An AI voice interaction-based pension service scheduling method and system

By integrating multimodal data fusion and spatiotemporal graph neural network decision-making, the shortcomings of existing home-based elderly care service systems in terms of the elderly's ambiguous expressions and dialect adaptability are addressed. This enables accurate identification of the elderly's needs and intelligent response to complex services, improving the efficiency of service resource scheduling and the ability to respond to emergency needs.

CN120748370BActive Publication Date: 2026-07-24HENGFENG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HENGFENG INFORMATION TECH CO LTD
Filing Date
2025-07-25
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing home-based elderly care service systems have poor adaptability to the ambiguous expressions and regional accents of the elderly in voice recognition, making it impossible to achieve proactive linkage between risk warning and services. This results in high-risk scenarios being missed and makes it difficult to cope with emergencies or complex service needs.

Method used

By combining multimodal data fusion and spatiotemporal graph neural network decision-making, user voice requests and environmental perception data are collected, and deep neural network multi-dialect speech recognition and intent parsing are performed. Combined with spatiotemporal consistency calibration processing, service type features and urgency features are extracted, a service decision graph is constructed, multi-granularity task decomposition is performed, and the final service task list is generated.

Benefits of technology

It enables accurate identification of the ambiguous needs of the elderly and intelligent response to complex service scenarios, improves the accuracy of matching service needs with environmental conditions, optimizes the scheduling efficiency of elderly care service resources, and ensures timely response to emergency needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748370B_ABST
    Figure CN120748370B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on AI voice interaction's old-age service scheduling method and system, method includes: user voice request and environmental perception data are collected, the joint processing of multi-dialect voice recognition and intention analysis based on deep neural network is carried out to voice request, obtains structured service demand;Environmental perception data are fused multi-sensor data and are handled with the spatiotemporal consistency calibration, obtain enhanced environmental data;Service type features and urgency degree features are extracted from structured service demand and are compensated based on attention mechanism prediction;User activity mode and environmental risk features are extracted from enhanced environmental data;Temporal-spatial graph neural network is used to construct service decision graph, realize demand-environment joint optimization;Finally, multi-granularity task decomposition generates final service task sheet.The present application realizes the accurate identification of the needs of the elderly and the intelligent response of the complex service scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent elderly care service technology, specifically to an elderly care service scheduling method and system based on AI voice interaction. Background Technology

[0002] With the accelerating aging of the population, the demand for home-based elderly care services is becoming increasingly complex, and existing systems face the core bottleneck of insufficient intelligence. Most current home-based elderly care service systems on the market are based on traditional telephone or online platforms, which suffer from inherent defects such as delayed response and low service matching accuracy. Although artificial intelligence technology has been gradually applied to this field in recent years, the voice recognition of these systems is poorly adapted to the ambiguous expressions and regional accents unique to the elderly, and cannot achieve proactive linkage between risk warnings and services. This can easily lead to the omission of high-risk scenarios and make it difficult to cope with emergencies or complex service needs. Summary of the Invention

[0003] In view of the above problems, the present invention provides an AI-based voice interaction-based elderly care service scheduling method and system. Through multimodal data fusion and spatiotemporal graph neural network decision-making, it realizes intelligent linkage analysis of voice demand and environmental risk, solving the problem that existing systems cannot accurately identify the ambiguous needs of the elderly and respond to complex service scenarios in a timely manner.

[0004] To achieve the above objectives, in a first aspect, the present invention provides a method for scheduling elderly care services based on AI voice interaction, comprising:

[0005] Collect user voice requests and environmental perception data. Voice requests include real-time voice signals and emotional feature parameters, while environmental perception data includes indoor and outdoor environmental status information collected by smart terminal devices.

[0006] The voice request undergoes a first preprocessing step to obtain structured service requirements. This first preprocessing step is configured as a joint processing of multi-dialect speech recognition and intent parsing based on a deep neural network.

[0007] In addition, the environmental perception data is subjected to a second preprocessing to obtain enhanced environmental data. The second preprocessing is configured as a spatiotemporal consistency calibration process that fuses multi-sensor data.

[0008] Service type features and urgency features are extracted from structured service requirements, and service type prediction compensation is performed on fuzzy requirements based on an attention mechanism to obtain compensated service requirements.

[0009] In addition, user activity pattern features and environmental risk features are extracted from enhanced environmental data;

[0010] A spatiotemporal graph neural network is used to construct a service decision graph by combining post-compensation service demand, user activity pattern characteristics and environmental risk characteristics, thereby obtaining enhanced service demand. The service decision graph is configured as a demand-environment joint optimization model.

[0011] The enhanced service requirements are decomposed into multi-granularity tasks to obtain the final service task list. The multi-granularity task decomposition is configured to generate end-to-end tasks that integrate voice intent, behavioral habits, and resource scheduling.

[0012] In some embodiments, the voice request undergoes a first preprocessing step to obtain a structured service requirement. This first preprocessing is configured as a joint processing of multi-dialect speech recognition and intent parsing based on a deep neural network, including:

[0013] The speech feature coding network extracts the time-frequency domain hybrid features of real-time speech signals. The speech feature coding network is configured as a dual-branch parallel network structure including a one-dimensional convolutional layer in the time domain and a two-dimensional convolutional layer in the frequency domain.

[0014] The time-frequency domain hybrid features are input into the dialect adaptation module, which outputs dialect-standardized speech features and dialect category identifiers. The dialect adaptation module is configured as a domain adversarial neural network based on a gradient inversion layer.

[0015] The dialect-standardized speech features are input into the intent parsing network to generate service type classification results and a set of demand parameters. The intent parsing network is configured as a hierarchical Transformer architecture based on a multi-head self-attention mechanism.

[0016] In addition, it processes sentiment feature parameters and outputs sentiment quantification indicators through a sentiment analysis network, which is configured as a temporal model containing long short-term memory units and attention pooling layers.

[0017] By integrating service type classification results, demand parameter sets, and sentiment quantification indicators, structured service requirements are generated, and the fusion processing is configured as a vector concatenation operation based on dynamic weight allocation.

[0018] In some embodiments, the environmental sensing data undergoes a second preprocessing to obtain enhanced environmental data. The second preprocessing is configured as a spatiotemporal consistency calibration process that fuses multi-sensor data, including:

[0019] The spatiotemporal alignment module performs timestamp synchronization and spatial coordinate unification of multi-source heterogeneous sensor data. The spatiotemporal alignment module is configured to perform time-series calibration based on dynamic time warping algorithm and spatial mapping based on coordinate system transformation matrix.

[0020] The calibrated multi-source heterogeneous sensor data is input into the feature extraction network to extract static environmental features and dynamic change features respectively. The feature extraction network is configured to have a static feature extraction branch containing three-dimensional convolutional kernels and a dynamic feature extraction branch containing recurrent neural units.

[0021] An attention fusion mechanism is used to weightedly fuse static and dynamic environmental features to generate a spatiotemporal correlation feature matrix. The attention fusion mechanism is configured as a feature importance evaluation model based on gated recurrent units.

[0022] Anomaly detection and missing data compensation are performed on the spatiotemporal correlation feature matrix to output enhanced environmental data. Anomaly detection is configured to identify outliers based on the isolated forest algorithm, and missing data compensation is configured to be a data reconstruction model based on generative adversarial networks.

[0023] In some embodiments, service type features and urgency features are extracted from structured service requirements, and service type prediction compensation is performed on fuzzy requirements based on an attention mechanism to obtain compensated service requirements, including:

[0024] The service feature extraction module parses structured service requirements and outputs a service type probability distribution vector and an urgency score. The service feature extraction module is configured as a classification and regression joint network containing a dual-channel fully connected layer.

[0025] The service type probability distribution vector is input into the fuzzy demand identifier to detect low-confidence service requests. The fuzzy demand identifier is configured as a probability distribution analyzer based on entropy threshold judgment.

[0026] Context-aware compensation is performed on the identified ambiguous demands to generate compensation service type features, including:

[0027] The temporal context features of historical service requests are extracted through a bidirectional LSTM network and denoted as historical features.

[0028] An attention mechanism is used to calculate the association weight between the current fuzzy demand and historical features, and a compensation prediction vector is generated, which is denoted as the compensation service type feature.

[0029] The enhanced service type features are obtained by fusing the compensated service type features with the original service type features, and the fusion process is configured as a feature interpolation algorithm based on confidence weighting.

[0030] The enhanced service type features are concatenated with the urgency score to form the compensated service requirements. The concatenation operation is configured as a vector merging process that includes a dimension alignment layer.

[0031] In some embodiments, extracting user activity pattern features and environmental risk features from enhanced environmental data includes:

[0032] The enhanced environment data is processed by the activity feature extraction module, which outputs user activity temporal features and environment state features. The activity feature extraction module is configured as a two-branch feature extraction network containing spatiotemporal convolutional layers.

[0033] User activity temporal features are input into a pattern recognizer to generate user activity pattern features. The pattern recognizer is configured as a sequence matching model based on a dynamic time warping algorithm, including:

[0034] The activity time-series features are segmented using a sliding window and denoted as window features;

[0035] Calculate the similarity score between each window feature and the preset activity template to generate user activity pattern features;

[0036] Furthermore, environmental state characteristics are input into a risk assessor, which outputs environmental risk characteristics. The risk assessor is configured as an anomaly detection network incorporating a multi-head attention mechanism, including:

[0037] The environmental state features are mapped to the risk feature space through the feature projection layer, and this is denoted as the initial risk feature.

[0038] An attention mechanism is used to calculate the contribution weight of each environmental dimension to the risk score, generating environmental risk characteristics;

[0039] Furthermore, the user activity pattern characteristics and environmental risk characteristics are normalized to obtain normalized user activity pattern characteristics and environmental risk characteristics. The normalization process is configured as a standardization algorithm based on sliding window statistics.

[0040] In some embodiments, a spatiotemporal graph neural network is used to construct a service decision graph from post-compensation service demand, user activity pattern characteristics, and environmental risk characteristics to obtain enhanced service demand. The service decision graph is configured as a demand-environment joint optimization model, including:

[0041] An initial decision graph is generated using the graph construction module, which is configured as follows:

[0042] The compensated service requirements are taken as requirement nodes, and the node characteristics include service type code and urgency value;

[0043] User activity pattern features are used as activity nodes, and the node features include an activity type identifier and a pattern strength vector.

[0044] Environmental risk characteristics are used as environmental nodes, and the node characteristics include risk type codes and risk level values;

[0045] The initial decision graph is enhanced with spatiotemporal features to obtain an enhanced decision graph, including:

[0046] The graph attention network is configured as a multi-head attention mechanism that includes spatiotemporal location encoding to calculate the association weights between demand nodes, activity nodes, and environment nodes.

[0047] The graph convolutional layer aggregates the node features of adjacent demand nodes, activity nodes, and environment nodes to generate enhanced node features. The graph convolutional layer is configured as a spectral graph convolution based on Chebyshev multinomials to obtain an enhanced decision graph.

[0048] The service optimizer processes the enhancement decision graph to generate enhancement service requirements, including:

[0049] The importance of the features of the augmentation nodes corresponding to the demand nodes in the augmentation decision graph is ranked to generate a service priority vector;

[0050] The enhanced node features corresponding to the activity nodes of the enhanced decision graph and the enhanced node features corresponding to the environment nodes of the enhanced decision graph are integrated to generate a service constraint matrix.

[0051] The service priority vector and service constraint matrix are jointly optimized to output enhanced service requirements containing the optimal service path.

[0052] In some embodiments, the association weights between demand nodes, activity nodes, and environment nodes are calculated using a graph attention network. The graph attention network is configured as a multi-head attention mechanism that includes spatiotemporal location encoding, comprising:

[0053] The spatiotemporal location features of the nodes are generated using a spatiotemporal encoder, which is configured as follows:

[0054] Spatial location codes are assigned to demand nodes, activity nodes, and environment nodes. These codes are generated based on the three-dimensional location vectors of the nodes' physical coordinates.

[0055] The demand node, activity node, and environment node are assigned time location codes, which are generated based on the sinusoidal location code of the event occurrence timestamp.

[0056] Spatial location codes and temporal location codes are concatenated to form a spatiotemporal location feature vector;

[0057] A multi-head attention mechanism is used to calculate the dynamic association weights between demand nodes, activity nodes, and environment nodes, including:

[0058] The attention head is obtained by adding the demand node, activity node, and environment node with the spatiotemporal location features;

[0059] Generate a query vector, key vector, and value vector for each attention head;

[0060] Calculate the query-key similarity score and overlay it with the spatiotemporal association constraint matrix;

[0061] The association weights of each attention head are obtained by normalization using the softmax function;

[0062] Aggregate the results of multi-head attention to generate the final node association weights, including:

[0063] Linear projection is performed on the output of each attention head;

[0064] A gating mechanism is used to fuse feature representations from different attention heads;

[0065] The output contains an edge weight matrix that includes spatiotemporal correlation information. The edge weight matrix includes the correlation weights between demand nodes, activity nodes, and environment nodes.

[0066] In some embodiments, the enhanced service requirements are decomposed into multi-granularity tasks to obtain a final service task list. This multi-granularity task decomposition is configured to generate end-to-end tasks that integrate voice intent, behavioral habits, and resource scheduling, including:

[0067] The task parsing module processes the enhanced service requirements and outputs a set of basic task units. The task parsing module is configured as a task decomposition network that includes a hierarchical attention mechanism.

[0068] The basic set of task units is input into the resource adapter to generate an executable task sequence. The resource adapter is configured as follows:

[0069] Information on currently available service resources is obtained through the resource status monitoring layer and recorded as resource characteristics;

[0070] A graph matching algorithm is used to calculate the fit score between task units and resource features;

[0071] A preliminary task allocation scheme is generated based on the suitability score;

[0072] Optimize the initial task allocation plan by incorporating behavioral habits, including:

[0073] Historical behavioral pattern features are extracted through the user profiling module and recorded as habit features;

[0074] The matching degree between the preliminary task allocation scheme and habitual features is calculated using a collaborative filtering algorithm, and is denoted as the behavior matching degree.

[0075] An optimized task allocation scheme is obtained by balancing resource suitability and behavior matching through a multi-objective optimization algorithm.

[0076] Generate the final service task order, including:

[0077] The optimized task allocation scheme is then time-series orchestrated.

[0078] Add execution constraints and quality evaluation metrics to each task unit;

[0079] The output includes a standardized task sheet containing the task path, resource binding, and execution sequence.

[0080] In a second aspect, the present invention also provides an elderly care service scheduling system based on AI voice interaction, which is applicable to the method described in the first aspect.

[0081] Unlike existing technologies, the above technical solution provides a method and system for scheduling elderly care services based on AI voice interaction. The method includes: collecting user voice requests and environmental perception data; performing joint processing of voice requests using multi-dialect speech recognition and intent parsing based on deep neural networks to obtain structured service requirements; performing spatiotemporal consistency calibration processing on the environmental perception data by fusing multi-sensor data to obtain enhanced environmental data; extracting service type features and urgency features from the structured service requirements and performing predictive compensation based on an attention mechanism; extracting user activity patterns and environmental risk features from the enhanced environmental data; constructing a service decision graph using a spatiotemporal graph neural network to achieve joint optimization of demand and environment; and finally, performing multi-granularity task decomposition to generate the final service task list. This invention achieves accurate identification of ambiguous needs of the elderly and intelligent response to complex service scenarios.

[0082] The above description of the invention is merely an overview of the technical solution of this application. In order to enable those skilled in the art to better understand the technical solution of this application and to implement it based on the description and drawings, and to make the above-mentioned objectives and other objectives, features and advantages of this application easier to understand, the following description is provided in conjunction with the specific embodiments and drawings of this application. Attached Figure Description

[0083] The accompanying drawings are only used to illustrate the principles, implementation methods, applications, features, and effects of specific embodiments of the present invention and other related contents, and should not be considered as limitations on this application.

[0084] In the accompanying drawings of the instruction manual:

[0085] Figure 1 This is a flowchart illustrating steps S101 to S105 of the scheduling method described in a specific implementation.

[0086] Figure 2 The method steps S201 to S204 of the scheduling method described in the specific implementation are shown in the figure. Detailed Implementation

[0087] To illustrate the possible application scenarios, technical principles, implementable specific solutions, and achievable objectives and effects of this application in detail, the following description, in conjunction with the listed specific embodiments and accompanying drawings, provides a detailed explanation. The embodiments described herein are merely illustrative of the technical solutions of this application and are therefore intended to limit the scope of protection of this application.

[0088] In this document, the term "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The term "embodiment" appearing in various places throughout the specification does not necessarily refer to the same embodiment, nor does it specifically limit its independence or connection with other embodiments. In principle, in this application, as long as there are no technical contradictions or conflicts, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.

[0089] Unless otherwise defined, the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the use of related terms herein is merely for the purpose of describing particular embodiments and is not intended to limit this application.

[0090] In the description of this application, the term "and / or" is used to describe the logical relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A exists, B exists, and A and B exist simultaneously. Additionally, the character " / " in this document generally indicates that the preceding and following objects have an "or" logical relationship.

[0091] In this application, terms such as “first” and “second” are used only to distinguish one entity or operation from another, and do not necessarily require or imply any actual quantity, hierarchy or order relationship between these entities or operations.

[0092] Without further limitations, the use of terms such as “comprising,” “including,” “having,” or other similar open-ended expressions in this application is intended to cover non-exclusive inclusion, which does not exclude the presence of additional elements in a process, method, or product that includes the stated elements, such that a process, method, or product that includes a list of elements may include not only those defined elements but also other elements not expressly listed, or elements inherent to such a process, method, or product.

[0093] Similar to the understanding in the Examination Guidelines, in this application, expressions such as "greater than," "less than," and "exceeding" are understood to exclude the stated number; expressions such as "above," "below," and "within" are understood to include the stated number. Furthermore, in the description of the embodiments in this application, "multiple" means two or more (including two), and similar expressions related to "multiple" are also understood in this way, such as "multiple groups" and "multiple times," unless otherwise explicitly specified.

[0094] In the description of the embodiments of this application, the space-related expressions used, such as "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "vertical," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential," indicate the orientation or positional relationship based on the orientation or positional relationship shown in the specific embodiments or drawings. They are only for the purpose of describing the specific embodiments of this application or for the reader's understanding, and do not indicate or imply that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this application.

[0095] Please see Figure 1 In a first aspect, this embodiment provides a method for scheduling elderly care services based on AI voice interaction, including:

[0096] S101. Collect user voice requests and environmental perception data. The voice requests include real-time voice signals and emotional feature parameters. The environmental perception data includes indoor and outdoor environmental status information collected by smart terminal devices.

[0097] S102. Perform a first preprocessing on the voice request to obtain structured service requirements. The first preprocessing is configured as a joint processing of multi-dialect speech recognition and intent parsing based on a deep neural network.

[0098] In addition, the environmental perception data is subjected to a second preprocessing to obtain enhanced environmental data. The second preprocessing is configured as a spatiotemporal consistency calibration process that fuses multi-sensor data.

[0099] S103. Extract service type features and urgency features from structured service requirements, and perform service type prediction compensation for fuzzy requirements based on attention mechanism to obtain compensated service requirements.

[0100] In addition, user activity pattern features and environmental risk features are extracted from enhanced environmental data;

[0101] S104. A service decision graph is constructed by using a spatiotemporal graph neural network to combine the compensated service demand, user activity pattern characteristics and environmental risk characteristics, thereby obtaining enhanced service demand. The service decision graph is configured as a demand-environment joint optimization model.

[0102] S105. Perform multi-granularity task decomposition on the enhanced service requirements to obtain the final service task list. The multi-granularity task decomposition is configured as an end-to-end task generation that integrates voice intent, behavioral habits, and resource scheduling.

[0103] In step S101, the user's voice request is collected through a smart terminal device equipped with a voice recognition module. This device employs multi-dialect voice processing technology, which can accurately identify the unique pronunciation characteristics and ambiguous expressions of the elderly. Real-time voice signals are acquired through a high-sensitivity microphone array, and emotional feature parameters are extracted by analyzing acoustic features such as intonation and speech rate. Environmental perception data is collected by a distributed sensor network deployed in the residence, including IoT devices such as temperature and humidity sensors and infrared sensors. These devices use spatial positioning technology to accurately acquire indoor and outdoor environmental status information. The voice emotion analysis module has a built-in emotion recognition model that can extract emotional feature parameters from the voice signal.

[0104] In step S102, the first preprocessing uses a deep neural network that integrates speech recognition and intent parsing functions. Through joint training and optimization, it achieves accurate understanding of the elderly's voice requests. The speech recognition module is specifically trained to strengthen common dialects, while the intent parsing module establishes a mapping relationship between service needs and semantic expressions. The second preprocessing uses a spatiotemporal calibration algorithm to fuse multi-source sensor data and adopts an adaptive sliding window mechanism to process data streams with different sampling frequencies to ensure the spatiotemporal consistency of environmental state representation. Spatial calibration is based on device positioning information, while time calibration is completed through unified clock synchronization.

[0105] In step S103, preferably, the service type feature extraction module has a built-in domain knowledge base, supporting automatic classification of common service types such as medical and daily life services. The classification process comprehensively considers keywords and sentence structure features in the voice text. The urgency assessment model performs multi-dimensional scoring by integrating emotional feature parameters and semantic keywords, where emotional feature parameters reflect the user's emotional state and semantic keywords indicate time sensitivity. An attention mechanism constructs a knowledge graph of the user's service history and achieves context-aware demand compensation through a graph neural network. This mechanism can identify potential associations between historical service records and current requests. User activity pattern features are derived by analyzing movement trajectories and device usage records in environmental sensor data, while environmental risk features identify abnormal environmental states based on preset safety thresholds.

[0106] In step S104, the service decision graph construction module employs a dynamic graph representation learning method. Node features include three types of information: compensated service demand, user activity pattern features, and environmental risk features. Edge weights are obtained through spatiotemporal correlation learning, reflecting the strength of the correlation between features. The demand-environment joint optimization model integrates graph attention mechanisms and reinforcement learning algorithms to achieve intelligent decision-making in service resource allocation. The graph attention mechanism dynamically adjusts the contribution of each feature, while the reinforcement learning algorithm optimizes long-term service performance.

[0107] In step S105, preferably, the multi-granularity task decomposition adopts a hierarchical processing architecture. The top layer is classified according to service domain, and the bottom layer is refined into executable atomic operations. The decomposition process comprehensively considers the results of voice intent understanding, user behavior habit analysis, and resource scheduling constraints. The end-to-end task generation system outputs the final service task list, which includes dynamic priority marking and optimal execution path planning. The priority is dynamically adjusted according to urgency characteristics and environmental risk characteristics, and the execution path is determined based on a resource scheduling optimization algorithm.

[0108] This embodiment collects voice requests containing emotional feature parameters and indoor / outdoor environmental status information. It then uses a deep neural network to jointly process multi-dialect speech recognition and intent parsing. Combined with spatiotemporally consistent calibrated environmental data, it constructs a service decision graph that integrates service type features, urgency features, user activity pattern features, and environmental risk features. This allows for accurate understanding of the ambiguous needs expressed by the elderly, context-aware demand compensation through an attention mechanism, and demand-environment joint optimization based on a spatiotemporal graph neural network. Ultimately, it generates a service task list that comprehensively considers voice intent, behavioral habits, and resource scheduling. This embodiment significantly improves the accuracy of recognizing elderly people's voice requests, achieves intelligent matching of service needs with environmental conditions, optimizes the scheduling efficiency of elderly care service resources, and ensures timely response to urgent needs through a dynamic priority adjustment mechanism.

[0109] Please see Figure 2 In some embodiments, the voice request undergoes a first preprocessing step to obtain a structured service requirement. This first preprocessing is configured as a joint processing of multi-dialect speech recognition and intent parsing based on a deep neural network, including:

[0110] S201. Extract time-frequency domain hybrid features of real-time speech signals through a speech feature coding network. The speech feature coding network is configured as a dual-branch parallel network structure including a one-dimensional convolutional layer in the time domain and a two-dimensional convolutional layer in the frequency domain.

[0111] S202. Input the time-frequency domain hybrid features into the dialect adaptation module, and output the dialect standardized speech features and dialect category identifier. The dialect adaptation module is configured as a domain adversarial neural network based on the gradient inversion layer.

[0112] S203. Input the dialect-standardized speech features into the intent parsing network to generate service type classification results and a set of demand parameters. The intent parsing network is configured as a hierarchical Transformer architecture based on a multi-head self-attention mechanism.

[0113] In addition, it processes sentiment feature parameters and outputs sentiment quantification indicators through a sentiment analysis network, which is configured as a temporal model containing long short-term memory units and attention pooling layers.

[0114] S204. By integrating the service type classification results, the set of demand parameters, and the sentiment quantification indicators, structured service requirements are generated. The fusion processing is configured as a vector concatenation operation based on dynamic weight allocation.

[0115] In step S201, the speech feature coding network extracts time-frequency domain hybrid features of the real-time speech signal through a dual-branch parallel structure consisting of a one-dimensional time-domain convolutional layer and a two-dimensional frequency-domain convolutional layer. The one-dimensional time-domain convolutional layer captures the dynamic temporal patterns of the speech signal, while the two-dimensional frequency-domain convolutional layer analyzes the spatial distribution characteristics of the speech spectrum. The dual-branch structure achieves complementary fusion of time-domain and frequency-domain information through feature concatenation. These time-frequency domain hybrid features serve as input to the dialect adaptation module, providing a feature representation containing complete speech characteristics for subsequent processing.

[0116] In step S202, the dialect adaptation module performs dialect feature standardization based on a domain adversarial neural network with a gradient inversion layer. This module eliminates the impact of dialect differences on speech recognition through adversarial training while preserving key features of semantic content. The dialect-standardized speech feature output is a unified feature representation with dialect invariance, while the dialect category identifier is used to record the dialect attribute information of the speech. Together, they provide standardized input for intent parsing.

[0117] In step S203, the intent parsing network employs a hierarchical Transformer architecture based on a multi-head self-attention mechanism to process dialect-standardized speech features. It captures long-distance dependencies in speech features through self-attention, and the hierarchical structure achieves a gradual abstraction from local features to global semantics. The service type classification result identifies the service category of the voice request, and the requirement parameter set contains the specific parameters required for service execution. The sentiment analysis network models the temporal changes of sentiment feature parameters through long short-term memory units, and the attention pooling layer focuses on key sentiment segments. The output sentiment quantification index is used to assess the user's emotional state.

[0118] In step S204, the vector concatenation operation with dynamic weight allocation automatically adjusts the fusion weights of each input feature according to the service type. The service type classification result determines the basic service framework, the requirement parameter set provides specific execution parameters, and the sentiment quantification index affects service priority and response method. The structured service requirements ultimately generate a standardized representation containing complete service information, providing accurate input for subsequent service scheduling.

[0119] This embodiment achieves accurate understanding of voice requests through multi-stage feature processing. First, it extracts the time-frequency mixed features of the speech, then performs dialect standardization processing, and finally combines intent parsing and sentiment analysis to generate structured service requirements. This effectively addresses challenges such as dialect differences, ambiguous expressions, and emotional factors in elderly people's voice interactions, ensuring that service requirements are accurately translated into executable tasks. For example, when an elderly person expresses discomfort and needs help in a dialect, the system can accurately identify dialect features, understand the intent to seek medical help, and assess the urgency based on anxiety levels, ultimately generating a structured medical service request.

[0120] In some embodiments, the environmental sensing data undergoes a second preprocessing to obtain enhanced environmental data. The second preprocessing is configured as a spatiotemporal consistency calibration process that fuses multi-sensor data, including:

[0121] The spatiotemporal alignment module performs timestamp synchronization and spatial coordinate unification of multi-source heterogeneous sensor data. The spatiotemporal alignment module is configured to perform time-series calibration based on dynamic time warping algorithm and spatial mapping based on coordinate system transformation matrix.

[0122] The calibrated multi-source heterogeneous sensor data is input into the feature extraction network to extract static environmental features and dynamic change features respectively. The feature extraction network is configured to have a static feature extraction branch containing three-dimensional convolutional kernels and a dynamic feature extraction branch containing recurrent neural units.

[0123] An attention fusion mechanism is used to weightedly fuse static and dynamic environmental features to generate a spatiotemporal correlation feature matrix. The attention fusion mechanism is configured as a feature importance evaluation model based on gated recurrent units.

[0124] Anomaly detection and missing data compensation are performed on the spatiotemporal correlation feature matrix to output enhanced environmental data. Anomaly detection is configured to identify outliers based on the isolated forest algorithm, and missing data compensation is configured to be a data reconstruction model based on generative adversarial networks.

[0125] In this embodiment, the spatiotemporal alignment module achieves timestamp synchronization of multi-source heterogeneous sensor data through a dynamic time warping algorithm, enabling adaptive processing of sensor data streams with different sampling frequencies. Spatial mapping based on the coordinate system transformation matrix unifies the data from various sensors distributed indoors to a standard coordinate system, ensuring spatial consistency. The calibrated multi-source heterogeneous sensor data provides spatiotemporal alignment input for subsequent feature extraction.

[0126] The static feature extraction branch of the feature extraction network, consisting of 3D convolutional kernels, captures the spatial layout features of the environment, such as fixed attributes like furniture positions and door / window status. The dynamic feature extraction branch, containing recurrent neural units, analyzes the temporal variation patterns of sensor data, such as temperature fluctuations and human movement trajectories. The two branches process data in parallel, outputting static features characterizing the inherent properties of the environment and dynamic features reflecting real-time changes, respectively.

[0127] The attention fusion mechanism evaluates the relative importance of static and dynamic features through a gated recurrent unit, generating a feature matrix with spatiotemporal correlation. This mechanism can automatically adjust feature weights based on the current environmental state; for example, it enhances the contribution of dynamic features when abnormal temperature changes are detected. The spatiotemporally correlated feature matrix comprehensively represents the complete state of the environment, providing multi-dimensional analytical basis for anomaly detection.

[0128] Anomaly detection employs the Isolation Forest algorithm to identify outliers in the spatiotemporal correlation feature matrix, and detects anomalous combinations of environmental states by constructing a random segmentation tree. Missing data compensation utilizes a generative adversarial network to reconstruct data lost due to sensor failures or communication interruptions, ensuring the continuity of environmental monitoring. The final output of the enhanced environmental data is a reliable environmental characterization that has been calibrated, fused, and repaired.

[0129] This embodiment achieves accurate fusion of environmental perception data through multi-level processing, unifying the spatiotemporal benchmark of multi-source data, extracting static and dynamic features, realizing intelligent fusion through an attention mechanism, and finally optimizing data quality. This embodiment can effectively solve the fusion challenge caused by the heterogeneity of multi-sensor data in elderly care service scenarios, providing accurate environmental status input for service decisions. For example, when the system detects an abnormally high bedroom temperature and that doors and windows have not been opened for a long time, it can accurately identify this risk combination and trigger corresponding care services.

[0130] In some embodiments, service type features and urgency features are extracted from structured service requirements, and service type prediction compensation is performed on fuzzy requirements based on an attention mechanism to obtain compensated service requirements, including:

[0131] The service feature extraction module parses structured service requirements and outputs a service type probability distribution vector and an urgency score. The service feature extraction module is configured as a classification and regression joint network containing a dual-channel fully connected layer.

[0132] The service type probability distribution vector is input into the fuzzy demand identifier to detect low-confidence service requests. The fuzzy demand identifier is configured as a probability distribution analyzer based on entropy threshold judgment.

[0133] Context-aware compensation is performed on the identified ambiguous demands to generate compensation service type features, including:

[0134] The temporal context features of historical service requests are extracted through a bidirectional LSTM network and denoted as historical features.

[0135] An attention mechanism is used to calculate the association weight between the current fuzzy demand and historical features, and a compensation prediction vector is generated, which is denoted as the compensation service type feature.

[0136] The enhanced service type features are obtained by fusing the compensated service type features with the original service type features, and the fusion process is configured as a feature interpolation algorithm based on confidence weighting.

[0137] The enhanced service type features are concatenated with the urgency score to form the compensated service requirements. The concatenation operation is configured as a vector merging process that includes a dimension alignment layer.

[0138] In this embodiment, the service feature extraction module processes structured service requirements through a classification and regression joint network composed of dual-channel fully connected layers. It can simultaneously extract two key features, service category and urgency, from the standardized input. The classification channel outputs a service type probability distribution vector, representing the probability that the request belongs to a certain type of service; the regression channel outputs an urgency score, quantifying the timeliness requirements of the request.

[0139] The fuzzy demand identifier analyzes the probability distribution vector of service types based on the entropy threshold determination method. When the entropy value of the probability distribution exceeds the preset threshold, it is determined to be a low-confidence service request, which effectively captures the fuzzy demands caused by the unclear expression of the elderly and provides a trigger signal for subsequent compensation processing.

[0140] Context-aware compensation extracts temporal patterns from historical service requests using a bidirectional LSTM network, generating historical features that include user behavior habits. An attention mechanism dynamically calculates the correlation between current ambiguous requests and historical features, producing a compensation prediction vector that reflects potential user needs. Furthermore, compensation service type features compensate for the deficiencies in current ambiguous representations by capturing patterns in historical service records.

[0141] The confidence-weighted feature interpolation algorithm adaptively adjusts the fusion ratio of the compensated service type features based on the certainty of the original service type features, outputting enhanced service type features that retain current request information while supplementing contextual knowledge. The vector merging process in the dimension alignment layer ensures the standardized concatenation of the enhanced service type features and the urgency score, forming a complete post-compensation service requirement.

[0142] This embodiment achieves intelligent compensation for ambiguous service needs of the elderly through multi-stage processing. First, it accurately extracts basic service features, then identifies low-confidence requests, and then combines historical context for predictive compensation. Finally, it generates an accurate and reliable representation of service needs. For example, when the elderly express themselves vaguely, it combines their historical medical help records to accurately compensate for the corresponding medical service needs, effectively solving the problem of service identification caused by unclear expression by the elderly.

[0143] In some embodiments, extracting user activity pattern features and environmental risk features from enhanced environmental data includes:

[0144] The enhanced environment data is processed by the activity feature extraction module, which outputs user activity temporal features and environment state features. The activity feature extraction module is configured as a two-branch feature extraction network containing spatiotemporal convolutional layers.

[0145] User activity temporal features are input into a pattern recognizer to generate user activity pattern features. The pattern recognizer is configured as a sequence matching model based on a dynamic time warping algorithm, including:

[0146] The activity time-series features are segmented using a sliding window and denoted as window features;

[0147] Calculate the similarity score between each window feature and the preset activity template to generate user activity pattern features;

[0148] Furthermore, environmental state characteristics are input into a risk assessor, which outputs environmental risk characteristics. The risk assessor is configured as an anomaly detection network incorporating a multi-head attention mechanism, including:

[0149] The environmental state features are mapped to the risk feature space through the feature projection layer, and this is denoted as the initial risk feature.

[0150] An attention mechanism is used to calculate the contribution weight of each environmental dimension to the risk score, generating environmental risk characteristics;

[0151] Furthermore, the user activity pattern characteristics and environmental risk characteristics are normalized to obtain normalized user activity pattern characteristics and environmental risk characteristics. The normalization process is configured as a standardization algorithm based on sliding window statistics.

[0152] In this embodiment, the activity feature extraction module processes the enhanced environment data through a dual-branch network composed of spatiotemporal convolutional layers, separating two key types of information from the calibrated environment data: user activity and environment state. The temporal convolutional branch extracts the temporal variation features of user activity, while the spatial convolutional branch captures the distribution characteristics of the environment state. The dual-branch parallel processing outputs a multi-dimensional feature representation containing user behavior and environment state.

[0153] The pattern recognizer employs a dynamic time warping algorithm to analyze the temporal characteristics of user activities. It segments continuous activity sequences into processable segments using a sliding window and calculates the similarity score between the features of each window and a preset activity template. This pattern recognizer can adapt to individual differences in the activity rhythms of the elderly, accurately identifying the pattern features of daily activities such as getting up and eating. The preset activity templates are trained using typical behavioral samples and cover common activity types in elderly care service scenarios.

[0154] The risk assessor transforms environmental state features into a risk feature space through a feature projection layer, and then uses a multi-head attention mechanism to analyze the risk contribution of each environmental dimension. This risk assessor can identify combinations of risk factors such as slippery ground and abnormally high temperatures, outputting comprehensive environmental risk characteristics. The attention mechanism enables the model to focus on the most dangerous environmental factors, improving the relevance of the risk assessment.

[0155] A standardization algorithm based on sliding window statistics normalizes user activity pattern features and environmental risk features, eliminating the influence of different feature units and numerical ranges while preserving the relative relationships of the original features. This ensures that subsequent modules can fairly utilize various feature information. The sliding window statistics dynamically adapt to changes in feature distribution, maintaining the real-time nature of the normalization effect.

[0156] This embodiment achieves a deep understanding of the environment and user behavior through hierarchical feature extraction. First, it separates activity and environmental features, then performs pattern recognition and risk assessment separately, ultimately outputting standardized feature representations. This embodiment can accurately capture the daily activity patterns of the elderly and environmental safety hazards, such as identifying the high-risk behavioral pattern of "prolonged bed rest" combined with the environmental risk of "insufficient nighttime lighting," providing a basis for decision-making in smart elderly care services.

[0157] In some embodiments, a spatiotemporal graph neural network is used to construct a service decision graph from post-compensation service demand, user activity pattern characteristics, and environmental risk characteristics to obtain enhanced service demand. The service decision graph is configured as a demand-environment joint optimization model, including:

[0158] An initial decision graph is generated using the graph construction module, which is configured as follows:

[0159] The compensated service requirements are taken as requirement nodes, and the node characteristics include service type code and urgency value;

[0160] User activity pattern features are used as activity nodes, and the node features include an activity type identifier and a pattern strength vector.

[0161] Environmental risk characteristics are used as environmental nodes, and the node characteristics include risk type codes and risk level values;

[0162] The initial decision graph is enhanced with spatiotemporal features to obtain an enhanced decision graph, including:

[0163] The graph attention network is configured as a multi-head attention mechanism that includes spatiotemporal location encoding to calculate the association weights between demand nodes, activity nodes, and environment nodes.

[0164] The graph convolutional layer aggregates the node features of adjacent demand nodes, activity nodes, and environment nodes to generate enhanced node features. The graph convolutional layer is configured as a spectral graph convolution based on Chebyshev multinomials to obtain an enhanced decision graph.

[0165] The service optimizer processes the enhancement decision graph to generate enhancement service requirements, including:

[0166] The importance of the features of the augmentation nodes corresponding to the demand nodes in the augmentation decision graph is ranked to generate a service priority vector;

[0167] The enhanced node features corresponding to the activity nodes of the enhanced decision graph and the enhanced node features corresponding to the environment nodes of the enhanced decision graph are integrated to generate a service constraint matrix.

[0168] The service priority vector and service constraint matrix are jointly optimized to output enhanced service requirements containing the optimal service path.

[0169] In this embodiment, the graph construction module converts three types of key features into graph nodes through structured processing: demand nodes represent the core attributes of service requests through service type encoding and urgency values; activity nodes record user behavior characteristics through activity type identifiers and pattern strength vectors; and environment nodes describe the environmental security status through risk type encoding and risk level values. The final constructed initial decision graph provides a topological foundation for subsequent joint optimization.

[0170] The spatiotemporal feature enhancement process computes dynamic relationships between nodes through a multi-head attention mechanism that incorporates spatiotemporal location encoding, enabling the graph attention network to simultaneously capture spatial proximity and temporal relevance. Chebyshev multinomial-based spectral graph convolution achieves intelligent aggregation of local neighborhood features, generating enhanced node features that retain original characteristics while incorporating contextual information. This processing allows the decision graph to reflect the complex interactions between needs, activities, and the environment.

[0171] The service optimizer analyzes the enhancement features of demand nodes using an importance ranking algorithm to generate a priority vector reflecting service urgency. Simultaneously, it integrates the enhancement features of active and environmental nodes to construct a service constraint matrix representing implementation limitations. The joint optimization process comprehensively considers service priority and implementation feasibility, ultimately outputting enhanced service requirements containing the optimal service path, which balances service urgency with execution safety.

[0172] This embodiment achieves the organic integration of multi-source information through graph structure modeling. First, a basic decision graph is constructed, then the spatiotemporal correlation between nodes is enhanced, and finally, multi-objective optimization decision-making is carried out. It can intelligently coordinate service needs and implementation conditions. For example, when it is found that "emergency medical needs" and "high-risk fall environment" coexist, an optimal service plan that includes emergency calls and fall prevention assistance is generated in parallel to ensure that core needs are safely met in risky environments.

[0173] In some embodiments, the association weights between demand nodes, activity nodes, and environment nodes are calculated using a graph attention network. The graph attention network is configured as a multi-head attention mechanism that includes spatiotemporal location encoding, comprising:

[0174] The spatiotemporal location features of the nodes are generated using a spatiotemporal encoder, which is configured as follows:

[0175] Spatial location codes are assigned to demand nodes, activity nodes, and environment nodes. These codes are generated based on the three-dimensional location vectors of the nodes' physical coordinates.

[0176] The demand node, activity node, and environment node are assigned time location codes, which are generated based on the sinusoidal location code of the event occurrence timestamp.

[0177] Spatial location codes and temporal location codes are concatenated to form a spatiotemporal location feature vector;

[0178] A multi-head attention mechanism is used to calculate the dynamic association weights between demand nodes, activity nodes, and environment nodes, including:

[0179] The attention head is obtained by adding the demand node, activity node, and environment node with the spatiotemporal location features;

[0180] Generate a query vector, key vector, and value vector for each attention head;

[0181] Calculate the query-key similarity score and overlay it with the spatiotemporal association constraint matrix;

[0182] The association weights of each attention head are obtained by normalization using the softmax function;

[0183] Aggregate the results of multi-head attention to generate the final node association weights, including:

[0184] Linear projection is performed on the output of each attention head;

[0185] A gating mechanism is used to fuse feature representations from different attention heads;

[0186] The output contains an edge weight matrix that includes spatiotemporal correlation information. The edge weight matrix includes the correlation weights between demand nodes, activity nodes, and environment nodes.

[0187] In this embodiment, the spatiotemporal encoder generates spatial location codes using three-dimensional location vectors to accurately represent the distribution relationships of demand nodes, activity nodes, and environmental nodes in physical space. Simultaneously, it generates temporal location codes based on sinusoidal location codes, effectively capturing the temporal sequence characteristics of events. The features from the spatial and temporal location codes are concatenated to form a complete spatiotemporal location feature vector, providing a spatiotemporal benchmark for subsequent correlation analysis.

[0188] The multi-head attention mechanism forms an attention head by adding node features to spatiotemporal location features. Each attention head independently generates a query vector, key vector, and value vector, enabling multi-perspective feature interaction analysis. The query-key similarity score, superimposed on the spatiotemporal association constraint matrix, is then normalized using softmax to generate dynamic association weights reflecting spatiotemporal dependencies. This mechanism can adaptively capture node association patterns at different spatiotemporal scales.

[0189] The aggregation process of multi-head attention results is achieved through linear projection and gating mechanisms. Linear projection unifies the feature representation space, while the gating mechanism dynamically adjusts the contribution ratio of different attention heads, ultimately outputting an edge weight matrix containing complete spatiotemporal correlation information. This matrix accurately quantifies the interaction strength between demand nodes, activity nodes, and environment nodes, providing a relational basis for service decisions.

[0190] This embodiment achieves accurate modeling of node associations through spatiotemporal coding and multi-head attention mechanisms. First, it constructs spatiotemporal baseline features, then analyzes multi-dimensional interaction relationships, and finally generates a quantified association matrix. This technical solution can intelligently identify the spatiotemporal correlation between service demands and user activities and environmental states, providing crucial evidence for service optimization. By accurately modeling the spatiotemporal relationships between nodes, this embodiment enables service decisions to fully consider the spatiotemporal context of demand occurrence, improving the accuracy and timeliness of service matching, while also enhancing the system's adaptability to complex environmental changes.

[0191] In some embodiments, the enhanced service requirements are decomposed into multi-granularity tasks to obtain a final service task list. This multi-granularity task decomposition is configured to generate end-to-end tasks that integrate voice intent, behavioral habits, and resource scheduling, including:

[0192] The task parsing module processes the enhanced service requirements and outputs a set of basic task units. The task parsing module is configured as a task decomposition network that includes a hierarchical attention mechanism.

[0193] The basic set of task units is input into the resource adapter to generate an executable task sequence. The resource adapter is configured as follows:

[0194] Information on currently available service resources is obtained through the resource status monitoring layer and recorded as resource characteristics;

[0195] A graph matching algorithm is used to calculate the fit score between task units and resource features;

[0196] A preliminary task allocation scheme is generated based on the suitability score;

[0197] Optimize the initial task allocation plan by incorporating behavioral habits, including:

[0198] Historical behavioral pattern features are extracted through the user profiling module and recorded as habit features;

[0199] The matching degree between the preliminary task allocation scheme and habitual features is calculated using a collaborative filtering algorithm, and is denoted as the behavior matching degree.

[0200] An optimized task allocation scheme is obtained by balancing resource suitability and behavior matching through a multi-objective optimization algorithm.

[0201] Generate the final service task order, including:

[0202] The optimized task allocation scheme is then time-series orchestrated.

[0203] Add execution constraints and quality evaluation metrics to each task unit;

[0204] The output includes a standardized task sheet containing the task path, resource binding, and execution sequence.

[0205] In this embodiment, the task parsing module decomposes enhanced service requirements into a set of basic task units through a hierarchical attention mechanism. This allows it to identify the core and auxiliary elements within the service requirements, achieving a precise conversion from macro-level service objectives to micro-level operational units. The hierarchical attention mechanism captures the inclusion relationships between tasks through multi-level feature extraction, ensuring that the decomposed task units retain the integrity of the original requirements.

[0206] The resource adapter obtains service resource information in real time through the resource status monitoring layer, and uses a graph matching algorithm to quantify the suitability between task units and resources. It comprehensively considers dimensions such as resource type, location distribution, and service capabilities to generate a preliminary allocation scheme that meets both technical requirements and the current resource situation. The graph matching algorithm calculates the optimal task-resource correspondence by constructing a bipartite graph model, thereby maximizing resource utilization.

[0207] The behavior habit optimization process extracts historical behavior pattern features through the user profiling module and uses a collaborative filtering algorithm to evaluate the fit between the task allocation scheme and user habits. A multi-objective optimization algorithm seeks a balance between resource suitability and behavior matching to ensure that the final solution is both efficient and feasible while conforming to user preferences. Preferably, this optimization process uses Pareto front analysis to determine the optimal trade-off point, retaining non-dominated solutions that meet the constraints.

[0208] This embodiment achieves refined decomposition of service requirements through multi-stage processing. First, it analyzes the requirement structure; then, it matches resource capabilities; and finally, it incorporates user habits to generate a standardized task list containing complete execution elements. For example, "health check service" is decomposed into atomic tasks such as "booking a doctor" and "preparing equipment," and the execution sequence is optimized based on medical staff scheduling and the elderly's daily routines. This embodiment, through hierarchical task decomposition and resource-habit dual-objective optimization, ensures that the service task list not only conforms to actual resource conditions but also respects user behavior patterns, improving task execution efficiency while enhancing service acceptance, and achieving precise matching between the supply and demand of elderly care services.

[0209] In a second aspect, this embodiment also provides an elderly care service scheduling system based on AI voice interaction, which is applicable to the method described in the first aspect.

[0210] Unlike existing technologies, the above technical solution has the following beneficial effects: This invention achieves precision and personalization in elderly care service scheduling through multimodal data fusion and intelligent decision-making technology. In terms of voice interaction, it employs multi-dialect speech recognition and intent parsing for joint processing, effectively solving the problems of ambiguous expression and dialect differences among the elderly. Regarding environmental perception, it achieves accurate identification of user activity patterns and environmental risks through spatiotemporal consistency calibration of multi-sensor data. In terms of service decision-making, it constructs a spatiotemporal graph neural network that integrates needs, activities, and environmental characteristics, achieving intelligent matching of service needs with environmental conditions. In terms of task decomposition, it adopts a hierarchical attention mechanism and a multi-objective optimization algorithm to generate service task lists that both meet resource conditions and respect user habits.

[0211] This invention significantly improves the accuracy of recognizing elderly people's voice requests. Through emotional feature analysis and context-aware compensation, it can accurately understand ambiguous expressions. It achieves dynamic matching of service needs with environmental conditions, optimizing resource allocation efficiency by modeling complex relationships through spatiotemporal graph neural networks. It also improves the execution effect of service tasks, ensuring that service solutions are both efficient and feasible while conforming to user preferences through multi-granularity decomposition and behavioral habit optimization. These technical solutions construct a complete intelligent chain from needs understanding to task execution, effectively improving the response speed and quality of elderly care services.

[0212] Finally, it should be noted that although the above embodiments have been described in the text and drawings of this application, this should not limit the scope of patent protection of this application. Any technical solutions that are based on the essential concept of this application and utilize the content described in the text and drawings of this application, resulting in equivalent structural or procedural substitutions or modifications, as well as the direct or indirect application of the technical solutions of the above embodiments to other related technical fields, are all included within the scope of patent protection of this application.

Claims

1. A method for scheduling elderly care services based on AI voice interaction, characterized in that, include: Collect user voice requests and environmental perception data. The voice requests include real-time voice signals and emotional feature parameters, and the environmental perception data includes indoor and outdoor environmental status information collected by smart terminal devices. The voice request is subjected to a first preprocessing to obtain a structured service requirement. The first preprocessing is configured as a joint processing of multi-dialect speech recognition and intent parsing based on a deep neural network. In addition, the environmental perception data is subjected to a second preprocessing to obtain enhanced environmental data, wherein the second preprocessing is configured as a spatiotemporal consistency calibration process for fusing multi-sensor data; Service type features and urgency features are extracted from the structured service requirements, and service type prediction compensation is performed on fuzzy requirements based on an attention mechanism to obtain compensated service requirements. In addition, user activity pattern features and environmental risk features are extracted from the enhanced environmental data; A service decision graph is constructed by using a spatiotemporal graph neural network to combine the compensated service demand, user activity pattern characteristics and environmental risk characteristics to obtain enhanced service demand. The service decision graph is configured as a demand-environment joint optimization model. The enhanced service requirements are decomposed into multi-granularity tasks to obtain the final service task list. The multi-granularity task decomposition is configured to generate end-to-end tasks that integrate voice intent, behavioral habits, and resource scheduling.

2. The method for scheduling elderly care services based on AI voice interaction according to claim 1, characterized in that, The voice request undergoes a first preprocessing step to obtain a structured service requirement. This first preprocessing step is configured as a joint processing of multi-dialect speech recognition and intent parsing based on a deep neural network, including: The speech feature coding network extracts the time-frequency domain hybrid features of the real-time speech signal through a speech feature coding network, which is configured as a dual-branch parallel network structure including a one-dimensional convolutional layer in the time domain and a two-dimensional convolutional layer in the frequency domain. The time-frequency domain hybrid features are input into the dialect adaptation module, which outputs dialect-standardized speech features and dialect category identifiers. The dialect adaptation module is configured as a domain adversarial neural network based on a gradient inversion layer. The dialect-standardized speech features are input into the intent parsing network to generate service type classification results and a set of demand parameters. The intent parsing network is configured as a hierarchical Transformer architecture based on a multi-head self-attention mechanism. Furthermore, the emotional feature parameters are processed, and an emotional quantification index is output through an emotional analysis network, which is configured as a temporal model containing long short-term memory units and attention pooling layers. The structured service requirements are generated by integrating the service type classification results, the set of requirement parameters, and the sentiment quantification indicators. The integration is configured as a vector concatenation operation based on dynamic weight allocation.

3. The method for scheduling elderly care services based on AI voice interaction according to claim 1, characterized in that, The environmental perception data undergoes a second preprocessing step to obtain enhanced environmental data. This second preprocessing is configured as a spatiotemporal consistency calibration process that fuses multi-sensor data, including: The spatiotemporal alignment module performs timestamp synchronization and spatial coordinate unification of multi-source heterogeneous sensor data. The spatiotemporal alignment module is configured to perform time-series calibration based on dynamic time warping algorithm and spatial mapping based on coordinate system transformation matrix. The calibrated multi-source heterogeneous sensor data is input into the feature extraction network to extract static environmental features and dynamic change features respectively. The feature extraction network is configured to include a static feature extraction branch containing three-dimensional convolutional kernels and a dynamic feature extraction branch containing recurrent neural units. An attention fusion mechanism is used to weightedly fuse the static and dynamic features of the environment to generate a spatiotemporal correlation feature matrix. The attention fusion mechanism is configured as a feature importance evaluation model based on a gated recurrent unit. The spatiotemporal correlation feature matrix is ​​subjected to anomaly detection and missing data compensation to output the enhanced environment data. The anomaly detection is configured as outlier identification based on the isolated forest algorithm, and the missing data compensation is configured as a data reconstruction model based on generative adversarial networks.

4. The method for scheduling elderly care services based on AI voice interaction according to claim 1, characterized in that, Service type features and urgency features are extracted from the structured service requirements, and service type prediction compensation is performed on fuzzy requirements based on an attention mechanism to obtain compensated service requirements, including: The structured service requirements are analyzed by the service feature extraction module, which outputs a service type probability distribution vector and an urgency score. The service feature extraction module is configured as a classification and regression joint network containing a dual-channel fully connected layer. The service type probability distribution vector is input into the fuzzy demand identifier to detect low-confidence service requests. The fuzzy demand identifier is configured as a probability distribution analyzer based on entropy threshold determination. Context-aware compensation is performed on the identified ambiguous demands to generate compensation service type features, including: The temporal context features of historical service requests are extracted through a bidirectional LSTM network and denoted as historical features. An attention mechanism is used to calculate the association weight between the current fuzzy demand and historical features, and a compensation prediction vector is generated, which is denoted as the compensation service type feature. The enhanced service type features are obtained by fusing the compensated service type features with the original service type features, and the fusion is configured as a feature interpolation algorithm based on confidence weighting. The enhanced service type features are concatenated with the urgency score to form the compensated service requirement. The concatenation is configured as a vector merging process that includes a dimension alignment layer.

5. The method for scheduling elderly care services based on AI voice interaction according to claim 1, characterized in that, Extracting user activity pattern features and environmental risk features from the enhanced environment data includes: The enhanced environment data is processed by the activity feature extraction module, which outputs user activity temporal features and environment state features. The activity feature extraction module is configured as a dual-branch feature extraction network containing spatiotemporal convolutional layers. The user activity time-series features are input into a pattern recognizer to generate user activity pattern features. The pattern recognizer is configured as a sequence matching model based on a dynamic time warping algorithm, including: The activity time-series features are segmented using a sliding window and denoted as window features; Calculate the similarity score between each window feature and the preset activity template to generate user activity pattern features; Furthermore, the environmental state characteristics are input into a risk evaluator, which outputs environmental risk characteristics. The risk evaluator is configured as an anomaly detection network incorporating a multi-head attention mechanism, including: The environmental state features are mapped to the risk feature space through the feature projection layer, and this is denoted as the initial risk feature. An attention mechanism is used to calculate the contribution weight of each environmental dimension to the risk score, generating environmental risk characteristics; Furthermore, the user activity pattern features and environmental risk features are normalized to obtain the normalized user activity pattern features and environmental risk features. The normalization process is configured as a standardization algorithm based on sliding window statistics.

6. The method for scheduling elderly care services based on AI voice interaction according to claim 1, characterized in that, A service decision graph is constructed using a spatiotemporal graph neural network to integrate the compensated service demand, user activity pattern characteristics, and environmental risk characteristics, thereby obtaining enhanced service demand. This service decision graph is configured as a demand-environment joint optimization model, including: An initial decision graph is generated through a graph construction module, which is configured as follows: The compensated service requirements are taken as requirement nodes, and the node characteristics include service type code and urgency value; The user activity pattern features are used as activity nodes, and the node features include an activity type identifier and a pattern strength vector; The environmental risk characteristics are used as environmental nodes, and the node characteristics include risk type codes and risk level values; The initial decision graph is enhanced with spatiotemporal features to obtain an enhanced decision graph, including: The graph attention network is configured as a multi-head attention mechanism that includes spatiotemporal location encoding to calculate the association weights between demand nodes, activity nodes, and environment nodes. The node features of adjacent demand nodes, activity nodes, and environment nodes are aggregated using graph convolutional layers to generate enhanced node features. The graph convolutional layers are configured as spectral graph convolutions based on Chebyshev multinomials to obtain the enhanced decision graph. The enhanced service requirements are generated by processing the enhanced decision graph through a service optimizer, including: The importance of the features of the augmented nodes corresponding to the demand nodes in the augmented decision graph is ranked to generate a service priority vector; The enhanced node features corresponding to the activity nodes of the enhanced decision graph and the enhanced node features corresponding to the environment nodes of the enhanced decision graph are integrated to generate a service constraint matrix. The service priority vector and the service constraint matrix are jointly optimized to output enhanced service requirements containing the optimal service path.

7. The method for scheduling elderly care services based on AI voice interaction according to claim 6, characterized in that, The graph attention network is configured as a multi-head attention mechanism incorporating spatiotemporal location encoding to calculate the association weights between demand nodes, activity nodes, and environment nodes. The spatiotemporal location features of the nodes are generated by a spatiotemporal encoder, which is configured as follows: Demand nodes, activity nodes, and environment nodes are assigned spatial location codes, which are generated based on the three-dimensional location vectors of the nodes' physical coordinates; The demand node, activity node, and environment node are assigned time location codes, which are generated based on the sinusoidal location code of the event occurrence timestamp. Spatial location codes and temporal location codes are concatenated to form a spatiotemporal location feature vector; A multi-head attention mechanism is used to calculate the dynamic association weights between demand nodes, activity nodes, and environment nodes, including: The attention head is obtained by adding the demand node, activity node, and environment node with the spatiotemporal location features; Generate a query vector, key vector, and value vector for each attention head; Calculate the query-key similarity score and overlay it with the spatiotemporal association constraint matrix; The association weights of each attention head are obtained by normalization using the softmax function; Aggregate the results of multi-head attention to generate the final node association weights, including: Linear projection is performed on the output of each attention head; A gating mechanism is used to fuse feature representations from different attention heads; The output contains an edge weight matrix that includes spatiotemporal correlation information. The edge weight matrix includes the correlation weights between demand nodes, activity nodes, and environment nodes.

8. The method for scheduling elderly care services based on AI voice interaction according to claim 1, characterized in that, The enhanced service requirements are decomposed into multi-granularity tasks to obtain the final service task list. This multi-granularity task decomposition is configured as an end-to-end task generation that integrates voice intent, behavioral habits, and resource scheduling, including: The enhanced service requirements are processed by the task parsing module, which outputs a set of basic task units. The task parsing module is configured as a task decomposition network containing a hierarchical attention mechanism. The set of basic task units is input into a resource adapter to generate an executable task sequence. The resource adapter is configured as follows: Information on currently available service resources is obtained through the resource status monitoring layer and recorded as resource characteristics; A graph matching algorithm is used to calculate the fit score between task units and resource features; A preliminary task allocation scheme is generated based on the suitability score; Optimize the behavior and habits of the preliminary task allocation scheme, including: Historical behavioral pattern features are extracted through the user profiling module and recorded as habit features; The matching degree between the preliminary task allocation scheme and habitual features is calculated using a collaborative filtering algorithm, and is denoted as the behavior matching degree. An optimized task allocation scheme is obtained by balancing resource suitability and behavior matching through a multi-objective optimization algorithm. Generating the final service task order includes: The optimized task allocation scheme is then time-series orchestrated. Add execution constraints and quality evaluation metrics to each task unit; The output includes a standardized task sheet containing the task path, resource binding, and execution sequence.

9. A dispatching system for elderly care services based on AI voice interaction, characterized in that, The method applicable to any one of claims 1 to 8.