A method, system, and device for ordering home delivery services based on AI voice.

CN122575368APending Publication Date: 2026-08-14HUNAN NATIONAL HOME SERVICE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]为了解决现有的基于AI语音的下单到家服务无法确保服务交付准确性和用户满意度的技术问题,本申请提供了一种基于AI语音下单到家服务方法、系统及设备

Benefits of technology

本申请提供了一种基于AI语音下单到家服务的方法,包括:基于接收到的用户语音指令和识别到的环境传感数据,进行自然语言处理和上下文理解,得到增强语义信息;根据增强语义信息进行服务意图识别和情感分析处理,生成服务请求和用户偏好信息;依据服务请求和用户偏好信息进行服务匹配操作和服务调度操作,生成匹配调度信息;获取匹配调度信息的反馈信息,并依据反馈信息进行服务确认和环境感知,以优化到家服务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575368A_ABST
    Figure CN122575368A_ABST
Patent Text Reader

Abstract

This invention relates to the field of voice interaction technology, specifically to a method, system, and device for ordering home delivery services based on AI voice. The method includes: performing natural language processing and contextual understanding based on received user voice commands and identified environmental sensor data to obtain enhanced semantic information; performing service intent recognition and sentiment analysis based on the enhanced semantic information to generate service requests and user preference information; performing service matching and scheduling operations based on the service requests and user preference information to generate matching and scheduling information; obtaining feedback information from the matching and scheduling information, and performing service confirmation and environmental awareness based on the feedback information to optimize the home delivery service. This solves the problem that existing AI voice-based home delivery services cannot ensure service delivery accuracy and user satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice interaction technology, specifically to a method, system, and device for ordering home delivery services based on AI voice. Background Technology

[0002] With the deep integration of artificial intelligence and Internet of Things technologies, smart home systems based on AI voice interaction are gradually becoming more widespread, providing users with a new way to conveniently book various home services through natural voice commands.

[0003] Currently, services typically rely on speech recognition and natural language processing of user voice commands to convert them into text commands. These text commands are then matched against a pre-defined service keyword database to generate a standardized service order. This order is subsequently sent to the corresponding service platform or dispatch system for service matching and notification. However, the logic behind this service generation heavily depends on the explicit text content of the voice commands, completely ignoring the specific physical environment in which the user issued the command, as well as the user's emotional state and potential preferences implied in the command. This results in rigid service requests that are detached from real-world scenarios, failing to deeply understand the user's real-time, authentic, and even unspoken intentions in the current environment. For example, it cannot recognize that a user's anxiety due to unexpected damage to household items might lead to a higher demand for repair services. Consequently, subsequent service matching and dispatch lack precise and personalized decision-making basis, affecting the accuracy of service delivery and user satisfaction. Summary of the Invention

[0004] To address the technical issues that existing AI voice-based ordering and home delivery services cannot ensure service delivery accuracy and user satisfaction, this application provides a method, system, and device for AI voice-based ordering and home delivery services.

[0005] This application provides a method, system, and device for ordering and delivering home services based on AI voice, which adopts the following technical solution: A method for ordering home delivery services based on AI voice includes: Based on the received user voice commands and the identified environmental sensor data, natural language processing and contextual understanding are performed to obtain enhanced semantic information; Based on enhanced semantic information, service intent recognition and sentiment analysis are performed to generate service requests and user preference information; Based on service requests and user preference information, perform service matching and scheduling operations to generate matching and scheduling information; Obtain feedback information on matching and scheduling information, and perform service confirmation and environmental awareness based on the feedback information to optimize home delivery services.

[0006] Furthermore, based on the received user voice commands and the identified environmental sensor data, the steps of performing natural language processing and contextual understanding to obtain enhanced semantic information include: Based on user voice commands and environmental sensor data, speech-to-text conversion and emotion state recognition are performed to obtain speech-to-text information and user emotion data. Anomaly detection processing is performed on environmental sensor data to obtain environmental event information. Based on user emotion data, the weights of user voice commands and environmental sensor data are calculated to obtain modal weight coefficients. Based on the modal weight coefficients, weighted features and cross-modal alignment operations are performed on speech text information and environmental event information to obtain semantic fusion features; Based on semantic fusion features, contextual association and information completion processing are performed by combining historical service records to obtain enhanced semantic information.

[0007] Furthermore, the steps for calculating the modal weight coefficients by weighting user voice commands and environmental sensor data based on user emotion data include: Based on user emotion data, feature extraction is performed on emotion state, emotion intensity, and emotion stability to obtain emotion feature parameters; Based on the emotional feature parameters, attention allocation pattern analysis is performed to generate a weight allocation strategy for user voice commands and environmental sensor data. Based on the weight allocation strategy, the weight values ​​of user voice commands and environmental sensor data are mapped to obtain the voice command weight value and the environmental sensor weight value. The voice command weight values ​​and environmental sensor weight values ​​are integrated and normalized to obtain the modal weight coefficients.

[0008] Furthermore, the steps of performing service intent recognition and sentiment analysis based on enhanced semantic information to generate service request and user preference information include: Based on enhanced semantic information, semantic role labeling, dependency parsing, and multimodal information alignment operations are performed to obtain structured semantic units; Based on structured semantic units, implicit intention reasoning and emotional state transition analysis are performed to obtain multi-granularity intention classification and emotional evolution trajectory; Based on multi-granularity intent classification and sentiment evolution trajectory, and combined with user historical behavior characteristics, intent decomposition and sentiment-guided preference identification are performed to obtain service tasks and preference constraints. The service tasks and preference constraints are parameterized and prioritized to generate service requests and user preference information.

[0009] Furthermore, based on structured semantic units, the steps for implicit intention reasoning and sentiment state transition analysis to obtain multi-granularity intention classification and sentiment evolution trajectory include: The service timing information and service entity relationships included in the structured semantic units are extracted and aligned to obtain the service operation sequence; Based on the service operation sequence, and combined with historical service records, sequence matching and operation difference comparison are performed to identify explicit intent sequences and implicit intent information; Based on explicit intention sequences and implicit intention information, an intention state transition diagram is constructed and a time series analysis of emotional intensity is performed to obtain the intention evolution path and the emotional intensity change curve. Multi-level aggregation and trajectory smoothing are performed on the intention evolution path and emotion intensity change curve to obtain multi-granularity intention classification and emotion evolution trajectory.

[0010] Furthermore, the steps for generating matching and scheduling information, based on service requests and user preference information, include: Based on service request and user preference information, and combined with real-time status and location information of service resources, conditional filtering and scoring are performed to obtain multiple candidate service resource items; Based on multiple candidate service resource items, task decomposition and scheduling strategy calculation are performed on multiple service tasks included in the service request and preference constraints in user preference information to obtain an initial scheduling scheme; Based on the initial scheduling scheme, changes in the real-time status of service resources, and external environmental factors, conflict detection and optimization adjustments are performed to generate optimized scheduling information, including the main execution scheme and at least one backup execution scheme. The optimized scheduling information is encoded with instructions and resource locking is confirmed to generate matching scheduling information.

[0011] Furthermore, based on multiple candidate service resource items, the steps for decomposing tasks and calculating scheduling strategies for multiple service tasks included in the service request and preference constraints in the user preference information to obtain the initial scheduling scheme include: Based on the logical relationships and preference constraints among multiple candidate service resource items, task dependency analysis and constraint satisfaction verification are performed to obtain a task dependency graph. Based on the task dependency graph and multiple candidate service resource items, service capability matching and parallel execution path planning are performed to obtain a feasible execution sequence; Based on the feasible execution sequence, multiple service task subsequences that can be processed by the same candidate service resource item are aggregated to generate an aggregated task block with a time window; Based on the aggregated task block with time windows, the timeline is filled and conflict avoidance calculations are performed by combining the real-time location and estimated ready time of each candidate service resource item to generate an initial scheduling scheme.

[0012] This application also provides a system for ordering home delivery services based on AI voice, which is applied to the method for ordering home delivery services based on AI voice as described above, including: The data processing module is used to perform natural language processing and contextual understanding based on the received user voice commands and the recognized environmental sensor data to obtain enhanced semantic information; The information analysis module is used to perform service intent recognition and sentiment analysis based on enhanced semantic information, and to generate service request and user preference information. The matching and scheduling module is used to perform service matching and scheduling operations based on service requests and user preference information, and to generate matching and scheduling information. The service optimization module is used to obtain feedback information from matching scheduling information, and to perform service confirmation and environmental awareness based on the feedback information in order to optimize the home delivery service.

[0013] This application also provides a device for AI voice-based ordering and home delivery services. The device includes the system for AI voice-based ordering and home delivery services as described above, a memory storing executable instructions, and a processor configured to execute the executable instructions in the memory to implement the method for AI voice-based ordering and home delivery services as described above.

[0014] Beneficial effects achieved: This application provides a method for ordering home delivery services based on AI voice, including: performing natural language processing and contextual understanding based on received user voice commands and identified environmental sensor data to obtain enhanced semantic information; performing service intent recognition and sentiment analysis based on the enhanced semantic information to generate service requests and user preference information; performing service matching and service scheduling operations based on the service requests and user preference information to generate matching and scheduling information; obtaining feedback information on the matching and scheduling information, and performing service confirmation and environmental awareness based on the feedback information to optimize the home delivery service.

[0015] In this application, by fusing user voice commands with environmental sensing data, the system can perceive the user's current situation and obtain enhanced semantic information that includes environmental context and is closer to the actual situation. Then, based on this enhanced semantic information, service intent recognition and sentiment analysis are performed to generate more refined and personalized service requests and user preference information. This makes subsequent matching and scheduling no longer fixed orders but decisions based on a solid foundation. Next, service matching and scheduling are performed based on the service requests and user preference information, taking into account service personnel skills, real-time location, current load, and the user's emotional needs, thereby generating matching and scheduling information that is more relevant to the scenario and better meets user expectations. Finally, the system obtains execution feedback and performs service confirmation and environmental perception based on the feedback, forming a closed-loop optimization mechanism. This mechanism can adjust service execution details in real time to cope with changes in the situation, ensuring that the entire process from service generation and matching to final delivery closely revolves around the dynamically changing real user intent and scenario state. This fundamentally guarantees that the final service delivery result is highly consistent with the user's complex and real-time expectations, thus significantly improving service accuracy and user satisfaction. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the steps of a method for ordering home delivery services based on AI voice, as described in this application. Figure 2 This is a flowchart illustrating the steps involved in obtaining enhanced language information in this application; Figure 3 A flowchart illustrating the steps involved in generating service requests and user preference information for this application; Figure 4 A flowchart illustrating the steps involved in generating matching scheduling information for this application; Figure 5 This is a schematic diagram of a system for ordering and delivering services based on AI voice, as described in this application.

[0017] Explanation of reference numerals in the attached figures: 10. Data processing module; 20. Information analysis module; 30. Matching and scheduling module; 40. Service optimization module. Detailed Implementation

[0018] The following combination Figures 1 to 5 This application will be described in further detail.

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0020] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative positional relationship and movement of the components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.

[0021] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the use of "and / or" or "and / or" throughout the text includes three parallel solutions. For example, "A and / or B" includes solution A, solution B, or a solution where both A and B are satisfied simultaneously. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0022] This application discloses a method for ordering home delivery services based on AI voice commands.

[0023] Please refer to Figure 1 The method for ordering home delivery service based on AI voice proposed in this embodiment includes steps S10 to S40: Step S10: Based on the received user voice commands and the identified environmental sensor data, perform natural language processing and contextual understanding to obtain enhanced semantic information.

[0024] Based on the received user voice commands and the identified environmental sensor data, natural language processing and contextual understanding are performed. The fundamental purpose is to break through the limitations of the traditional processing mode that only converts voice signals into text. It aims to deeply integrate and cross-validate the user's subjective verbal expression with the real-time environmental status, thereby building a more realistic foundation for decision-making information.

[0025] By parsing the explicit meaning of user voice commands through natural language processing, and transforming the scene information reflected by environmental sensor data into a supplementary context that can be semantically interpreted through contextual understanding, the resulting enhanced semantic information is no longer an isolated text command, but a comprehensive semantic representation that integrates the user's explicit demands, implicit physical scene constraints, and real-time environmental background. This provides a solid and rich input for subsequent steps to understand the user's true service intentions and emotional preferences.

[0026] Step S20: Based on the enhanced semantic information, perform service intent recognition and sentiment analysis to generate service request and user preference information.

[0027] Service intent recognition and sentiment analysis based on enhanced semantic information involve in-depth analysis and mining of enhanced semantic information. The aim is to penetrate the surface meaning of user voice commands and capture their potential service needs and personalized tendencies in specific environments and emotional states. In this way, the context-rich enhanced semantic information is transformed into service requests and user preference information that can be directly used for decision-making and scheduling.

[0028] By extracting core service actions, target objects, and constraints from enhanced semantic information through service intent recognition, and judging the user's current emotional state, urgency, and satisfaction preferences through sentiment analysis, a user's voice command, which may be vague, general, or emotionally charged, is deconstructed and transformed into a service request with clear objectives, distinct elements, and personalized parameters, along with user preference information. The generated service request is no longer a standardized service type code, but a complete task package that includes "what to do," "what are the specific requirements," and "why the user needs it." The user preference information further quantifies the expected standards for service execution, providing decision input for highly accurate and personalized service resource matching and scheduling in subsequent steps. This is a key link in ensuring that service supply closely matches the dynamic and complex real needs of users.

[0029] Step S30: Perform service matching and service scheduling operations based on service requests and user preference information to generate matching and scheduling information.

[0030] The service matching and scheduling operations based on service requests and user preference information are transformed into a specific action plan, namely matching and scheduling information. This solves the problems of coarse resource matching and rigid scheduling caused by the lack of single decision-making information in traditional dispatch systems.

[0031] By using abstract service requests and user preference information as the core decision-making basis, the system drives multi-dimensional and dynamic screening and weighing of massive service resources. The service matching operation searches for the most suitable candidate service based on multiple factors such as service request content, user preference information, and service provider skills, reputation, and real-time location. The service scheduling operation further considers task sequence, time window, path planning, and resource load, and arranges the matching and scheduling results in a time sequence and resolves conflicts. This ensures that the generated matching and scheduling information is not a simple personnel assignment notification, but an optimized matching and scheduling information that integrates "who," "when," "where," "in what order," and "what specific task" to execute. This ensures that the most suitable service resources can be delivered to the service site at the most appropriate time and in a way that best meets user expectations, thereby effectively transforming the service requests and user preference information obtained in the previous steps into high-satisfaction physical service delivery.

[0032] Step S40: Obtain feedback information on the matching scheduling information, and perform service confirmation and environmental awareness based on the feedback information to optimize the home delivery service.

[0033] By acquiring feedback information from matching scheduling data and performing service confirmation and environmental awareness based on this feedback, an optimized closed loop based on real-time execution feedback is established for the entire service process. This aims to solve the static scheduling defects in traditional service processes, such as the disconnect between planning and execution and the inability to cope with sudden changes in the field. This ensures that the final service delivery result continuously reflects the dynamic changes in user needs and the actual situation on site. By placing the matching scheduling information generated in the preceding steps into a real service execution environment for verification and calibration, and by proactively acquiring execution feedback from service providers and sensing real-time environmental changes in the service field, the system provides real-time decision-making basis. This allows the system to promptly confirm or adjust services based on feedback information and update its understanding of the scenario using the latest environmental awareness. Thus, the "GetHome" service is no longer a one-way, one-time dispatch of instructions, but an adaptive process that can dynamically optimize based on execution feedback and on-site changes. This fundamentally improves the flexibility, accuracy, and ultimate user satisfaction of service delivery, completing a closed loop from accurate demand understanding and intelligent scheduling to reliable final delivery.

[0034] Specifically, after the service is started based on the matching scheduling information, the terminal carried by the service resource item actively reports its execution status, and the latest environmental perception data of the service site is continuously collected through the environmental sensor network to obtain feedback information on the execution status of the scheduling plan and changes in the site environment in real time.

[0035] The system compares the feedback information reported by the service resource items with the plan in the matching scheduling information to confirm whether the task is progressing according to the predetermined sequence, time window and standard. If a delay, interruption or failure to meet the execution conditions is found, a confirmation exception is triggered and service confirmation is performed.

[0036] The system analyzes the newly acquired environmental sensors in real time and compares them with the context understanding in step S10 to detect whether any environmental events affecting the scheduling plan have occurred, thereby performing environmental perception.

[0037] Based on the service confirmation and environmental awareness results, the home delivery service is dynamically optimized. When the execution is confirmed to be normal and there are no significant changes in the environment, the original scheduling is maintained and the service is confirmed to proceed as planned. When the execution is confirmed to be abnormal or new relevant environmental events are detected, the real-time replanning of the scheduling information is triggered. Based on the latest feedback information and the original service requests and user preference information, local or global service matching and scheduling calculations are re-performed, and adjusted matching and scheduling information is generated and issued to ensure that the final delivery of the service can adapt to the dynamic changes in the execution site.

[0038] In one feasible implementation, refer to Figure 2 It can be seen that step S10 may specifically include steps S11 to S14: Step S11: Based on user voice commands and environmental sensor data, perform voice-to-text conversion and emotion state recognition to obtain voice-to-text information and user emotion data, and perform anomaly detection processing on environmental sensor data to obtain environmental event information.

[0039] Specifically, the processing method based on user voice commands and environmental sensor data involves executing two independent processing streams in parallel: For user voice commands, continuous analog speech signals are converted into discrete digital audio signals using sound acquisition devices such as microphones. Sampling captures sound amplitude at a fixed frequency (e.g., 16kHz), quantization maps the amplitude value of each sample point to a finite-precision digital representation (e.g., a 16-bit integer), and encoding generates a digital sequence according to a specific format (e.g., PCM). Noise reduction is then applied to the digital audio signal, typically using digital filtering (e.g., high-pass filtering to remove low-frequency noise) or spectral subtraction based on spectral analysis. The signal-to-noise ratio is improved by estimating and subtracting the spectral components of steady-state noise in the signal. The denoised digital audio signal is then framed, dividing the continuous digital audio signal into a series of short time segments (frames). Each frame is typically 20-40 milliseconds long, with overlapping regions (e.g., 10-millisecond overlap) between frames to maintain continuity. Windowing is then applied to each frame of the digital audio signal, typically using a Hamming or Hanning window function multiplied point-by-point with the frame data to reduce spectral leakage caused by signal truncation. The process involves several steps: First, a series of preprocessed valid speech segments are obtained. Next, acoustic features are extracted from these segments, typically using feature vectors such as Mel-frequency cepstral coefficients to characterize the time-frequency properties of the speech signal. Then, the extracted acoustic features are input into a pre-trained speech recognition model, which usually consists of an acoustic model, a pronunciation dictionary, and a language model. The acoustic model (e.g., a deep neural network-based model) maps acoustic features to phoneme probability distributions. The pronunciation dictionary provides the correspondence between phonemes and words, while the language model (e.g., a Transformer-based model) is trained on a large-scale text corpus to evaluate the probability of word sequences, thus constraining the decoding process. Finally, a decoder (e.g., a weighted finite-state converter) dynamically searches for the optimal path, combining the probability sequence output by the acoustic model with the prior knowledge provided by the language model and the constraints of the pronunciation dictionary to find the word sequence with the highest probability, thereby generating the corresponding speech-text information.

[0040] Simultaneously, acoustic feature vectors such as pitch, speech rate, and energy are extracted from the preprocessed effective speech segments. Pitch reflects the basic frequency changes of speech, speech rate is calculated by the number of speech segments per unit time, and energy represents the amplitude intensity of the speech signal. Next, these acoustic feature vectors are input into a pre-trained emotion recognition model. This emotion recognition model is usually built based on machine learning or deep learning methods (such as convolutional neural networks or recurrent neural networks), which learns the mapping relationship between acoustic features and emotional states through training on a large amount of speech data with emotion labels. Then, the emotion recognition model performs forward propagation calculation on the input acoustic features, outputting probability distribution predictions for various emotion categories (such as calm, anxiety, joy, etc.), and estimates the intensity of emotion through internal mechanisms of the emotion recognition model (such as regression layers or attention weights), thereby obtaining quantified numerical values ​​of emotion intensity and emotion category labels. Finally, these outputs are integrated into structured user emotion data, including specific emotion categories, corresponding confidence or probability values, and continuous or discrete emotion intensity parameters, thus obtaining user emotion data.

[0041] For environmental sensing data, the system receives continuous data streams from environmental sensors in real time and compares them with preset normal state thresholds or historical data baselines. When data is detected to deviate from the normal range continuously or momentarily, it is determined to be an abnormal event. The system then structures information such as the type, location, and severity of the event to obtain environmental event information. Each piece of information in this environmental event information is encoded.

[0042] This technology enables the parallel conversion and extraction of user voice commands and environmental sensor data into three structured, quantifiable information types: voice-text information, user emotion data, and environmental event information. This provides an input basis for subsequent steps of adaptive weighted fusion guided by emotion data, ensuring that the fusion process is based on reliable information and supporting a deeper understanding of complex user intentions.

[0043] Step S12: Based on the user's emotional data, calculate the weights of the user's voice commands and environmental sensor data to obtain the modal weight coefficients.

[0044] This step introduces a dynamic decision-making mechanism driven by the user's real-time emotional state to quantitatively evaluate and dynamically allocate the relative importance and credibility of two different modalities, speech and environment, in semantic understanding. This solves the problem of fusion result deviation caused by the inability to adapt to complex and variable user states due to the use of fixed weights in modal information fusion.

[0045] By transforming the user's emotional data, which reflects the user's subjective psychological state, identified in step S11 into an operable emotional feature parameter, and analyzing the emotional feature parameter, the reliability and accuracy of the user's speech expression under the current emotional state, as well as the degree of dependence on the objective state of the environment, are inferred. This generates a weight allocation strategy, which determines the contribution ratio of speech text information and environmental event information to the enhanced semantic information in subsequent fusion steps. This enables dynamic adjustment of the information fusion balance according to the user's different emotional states, making the subsequent feature fusion process personalized and context-adaptive, and providing a key guarantee for the accuracy and robustness of core semantic understanding.

[0046] Furthermore, step S12 also includes steps S121 to S124: Step S121: Based on user emotion data, extract features of emotion state, emotion intensity and emotion stability to obtain emotion feature parameters.

[0047] The emotional state is extracted from the user emotional data obtained in step S11. The emotional state is a categorical variable and is converted into a vector as a feature representation of the emotional state through a preset mapping table. For example, calm is mapped to vector [1,0,0,0,0], and anxious is mapped to vector [0,1,0,0,0], etc.

[0048] Simultaneously, the numerical values ​​of quantified emotional intensity are directly extracted from the user's emotional data. These values ​​are typically continuous scalars or discrete level values, and are directly used as emotional intensity features. Based on a pre-defined threshold range corresponding to the value type (for example, for levels 1-5, the threshold range is set to [1,5]), a linear normalization formula is applied for calculation. The normalization formula is: Normalized emotional intensity = (value - lower limit of threshold range) / (upper limit of threshold range - lower limit of threshold range). This maps the value to a continuous scale of [0,1]. When the value is equal to or lower than the lower limit of the threshold range, the normalized emotional intensity is set to 0; when the value is equal to or higher than the upper limit of the threshold range, the normalized emotional intensity is set to 1. This ensures that all emotional intensity features are within a uniform and comparable numerical range.

[0049] In addition, feature extraction of emotional stability is performed by calculating the change characteristics between the current emotional intensity value and the numerical sequence of multiple emotional intensities extracted in a previous continuous time period (such as the previous 5 seconds). Specifically, the statistical characteristics of the numerical sequence are calculated, such as variance to measure the degree of fluctuation or the average of the differences between adjacent points to measure the severity of the change trend, thereby obtaining one or more parameters characterizing emotional stability.

[0050] Finally, the vector of the transformed emotional state, the normalized emotional intensity value, and the calculated emotional stability parameter are combined into an emotional feature parameter, which provides a quantitative input for generating attention allocation patterns based on data analysis in subsequent steps.

[0051] Step S122: Based on the emotional feature parameters, perform attention allocation pattern analysis to generate a weight allocation strategy for user voice commands and environmental sensor data.

[0052] The system receives the emotional feature parameters from the previous step and inputs them into a predefined attention allocation decision mechanism. This mechanism first matches a basic attention allocation pattern based on the specific emotional category represented by the vector of emotional state (such as calm, anxious, happy, etc.). For example, it matches a basic attention allocation pattern of "high voice weight and low environment weight" for calmness and a basic attention allocation pattern of "moderately reducing voice weight and increasing environment weight" for anxiety.

[0053] Next, the attention allocation decision mechanism adjusts the basic attention allocation pattern based on the normalized emotion intensity value. The higher the emotion intensity value, the greater the adjustment of the preset weight value. For example, a high value in the "anxious" state will further reduce the preset weight value of the user's voice commands and increase the preset weight value of the environmental sensor data. At the same time, the attention allocation decision mechanism also incorporates the emotion stability parameter. If the emotion stability parameter is low (i.e., large emotion fluctuations), a decay factor is applied to the adjustment range driven by emotion intensity to address the inconsistency or incompleteness of user voice commands that may be caused by emotional fluctuations.

[0054] Finally, the attention allocation decision mechanism integrates the above calculations to generate a clear weight allocation strategy. This weight allocation strategy includes a basic attention allocation pattern based on vector matching of emotional state, and the output of the weight values ​​of user voice commands and environmental sensor data after emotional intensity adjustment and emotional stability correction.

[0055] Step S123: Based on the weight allocation strategy, the weight values ​​of user voice commands and environmental sensor data are mapped to obtain the voice command weight value and the environmental sensor weight value.

[0056] Based on the fundamental attention allocation pattern in the weighting strategy, initial weight values ​​for user voice commands and environmental sensor data are determined. For example, for a "calm" state, the initial weight value for user voice commands is set to a higher value, such as 0.7, and the initial weight value for environmental sensor data is set to a lower value, such as 0.3. Next, the initial weight values ​​for user voice commands are adjusted according to emotional intensity. The adjustment magnitude is usually proportional to the emotional intensity value. For example, when the emotional intensity value is 0.8, the initial weight value for user voice commands is decreased by 0.1, and the initial weight value for environmental sensor data is increased by 0.1.

[0057] Simultaneously, an emotion stability correction factor is applied. If the emotion stability parameter is low, the adjustment amplitude is attenuated, for example, by multiplying by 0.5, to smooth the changes in the initial user voice command weight value and the initial environmental sensor data weight value, thus obtaining the voice command weight value and environmental sensor weight value. This transforms the abstract weight allocation strategy into operable specific values, enabling the user voice command and environmental sensor data to obtain accurate and adaptive weight allocation based on the user's real-time emotional state in subsequent modal fusion steps, thereby optimizing the accuracy of fusion and the responsiveness to dynamic interaction scenarios.

[0058] Step S124: Integrate and normalize the voice command weight value and the environmental sensing weight value to obtain the modal weight coefficient.

[0059] After combining the voice command weights and environmental sensor weights calculated in the above steps into a weight pair, this weight pair is normalized. Specifically, the sum of the voice command weights and environmental sensor weights is calculated, and each voice command weight and environmental sensor weight is divided by this sum to ensure that the sum of the processed voice command weights and environmental sensor weights is 1. This yields normalized voice command weight coefficients and normalized environmental sensor weight coefficients, which together constitute the modal weight coefficients. This transforms the voice command weights and environmental sensor weights calculated in the previous step into a standardized modal weight coefficient that sums to one. This ensures that in the subsequent weighted feature steps, the contributions of the two sources of information from user voice commands and environmental sensor data are standardized into a unified and comparable probabilistic framework, making the fusion process consistent and stable. This guarantees that the semantic fusion features output later will not be biased due to the arbitrariness of the weight scale, laying a mathematical foundation for generating accurate and reliable enhanced semantic information.

[0060] Step S13: Based on the modal weight coefficients, perform weighted feature and cross-modal alignment operations on the speech text information and environmental event information to obtain semantic fusion features.

[0061] The feature vectors of the speech-text information are multiplied by the speech command weight coefficient in the modal weighting coefficient, and the feature vectors of the environmental event information are multiplied by the environmental sensing weight coefficient in the modal weighting coefficient. After obtaining the weighted speech-text feature vectors and weighted environmental sensing feature vectors, a cross-modal alignment operation is performed. By calculating the correlation between the weighted speech-text feature vectors and weighted environmental sensing feature vectors, the speech-text feature vectors representing the same entity or concept and the environmental sensing feature vectors are aligned and bound in the vector space. For example, the feature vectors of "water cup" mentioned in the speech-text information and "cylindrical object" detected in the environmental event information are brought closer together, while irrelevant ones are downplayed. The connections between feature vectors eliminate information redundancy and conflict between different modalities and strengthen complementary information. Finally, the features after weighting and alignment are integrated into a semantic fusion feature with internal semantic consistency and unified structure. This generates a feature representation that is deeply integrated and calibrated to reflect the user's subjective expression and the objective environment. It not only balances the relative credibility and importance of the two modal information in the current user's emotional state through corresponding weight coefficients, but also ensures that the information of different modalities points to consistent real-world entities and events at the semantic level through cross-modal alignment. This provides input for subsequent intent recognition and greatly improves the accuracy of understanding user intent in complex real-world scenarios.

[0062] The feature vector of the speech text information is obtained by inputting the speech text information into a pre-trained deep neural network such as BERT. This pre-trained deep neural network models the contextual semantics of the speech text information through its multi-layer Transformer encoder and outputs a dense vector of fixed dimensions as the feature vector representation of the speech text information.

[0063] As shown in step S11, each piece of information in the environmental event information is encoded. Each encoded piece of information is then mapped / transformed. For example, discrete event types are converted into unique integer indices, which are then searched using a pre-trained embedding matrix. Each row of this embedding matrix corresponds to an embedding vector for an event type, thus mapping the index to a type feature vector (i.e., the event type is mapped to a type feature vector through the embedding layer). The specific process of converting the absolute coordinates of the occurrence location into a relative position vector through normalization involves obtaining the absolute coordinates of the event occurrence (e.g., latitude and longitude or image pixel coordinates), and then scaling the coordinate values ​​to a range between 0 and 1 through a linear transformation. Alternatively, the relative offset can be calculated using a specific reference point in the scene (such as the center of the room or the center of the camera's field of view) as the origin and scaled accordingly to obtain a numerical vector representing the relative position (i.e., the location of occurrence is converted from absolute coordinates to a relative position vector through normalization). The severity value, such as discrete levels (e.g., 1 to 5) k or continuous values ​​(e.g., 0.0 to 1.0) s, is treated as a scalar input x. Then, a linear transformation layer is defined, parameterized by a learnable weight matrix W and a learnable bias vector b. The weight matrix W has a dimension of d*1, where d is a fixed dimension of the output feature vector, and the bias vector b has a dimension of d. Then, the calculation... ,in, This represents the multiplication of the weight matrix with the scalar input. Essentially, x is treated as a 1-dimensional vector, resulting in an output vector y, which is a d-dimensional real vector, i.e., a severity feature vector (i.e., the severity is converted into a severity feature vector through numerical encoding).

[0064] Step S14: Based on the semantic fusion features, context association and information completion processing are performed by combining historical service records to obtain enhanced semantic information.

[0065] First, retrieve the historical service records associated with the user from the local or cloud database. In this embodiment, the historical service records refer to the structured log data stored in time sequence that was recorded during the user's past interactions with the system. These records typically include historical interaction times, historical voice command texts, historical environmental event information, service actions performed by the system, and their results.

[0066] When performing context association, the similarity between semantic fusion features and various historical service records is calculated. Specifically, the cosine similarity method can be used. This involves calculating the cosine similarity by dividing the dot product of the weighted speech-text feature vector and the weighted environmental sensing feature vector in the semantic fusion features by the product of their respective moduli with each feature vector stored in the historical service records. The historical service records are then sorted in descending order based on their cosine similarity scores, and the top few historical service records with the highest similarity scores are selected as the most relevant historical service records to the current context. Finally, these retrieved relevant historical service records are input into an encoder, typically a pre-trained neural network such as a multilayer perceptron or Transformer encoder. This encoder encodes the text, event tags, or multimodal data included in the selected historical service records into a fixed-dimensional dense vector, thus obtaining the corresponding... After generating historical service feature vectors, information completion processing is performed. The historical service feature vectors are then fused with semantic fusion features in a secondary process. This process is achieved through an attention mechanism, which evaluates the importance of historical service records in completing the semantic fusion features and dynamically integrates missing context from historical service information (such as omitted objects in user references, recurring personalized preferences, and unspecified task preconditions) into the current semantic fusion features. The result is an enhanced semantic information that integrates real-time user emotional commands, environmental sensor data, and relevant historical service record context. This allows the understanding of user intent to move beyond isolated information from a single interaction. Instead, it enables the parsing of ambiguous user voice commands, completion of missing information, and prediction of potential needs by leveraging the user's historical behavior and service context. This significantly improves the ability to understand semantics and provides a solid foundation for generating more accurate and appropriate service responses.

[0067] In one feasible implementation, refer to Figure 3 It can be seen that step S20 may specifically include steps S21 to S24: Step S21: Based on the enhanced semantic information, perform semantic role labeling, dependency parsing and multimodal information alignment operations to obtain structured semantic units.

[0068] First, semantic role labeling is performed on the text portion of the enhanced semantic information. A pre-trained semantic role labeling model is used to identify the arguments governed by the predicates in the text portion and their roles, such as agent, patient, time, and place. This decomposes the text portion into a structured role framework of "who-to what-when and where-did what".

[0069] Meanwhile, dependency parsing is performed on the text portion, and the dependency parser is used to analyze the syntactic modification relationships between words, constructing a tree structure with the core verb as the root, and clarifying the grammatical dependency relationships between words.

[0070] Next, the structured role framework and syntactic dependency relations obtained from the above analysis are associated and matched with the environmental event information in the enhanced semantic information in the same semantic space. For example, the text entity "water cup" in the structured role framework and the visual entity "cylindrical object on the table" detected by the environmental event information are similar to each other through their feature vectors to establish cross-modal entity association. The attributes of the environmental event (such as location and state) are bound to the corresponding text entity as supplementary information, and the timestamp or event occurrence order is extracted from the environmental event information to construct service time sequence information. At the same time, the entity associations revealed by the structured role framework and syntactic dependency relations are integrated to form service entity relations.

[0071] Finally, service time sequence information and service entity relationships are integrated to generate structured semantic units. This transforms unstructured enhanced semantic information into a structured representation labeled with "action-entity-attribute-relationship" and its cross-modal correspondences, laying the semantic analysis foundation for subsequent inference of the user's explicit and implicit intent information.

[0072] Step S22: Based on the structured semantic units, perform implicit intention reasoning and emotional state transition analysis to obtain multi-granularity intention classification and emotional evolution trajectory.

[0073] Implicit intent reasoning and emotional state transition analysis are based on structured semantic units to avoid merely understanding the surface meaning of users' explicit instructions. Instead, it can delve deeper into their unexpressed potential intentions and dynamically changing emotional states, thereby achieving a more fundamental insight into users' needs.

[0074] This step achieves the effect of moving from "understanding what the user said" to "understanding what the user wants and why." By utilizing the structured role framework and service entity relationships provided by structured semantic units, the system identifies the logically implicit goals in the user's voice commands that are not directly stated. At the same time, by analyzing the patterns of emotional states during dialogue or service processes, the system captures the user's emotional cues and changes in satisfaction, enabling the system to have deep understanding and predictive capabilities. It can not only respond to current direct requests, but also proactively identify the user's long-term intention evolution trajectory and emotional driving factors, thereby providing a decision-making basis for generating truly personalized services with emotional interaction capabilities and the ability to dynamically adapt to changes in the user's state.

[0075] Furthermore, step S22 also includes steps S221 to S224: Step S221: Extract and align the service timing information and service entity relationships included in the structured semantic unit to obtain the service operation sequence.

[0076] First, service time sequence information is parsed from structured semantic units. This service time sequence information exists in the form of timestamp sequences or event occurrence sequences. At the same time, service entity relationships are extracted. These service entity relationships are represented in the form of graph structures or sets of relation triples, showing the interactions, dependencies, or subordinate associations between service entities.

[0077] Next, an extraction and alignment operation is performed. This operation matches and binds each relation instance in the service entity relationship with its corresponding time point or sequential position in the service time sequence information. For example, it compares the time attribute of the relation instance with the time interval in the service time sequence information, or sorts the relation instances according to the order in the service time sequence information. Finally, these aligned relation instances are organized into an ordered list according to the time order or logical order, where each element represents a service action involving the interaction of a specific service entity that occurs at a specific time or step, thus forming a service operation sequence. This integrates discrete time sequence information and service entity relationships into a time-sequential and structured service action flow, providing an input basis for sequence matching and operation difference comparison in subsequent steps.

[0078] Step S222: Based on the service operation sequence, perform sequence matching and operation difference comparison by combining historical service records to identify explicit intent sequences and implicit intent information.

[0079] The service operation sequence obtained in the previous step is aligned and compared with the historical service operation sequences in the historical service record library. Specifically, a distance matrix is ​​constructed, where the rows of the distance matrix correspond to each operation point of the service operation sequence, and the columns correspond to each operation point of the historical service operation sequence. The local distance between the corresponding two points (usually based on operation type, parameters, and other features) is calculated for each cell in the distance matrix. A target path with the smallest cumulative distance is found from the lower left corner to the upper right corner of the distance matrix. This target path allows the sequence to be non-linearly stretched or compressed on the time axis to achieve the best alignment. The cumulative distance value of the final target path is the DTW distance between the two sequences. The smaller this distance, the higher the similarity. A list of similarity values ​​between the service operation sequence and all historical service operation sequences is obtained. The similarity values ​​are sorted from high to low according to the list of similarity values, thereby selecting one or more historical service operation sequences with the highest similarity values.

[0080] Based on the historical service operation sequences with the highest similarity scores obtained through the above screening, the operation differences are compared. That is, the service operation sequence is time-aligned with each of the screened similar historical service operation sequences, and then each operation point in the sequence is traversed item by item. The operation type at the corresponding position is compared to see if the operation parameter values ​​or settings are consistent, the operation entity (such as target object or tool) is matched, and the order of the operation in the overall sequence is consistent. If the service operation sequence is consistent with a certain similar historical sequence in most operation types, operation parameters, operation entities, and operation order, and the corresponding service operation sequence appears frequently in the history record or directly corresponds to the service action explicitly stated in the user's voice command, then the corresponding service operation sequence is determined to be a stable and repetitive regular operation combination executed by the user, and it is identified as an explicit intent sequence, such as the fixed process of "turning on the lights first and then adjusting the air conditioner after returning home each time".

[0081] Simultaneously, the identified discrepancies (such as omitted, added, or parameter-changed service operations) are categorized, and complete historical context information associated with similar historical service operation sequences matching the current service operation sequence is retrieved. This historical context information includes the original user command, the environmental state during service execution, user historical preferences, and service feedback. For omitted operations, the reason for omission is inferred by considering the context in which the operation typically occurs in historical service operation sequences. For example, if the operation is "detailed check" and is often performed when the user has ample time in the history, while the current context indicates the user is anxious, it can be inferred that the user has a new need for "speed and simplification of the process." For added operations, the correlation between the operation and the current context and historical preferences is analyzed. For example, if a "disinfection" operation is added to a cleaning task sequence, and the historical record uses... Even if a user never explicitly requests disinfection, but current environmental sensor data shows "recent flu outbreak" or the user has frequently searched for disinfection products recently, it can be inferred that the user has an implicit preference adjustment to "enhance hygiene protection." For operations involving parameter changes, comparing the new parameters with historical parameters and current constraints—for example, changing the detergent dosage parameter from "standard" to "small amount"—while there is no related record in the user's preference constraints, but environmental sensor data shows "there are pets indoors," it can be inferred that the parameter adjustment is due to the external constraint of "avoiding pet allergies." This information constitutes implicit intent information, thus going beyond the understanding of the service operation sequence. By identifying stable and repetitive behavioral patterns of users and the change signals contained in the current service operation sequence through historical context, it provides input including explicit patterns and implicit cues for subsequent construction of intent state transition maps and sentiment analysis.

[0082] Step S223: Based on the explicit intention sequence and implicit intention information, construct the intention state transition diagram and perform time series analysis of emotional intensity to obtain the intention evolution path and emotional intensity change curve.

[0083] Using regular operation combinations identified in explicit intent sequences as basic state nodes, each node represents a specific intent state (such as "cleaning start," "item organization," etc.). Transition edges between states are constructed based on the sequence and dependencies of service operations within these combinations, forming the basic framework of the intent state transition graph. Then, combining implicit intent information, state nodes and transition edges are added or adjusted in the graph by analyzing the new needs or preference changes implied by differences. For example, new states are created for new operations, or transition conditions are modified to reflect potential intents. Simultaneously, all historical service operation sequences are traversed, and the frequency of direct transitions from intent state a to intent state b is counted. This frequency is divided by the total frequency of transitions originating from intent state a to calculate the initial weights based on historical service operation sequences. These initial weights reflect the statistical regularity of intent state b appearing after intent state a in the past. Finally, the initial weights are dynamically adjusted using real-time context information, including the user's current emotional intensity, emotional evolution trends, new environmental events, and services. The urgency of the task is considered, and influence coefficients are assigned to these real-time contexts. For example, when the user's current emotional intensity is displayed as "highly anxious," the weight of transition edges pointing to states like "task completed" or "problem solved" is increased. Or, when a new environmental event, "item damaged," is detected, the weight of transitioning from the current state to the "repair processing" state is increased. An influence coefficient is predefined for each real-time context factor, which quantifies the relative contribution of each real-time context factor to the probability of state transition. This coefficient is set based on historical experience. The specific values ​​of these real-time context factors at the current moment are collected (for example, an emotional intensity of "highly anxious" can be mapped to a value of 0.9, and a new environmental event "item damaged" can be mapped to a Boolean value of 1 or 0). The sum of the products of the values ​​of each real-time context factor and their influence coefficients is calculated to obtain an adjustment factor, i.e., adjustment factor = ∑(value × influence coefficient). This adjustment factor is multiplied by the initial weight to obtain the dynamic weight of the corresponding transition edge. This dynamic weight quantifies the probability of the state transition occurring in the current specific context.

[0084] While constructing the intent state transition graph, a time series analysis of sentiment intensity is performed. By extracting the sentiment intensity value of the timestamp corresponding to each operation from the service operation sequence (derived from the user sentiment data in step S11), a time series data is formed. Then, each data point in the time series data is assigned a weight related to distance to other data points within the adjacent time window. The closer the distance, the greater the weight. The data points are weighted to obtain a smoothed estimate of each data point. By performing the above operation on all data points, a continuous fitting curve, i.e., the sentiment intensity change curve, can be generated.

[0085] Finally, when calculating the intent evolution path, based on the real-time data of the current interaction and the intent state transition graph, all possible service operation states are traversed in chronological order. For each time step, during the initialization phase, the maximum cumulative probability value of the initial state is set to 1, and the maximum cumulative probability value of other states is set to 0 or a very small negative value to indicate initial unreachability. For each subsequent time step or state transition opportunity, all possible current intent state nodes are traversed. For each current intent state node, all possible previous intent state nodes are considered, and the transition probability (i.e., the weight of the corresponding transition edge) from the previous intent state node to the current intent state node is calculated. This probability is multiplied by the maximum cumulative probability value already calculated for the previous intent state node. From all possible previous intent state nodes, the value that maximizes this product is selected as the maximum cumulative probability value for the current intent state node. The probability value is calculated, and the previous intent state node that provides the maximum value is recorded as the backtracking point. This process is iterated in chronological order until all relevant time steps are processed or the intent state node corresponding to the current interaction is reached. Then, the backtracking mechanism is used to trace back from the current intent state node to find the previous intent state node of each intent state node according to the backtracking point of each state record until the initial state is reached. This results in the maximum cumulative probability path from the initial state to the current intent state node. This maximum cumulative probability path is the intent evolution path. The emotion intensity change curve is aligned with time. In this way, the evolution logic of user intent and the dynamic fluctuation of emotion are captured in a structured and quantitative way. This provides temporal and state information for intent decomposition and preference recognition in subsequent steps, thereby supporting more accurate and adaptive service planning.

[0086] Step S224 involves multi-level aggregation and trajectory smoothing of the intention evolution path and emotion intensity change curve to obtain multi-granularity intention classification and emotion evolution trajectory.

[0087] The system merges adjacent or semantically similar intent state nodes in the intent evolution path into higher-level abstract intent categories. For example, intent state nodes such as "pick up a rag" and "wipe the table" are aggregated into a medium-granularity intent called "clean furniture." Further, multiple medium-granularity intents are aggregated into coarse-granular intents such as "household chores," thus forming a multi-granularity intent classification from fine to coarse. Simultaneously, the emotion intensity change curve is smoothed using Kalman filtering to remove high-frequency noise and abnormal fluctuations while preserving the overall trend and key inflection points of emotion change. This results in a smooth and continuous emotion evolution trajectory, enabling a structured and hierarchical understanding of the abstract levels and evolutionary context of user intents. It also provides a more reliable dynamic representation of emotions after denoising, offering multi-dimensional contextual information for subsequent intent decomposition and preference recognition. This allows the system to flexibly select appropriate granularity intents to respond to based on different scenario requirements and optimize interaction strategies based on the emotion trajectory.

[0088] Step S23: Based on multi-granularity intent classification and emotional evolution trajectory, and combined with user historical behavior characteristics, intent decomposition and emotional guidance preference identification are performed to obtain service tasks and preference constraints.

[0089] Based on the tree-like hierarchical structure of multi-granularity intent classification, the coarse-grained intents identified are decomposed step by step into a series of specific sub-tasks with logical order and execution dependencies, such as cleaning the desktop, tidying up clutter, and disposing of trash, forming service tasks.

[0090] Simultaneously, the system analyzes the trajectory of emotional evolution, extracting emotional features such as emotional trends (positive increase or negative decrease), peak intensity points, and emotional baselines at different intention stages. These emotional features are then time-aligned and sequence-matched with user historical behavior features stored in the user historical behavior feature database (such as the service options ultimately accepted by the user under similar emotional intensities, the satisfaction rating given, and parameters frequently modified during interactions). This identifies the stable constraints implicit in specific emotional features (for example, when the emotional trajectory shows "increasing anxiety," historical records indicate that the user has a 95% probability of prioritizing the "expedited processing" option and will use "service personnel operating quietly" as a constraint). This transforms service planning from a mechanical task list into a personalized solution that deeply integrates user intention logic, real-time emotional state, and historical behavioral habits. This dynamically responds to changes in user emotions, while the generated preference constraints are directly related to emotional drivers, greatly improving the accuracy of service generation.

[0091] Step S24: Parameterize and prioritize the service tasks and preference constraints to generate service requests and user preference information.

[0092] Each subtask in the service task is parameterized, meaning a set of structured parameters is defined for each subtask, including fields such as execution action, target object, required tools, and expected output. The sequential / parallel execution between subtasks is also encoded as specific relational parameters, thereby transforming the unstructured service task description into a machine-parseable parameterized task object. At the same time, preference constraints are parameterized, quantifying constraints (such as timeliness requirements, service methods, and personnel attributes) into specific parameter key-value pairs or numerical ranges. For example, "latest completion deadline" is converted into a specific deadline timestamp, and "preferred service personnel type" is mapped to a list of skill tags and confidence scores.

[0093] Next, priority ranking is performed. This priority ranking is based on the real-time emotional intensity and trend in the emotional evolution trajectory, as well as the success rate of similar tasks and user satisfaction feedback extracted from the user's historical behavioral characteristics. A weighted scoring algorithm is used to calculate a dynamic priority score for each parameterized task. This algorithm typically assigns different weights to factors such as real-time emotional intensity, emotional trend, task urgency, and success rate of similar tasks and calculates a weighted sum, thereby generating an execution sequence for all service tasks in descending order of score. It should be noted that the weight allocation is set based on the actual application requirements.

[0094] Finally, the parameterized subtasks and their priority sequences are packaged to generate service requests, and the parameterized preference constraints are encapsulated into user preference information. This results in a set of quantified service requests and user preference information, ensuring that subsequent service matching and scheduling operations can perform machine-executable computations and resource allocation.

[0095] In one feasible implementation, refer to Figure 4 It can be seen that step S30 may specifically include steps S31 to S34: Step S31: Based on the service request and user preference information, and combined with the real-time status of service resources and service resource location information, condition filtering and scoring are performed to obtain multiple candidate service resource items.

[0096] First, based on the fields of the parameterized subtasks required in the service request, such as the execution action, target object, and required tools, they are matched with the static attributes of all available service resource items, such as skill tags, tool configuration, and capability descriptions. Then, based on the parameterized constraints in the user preference information, all available service resource items are filtered, and only available service resource items that fully meet the user preference information are retained to enter the candidate pool.

[0097] Next, combining the real-time status of service resources, such as whether service personnel are currently available, the estimated remaining time of service personnel performing tasks, the current battery level of service equipment, the level of service materials, and the location information of service resources, a comprehensive score is calculated for each available service resource item in the candidate pool. This comprehensive score is usually based on a preset weighted scoring algorithm. Its input variables include the real-time distance between the available service resource item and the service task location, the real-time readiness time of the available service resource item, the historical service success rate and user rating of the available service resource item, and the degree of conformity between the available service resource item and the soft preferences in user preference information (such as "prioritizing service providers with a rating higher than 4.5 stars"). After assigning a predetermined variable weight to each input variable and calculating the weighted sum, all available service resource items in the candidate pool are sorted in descending order according to the comprehensive score, and a certain number of the top-ranked available service resource items are selected as multiple candidate service resource items for output. This provides a decision input set for subsequent scheduling and path planning steps, greatly improving the efficiency of the overall scheduling process and the quality of the final matching result.

[0098] Step S32: Based on multiple candidate service resource items, perform task decomposition and scheduling strategy calculation on multiple service tasks included in the service request and preference constraints in user preference information to obtain an initial scheduling scheme.

[0099] Based on multiple candidate service resource items, task decomposition and scheduling strategy calculation are performed on multiple service tasks in the service request and preference constraints in the user preference information. This integrates and spatiotemporally plans the candidate service resource items, service tasks, and user preference information selected in the previous step, transforming the matching problem between a candidate service resource item and a service task into an executable, time-sequential action plan. This plan acts as the execution step from "which service resources are available" to "how the service resources are specifically used." The generated initial scheduling plan not only ensures that all service tasks have qualified candidate service resource items to execute, but also optimizes the continuous utilization rate and overall execution efficiency of candidate service resource items through task aggregation, parallel path planning, and time window scheduling, laying the core foundation for generating a theoretically feasible matching and scheduling information.

[0100] Furthermore, step S32 also includes steps S321 to S324: Step S321: Based on the logical relationships and preference constraints among multiple candidate service resource items, perform task dependency analysis and constraint satisfaction verification to obtain a task dependency graph.

[0101] Dependency analysis is performed on multiple parameterized subtasks included in the service request. By parsing the input-output relationships, sequential requirements, and physical and logical constraints defined in the parameterized subtasks (e.g., wiping the floor must be done after washing the mop), a directed acyclic graph is constructed to represent the dependencies between parameterized subtasks. In this graph, task nodes represent tasks, and directed edges represent hard dependencies that "must be executed after...".

[0102] At the same time, each parameterized subtask is labeled according to the parameterized preference constraints in the user preference information, and these parameterized preference constraints are matched and verified with the real-time capabilities, attributes and locations of candidate service resource items. For example, it verifies whether a candidate service resource item has the skills required to perform a specific task, whether it can arrive at the scene within the specified time window, and whether its attributes (such as gender and language) meet the user preferences. Task nodes that cannot meet any candidate service resource item are eliminated or dependencies are adjusted.

[0103] Finally, the verified task nodes and directed edges, along with the parameterized preference constraints bound to the task nodes that have been verified to be satisfied by at least one candidate service resource, are integrated into a task dependency graph. This ensures that each task in the graph has at least one available candidate service resource that can satisfy all its constraints, thus providing a logically correct and resource-feasible data foundation for subsequent path planning and resource allocation. This avoids scheduling failures caused by task logic conflicts or hard resource mismatches from the source.

[0104] Step S322: Based on the task dependency graph and multiple candidate service resource items, perform service capability matching and parallel execution path planning to obtain a feasible execution sequence.

[0105] First, each task node in the task dependency graph is traversed. Based on the parameterized preference constraints bound to each task node, it is matched with all candidate service resource items, generating a list of candidate service resource items that can execute each task node. On this basis, parallel execution path planning is performed. The core of this process is based on the task order constraints defined in the task dependency graph, utilizing graph traversal and scheduling lists. Under the premise of allowing parallel execution, each task node selects a suitable candidate service resource item from its list and determines its relative execution time order. During the planning process, task nodes without direct dependencies are assigned to different candidate service resource items as much as possible to enable simultaneous execution, thereby shortening the overall completion time and obtaining a feasible execution sequence. Each element in this feasible execution sequence is a tuple that explicitly defines the execution order of tasks and the specific candidate service resource item assigned to execute that task. This transforms a static task dependency graph into a preliminary scheduling scheme, which not only strictly follows the logical dependencies between all task nodes but also maps tasks to specific candidate service resource items and explores the potential for parallel execution, laying the foundation for generating an action plan that is feasible in terms of resource allocation and execution timing.

[0106] The specific implementation method for selecting suitable candidate service resource items from the candidate service resource item list for each task node is as follows: First, a set of evaluation attributes is defined for each candidate service resource item, which typically includes execution cost, estimated time, historical reliability score, and current load status. Preset weights are assigned to these evaluation attributes to reflect the priority of different scheduling objectives (such as cost minimization, shortest time, or highest reliability). Subsequently, the resource selection problem of each task node is encoded as a decision variable of the optimization problem. Different resource allocation schemes are explored through iterative search, and an allocation evaluation value is calculated in each generation to evaluate each resource classification scheme. The allocation evaluation value is obtained by multiplying each evaluation attribute (i.e., the input variable in step S31) of the candidate service resource items allocated to each task node by the weight corresponding to each evaluation attribute and then summing them up. The weights corresponding to each evaluation attribute reflect the relative importance of each attribute in the current scheduling objective. For example, the smallest real-time distance has the highest priority and a higher weight. The resource classification scheme with the highest allocation evaluation value is selected as the candidate service resource item for the corresponding task node.

[0107] Step S323: Based on the feasible execution sequence, aggregate multiple service task sub-sequences that can be processed by the same candidate service resource item to generate an aggregated task block with a time window.

[0108] Traverse feasible execution sequences to identify multiple service task subsequences that are assigned to the same candidate service resource item and are consecutive in execution order (i.e., no tasks are inserted in the middle that are assigned to other candidate service resource items). For each service task subsequence, check whether there are only simple sequential dependencies such as "completed-started" between these service task subsequences, and whether there are complex dependencies that require the intervention of other candidate service resource items. Also, check whether the geographical locations of the task execution are close or the movement paths are reasonable, and whether all relevant constraints in the user preference information are met. When the conditions are met, logically merge these consecutive tasks into a larger work unit, i.e., aggregate task block.

[0109] Next, a "time window" is calculated for the aggregated task block. Based on the earliest completion time of all tasks of the first task node in the task dependency graph, the earliest start time of the aggregated task block is determined. Based on the user preference constraints of the last task in the aggregated task block, its latest end time is determined, thus forming a suggested execution time interval. This time window reduces the number of service entities that need to be directly scheduled and optimized in subsequent timeline filling and conflict avoidance calculations, reducing scheduling complexity. Furthermore, by ensuring that the same candidate service resource items can work continuously and in batches, it reduces the idle time of service resource items and the overhead of task switching, thus laying the foundation for generating a highly efficient and resource-utilization-efficient initial scheduling scheme.

[0110] Step S324: Based on the aggregated task block with time window, combine the real-time location and estimated ready time of each candidate service resource item to perform timeline filling and conflict avoidance calculations to generate an initial scheduling scheme.

[0111] Create a timeline with time points as scales to record the time intervals during which candidate service resources will be occupied in the future. Use the current system time as the logical starting point of the timeline, truncate all time periods before the logical starting point, and initialize all time periods after the logical starting point to an idle and available state.

[0112] Next, iterate through all aggregated task blocks with time windows. According to the global order determined by the task dependency graph and the optimization objective (such as minimizing the overall completion time), try to schedule a specific execution period for each aggregated task block within its time window. When scheduling, the start time of the candidate service resource item allocated to the aggregated task block needs to be calculated. This start time must meet three conditions simultaneously: First, it must be later than the latest of the end time of the last scheduled task on the timeline of the corresponding candidate service resource item, the estimated ready time (i.e., the estimated time for it to complete its existing tasks or arrive at the starting point), and the actual completion time of all tasks of the first task in the aggregated task block. Second, the time required for the candidate service resource item to move from its previous task position (or current position) to the execution location of the aggregated task block must be included in the calculation of the start time. Third, the execution duration of the aggregated task block (estimated based on all its subtasks) must be fully contained within the scheduled time period, and the time period must not exceed the time window of the aggregated task block itself.

[0113] During the scheduling process, conflict avoidance calculations are continuously performed. When it is detected that the time period attempted to be scheduled for the current aggregate task block overlaps with other aggregate task blocks already scheduled on the timeline of the same candidate service resource item, or that a subsequent aggregate task block cannot be scheduled due to time window limitations, the start time of the current aggregate task block is dynamically adjusted (sliding within its time window) or its candidate service resource item allocation is re-evaluated to resolve the conflict.

[0114] Finally, when all aggregated task blocks are successfully scheduled onto the timeline of each candidate service resource item, an initial scheduling scheme is generated. This transforms the logically feasible task-resource matching relationship into an operable schedule under strict time and space constraints, providing a reliable foundation for subsequent real-time conflict detection and dynamic adjustment.

[0115] Step S33: Based on the initial scheduling scheme, changes in the real-time status of service resources, and external environmental factors, conflict detection and optimization adjustments are performed to generate optimized scheduling information including the main execution scheme and at least one backup execution scheme.

[0116] The system monitors real-time changes in the status of candidate service resources (e.g., the actual ready time of the scheduled resource is later than the estimated time due to delays in the previous task, sudden failures, or withdrawals) and external environmental factors (e.g., changes in traffic conditions affecting travel time, weather changes causing service tasks to be unable to execute). This dynamic information is compared with the scheduled time, resources, and path of each arrangement in the initial scheduling plan to detect conflicts. Conflict types include, but are not limited to, time conflicts (e.g., resources cannot arrive at the scheduled time), resource conflicts (e.g., allocated resources suddenly become unavailable), and conditional conflicts (e.g., environmental changes cause task prerequisites to be invalid). Once a conflict is detected, optimization adjustments are immediately triggered. Based on the real-time status changes of service resources and external environmental factors, and under the premise of satisfying all task dependencies and preference constraints, local replanning is performed. That is, based on the initial scheduling plan, conflicts are resolved by adjusting the candidate service resources of the affected tasks (re-selecting from other candidate service resources), postponing execution time (sliding within a time window), or modifying the execution path, resulting in a main execution plan.

[0117] Next, after generating the primary execution plan, the current global scheduling state is copied, including the real-time status of all candidate service resource items, the dependency graph of incomplete tasks, and preference constraints. Based on this global scheduling state, alternative execution plans are generated. For example, the scheduling objective can be adjusted to "minimize the total service cost while ensuring the total completion time does not exceed 120% of the primary execution plan." Driven by this scheduling objective, candidate service resource items with lower prices but potentially lower efficiency are selected, or tasks can be scheduled during off-peak hours to obtain discounts, thereby generating a lower-cost but longer-lasting alternative execution plan. Alternatively, the scheduling objective can be adjusted to "maximize the use of alternative candidate service resource items that are different from the primary execution plan to diversify risk." Driven by this scheduling objective, alternative execution plans are generated by selecting from unused but equally qualified candidate service resource items in the primary execution plan, and then re-performing conflict detection and optimization adjustments.

[0118] By encapsulating the backup execution plan with the main execution plan, optimized scheduling information is obtained. This allows for rapid scheduling based on the backup execution plan when the main execution plan partially fails due to uncertainties in the actual execution process. This ensures that the service execution process can continue reliably in the face of various unexpected disturbances, ultimately guaranteeing the success rate of service delivery and the stability of user experience.

[0119] Step S34: Encode the optimized scheduling information into instructions and confirm resource locking to generate matching scheduling information.

[0120] The main execution plan in the optimized scheduling information is parsed and transformed according to a predefined machine-readable instruction protocol format (such as a specific data structure based on JSON or Protocol Buffers) to generate a series of scheduling instructions with clear operational semantics. The candidate service resource items, allocated execution time windows, and movement path planning corresponding to each aggregate task block in the main execution plan are formatted and transformed according to the predefined machine-readable instruction protocol format (such as a specific data structure based on JSON or Protocol Buffers).

[0121] Simultaneously, a lock request is sent to the resource management service for all candidate service resource items involved in the scheduling instruction. The lock request includes the resource identifier and the time interval to be locked. After receiving the lock request, the resource management service verifies in real time whether these candidate service resource items are still available at the current time and do not conflict with the lock interval. After the verification is successful, their status is marked as "pre-occupied" and a successful lock confirmation is sent back.

[0122] After receiving successful confirmation of all necessary resource locks, the coded scheduling instructions are bound to the resource lock credentials, and matching scheduling information is generated. The instruction encoding ensures that the execution end can understand the task unambiguously, and the resource lock confirmation reserves the necessary resources in advance. Thus, the exclusivity and executability of the scheduling are established the moment the scheduling instruction is issued, providing a crucial final guarantee for the seamless and reliable transition of the service from the planning stage to the actual execution stage.

[0123] This application also provides a system for ordering home delivery services based on AI voice, see reference. Figure 5 As shown, the system for ordering home delivery services based on AI voice includes: Data processing module 10 is used to perform natural language processing and contextual understanding based on the received user voice commands and the recognized environmental sensor data to obtain enhanced semantic information; Information analysis module 20 is used to perform service intent recognition and sentiment analysis based on enhanced semantic information, and generate service request and user preference information; The matching and scheduling module 30 is used to perform service matching and service scheduling operations based on service requests and user preference information, and generate matching and scheduling information. The service optimization module 40 is used to obtain feedback information on matching scheduling information, and to perform service confirmation and environmental awareness based on the feedback information in order to optimize the home delivery service.

[0124] Optionally, the data processing module 10 is also used for: Based on user voice commands and environmental sensor data, speech-to-text conversion and emotion state recognition are performed to obtain speech-to-text information and user emotion data. Anomaly detection processing is performed on environmental sensor data to obtain environmental event information. Based on user emotion data, the weights of user voice commands and environmental sensor data are calculated to obtain modal weight coefficients. Based on the modal weight coefficients, weighted features and cross-modal alignment operations are performed on speech text information and environmental event information to obtain semantic fusion features; Based on semantic fusion features, contextual association and information completion processing are performed by combining historical service records to obtain enhanced semantic information.

[0125] Optionally, the data processing module 10 is also used for: Based on user emotion data, feature extraction is performed on emotion state, emotion intensity, and emotion stability to obtain emotion feature parameters; Based on the emotional feature parameters, attention allocation pattern analysis is performed to generate a weight allocation strategy for user voice commands and environmental sensor data. Based on the weight allocation strategy, the weight values ​​of user voice commands and environmental sensor data are mapped to obtain the voice command weight value and the environmental sensor weight value. The voice command weight values ​​and environmental sensor weight values ​​are integrated and normalized to obtain the modal weight coefficients.

[0126] Optionally, the information analysis module 20 is also used for: Based on enhanced semantic information, semantic role labeling, dependency parsing, and multimodal information alignment operations are performed to obtain structured semantic units; Based on structured semantic units, implicit intention reasoning and emotional state transition analysis are performed to obtain multi-granularity intention classification and emotional evolution trajectory; Based on multi-granularity intent classification and sentiment evolution trajectory, and combined with user historical behavior characteristics, intent decomposition and sentiment-guided preference identification are performed to obtain service tasks and preference constraints. The service tasks and preference constraints are parameterized and prioritized to generate service requests and user preference information.

[0127] Optionally, the information analysis module 20 is also used for: The service timing information and service entity relationships included in the structured semantic units are extracted and aligned to obtain the service operation sequence; Based on the service operation sequence, and combined with historical service records, sequence matching and operation difference comparison are performed to identify explicit intent sequences and implicit intent information; Based on explicit intention sequences and implicit intention information, an intention state transition diagram is constructed and a time series analysis of emotional intensity is performed to obtain the intention evolution path and the emotional intensity change curve. Multi-level aggregation and trajectory smoothing are performed on the intention evolution path and emotion intensity change curve to obtain multi-granularity intention classification and emotion evolution trajectory.

[0128] Optionally, the matching and scheduling module 30 is also used for: Based on service request and user preference information, and combined with real-time status and location information of service resources, conditional filtering and scoring are performed to obtain multiple candidate service resource items; Based on multiple candidate service resource items, task decomposition and scheduling strategy calculation are performed on multiple service tasks included in the service request and preference constraints in user preference information to obtain an initial scheduling scheme; Based on the initial scheduling scheme, changes in the real-time status of service resources, and external environmental factors, conflict detection and optimization adjustments are performed to generate optimized scheduling information, including the main execution scheme and at least one backup execution scheme. The optimized scheduling information is encoded with instructions and resource locking is confirmed to generate matching scheduling information.

[0129] Optionally, the matching and scheduling module 30 is also used for: Based on the logical relationships and preference constraints among multiple candidate service resource items, task dependency analysis and constraint satisfaction verification are performed to obtain a task dependency graph. Based on the task dependency graph and multiple candidate service resource items, service capability matching and parallel execution path planning are performed to obtain a feasible execution sequence; Based on the feasible execution sequence, multiple service task subsequences that can be processed by the same candidate service resource item are aggregated to generate an aggregated task block with a time window; Based on the aggregated task block with time windows, the timeline is filled and conflict avoidance calculations are performed by combining the real-time location and estimated ready time of each candidate service resource item to generate an initial scheduling scheme.

[0130] This application also provides a device for AI voice-based ordering and home delivery services. The device includes: a system for AI voice-based ordering and home delivery services, a memory storing executable instructions, and a processor configured to execute the executable instructions in the memory to implement a method for AI voice-based ordering and home delivery services.

[0131] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A method for ordering home delivery services based on AI voice, characterized in that, include: Based on the received user voice commands and the identified environmental sensor data, natural language processing and contextual understanding are performed to obtain enhanced semantic information; Based on the enhanced semantic information, service intent recognition and sentiment analysis are performed to generate service requests and user preference information; Based on the service request and the user preference information, perform service matching and service scheduling operations to generate matching and scheduling information; The system obtains feedback information from the matching and scheduling information, and performs service confirmation and environmental awareness based on the feedback information to optimize the home delivery service.

2. The method for ordering home delivery service based on AI voice according to claim 1, characterized in that, The step of performing natural language processing and contextual understanding based on the received user voice commands and recognized environmental sensor data to obtain enhanced semantic information includes: Based on the user's voice commands and the environmental sensor data, speech-to-text conversion and emotion state recognition are performed to obtain speech-to-text information and user emotion data. Anomaly detection processing is also performed on the environmental sensor data to obtain environmental event information. Based on the user emotion data, the weights of the user voice commands and the environmental sensor data are calculated to obtain modal weight coefficients. Based on the modal weight coefficients, weighted features and cross-modal alignment operations are performed on the speech text information and the environmental event information to obtain semantic fusion features; Based on the semantic fusion features, context association and information completion processing are performed by combining historical service records to obtain the enhanced semantic information.

3. The method for ordering home delivery service based on AI voice according to claim 2, characterized in that, The step of calculating the weights of the user's voice commands and the environmental sensor data based on the user's emotional data to obtain the modal weight coefficients includes: Based on the user's emotional data, features of emotional state, emotional intensity, and emotional stability are extracted to obtain emotional feature parameters; Based on the emotional feature parameters, attention allocation pattern analysis is performed to generate a weight allocation strategy for the user's voice command and the environmental sensing data. Based on the weight allocation strategy, the weight values ​​of the user voice command and the environmental sensing data are mapped to obtain the voice command weight value and the environmental sensing weight value. The voice command weight value and the environmental sensing weight value are integrated and normalized to obtain the modal weight coefficient.

4. The method for ordering home delivery service based on AI voice according to claim 1, characterized in that, The step of performing service intent recognition and sentiment analysis based on the enhanced semantic information to generate service request and user preference information includes: Based on the enhanced semantic information, semantic role labeling, dependency parsing, and multimodal information alignment operations are performed to obtain structured semantic units; Based on the structured semantic units, implicit intention reasoning and emotional state transition analysis are performed to obtain multi-granularity intention classification and emotional evolution trajectory; Based on the multi-granularity intent classification and the emotion evolution trajectory, and combined with the user's historical behavior characteristics, intent decomposition and emotion-guided preference identification are performed to obtain service tasks and preference constraints. The service task and the preference constraint are parameterized and prioritized to generate the service request and the user preference information.

5. The method for ordering home delivery service based on AI voice according to claim 4, characterized in that, The steps of performing implicit intention reasoning and sentiment state transition analysis based on the structured semantic units to obtain multi-granularity intention classification and sentiment evolution trajectory include: The service timing information and service entity relationships included in the structured semantic unit are extracted and aligned to obtain the service operation sequence; Based on the service operation sequence, sequence matching and operation difference comparison are performed by combining historical service records to identify explicit intent sequences and implicit intent information; Based on the explicit intention sequence and the implicit intention information, an intention state transition diagram is constructed and a time series analysis of emotional intensity is performed to obtain the intention evolution path and the emotional intensity change curve. The intention evolution path and the emotion intensity change curve are subjected to multi-level aggregation and trajectory smoothing to obtain the multi-granularity intention classification and the emotion evolution trajectory.

6. The method for ordering home delivery service based on AI voice according to claim 1, characterized in that, The step of performing service matching and service scheduling operations based on the service request and the user preference information, and generating matching and scheduling information, includes: Based on the service request and the user preference information, and combined with the real-time status of service resources and the location information of service resources, condition filtering and scoring are performed to obtain multiple candidate service resource items; Based on the multiple candidate service resource items, task decomposition and scheduling strategy calculation are performed on the multiple service tasks included in the service request and the preference constraints in the user preference information to obtain an initial scheduling scheme; Based on the initial scheduling scheme, the real-time changes in the status of the service resources, and external environmental factors, conflict detection and optimization adjustments are performed to generate optimized scheduling information including the main execution scheme and at least one backup execution scheme. The optimized scheduling information is encoded with instructions and resource locking is confirmed to generate the matching scheduling information.

7. The method for ordering home delivery service based on AI voice according to claim 6, characterized in that, The step of performing task decomposition and scheduling strategy calculation on multiple service tasks included in the service request and preference constraints in the user preference information based on the multiple candidate service resource items to obtain an initial scheduling scheme includes: Based on the logical relationships between the multiple candidate service resource items and the preference constraints, task dependency analysis and constraint satisfaction verification are performed to obtain a task dependency graph. Based on the task dependency graph and the multiple candidate service resource items, service capability matching and parallel execution path planning are performed to obtain a feasible execution sequence; Based on the feasible execution sequence, multiple service task sub-sequences that can be processed by the same candidate service resource item are aggregated to generate an aggregated task block with a time window; Based on the aggregated task block with time windows, and combining the real-time location and estimated ready time of each candidate service resource item, timeline filling and conflict avoidance calculations are performed to generate the initial scheduling scheme.

8. A system for ordering home delivery services based on AI voice, characterized in that, The system for ordering home delivery services based on AI voice is applied to the method for ordering home delivery services based on AI voice as described in any one of claims 1 to 7, wherein the system for ordering home delivery services based on AI voice includes: The data processing module is used to perform natural language processing and contextual understanding based on the received user voice commands and the recognized environmental sensor data to obtain enhanced semantic information; The information analysis module is used to perform service intent recognition and sentiment analysis processing based on the enhanced semantic information, and to generate service request and user preference information; The matching and scheduling module is used to perform service matching and service scheduling operations based on the service request and the user preference information, and generate matching and scheduling information. The service optimization module is used to obtain feedback information from the matching and scheduling information, and to perform service confirmation and environmental awareness based on the feedback information in order to optimize the home delivery service.

9. A device for ordering home delivery services based on AI voice, characterized in that, The device for AI voice-based ordering and home delivery service includes the system for AI voice-based ordering and home delivery service as described in claim 8, a memory storing executable instructions, and a processor configured to execute the executable instructions in the memory to implement the method for AI voice-based ordering and home delivery service as described in any one of claims 1 to 7.