Passenger cabin service intelligent auxiliary system based on multi-mode interaction technology

Through multimodal information collection, intelligent service scheduling and multimodal feedback technology, combined with blockchain security mechanism, the problems of single interaction mode, poor environmental adaptability and inaccurate service matching of traditional cabin service systems are solved, and an efficient and safe cabin service system is achieved.

CN120338799AInactive Publication Date: 2025-07-18XINJIANG JIAOTONG VOCATIONAL & TECHNICAL UNIVERSITY

Patent Information

Application Number
CN202510408328.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional cabin service systems have problems such as single interaction methods, poor environmental adaptability, inaccurate service matching, and low efficiency of flight attendants. The existing multimodal interaction technology has failed to effectively solve technical bottlenecks such as dynamic fusion of multi-source data, scenario adaptive decision-making and privacy protection.

Method used

The multimodal information acquisition module is used to obtain voice, vision, touch and physiological signals, combine the space-time encoding model and knowledge graph to analyze passenger needs, dynamically optimize resource allocation through the intelligent service scheduling module, and use the multimodal feedback module to provide voice, vision, touch and environmental linkage response, and use blockchain technology to ensure data security through the data management and self-learning module.

Benefits of technology

It realizes accurate identification and rapid response to passenger needs, improves service efficiency and passenger satisfaction, reduces the workload of flight attendants, ensures data integrity and privacy security, and enhances the environmental adaptability and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338799A_ABST
    Figure CN120338799A_ABST
Patent Text Reader

Abstract

The invention provides a passenger cabin service intelligent auxiliary system based on a multi-modal interaction technology. The passenger cabin service intelligent auxiliary system comprises a multi-modal information acquisition module, a passenger intention identification module, an intelligent service scheduling module, a multi-modal feedback module and a data management and self-learning module. According to the passenger cabin service intelligent auxiliary system based on the multi-modal interaction technology, the demand identification accuracy is high, the emergency response speed is extremely high, the workload of a steward is greatly reduced, and the service efficiency is greatly improved; health monitoring and early warning accuracy is extremely high, block chain evidence storage guarantees data integrity and traceability, and safety is enhanced; the federated learning-driven recommendation system greatly improves the satisfaction degree of passengers, the holographic projection interaction naturalness score is high, the personalized experience is optimized, various complex scenes such as day and night, bumping and strong light are supported, the system availability is extremely high, and the environmental adaptability is good; and multi-modal deep fusion is realized, and high-precision interaction is still kept under environmental interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent services, and specifically relates to a cabin service intelligent assistance system based on multi-modal interaction technology, which is applicable to the cabin environment of transportation means such as aviation and high-speed rail. By integrating multi-modal interaction methods such as voice, gesture, vision, and touch, and combining dynamic scene perception and intelligent decision-making technology, it realizes the accurate recognition of passenger needs, the optimal scheduling of service resources, and personalized service recommendations, significantly improving the cabin service efficiency and passenger experience. Background Art

[0002] Traditional cabin services rely on manual responses to passenger needs, suffering from problems such as a single interaction method, low efficiency, response latency, and insufficient service standardization. In the prior art, some systems adopt single-modal interaction (such as voice or touch), but still have the following defects:

[0003] Poor environmental adaptability: Voice recognition is easily interfered by cabin noise, and touch screen operations are inconvenient to use in a bumpy environment;

[0004] Inaccurate service matching: It is unable to provide proactive services in combination with passengers' personalized needs (such as special diets and health status);

[0005] Insufficient crew coordination: Task allocation relies on manual coordination, and the resource allocation efficiency is low in emergency scenarios.

[0006] Although existing patents (such as a multi-modal interaction intelligent control system disclosed in a Chinese patent with the application number 202410656932.5) have proposed multi-modal interaction solutions, they have not solved technical bottlenecks such as dynamic fusion of multi-source data, scene adaptive decision-making, and privacy protection. Therefore, there is an urgent need for a cabin service system that deeply integrates multi-modal interaction technology and supports real-time intent recognition and intelligent scheduling. Summary of the Invention

[0007] The purpose of the present invention is to solve the problems raised in the background art, and provides a cabin service intelligent assistance system based on multi-modal interaction technology. Through multi-modal data collection, dynamic scene perception, intelligent decision-making, and multi-channel feedback technology, it realizes the accurate recognition and rapid response of passenger needs, optimizes the crew resource scheduling, and improves service efficiency and passenger satisfaction.

[0008] The specific technical solutions are as follows:

[0009] A cabin service intelligent assistance system based on multi-modal interaction technology, comprising:

[0010] A multi-modal information collection module, used to obtain voice, vision, touch, and physiological signal inputs;

[0011] The passenger intention recognition module analyzes explicit and implicit service demands by fusing multi-source data through a spatio-temporal coding model and a knowledge graph;

[0012] The intelligent service scheduling module dynamically optimizes task priorities and resource allocation strategies based on federated learning;

[0013] The multi-modal feedback module provides voice, visual, tactile, and environmental linkage responses;

[0014] The data management and self-learning module encrypts and stores service records using blockchain technology.

[0015] The above intelligent cabin service assistant system based on multi-modal interaction technology, wherein: the multi-modal information acquisition module includes a voice unit, a visual unit, a touch unit, and a physiological signal unit, wherein:

[0016] The voice unit includes a microphone array deployed on the seat, which combines beamforming technology and noise reduction algorithms to suppress engine noise and extract passenger voice commands;

[0017] The visual unit includes in-cabin cameras, infrared sensors, thermal imaging cameras, and image recognition devices for capturing passenger gestures, expressions, physical signs, and night-time gesture trajectories;

[0018] The touch unit includes a touch screen integrated into the seat armrest and a high-precision pressure sensor array, which supports gesture sliding input and emergency calls, and includes a pressure distribution sensor array to analyze passenger comfort through sitting posture changes and trigger adaptive seat adjustment;

[0019] The physiological signal unit includes a non-invasive electroencephalogram sensor and a seat pressure distribution sensor to collect passenger intention prediction data and sitting posture comfort information, wherein the non-invasive electroencephalogram sensor predicts passenger intentions through an LSTM neural network and cross-verifies with voice commands.

[0020] The above intelligent cabin service assistant system based on multi-modal interaction technology, wherein: the passenger intention recognition module fuses multi-source data of voice, vision, touch, and physiological signals based on a deep learning model to analyze passengers' explicit and implicit demands; at the same time, it combines a knowledge graph to associate passengers' historical data with real-time situations, generates personalized service suggestions, and integrates a voice emotion analysis sub-module to judge the emotional state through MFCC feature extraction and a convolutional neural network.

[0021] The above intelligent cabin service assistant system based on multi-modal interaction technology, wherein: the intelligent service scheduling module is used for dynamic priority division and resource optimization scheduling, wherein:

[0022] Dynamic priority division, allocates service priorities according to the urgency of demand and the passenger's state;

[0023] Resource optimization scheduling, constructing a real-time three-dimensional in-cabin map based on LiDAR and SLAM technologies, tracking the positions of flight attendants and equipment status, pushing the optimal task path through AR glasses, and using a vibrotactile device to remind of emergency tasks, with the vibration intensity being positively correlated with the task urgency.

[0024] The above-mentioned intelligent cabin service assistance system based on multimodal interaction technology, wherein: the multimodal feedback module is used for voice feedback, visual feedback, tactile feedback, and environmental linkage feedback, where:

[0025] Voice feedback, broadcasting service confirmation information to designated seats through directional sound field technology;

[0026] Visual feedback, the seat screen displays the service progress, and the color of the cabin ceiling LED light strip indicates the service status;

[0027] Tactile feedback, the intelligent bracelet worn by the flight attendant prompts task details through vibration, and synchronously triggers seat vibration and AR-HUD warning in emergency scenarios;

[0028] Environmental linkage feedback, the odor release device automatically adjusts the type and concentration of fragrance according to the passenger emotion index, and the air conditioning system is linked to adjust the temperature.

[0029] The above-mentioned intelligent cabin service assistance system based on multimodal interaction technology, wherein: the data management and self-learning module is used to store passengers' historical service data and preference information, support personalized recommendation; adopt federated learning technology to update the intent recognition model to ensure data privacy and security; adopt blockchain encryption to store and prove key service records to ensure data immutability and traceability, and support multi-node distributed storage and audit traceability.

[0030] The above-mentioned intelligent cabin service assistance system based on multimodal interaction technology, wherein: the physiological signal unit in the multimodal information collection module integrates an electrocardiogram sensor and a blood oxygen sensor to continuously monitor the passenger's heart rate variability and blood oxygen saturation, and triggers a health warning and preferentially schedules medical assistance when abnormal values are detected.

[0031] The above-mentioned intelligent cabin service assistance system based on multimodal interaction technology, wherein: the passenger intent recognition module includes an environment adaptive sub-module that dynamically adjusts the interaction modality weight according to the real-time in-cabin environment parameters.

[0032] The above-mentioned intelligent cabin service assistance system based on multimodal interaction technology, wherein: the multimodal feedback module integrates a holographic projection virtual assistant, generates a three-dimensional virtual image through AID technology, supports multilingual interaction, entertainment service recommendation, and emergency operation guidance, and adjusts the projection position based on the passenger's line of sight direction.

[0033] The above intelligent cabin service assistance system based on multimodal interaction technology, wherein: it supports cross-device collaborative interaction, passengers can initiate service requests through personal intelligent terminals and synchronously receive multimodal feedback, and the multimodal feedback module includes redundant design, and in emergency scenarios, AR-HUD warnings, seat vibrations, and directional voice broadcasts are triggered simultaneously

[0034] The present invention has the following beneficial effects:

[0035] The intelligent cabin service assistance system based on multimodal interaction technology provided by the present invention has a relatively high demand recognition accuracy, an extremely fast emergency response speed, a significant reduction in the workload of flight attendants, and a significant improvement in service efficiency; the health monitoring and early warning accuracy is extremely high, and blockchain evidence storage ensures data integrity and traceability, enhancing security; the recommendation system driven by federated learning greatly improves passenger satisfaction, the holographic projection interaction has a relatively high naturalness score, the personalized experience is optimized, it supports various complex scenarios such as day and night, turbulence, and strong light, the system availability is extremely high, and the environmental adaptability is good; multimodal deep fusion, voice, vision, touch, and physiological signals cooperate, and high-precision interaction is still maintained under environmental interference, dynamic intelligent decision-making, federated learning optimizes the scheduling strategy, LiDAR three-dimensional maps improve resource utilization, full-scene adaptability, technologies such as thermal imaging and EEG prediction cover complex environments, the system robustness is significantly enhanced, it has a secure and trustworthy architecture, the combination of blockchain evidence storage and federated learning realizes "usable but invisible" data, and the privacy protection level is high Brief Description of the Drawings

[0036] Figure 1 It is a schematic diagram of the architecture of the intelligent cabin service assistance system based on multimodal interaction technology provided by the embodiment of the present invention

[0037] Figure 2 It is a bar chart of the key indicators of the intelligent cabin service assistance system based on multimodal interaction technology provided by the embodiment of the present invention

[0038] Figure 3 It is a line chart of the key indicators of the intelligent cabin service assistance system based on multimodal interaction technology provided by the embodiment of the present invention

[0039] Figure 4 It is a bar graph of the key indicators of the intelligent cabin service assistance system based on multimodal interaction technology provided by the embodiment of the present invention Detailed Embodiments

[0040] The technical solution of the present invention will be further described below in conjunction with the drawings and through specific embodiments

[0041] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams rather than physical diagrams, and should not be construed as a limitation to this patent; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.

[0042] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if terms such as "upper", "lower", "left", "right", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the attached drawings, and it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms used to describe the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as a limitation to this patent. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0043] In the description of the present invention, unless otherwise clearly defined and limited, if terms such as "connection" are used to indicate the connection relationship between components, this term should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0044] The in-cabin service intelligent assistance system based on multi-modal interaction technology provided in this embodiment, as Figures 1 - 4 shown, includes: a multi-modal information acquisition module, a passenger intention recognition module, an intelligent service scheduling module, a multi-modal feedback module, and a data management and self-learning module.

[0045] Among them, Figure 2 the bar chart shown presents the numerical values of key indicators such as the health monitoring accuracy, response time, and passenger satisfaction of this system, facilitating an intuitive comparison of the performance of different indicators; Figure 3 the line chart shown presents the trends of the key indicators of the system, facilitating an understanding of the performance of this system in different aspects; Figure 4 the bar graph shown presents the key indicators of the system in the form of horizontal bars, providing a visual perspective for an intuitive comparison of the performance of different indicators.

[0046] Among them, the multi-modal information acquisition module is used to obtain voice, visual, touch, and physiological signal inputs;

[0047] Among them, the passenger intention recognition module fuses multi-source data through a spatio-temporal encoding model and a knowledge graph to analyze explicit and implicit service requirements. Among them, the spatio-temporal encoding model adopts a multi-modal spatio-temporal fusion network (MSTFN) to achieve multi-source data fusion in the following ways:

[0048]

[0049] Among them, x i , x j are different modal data streams (such as voice stream, visual stream), δ is the time offset used to align the timestamps of multi-modal data; k is the time step index representing discrete moments in the data stream; p(x i , x j ) is the joint probability distribution that measures the correlation between the two modal data; p(x i ), p(x j ) are the marginal probability distributions of single-modal data;

[0050] Align the timestamp offsets of multi-modal data streams and construct a joint feature vector F:

[0051] F = α·GCN(G spatial ) + β·TCN(S temporal );

[0052] Among them, Gspatial is the gesture trajectory graph structure (nodes are joint points, edges are motion trajectories), GCN is the graph convolutional network; Stemporal is the voice time series signal, TCN is the temporal convolutional network; α, β are modal credibility weights that are dynamically adjusted according to the environmental light intensity and noise level;

[0053]

[0054] γ is the adjustment coefficient that controls the influence intensity of environmental parameters on the weights;

[0055] LightIntensity is the environmental light intensity (unit: lux);

[0056] NoiseLevel is the environmental noise level (unit: dB).

[0057] The proposed spatio-temporal encoding joint loss function can solve the problems of timestamp alignment and spatial feature association of multi-modal data, and introduce modal credibility weights to dynamically adjust the contribution degree of each modality in the fusion.

[0058] Among them, the intelligent service scheduling module dynamically optimizes the task priority and resource allocation strategy based on federated learning; the federated learning adopts a dynamic weight federated learning framework (DW-FL), and the client weight w k is determined by the data quality metric Qk Calculation:

[0059] Where:

[0060] w k : The federated learning weight of client k;

[0061] Q k : The client data quality metric;

[0062] SNR (voice): Signal-to-noise ratio of the voice signal (unit: dB);

[0063] IoU (gesture annotation): Intersection over Union of gesture annotation (range: 0 - 1);

[0064] N: The total number of clients participating in federated learning.

[0065] And inject Laplace noise L(0,λ) during gradient aggregation to ensure privacy;

[0066] Among them, the multimodal feedback module provides voice, vision, touch, and environmental linkage responses; among them, in emergency scenarios, the redundant feedback control algorithm (RFC) is adopted, and the feedback channel is selected and the delay is optimized through the priority matrix P i,j Select the feedback channel and optimize the delay:

[0067]

[0068] Where:

[0069] Cselected: The set of selected feedback channels (such as voice, vision, touch);

[0070] Dc: The feedback delay of channel c (unit: second);

[0071] Tmax: The maximum delay allowed by the scenario (such as Tmax = 0.3s in emergency scenarios);

[0072] Dynamically allocate federated learning weights based on client data quality, suppress interference from low-quality data, and ensure the irreversibility of user data by introducing differential privacy noise.

[0073] The algorithm process is as follows:

[0074] 1. Client weight calculation:

[0075] The weight w of each client k is determined by the data quality metric Q k (such as signal-to-noise ratio, annotation consistency)) and calculated as:

[0076]

[0077] 2. Gradient Aggregation and Noise Injection:

[0078] When updating the global model, Laplace noise L(0, λ) is added to the gradient:

[0079]

[0080] where λ is dynamically adjusted according to the privacy budget ∈:

[0081] Global model gradient;

[0082] Local gradient of client k;

[0083] L(0, λ): Laplace noise;

[0084] Δf: Gradient sensitivity (maximum gradient change);

[0085] ∈: Privacy budget (the smaller the value, the stronger the privacy protection).

[0086] Among them, the data management and self-learning module uses blockchain technology to encrypt and record service records.

[0087] Through the collaborative work of the five core modules, multi-modal data collection, intelligent intent parsing, dynamic resource scheduling, multi-channel feedback, and secure data management are realized, forming a closed-loop service system, significantly improving the efficiency of cabin service and system reliability.

[0088] More comprehensively, the multi-modal information collection module includes a voice unit, a visual unit, a touch unit, and a physiological signal unit, where:

[0089] The voice unit includes a microphone array deployed on the seat. Combining beamforming technology and noise reduction algorithms, it suppresses engine noise and extracts passenger voice commands (such as "Need a blanket", "Order food");

[0090] The visual unit includes in-cabin cameras, infrared sensors, thermal imaging cameras, and image recognition devices, which are used to capture passenger gestures (such as raising hands), expressions, physical signs (such as abnormal heart rate), and night gesture trajectories, and analyze requirements through pose estimation algorithms and image recognition technologies;

[0091] The touch unit includes a touch screen integrated on the seat armrest and a high-precision pressure sensor array, which supports gesture sliding input and emergency calls, and includes a pressure distribution sensor array, which analyzes passenger comfort through sitting posture changes and triggers adaptive seat adjustment;

[0092] The physiological signal unit includes a non-invasive electroencephalogram sensor and a seat pressure distribution sensor, which collect data for predicting passenger intentions and information on sitting posture comfort. The non-invasive electroencephalogram sensor predicts passenger intentions through an LSTM neural network and cross-verifies with voice commands.

[0093] Voice unit: In the cabin noise environment, beamforming and noise reduction technologies significantly improve the accuracy of voice command recognition.

[0094] Vision unit: The thermal imaging camera supports night gesture recognition with high accuracy and can accurately detect abnormal body temperatures.

[0095] Touch unit: The pressure distribution sensor triggers seat adjustment through sitting posture analysis, significantly improving passenger comfort.

[0096] Physiological signal unit: The EEG sensor cross-verifies predicted intentions with voice commands, significantly improving the recognition rate of implicit demands.

[0097] More comprehensively, the passenger intention recognition module fuses multi-source data such as voice, vision, touch, and physiological signals based on deep learning models (such as Transformer and LSTM neural networks), and introduces an adaptive multi-modal attention mechanism (AMA):

[0098] Environmental parameter encoding: Encodes the light intensity L and the bump index B into an environmental vector E:

[0099] E = ReLU(W e ·[L, B] + b e );

[0100] Where: L: Environmental light intensity (unit: lux); B: Bump index (range: 0 - 5, 5 being the most severe); W e , b e : Learnable parameter matrix and bias term; ReLU: Rectified linear unit activation function.

[0101] Modal attention score calculation:

[0102] a i = σ(W a ·[F i ||E] + b a );

[0103] Where: a i : Attention score of modality i; Fi: Feature vector of modality i; ||: Vector concatenation operation; Wa, ba: Learnable parameter matrix and bias term; σ: Sigmoid function, output range (0, 1).

[0104] The attention score a of each modality feature F i ​i controlled by the environmental vector E;

[0105] where σ is the Sigmoid function, and ‖ represents vector concatenation;

[0106] Dynamically adjust the weights of each modality, and the final intention prediction is:

[0107] where: y intent : intention prediction probability distribution; Wy: learnable parameter matrix;

[0108] Analyze the explicit needs (such as catering requests) and implicit needs (such as emotional anxiety) of passengers; at the same time, combine the knowledge graph to associate the passenger's historical data (such as flight records, allergy history) with the real-time situation (such as in-cabin temperature, flight phase), generate personalized service suggestions, and integrate the voice emotion analysis sub-module to judge the emotional state through MFCC feature extraction and convolutional neural network.

[0109] The deep learning model (Transformer + LSTM) fuses multi-source data, and the accuracy of explicit need recognition is greatly improved; the knowledge graph associates historical data to generate personalized suggestions, and the recommendation matching degree is greatly enhanced; the voice emotion analysis sub-module has extremely high accuracy in judging the emotional state.

[0110] More perfectly, the intelligent service scheduling module is used for dynamic priority division and resource optimization scheduling, where:

[0111] Dynamic priority division, allocate service priorities according to the urgency of needs (such as medical assistance > regular services) and passenger status (such as VIP level);

[0112] Resource optimization scheduling, build a real-time three-dimensional in-cabin map based on LiDAR and SLAM technologies, and adopt an improved three-dimensional path planning algorithm (LSE-A*), whose cost function is:

[0113] f(n) = g(n) + h(n) + λ·R(n), where: g(n): the actual movement cost from the starting point to node n (such as distance, time); h(n): heuristic function, estimating the cost from node n to the end point; λ: dynamic risk weight coefficient (range: 0 - 1); R(n): dynamic risk weight;

[0114] where, d obs : the distance between node n and the nearest obstacle (unit: meter); d safe : preset safety distance threshold (unit: meter); k: adjustment coefficient, controlling the steepness of the risk curve;

[0115] Track the position of the flight attendants and the status of the equipment, push the optimal task path through AR glasses, and use a vibrotactile device to remind of urgent tasks, with the vibration intensity being positively correlated with the task urgency.

[0116] Dynamic priority division significantly shortens the response time for urgent needs; LiDAR and SLAM technologies are used to construct a three-dimensional map, greatly improving the path planning efficiency of flight attendants and significantly reducing the task completion time.

[0117] More comprehensively, the multimodal feedback module is used for voice feedback, visual feedback, tactile feedback, and environmental linkage feedback, where:

[0118] Voice feedback: Broadcast service confirmation information to the designated seat through directional sound field technology;

[0119] Visual feedback: The seat screen displays the service progress, and the color of the LED light strip on the cabin ceiling indicates the service status (e.g., green means "responded", red means "urgent");

[0120] Tactile feedback: The smart bracelet worn by the flight attendant uses vibration to prompt task details, and in emergency scenarios, the seat vibration (frequency range 1 - 50Hz) and AR-HUD warning are triggered synchronously;

[0121] Environmental linkage feedback: The odor release device automatically adjusts the type and concentration of the fragrance according to the passenger's emotion index (such as anxiety), and the air conditioning system is linked to adjust the temperature.

[0122] The multimodal feedback module can effectively improve the naturalness of interaction through voice directional broadcast, LED status prompt, tactile vibration, and fragrance adjustment.

[0123] Specifically, the data management and self-learning module is used to store the historical service data and preference information of passengers, support personalized recommendations (such as actively pushing low-sugar beverages); use federated learning technology to update the intent recognition model, inject Laplace noise L(0,λ) during gradient aggregation, and the privacy budget ∈ = 0.1; ensure data privacy and security; use blockchain encryption to deposit and prove key service records (such as medical events), ensure the data is immutable and traceable, and support multi-node distributed deposit and audit traceability.

[0124] Federated learning technology significantly shortens the update cycle of the intent recognition model and greatly reduces the risk of data privacy leakage; Blockchain deposit and proof ensure that service records are immutable and significantly improve the audit traceability efficiency.

[0125] Specifically, the physiological signal unit in the multimodal information collection module integrates an electrocardiogram (ECG) sensor and a blood oxygen sensor to continuously monitor the passenger's heart rate variability and blood oxygen saturation. When abnormal values are detected, a health warning is triggered and medical assistance is prioritized for scheduling.

[0126] The ECG and blood oxygen sensors monitor health indicators in real time, significantly improving the accuracy of early warning for sudden diseases and greatly shortening the response time of medical assistance.

[0127] Specifically, the passenger intention recognition module includes an environment adaptive sub-module, which dynamically adjusts the interaction modality weights according to the real-time in-cabin environment parameters (light intensity, temperature, bump index), and is specifically implemented through the AMA mechanism. For example, in a strong light environment, the visual interaction weight is reduced, and the priorities of voice and touch interactions are increased.

[0128] The environment adaptive sub-module dynamically adjusts the interaction weights, significantly reducing the mis-touch rate of visual interaction in a strong light environment and greatly improving the touch success rate in a bumpy scenario.

[0129] More perfectly, the multi-modal feedback module integrates a holographic projection virtual assistant, which generates a three-dimensional virtual image through AID (holographic aerial intelligent display) technology, supports multi-language interaction, entertainment service recommendation, and emergency operation guidance, and adjusts the projection position based on the passenger's line of sight direction.

[0130] The holographic projection virtual assistant supports multi-language interaction, significantly increasing the click-through rate of entertainment service recommendations and greatly improving the efficiency of emergency operation guidance.

[0131] More perfectly, cross-device collaborative interaction is supported. Passengers can initiate service requests through their personal smart terminals and synchronously receive multi-modal feedback. Moreover, the multi-modal feedback module includes a redundant design. In an emergency scenario, AR-HUD warnings, seat vibrations, and directional voice broadcasts are triggered simultaneously. The cross-device collaborative interaction has a wide coverage, and the service request synchronization delay is short; the redundant design significantly reduces the mis-operation rate in an emergency scenario.

[0132] Example 1: Health Monitoring and Emergency Medical Response

[0133] Scene Description

[0134] During the flight, passenger D suddenly experienced chest pain, sweating on the forehead, and pressed the "Medical Help" button on the touch screen of the seat armrest.

[0135] System Operation Process

[0136] 1. Data Collection:

[0137] 1.1 The touch unit records the button press signal;

[0138] 1.2 The physiological signal unit (ECG sensor) detects an abnormal heart rate variability (HRV decreased by 50%), and the blood oxygen saturation drops to 88%;

[0139] 1.3 The visual unit captures the passenger's pale face and rapid breathing through an infrared camera.

[0140] 2. Intention Analysis:

[0141] 2.1 The intention recognition module determines it as a "heart discomfort emergency event", and the emergency index reaches the highest level.

[0142] 2.2 The knowledge graph associates with the passenger's medical history (such as hypertension) and generates priority medical assistance suggestions.

[0143] 3. Intelligent scheduling:

[0144] 3.1 The scheduling module assigns the nearest flight attendant to carry an AED device to go there, and the AR glasses display the optimal path (avoiding the obstacle of the food cart).

[0145] 3.2 Synchronously notify the captain to contact the ground medical support.

[0146] 4. Multimodal feedback:

[0147] 4.1 The passenger screen shows "Medical assistance has departed", and the LED light strip on the cabin ceiling turns red.

[0148] 4.2 The flight attendant's bracelet reminds the urgency of the task with high-frequency vibration (50Hz).

[0149] 4.3 The air conditioning system lowers the temperature to 20°C and releases a mint fragrance to relieve the tense mood.

[0150] Experimental data

[0151] Health monitoring accuracy: The detection accuracy of abnormal conditions by the ECG and blood oxygen sensors is ≥98%.

[0152] Response time: It takes 45 seconds from the detection of abnormalities to the arrival of the flight attendant (traditional manual response requires ≥3 minutes).

[0153] Passenger satisfaction: The satisfaction score for emergency event handling is 4.9 / 5.0.

[0154] Example 2: Optimization of multimodal interaction in a bumpy environment at night

[0155] Scene description

[0156] During the flight at night, the flight encounters strong air currents and turbulence. Passenger E tries to adjust the seat angle through gestures, but the cabin lights are dim and the turbulence is intense.

[0157] System operation process

[0158] 1. Data collection:

[0159] 1.1 The visual unit (thermal imaging camera) captures the passenger's gesture trajectory (the waving amplitude ≥30 cm).

[0160] 1.2 The touch unit detects the pressure fluctuation of the seat armrest (the mis-touch rate increases due to turbulence).

[0161] 1.3 The environmental sensor detects that the bump index is Level 3 (moderate bump).

[0162] 2 Intention parsing:

[0163] 2.1 The environmental adaptation sub-module reduces the visual interaction weight (from 70% to 30%) and enhances the priority of voice and touch control.

[0164] 2.2 The intention recognition module combines the gesture trajectory and the voice command "straighten the seat" and parses it as "seat angle adjustment requirement".

[0165] 3 Intelligent scheduling:

[0166] 3.1 The scheduling module directly drives the seat motor to complete the angle adjustment, avoiding the movement of the flight attendant during the bump.

[0167] 4 Multimodal feedback:

[0168] 4.1 Directional voice broadcast "The seat has been straightened".

[0169] 4.2 The seat vibrates to indicate that the operation is completed (frequency 20Hz, duration 0.5 seconds).

[0170] Experimental data

[0171] Gesture recognition accuracy: The accuracy rate of thermal imaging gesture recognition in the night environment is 97% (only 65% for traditional cameras).

[0172] Effect of accidental touch suppression: The accidental touch operation rate in the bump scenario is reduced from 35% to 8%.

[0173] Interaction efficiency: The seat adjustment response time ≤ 0.3 seconds, and the passenger operation convenience score is 4.8 / 5.0.

[0174] Example 3: Multilingual service of holographic projection virtual assistant

[0175] Scenario description

[0176] Foreign passenger F needs to inquire about transfer information but doesn't understand the language and tries to wake up the holographic assistant through gestures.

[0177] System operation process

[0178] 1 Data collection:

[0179] 1.1 The visual unit captures the passenger's waving action (lasting for 2 seconds).

[0180] 1.2 The voice unit detects the non-native keyword "Transfer".

[0181] 1.3 The passenger's personal mobile phone is connected to the system via Bluetooth.

[0182] 2 Intention parsing:

[0183] 2.1 The holographic projection module is activated to generate a 3D image of the virtual assistant (supporting 12 languages such as English and French);

[0184] 2.2 The knowledge graph is associated with the passenger's flight number to extract the transfer counter and time information.

[0185] 3 Intelligent scheduling:

[0186] 3.1 The virtual assistant projects the transfer route onto the passenger's field of vision through AR-HUD;

[0187] 3.2 Synchronously push the electronic guide to the passenger's mobile phone.

[0188] 4 Multimodal feedback:

[0189] 4.1 The holographic assistant announces in French "Your transfer counter is B12, and the estimated walking time is 5 minutes";

[0190] 4.2 The LED light strip on the cabin ceiling shows blue (information service logo).

[0191] Experimental data

[0192] Language support ability: Multilingual interaction accuracy ≥ 95%;

[0193] Projection positioning accuracy: Projection position error based on line-of-sight tracking ≤ 1 cm;

[0194] User acceptance: Service satisfaction of foreign passengers increased by 55%, and click-through rate of entertainment recommendations increased by 40%.

[0195] In summary, the intelligent cabin service assistance system based on multimodal interaction technology provided in this embodiment has the following advantages:

[0196] 1. Greatly improved service efficiency: High demand recognition accuracy, extremely fast emergency response speed, and significantly reduced workload of flight attendants;

[0197] 2. Enhanced safety: Extremely high accuracy of health monitoring and early warning, and blockchain evidence storage to ensure data integrity and traceability;

[0198] 3. Optimized personalized experience: The recommendation system driven by federated learning greatly improves passenger satisfaction, and the holographic projection interaction has a high naturalness score;

[0199] 4. Good environmental adaptability: Supports various complex scenarios such as day and night, turbulence, and strong light, and the system availability is extremely high.

[0200] 5. Deep multimodal integration: Voice, vision, touch, and physiological signals cooperate, and high-precision interaction is still maintained under environmental interference;

[0201] 6. Dynamic intelligent decision-making: The federated learning optimizes the scheduling strategy, and the LiDAR 3D map improves the resource utilization rate;

[0202] 7. Full-scenario adaptability: Technologies such as thermal imaging and EEG prediction cover complex environments, and the system robustness is significantly enhanced;

[0203] 8. Secure and trustworthy architecture: The combination of blockchain evidence storage and federated learning realizes "usable but invisible" data, with a high level of privacy protection.

[0204] Workflow

[0205] S1. Data collection: Passengers initiate interactions through voice, gestures, touch, or physiological signals, and multi-modal sensors collect data in real time;

[0206] S2. Intention parsing: The deep learning model fuses multi-source data, and combines the knowledge graph and historical records to parse the demand type and priority;

[0207] S3. Intelligent scheduling: Dynamically allocate tasks to the nearest steward or automatic device, and the AR glasses push the optimal path;

[0208] S4. Multi-modal feedback: Voice, vision, touch, and environmental linkage responses, and synchronously update the service status;

[0209] S5. Data management: The federated learning updates the model, and the blockchain encrypts and stores key records to support long-term personalized services.

[0210] The above are only the preferred embodiments of the present invention, and do not limit the implementation manners and protection scope of the present invention. For those skilled in the art, it should be able to realize that all equivalent replacements and obvious changes made by using the description and illustrations of the present invention should be included in the protection scope of the present invention.

Claims

1. An intelligent assistant system for cabin service based on multimodal interaction technology, characterized in that, Including: A multi-modal information collection module for obtaining voice, vision, touch, and physiological signal inputs; A passenger intention recognition module that fuses multi-source data through a spatio-temporal coding model and a knowledge graph to analyze explicit and implicit service requirements; An intelligent service scheduling module that dynamically optimizes task priorities and resource allocation strategies based on federated learning; A multi-modal feedback module that provides voice, vision, touch, and environmental linkage responses; A data management and self-learning module that uses blockchain technology to encrypt and store service records.

2. The intelligent cabin service assistance system based on multimodal interaction technology according to claim 1, characterized in that, The multi-modal information collection module includes a voice unit, a vision unit, a touch unit, and a physiological signal unit, where: The voice unit includes a microphone array deployed on the seat, which combines beamforming technology and a noise reduction algorithm to suppress engine noise and extract passenger voice commands; The vision unit includes in-cabin cameras, infrared sensors, thermal imaging cameras, and image recognition devices for capturing passenger gestures, expressions, physical signs, and night-time gesture trajectories; The touch unit includes a touch screen integrated into the seat armrest and a high-precision pressure sensor array, which supports gesture sliding input and emergency calls, and includes a pressure distribution sensor array that analyzes passenger comfort through sitting posture changes and triggers adaptive seat adjustment; The physiological signal unit includes a non-invasive electroencephalogram sensor and a seat pressure distribution sensor for collecting passenger intention prediction data and sitting posture comfort information. The non-invasive electroencephalogram sensor predicts passenger intentions through an LSTM neural network and cross-verifies with voice commands.

3. The intelligent cabin service assistance system based on multi-modal interaction technology according to claim 2, wherein The passenger intention recognition module fuses multi-source data of voice, vision, touch, and physiological signals based on a deep learning model to analyze passengers' explicit and implicit needs; at the same time, it combines a knowledge graph to associate passengers' historical data with real-time situations, generates personalized service recommendations, and integrates a voice emotion analysis sub-module to judge the emotional state through MFCC feature extraction and a convolutional neural network.

4. The intelligent cabin service assistance system based on multimodal interaction technology according to claim 1, wherein The intelligent service scheduling module is used for dynamic priority division and resource optimization scheduling, where: Dynamic priority division, which assigns service priorities according to the urgency of the demand and the passenger's state; Resource optimization scheduling, which constructs a real-time three-dimensional in-cabin map based on LiDAR and SLAM technologies, tracks the positions of flight attendants and the status of equipment, pushes the optimal task path through AR glasses, and uses a vibration tactile device to remind of urgent tasks, and the vibration intensity is positively correlated with the urgency of the task.

5. The intelligent cabin service assistance system based on multimodal interaction technology according to claim 1, wherein The multi-modal feedback module is used for voice feedback, vision feedback, touch feedback, and environmental linkage feedback, where: Voice feedback, which broadcasts service confirmation information to the designated seat through directional sound field technology; Vision feedback, where the seat screen displays the service progress, and the color of the ceiling LED light strip indicates the service status; Touch feedback, where the intelligent bracelet worn by the flight attendant vibrates to prompt task details, and in emergency scenarios, the seat vibration and AR-HUD warning are triggered synchronously; Environmental linkage feedback, where the odor release device automatically adjusts the type and concentration of the fragrance according to the passenger's emotion index, and the air conditioning system is linked to adjust the temperature.

6. The intelligent cabin service assistance system based on multi-modal interaction technology according to claim 1, characterized in that The data management and self-learning module is used to store passengers' historical service data and preference information to support personalized recommendations; The intention recognition model is updated using federated learning technology to ensure data privacy and security; blockchain encryption is used to store key service records to ensure data immutability and traceability, and support multi-node distributed storage and audit traceability.

7. The intelligent cabin service assistance system based on multimodal interaction technology according to claim 3, characterized in that, The physiological signal unit in the multi-modal information collection module integrates an electrocardiogram sensor and a blood oxygen sensor to continuously monitor the passenger's heart rate variability and blood oxygen saturation, and trigger a health warning and give priority to scheduling medical assistance when abnormal values are detected.

8. The intelligent cabin service assistance system based on multimodal interaction technology according to claim 2, wherein The passenger intention recognition module includes an environment adaptation sub-module that dynamically adjusts the interaction modality weights according to the real-time in-cabin environment parameters.

9. The intelligent cabin service assistance system based on multimodal interaction technology according to claim 2, wherein The multi-modal feedback module integrates a holographic projection virtual assistant that generates a three-dimensional virtual image through AID technology, supports multi-language interaction, entertainment service recommendation, and emergency operation guidance, and adjusts the projection position based on the passenger's line of sight.

10. The intelligent cabin service assistance system based on multimodal interaction technology according to any one of claims 1-9, characterized in that Cross-device collaborative interaction is supported. Passengers can initiate service requests through their personal smart terminals and receive multi-modal feedback synchronously. The multi-modal feedback module includes a redundant design that triggers AR-HUD warnings, seat vibrations, and directional voice announcements simultaneously in emergency scenarios.

Citation Information

Patent Citations

  • Multi-mode interactive intelligent control system

    CN118226967A

Cited By

  • Intelligent network connection automobile information interaction method and device based on category personnel judgment

    CN120909432A

  • Dynamic confrontation simulation system and method based on intelligent agent

    CN121257326A

  • Control method and artificial intelligence experiment system

    CN121354556A

  • Safe driving voice visual interaction control method

    CN122275933A