A Smart Medical Guidance Method and System for the Elderly Based on Multimodal Interaction
The multimodal interactive intelligent medical guidance system for the elderly, combined with smart bracelets and hospital servers, collects and integrates user and environmental data in real time, and dynamically adjusts route planning. This solves the problems of poor adaptability, difficulty in information acquisition, and low security of traditional medical guidance systems, and enables elderly patients to have an efficient and safe medical experience.
Patent Information
- Application Number
- CN202511612179.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-11-06
AI Technical Summary
Traditional hospital guidance systems cannot dynamically adjust route planning, elderly patients have difficulty obtaining information, have low interaction accuracy, and lack real-time physiological status monitoring and security guarantees.
The intelligent medical guidance system for the elderly, which adopts multimodal interaction, collects user status data and environmental data through smart bracelets, integrates various information for path planning, and provides real-time navigation using multi-stage image registration. Combined with voice and tactile feedback, it dynamically adjusts the path to adapt to changes in the user and environment.
It improves the efficiency and safety of medical treatment for elderly patients by using dynamic path planning and multimodal information fusion to achieve real-time information synchronization, natural and convenient interaction, and full-process safety protection.
Smart Images

Figure CN121075594B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of smart healthcare technology, and in particular to a smart medical guidance method and system for the elderly based on multimodal interaction. Background Technology
[0002] As the population ages, the demand for medical care among the elderly increases significantly. However, traditional hospital guidance systems suffer from the following pain points, severely impacting the efficiency and safety of elderly patients seeking medical treatment:
[0003] Poor adaptability: Traditional medical guides do not dynamically consider changes in the physiological state and environment of elderly patients in their pathway planning, and cannot make real-time adjustments to the guide pathways.
[0004] Difficulty in obtaining information: Traditional medical guidance relies on fixed signs or touch screen navigation. Elderly patients often cannot quickly obtain real-time information (such as changes in examination rooms or examination queue progress) due to declining vision or unfamiliarity with touch screen operation.
[0005] Poor interaction adaptation: Existing voice interaction systems have low recognition accuracy in noisy hospital environments, and commands must strictly follow fixed sentence patterns (such as "I want to go to the radiology department"). Due to differences in expression habits (such as dialects and colloquial expressions), interaction often fails for the elderly; touch screen interaction requires multiple clicks, which is not compatible with the reaction speed of the elderly.
[0006] Lack of safety and health monitoring: Elderly patients often have cardiovascular and cerebrovascular diseases, and are prone to falls or sudden health events (such as a sudden increase in heart rate) during the medical process. However, traditional medical guidance systems do not monitor physiological status in real time, making it difficult for medical staff to respond in a timely manner. Summary of the Invention
[0007] This invention provides an intelligent medical guidance method and system for the elderly based on multimodal interaction, in order to solve at least one of the above-mentioned technical problems.
[0008] In a first aspect, embodiments of the present invention provide an intelligent medical guidance method for the elderly based on multimodal interaction, comprising:
[0009] Obtain the destination for elderly users seeking medical care;
[0010] The path planning algorithm is run with the destination of the triage service as the endpoint, and the following operations are performed in each step of the path planning:
[0011] S1. Real-time updates of user status data, voice commands, and environmental data. User status data includes the user's location, posture, and physiological status indicators, while environmental data includes crowd density, environmental noise, and waiting time in the examination room.
[0012] S2. The user's basic information, user location, posture, voice commands and environmental data from multiple recent time steps are fused to obtain contextual features that comprehensively reflect the current medical environment.
[0013] S3. Update the user's triage scenario based on the aforementioned contextual features;
[0014] S4. Based on the triage scenario, the medical rules triggered by the physiological state indicators, and environmental risks, adjust the cost function of each candidate node in the path planning, and select the next node of the triage path according to the adjusted cost function.
[0015] S5. Based on the pre-constructed registration transformation matrix, the real-time images captured by multiple cameras closest to the user's future entry position at the next node are registered to the user's future entry view. Then, the registered images are registered a second time to obtain a final image consistent with the entry view, which is used by the user to identify the next node.
[0016] Secondly, embodiments of the present invention also provide an intelligent medical guidance system for the elderly based on multimodal interaction, comprising:
[0017] The intelligent medical guidance terminal is worn on the wrist of elderly users to obtain the user's destination for medical treatment and to monitor the user's status data in real time.
[0018] The hospital server is used to monitor hospital environmental data in real time.
[0019] A multimodal interaction engine is used to run a path planning algorithm with the triage destination as the endpoint, and performs the following operations at each step of the path planning:
[0020] S1. Real-time updates of user status data, voice commands, and environmental data. User status data includes the user's location, posture, and physiological status indicators, while environmental data includes crowd density, environmental noise, and waiting time in the examination room.
[0021] S2. The user's basic information, user location, posture, voice commands and environmental data from multiple recent time steps are fused to obtain contextual features that comprehensively reflect the current medical environment.
[0022] S3. Update the user's triage scenario based on the aforementioned contextual features;
[0023] S4. Based on the triage scenario, the medical rules triggered by the physiological state indicators, and environmental risks, adjust the cost function of each candidate node in the path planning, and select the next node of the triage path according to the adjusted cost function.
[0024] S5. Based on the pre-constructed registration transformation matrix, the real-time images captured by multiple cameras closest to the user's future entry position at the next node are registered to the user's future entry view. Then, the registered images are registered a second time to obtain a final image consistent with the entry view, which is used by the user to identify the next node.
[0025] In summary, the embodiments of the present invention provide a multimodal interactive intelligent medical guidance method and system for the elderly, offering elderly patients a medical experience characterized by "real-time information synchronization, natural and convenient interaction, and comprehensive safety escort," achieving the following beneficial effects:
[0026] 1. In path planning, the system dynamically determines the user's triage scenario based on user information and the surrounding environment. It also dynamically adjusts the cost function of candidate nodes in the path planning function based on the triage scenario, user physiological indicators, and environmental risks, changing the focus of each path planning iteration in real time to improve adaptability to changes in user status and environment. Simultaneously, it provides real-time image information for the user to identify the next node through multi-stage registration of real-time guidance images consistent with the user's perspective.
[0027] 2. Improve information acquisition efficiency: By integrating and utilizing multimodal information, provide better patient guidance paths and shorten the average waiting time for elderly patients;
[0028] 3. Improved Interaction Accuracy: Enhanced command recognition accuracy through multimodal fusion;
[0029] 4. Enhance safety assurance: Shorten fall response time and improve early warning coverage of sudden health events through multimodal interaction and dynamic path planning. Attached Figure Description
[0030] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0031] Figure 1 This is an architecture diagram of an intelligent medical guidance system for the elderly based on multimodal interaction, provided by an embodiment of the present invention.
[0032] Figure 2 This is a flowchart of an intelligent medical guidance method for the elderly based on multimodal interaction, provided by an embodiment of the present invention;
[0033] Figure 3 This is a schematic diagram of a multimodal interactive interface provided in an embodiment of the present invention;
[0034] Figure 4 This is a spatial top view of a user entering the next node provided by an embodiment of the present invention;
[0035] Figure 5 This is a flowchart of another intelligent medical guidance method for the elderly based on multimodal interaction provided in an embodiment of the present invention;
[0036] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0038] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0039] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0040] This invention provides a method for intelligent medical guidance for the elderly based on multimodal interaction. To illustrate this method, a multimodal interaction-based intelligent medical guidance system for the elderly that supports this method will be introduced first. Figure 1This is an architecture diagram of a multimodal interactive intelligent medical guidance system for the elderly, provided by an embodiment of the present invention. The system achieves information synchronization, convenient interaction, and safety protection for elderly patients during their hospital visits through the deep integration of smart bracelets, HIS (Hospital Information System), and multimodal interaction technology. Figure 1 As shown, the system includes: a smart medical guidance terminal (i.e., a smart bracelet), a hospital server, and a multimodal interaction engine. These components work together to achieve the medical guidance function.
[0041] The intelligent medical guidance terminal is a smart bracelet worn on the wrist of elderly patients. The bracelet integrates multiple sensors and communication modules, responsible for collecting user status data, location tracking, interactive feedback, and emergency response. For example, the specific hardware configuration of the smart bracelet is shown in the table below:
[0042]
[0043] For example, the millimeter-wave radar of the smart bracelet works in conjunction with the hospital's Bluetooth beacon (deployment density ≥ 1 per 100㎡) for positioning, with a positioning accuracy of ±0.5m; the triaxial accelerometer triggers a fall warning through an attitude detection algorithm (threshold: acceleration change > 2g and lasts > 2s, where g represents gravitational acceleration).
[0044] The hospital server is deployed on the hospital's local server or in the cloud, responsible for connecting with HIS and IoT (Internet of Things) devices to provide real-time data support for patient guidance. Optional core functional modules of the hospital server include:
[0045] HIS Interface Module: Connects with HIS via the HL7 (Health Level Seven) protocol to obtain registration information (patient name, department, doctor information), clinic status (vacant / occupied / closed), examination queue progress (current number of people waiting, estimated waiting time), and doctor's closure notices;
[0046] IoT device access module: Connects to hospital IoT devices (such as elevator sensors, corridor cameras, and smoke detectors) via MQTT (Message Queuing Telemetry Transport) protocol to collect real-time data on elevator operation status (floor and direction of travel), corridor pedestrian density (through camera image recognition with an accuracy of ±0.5 people / m²), environmental noise (collected through a microphone array), fire exit status (whether it is unobstructed), and fire alarm signals.
[0047] Electronic medical record storage module: After patient authorization, it stores the patient's basic information (age, gender), medical history (hypertension, diabetes, etc.), allergy history and past medical records, and supports personalized service recommendations (such as prioritizing diabetic patients to the fast examination channel).
[0048] The multimodal interaction engine is deployed on the hospital server or in the cloud to plan routes based on the patient's destination and specific scenario, and dynamically adjusts the planning strategy according to the real-time situation of the patient and the environment. Optionally, the core modules of the multimodal interaction engine include:
[0049] Data fusion module: Adopting an adaptive attention mechanism, it extracts and weights the features of patient basic information, millimeter-wave radar point cloud data, voice commands, physiological signals and hospital environment data to generate a unified scene representation and determine the type of triage scene;
[0050] Dynamic Path Planning Module: Based on the A* algorithm improved by reinforcement learning, it generates the optimal path that takes into account user status, medical rules and environmental risks. The path planning takes into account the user's current triage scenario and physiological state in real time and dynamically adjusts the path planning strategy.
[0051] Age-friendly output module: Outputs path planning results to users, guiding them to follow the triage path; the module supports dialect voice interaction, 3D spatial audio navigation, and large font high-contrast interface display.
[0052] Based on the above systems, Figure 2 This is a flowchart of a multimodal interactive intelligent medical guidance method for the elderly, provided by an embodiment of the present invention. This method is executed collaboratively by the elderly patient and various modules within the aforementioned system, covering the complete process from user request initiation to completion of the consultation. Figure 2 As shown, the method specifically includes:
[0053] S110, User Identity Binding.
[0054] Elderly patients can bind their smart bracelets to their electronic medical records via hospital self-service machines (NFC card swiping) or mobile apps (facial recognition). The system then simultaneously obtains the patient's basic information and current appointment registration information.
[0055] In one specific implementation, after arriving at the hospital, elderly patients can bind their smart bracelet to their electronic medical record using any of the following methods:
[0056] Self-service machine binding: Swipe your medical insurance card or ID card on the hospital's self-service machine, and the system will automatically read the patient's information and synchronize it to the wristband (the wristband needs to be brought close to the self-service machine to complete NFC pairing).
[0057] Mobile App Binding: Scan the wristband's QR code via the hospital's official app, and enter the patient's name and ID number to complete the binding (facial recognition verification is supported). Optionally, the mobile app / wristband screen features a high contrast ratio (≥4.5:1) and large font (18pt or higher), with information presented following a "progressive" principle (only one core piece of information is displayed per screen; details are viewed by swiping). Figure 3 This is a schematic diagram of a multimodal interactive interface provided by an embodiment of the present invention, demonstrating the large font, high contrast design and three-dimensional audio visualization effect of the mobile APP.
[0058] After the binding is completed, the system will automatically synchronize the patient's current medical information and obtain the destination for triage. For example, if the medical information is "CT scan, radiology room 3 on the 2nd floor, 12 people in line, estimated wait time 25 minutes", then the destination for triage will be radiology room 3 on the 2nd floor.
[0059] S120, Multimodal Requirements Acquisition.
[0060] Once the connection is established, patients can initiate triage requests through the following multimodal methods:
[0061] Voice commands: Speak natural sentences such as "I want to go to the radiology department" or "Find the nurses' station" (supports dialects, recognition accuracy ≥95%). Optionally, the system supports dialect voice interaction, using a pre-trained dialect voice recognition model to convert user voice commands into text, which is then parsed by natural language understanding to generate triage requests.
[0062] Touch button: Short press the physical button on the left side of the bracelet (preset "Guidance" function);
[0063] Posture trigger: Raise your wrist twice in a row (the wristband recognizes the posture change through the accelerometer and triggers the triage function).
[0064] S140, Dynamic Path Planning and Triage Output.
[0065] In response to users' guidance needs, the multimodal interaction engine runs a path planning algorithm based on the destination, dynamically adjusting the planning strategy at each step based on real-time user and environmental data. Optionally, the path planning algorithm can employ an improved A* algorithm, which uses key locations within the hospital as path nodes and, starting from the current position, progressively determines the next node with the lowest cost, guiding the user to that node. It should be noted that each step of path planning in this embodiment refers to the process of determining the next node; once the user moves to that next node, the next step of path planning begins.
[0066] In one specific implementation, the multimodal interaction engine performs the following operations at each step of path planning:
[0067] S1. Real-time updates of user status data, voice commands, and environmental data. User status data includes the user's location, posture, and physiological status indicators, while environmental data includes crowd density and waiting time in the clinic.
[0068] Optional, real-time updated data includes:
[0069] User location and posture: The user's point cloud data is obtained using millimeter-wave radar to obtain the user's location and posture (such as standing, walking, falling, etc.); at the same time, the user's location information and the sequence of locations that have been visited are collected by fusing positioning with the hospital's Bluetooth beacon and the wristband's millimeter-wave radar (positioning accuracy ±0.5m).
[0070] User's physiological indicators: Heart rate (60-100 beats / minute is normal) and blood oxygen (≥95% is normal) are monitored in real time by PPG sensor, and gait stability is analyzed by triaxial accelerometer (acceleration variance <0.5g is stable).
[0071] User voice commands: Various voice commands issued by the user in real time while walking.
[0072] Hospital environmental data: Collect hospital data through in-hospital equipment, such as the number of people queuing in the target department (e.g., 12 people are currently waiting in the CT room, with an estimated wait time of 25 minutes), temporary closure notices for consultation rooms (e.g., consultation room 3 on the 2nd floor is temporarily closed due to equipment maintenance), elevator operation status (e.g., elevator E1 is on the 1st floor, with an estimated arrival time of 30 seconds), etc.
[0073] S2. The user's basic information, user location, posture, voice commands and environmental data from multiple recent time steps are fused to obtain contextual features that comprehensively reflect the current medical environment.
[0074] In this step, the multimodal interaction engine fuses the collected data to generate a scene representation in real time. Optionally, this process may include the following steps:
[0075] Step 1: Based on the user's basic information, as well as the location, posture, voice commands, and environmental data at the same time step, generate the total feature vector for that time step. Optionally, for each time step... Feature extraction is performed on millimeter-wave radar point cloud data to obtain feature vectors reflecting user posture (such as standing, walking, or falling). ; for time step The voice commands are converted into text and semantically analyzed to obtain feature vectors that reflect user needs (such as "go to the radiology department" or "find the toilet"). The environmental sensor data is normalized to obtain feature vectors reflecting environmental conditions (such as pedestrian density, ambient noise, and elevator location). Furthermore, a user's basic information (including age, electronic medical records, etc.) can be encoded as a feature vector and exist independently, or this vector can be concatenated with... Inside, utilizing To unify the representation of user basic information and radar information, this embodiment adopts the latter representation method.
[0076] Step 2: Using an adaptive attention mechanism, fuse the total feature vectors from the most recent multiple time steps to obtain the current context features. Optionally, calculate the attention weights for each time step using the following formula. and according to Weighted fusion generates the scene representation vector at the current time step. Used for subsequent path planning and decision-making:
[0077] (1)
[0078] in, Harmony This is the weight matrix. and For bias terms, and This represents the index of the most recent time steps (including the current time step). This represents the activation function. express The concatenated vector, This represents the attention score at each time step. The fused representation. It can effectively reflect the information most relevant to the patient guidance task (such as the patient's urgency level, current location and distance to the target department).
[0079] Furthermore, in the above formula, multimodal feature concatenation is performed first. (dimension is) ), speech features (dimension is) ), environmental characteristics (dimension is) ) concatenate to form the total feature vector [ (dimension) Then calculate Through activation function Obtain intermediate features ; then calculate Attention scores at each time step are obtained. Finally, attention weights are calculated using Softmax. And use fusion to obtain contextual representations .
[0080] S3. Update the user's triage scenario based on the context features.
[0081] As mentioned above, the representation vector It is a "scene snapshot" after multimodal data (basic information, radar, voice, environmental data, etc.) are fused through an adaptive attention mechanism. Its core function is to provide dynamic and comprehensive scene understanding for path planning, which will affect the specific logic of subsequent path planning.
[0082] Optionally, the contextual features can be input into a classifier to obtain the probability that the user's triage scenario belongs to a general scenario, an emergency scenario, an elderly patient scenario, or a noise-sensitive scenario. The classifier can be a multilayer perceptron, and the output result... This is a vector representing the probability of a user belonging to various triage scenarios. Furthermore, a probability greater than a set threshold indicates belonging to that scenario; a user can belong to multiple triage scenarios at the same time step. Among these, "elderly patients" refers to elderly patients whose age exceeds the set threshold; "noise-sensitive scenarios" refers to environments that are relatively quiet but where the user is sensitive to noise, such as a scenario where a hearing-impaired patient undergoes a hearing test.
[0083] S4. Based on the triage scenario, the medical rules triggered by the physiological state indicators, and environmental risks, adjust the cost function of each candidate node in the path planning, and select the next node of the triage path according to the adjusted cost function.
[0084] As "contextual input" for path evaluation, it integrates the patient's recent state (e.g., walking speed, posture), environmental information (e.g., crowd density, ambient noise, elevator status), and task objectives (e.g., emergency room visit vs. general visit). These serve as key inputs during path planning, assisting the algorithm in determining "which factors are more important in the current scenario." For example, if... When the data includes information such as "patient walking slowly (radar data) + voice expression 'I need to go to the emergency room' (voice data) + low current patient flow in the emergency room (environmental data)," the expected path planning algorithm will prioritize the "shortest path" and reduce the importance of "smoothness" (because emergency patients are more concerned about speed). To achieve this, this embodiment improves the cost function in the improved A* algorithm and adjusts it according to the data at each time step. The cost function is dynamically adjusted.
[0085] In one specific implementation, considering the characteristics of elderly patient guidance, the cost function in the improved A* algorithm can be modified to the following form:
[0086] (2)
[0087] Among them, among them, Indicates candidate nodes The total cost, Indicates the distance from the current node to the candidate node. The actual cost, Indicates candidate nodes The heuristic cost to the destination, This indicates the cost of medical regulations and environmental risks. express The weighted weights. Candidate nodes refer to the adjacent nodes or directly reachable nodes of the current node.
[0088] (3)
[0089] in, , , and These represent the distances from the current node (i.e., the current position) to the candidate node. The path length, number of steps, slope, and noise exposure risk (which can be quantified in decibels). , , and They represent , , and The weighting coefficients.
[0090] (4)
[0091] in, Indicates candidate nodes Spatial distance to the destination (such as straight-line distance or Euclidean distance). Indicates candidate nodes Environmental parameters related to the spatial distance to the destination (such as the average population density and obstacle density of nodes near the line segment representing the straight-line distance between two nodes). and They represent and The weighted weights. Due to the candidate nodes The specific path to the destination has not yet been planned, so only spatial distance and environmental parameters of the area near the line segment representing that distance are used here to characterize the approximate heuristic cost.
[0092] (5)
[0093] in, This indicates the number of medical rules triggered by the user's physiological state. Indicates the first The intensity of the adjustment of node costs by each medical rule. Indicates the first An indicator function to determine whether a medical rule has been triggered. ,otherwise ; Represents a node Environmental risks, such as nodes Population density, obstacle density, etc. This represents the weighted average of environmental risks.
[0094] Based on the above cost function, in this embodiment, the weights of each item in the cost function are dynamically adjusted according to the real-time triage scenario determined in S3.
[0095] Optionally, if the current triage scenario includes emergency scenarios, the weight of path length can be increased. For example, will From default value Increase by 30% to obtain the path length weight in emergency scenarios. :
[0096] (6)
[0097] In cases where the triage scenario includes elderly patients, the weight of path smoothness should be increased. and .by For example, From default value Increase by 20% to obtain the step number weight in the elderly scenario:
[0098] (7)
[0099] In cases where the triage scenario includes noise-sensitive areas, the weighting of noise exposure risk should be increased. For example, will From default value Increase by 20% to obtain the weight of noise exposure risk in noise-sensitive scenarios:
[0100] (8)
[0101] Furthermore, this embodiment adjusts the medical rules triggered by the user's current physiological state indicators. For example:
[0102] Medical Rule 1 ( If heart rate > 100 beats / min and candidate nodes If it is a staircase node (triggering condition), then Increase by 0.3 (i.e., node) of If heart rate > 100 beats / min and node If it is an elevator node (triggering condition), then Decrease by 0.3 (i.e., node) of );
[0103] Medical Rule 2 ( If blood oxygen saturation is <95% and the nodes If it is a climbing node (triggering condition), then Increase by 0.2 (i.e., node) of );
[0104] Meanwhile, this embodiment is adjusted based on current environmental risks. For example:
[0105] When node When a corridor node is selected, if the pedestrian density in the corridor is greater than 4 people / m², it indicates a high risk of overcrowding (such as collisions or delays). Therefore, the corridor node will be designated as a corridor node. Increase the risk by 25% (multiply the original risk value by 1.25) to penalize the node and guide the algorithm to prioritize paths with less traffic.
[0106] When node When the elevator node is selected, if the elevator waiting time is greater than 2 minutes, it indicates that the node may face even longer delays. In this case, the "time advantage" of the pedestrian staircase becomes apparent, thus reducing the elevator node's time advantage. Increase by 15% to penalize the node for risk, guiding the algorithm to prioritize paths with shorter waiting periods. Alternatively, when the node... When a staircase node is a pedestrian staircase node, if the elevator waiting time on the same floor is greater than 2 minutes, then the pedestrian staircase node will be... Reduce the number of steps by 15% to reward the node and guide the algorithm to prioritize walking stairs.
[0107] Of course, the above thresholds can be adjusted as needed, and this embodiment does not impose specific limitations.
[0108] S5. Based on the pre-constructed registration transformation matrix, the real-time images captured by multiple cameras closest to the user's future entry position at the next node are registered to the user's future entry view. Then, the registered images are registered a second time to obtain a final image consistent with the entry view, which is used by the user to identify the next node.
[0109] This embodiment outputs multimodal triage information after each path planning step, guiding the patient to the next node. One method of triage output is to display a real-time environmental image of the next node to the user (e.g., via a mobile app), allowing the user to identify whether they have approached or reached that node. In particular, this embodiment utilizes multi-stage image registration to ensure that the image displayed to the user is as close as possible to the user's perspective when entering the next node.
[0110] In one specific implementation, multi-stage registration can be performed using a mature neural network model. This model takes the two images to be registered as input and outputs the registration transformation matrix between the two images, such as the DeepImage Homography Estimation model or the HomographyNet model. The basic steps are: extracting Regions of Interest (ROIs) of the same size in both the reference image and the image to be registered (i.e., feature matching regions or intersection regions); using the four vertices of the ROI rectangle in the reference image as feature points, finding the coordinates of these four points on the ROI region of the image to be registered; using these four pairs of matching points, directly solving for the homography matrix H, and then obtaining the registration transformation matrix from the image to be registered to the reference image. After this matrix transformation, the ROI of the image to be registered will be aligned with the ROI in the reference image.
[0111] In this embodiment, firstly, the boundary between the channel from the current node to the next node and the next node is taken as the user's future entry position; that is, the user will enter the next node from this position. Simultaneously, the viewpoint along the channel pointing to the next node is taken as the user's future entry viewpoint; that is, the image the user sees upon entering the next node will be the image viewed from this viewpoint. Multiple cameras may exist at the next node. This embodiment selects multiple cameras closest to the entry position that all cover part (or all) of the field of view of the entry viewpoint to provide environmental images from the user's entry viewpoint. Combined with... Figure 4 The diagram shows a top-down view of the next node, where the area marked by the dashed circle represents the next node (e.g., a fork in a corridor). The user will reach the next node from the entrance passage shown in the diagram (e.g., a corridor following a certain direction from the fork). The user's future entry position is the intersection of this passage and the next node, and the future entry viewpoint is the direction indicated by the black dashed arrow in the diagram. Camera 1 and Camera 2 in the diagram are the two cameras closest to the entry position at the next node. The red dashed line indicates the shooting area of Camera 1, and the green dashed line indicates the shooting area of Camera 2; they each cover a portion of the field of view of the entry perspective. In practical applications, these cameras are often reused from existing surveillance cameras in the hospital, eliminating the need for reinstallation specific to the method of this embodiment.
[0112] Based on the above illustration, this embodiment can pre-set a temporary camera (or camera) at the entrance location to capture images of the next node from the entrance viewpoint. These images can cover different times, different pedestrian flows, different lighting conditions, and different weather conditions to maintain image diversity as much as possible. Then, images captured at the entrance location at the same time from the entrance viewpoint, and images captured at the same time by any one of the multiple cameras closest to the entrance location, are combined to form image pairs, resulting in multiple image pairs between the entrance viewpoint and any one of the cameras. These multiple image pairs together constitute an image pair set, where the two images in each image pair were captured at the same time, but from different angles and with slightly different content. Similarly, by constructing such image pairs for each of the multiple cameras, image pair sets between the entrance viewpoint and each of the surrounding cameras can be obtained. Extending the above operation to each viewpoint of a user entering multiple nodes, image pair sets between each entrance viewpoint of each node (determined by the entrance channel deployed at each node) and each of the surrounding cameras (i.e., each camera around the entrance channel corresponding to each entrance viewpoint) can be obtained.
[0113] Then, the mature neural network model can be fine-tuned using each of the aforementioned image pairs to make the registration matrices of image pairs from the same camera more consistent, while maximizing the differences in registration transformation matrices of image pairs from different cameras. Specifically, the two images in each image pair are input into the aforementioned neural network model, where the image from the incoming viewpoint is used as the reference image, and the other image is used as the image to be registered. The model extracts the feature matching points of the two images as the Region of Interest (ROI), and calculates the registration transformation matrix from the image to the incoming viewpoint image based on this ROI; and then uses the following loss function... To fine-tune the model parameters:
[0114]
[0115] in, and Each index represents an index of a set of images, that is, each index represents a combination of an incoming viewpoint of a node and a camera surrounding it; and Each index represents an image pair index within a set of image pairs, meaning each index represents an image taken at the same time under one of the above combinations. Indicates from the The first image pair in each set The registration transformation matrix corresponding to each image pair and The meaning is similar. In the loss function... Indicates all Down Summing the cases, when When minimized, This ensures that the registration matrix is identical between every pair of images taken simultaneously by the same camera around the same entry point of the same node from the same viewpoint. In other words, the neural network model filters out random factors related to the image content itself, ensuring that the model extracts fixed transformation elements related to the angle combination between image pairs. Simultaneously, the loss function... Indicates all Summing the cases, when When minimized, It can ensure that the fixed transformation elements related to the combination extracted from different viewpoints of different nodes and different cameras around them (if any one of the node, viewpoint, or camera is different, the combination will be different), so as to ensure the differences between the combinations (such as inherent differences in spatial location). and These are the weights corresponding to the two loss terms. Furthermore, during model parameter fine-tuning, when... It remains essentially unchanged across multiple calculations, and In these calculations, the result approaches 0, and If the results remain essentially unchanged across these calculations, then it can be considered that the desired outcome has been achieved. Minimize the conditions.
[0116] Finally, the fine-tuned neural network model outputs the registration transformation matrix from each camera to the entering viewpoint. Specifically, the neural network model fine-tuned through the above steps can filter out random factors in the image content and output a registration transformation matrix that is only related to the viewpoint of the image pair. (Continuing with...) Figure 4 For example, the finely tuned neural network model is input with images captured by camera 1 at the same moment along the user's future entry view. The image captured along the user's future entry view is designated as the reference image, and the image captured by camera 1 is designated as the image to be registered. The finely tuned neural network model then inputs a registration transformation matrix from the image to the reference image. After the image to be registered is transformed by this matrix, the Region of Interest (ROI) will align with the ROI in the reference image. The transformed image is then displayed on the same screen as the reference image according to the new coordinates (i.e., stitched together according to the ROI), thus providing a continuous, wide-field-of-view image from the user's future entry view. In this embodiment, this matrix is referred to as the registration transformation matrix from camera 1 to the entry view.
[0117] The registration transformation matrix for each combination is prepared in advance. During path planning, after obtaining the next node each time, the user's future position and viewpoint for entering the next node are determined based on the channel between the current node and the next node. Real-time images captured by multiple cameras near that position are retrieved. For each camera, the following operations are performed: the registration transformation matrix from the current camera to the entering viewpoint is retrieved, and the real-time image captured by the current camera is transformed using the registration transformation matrix to transform the real-time image to the entering viewpoint (i.e., how the current environment of the next node appears from the entering viewpoint). Performing the above transformation on the real-time images of each camera completes the first stage of registration, also known as primary registration.
[0118] Then, the transformed images corresponding to each camera are registered a second time. This registration can utilize either the neural network model before fine-tuning or the neural network model before fine-tuning. In each registration, the image corresponding to the camera closer to the entry viewpoint can be used as the reference image, and the remaining images are aligned with the reference image. This can obtain an environmental image that is consistent with the entry viewpoint and has the largest possible field of view, providing users with a basis for image recognition. In particular, a special case may exist in secondary registration: the images after primary registration may intersect not only within the user's field of vision but also outside of it. In this case, when the neural network model extracts the Region of Interest (ROI) for each image to be registered, it may extract the ROI outside the field of vision and then perform registration and stitching based on this ROI. Consequently, the registered and stitched image will no longer be an environmental image within the user's field of vision, failing to guide the user. Therefore, in this embodiment, when performing secondary registration on the images after primary registration, the intersection of the images after primary registration is first determined. If the intersection contains a portion outside the field of vision, this portion is deleted from each of the images after primary registration. The deleted images are then input into the neural network model for secondary registration. The resulting registered image is one with the region within the user's field of vision as the ROI.
[0119] In addition to image output, triage output can also include the following methods:
[0120] Voice prompts: Dialect navigation voice is played through bone conduction speakers (e.g., "Grandpa Zhang, you are currently in the lobby on the 1st floor. You need to go to the 2nd floor for a CT scan. Follow me~ turn right first").
[0121] Vibration alert: Vibrate twice (frequency 100Hz, duration 0.5s) every time a key node is reached (such as a turn).
[0122] 3D spatial audio: Dynamically play supplementary prompts based on the user's head turning direction (e.g., "The elevator is 5 meters ahead, the door is open"). Optionally, based on the user's head posture (pitch angle, yaw angle) obtained from the wristband gyroscope and the user's position obtained from the hospital's UWB positioning system, the sound field direction can be dynamically adjusted through a head tracking algorithm, so that the navigation prompts change synchronously with the user's head turning direction (e.g., when the user is facing left, the left-side prompts are enhanced, and the right-side prompts are weakened).
[0123] Simultaneously, real-time tracking and anomaly handling can be performed during this process. Optional methods include the following:
[0124] Position deviation monitoring: If the patient deviates from the planned path by more than 5 meters (which can be determined by millimeter-wave radar positioning), the path will be automatically replanned and a prompt will be given, such as "Grandpa, there is a small slope ahead, it will be safer to take two steps to the left."
[0125] Physiological abnormality alarm: If the heart rate is >120 beats / min and lasts for 30 seconds (such as due to waiting anxiety), or blood oxygen <90% (such as difficulty breathing), immediately send an alarm to the nearest nurse station. The alarm content includes the location, patient's name, and medical history, such as "Patient Zhang XX, 72 years old, with a history of hypertension, is currently in the elevator lobby on the 2nd floor, with a heart rate of 120 beats / min. Please check in time."
[0126] Environmental change response: If a fire alarm is detected (such as when triggered by a hospital IoT smoke sensor), it will automatically guide you to the nearest safety exit (e.g., the path priority will be adjusted to "safety exit > target department") and provide a voice prompt, such as "Grandpa, there is a fire. Let's go to the nearest safety exit, number 3. Follow me!"
[0127] The above method dynamically generates the optimal path based on the fused scene representation and real-time changes in users and the environment, improving the responsiveness of the triage path to changes in user status and environment, and providing a better triage experience for elderly users. For example, the final generated optimal path is "1st floor lobby A area → turn right to corridor 2 → walk straight for 30 meters → east side elevator E1 → 2nd floor → radiology room 3", with a total length of 150 meters.
[0128] Furthermore, it should be noted that the real-time environment, real-time status, real-time optimal path, and real-time image in this embodiment are all relative to each step of path planning. As the user moves from the current node to the next node, these real-time data may change, but this application does not consider them. The focus is on ensuring that all data are updated in real time for each path planning step.
[0129] Below is a complete example of the guidance process for elderly patients undergoing CT scans.
[0130] Scenario: Mr. Zhang, 72 years old (with a history of hypertension), needs to go to the radiology department on the 2nd floor for a CT scan. He is currently in the lobby on the 1st floor.
[0131] First, identity verification and request initiation were completed. After arriving at the hospital, Grandpa Zhang swiped his medical insurance card at the self-service machine to complete the wristband binding (the system automatically synchronized the registration information: "CT scan, 2nd floor, radiology department, room 3, currently 12 people in line, estimated wait time 25 minutes"). Then, he raised his wrist twice (the wristband recognized his posture through the accelerometer, triggering the triage function).
[0132] Then, multimodal data acquisition and environmental perception were performed. The wristband's millimeter-wave radar scanned the surrounding environment, combined with hospital Bluetooth beacon positioning (current location: Area A, 1st floor lobby, coordinates X=10.2m, Y=5.6m). The PPG sensor detected that Grandpa Zhang's heart rate was 85 beats / min (normal) and blood oxygen saturation was 97% (normal). The hospital server returned real-time data: "12 people are currently waiting in Radiology Room 3 on the 2nd floor, with an estimated wait time of 25 minutes; the flow density in Corridor 2 is 3 people / ㎡ (unobstructed); Elevator E1 (east side) is currently on the 1st floor and is operating normally."
[0133] Next, dynamic path planning is performed. At each node, steps S1-S4 are executed to dynamically determine the next node and guide the patient there. The next node is then used as the new current node to determine the node after that. This process is repeated until the patient reaches their destination. Because Mr. Zhang is an elderly patient, the smoothness weight in the cost function is increased at each step of the path planning to avoid staircase nodes.
[0134] In the multimodal triage output, each node provides voice prompts (dialect recognition supported) to Grandpa Zhang, such as, "Grandpa Zhang, you are now in the lobby on the 1st floor. You need to go to the 2nd floor for a CT scan. Follow me~ First turn right, walk straight along the blue corridor, and when you see the elevator door, go in. The elevator will take you to the 2nd floor." Vibration prompts are also provided: two vibrations at each key node (such as a turn); spatial audio functionality is also offered: when Grandpa Zhang turns his head to the left, the system detects no key markers in his line of sight and does not play any additional prompts; when he faces the elevator, it plays, "The elevator is 5 meters ahead, the door is open~" (volume automatically adjusted to 75dB according to ambient noise).
[0135] During Grandpa Zhang's walk, the system provided real-time tracking and anomaly handling. The wristband's millimeter-wave radar continuously monitored his position (deviation <0.3m). When he deviated 2 meters from the path, the system replanned: "Grandpa, there's a small slope ahead. Let's take two steps to the left for safety." Simultaneously, while waiting for his appointment, the wristband displayed the waiting progress in real-time ("Currently waiting for 10 people, estimated 20 minutes remaining") and provided reminders every 5 minutes via bone conduction speaker. In case of emergencies, such as Grandpa Zhang experiencing dizziness upon standing up in the elevator (heart rate rising to 110 beats / min for 40 seconds), the wristband immediately sent an alert to the second-floor nurses' station: "Patient Zhang XX, 72 years old, history of hypertension, currently in the second-floor elevator lobby, heart rate 110 beats / min, please check immediately." A nurse arrived within 5 minutes, assisted him to rest, and measured his blood pressure (135 / 85 mmHg, back to normal).
[0136] Furthermore, the aforementioned attention mechanism and various parameters in the classifier (such as...) , , , The weights (etc.) are all obtained in advance through training. The weights of each term in the cost function can be determined empirically, or they can be obtained through the entire training process along with the parameters of the attention mechanism and classification. In one specific implementation, the following training method is provided:
[0137] First, initialize the parameters to be trained. The dimension is ,in It is the intermediate dimension after projection (usually taken as...). of (To balance computational efficiency and expressive power). The initialization method uses Xavier / Glorot initialization (suitable for random initialization of linear layers, ensuring variance stability during forward propagation). The dimension is ( ,1), Initialization method and Similarly. Bias term. (dimension) ), (Dimension 1) Initialize to 0 or a small random number (such as a normal distribution) The other parameters can also be initialized using appropriate methods, which will not be elaborated here.
[0138] Then, taking the PyTorch framework as an example, the following training is performed: the preprocessed multimodal data and supervision signals are encapsulated into a DataLoader, and the batch size and random shuffling are set. Input multimodal features. The parameters are updated based on the accuracy of the triage path. The accuracy of the triage path refers to the matching rate between the path generated by the model and the actual optimal path (e.g., above 95%), which can be obtained through manual annotation.
[0139] After training, the parameters can be validated by assessing the accuracy of the referral path or user interaction satisfaction during actual patient guidance processes. For example, a questionnaire survey can be used to collect data on patient satisfaction with the referral service (e.g., 4.5 / 5 or higher).
[0140] Furthermore, regarding the parameters in the attention mechanism, the rationality of the attention mechanism parameters can be verified by considering the uniformity of the weight distribution among different modal features in ordinary scenarios and the emphasis of each modal feature weight in different triage scenarios. Optionally, in ordinary scenarios, the mean value of the attention weights of each modality can be calculated (e.g., the mean weight of the speech modality is 0.4, the mean weight of the radar is 0.3, and the weight of the environment is 0.3) to ensure that no obvious modality is ignored. Optionally, the adjustment of attention weights in different scenarios can be recorded (e.g., the weight of speech decreases and the weight of radar increases when the environment is noisy) to verify the dynamic adaptability of the attention weights to the scenario.
[0141] Furthermore, the weights of each modal feature can be obtained as follows: by adding zeros, each single modal feature is expanded into a total feature vector, and then substituted into formula (1) to calculate the weights. That is: to respectively , and As Substituting into the first expression in formula (1), we obtain the results for the three modes. , respectively denoted as , and Then the attention weights for each modality are:
[0142]
[0143] in, and All are modal indexes. , and Representing modes corresponding And attention weights.
[0144] It should be noted that all user data involved in this application is information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0145] In summary, the embodiments of the invention provide a method and system for intelligent medical guidance for the elderly based on multimodal interaction. Through collaborative innovation of "hardware-algorithm-service," it provides elderly patients with a medical experience characterized by "real-time information synchronization, natural and convenient interaction, and comprehensive safety protection," achieving the following beneficial effects:
[0146] 1. In path planning, the user's triage scenario is dynamically determined based on user information and the surrounding environment. Based on the triage scenario, user physiological state indicators, and environmental risks, the cost function of candidate nodes in the path planning function is dynamically adjusted to change the focus of each path planning in real time, thereby improving the adaptability to changes in user status and environment.
[0147] 2. Improve information acquisition efficiency: By integrating and utilizing multimodal information, better patient guidance paths are provided, reducing the average waiting time for elderly patients from 42 minutes to 18 minutes;
[0148] 3. Improved interaction accuracy: Through multimodal fusion, the command recognition accuracy reaches 96% (traditional systems ≤73%).
[0149] 4. Enhanced safety: Through multimodal interaction and dynamic path planning, the fall response time is reduced to 8 seconds (industry standard 30 seconds), and the coverage rate of early warning for sudden health events is ≥95%.
[0150] Figure 5 This is a flowchart of another intelligent medical guidance method for the elderly based on multimodal interaction provided in an embodiment of the present invention. This method corresponds to the portion executed solely by the multimodal interaction engine in any of the above embodiments, and can also be executed solely by other electronic devices. For example... Figure 5 As shown, the method specifically includes:
[0151] S210, Obtain the destination for elderly users' medical guidance;
[0152] S220. Run the path planning algorithm with the destination of the triage as the endpoint, and perform the following operations in each step of the path planning:
[0153] S1. Real-time updates of user status data, voice commands, and environmental data. User status data includes the user's location, posture, and physiological status indicators, while environmental data includes crowd density, environmental noise, and waiting time in the examination room.
[0154] S2. The user's basic information, user location, posture, voice commands and environmental data from multiple recent time steps are fused to obtain contextual features that comprehensively reflect the current medical environment.
[0155] S3. Update the user's triage scenario based on the aforementioned contextual features;
[0156] S4. Based on the triage scenario, the medical rules triggered by the physiological state indicators, and environmental risks, adjust the cost function of each candidate node in the path planning, and select the next node of the triage path according to the adjusted cost function.
[0157] The method in this embodiment is based on the same inventive concept as the methods in any of the above embodiments, and the limitations of any of the above embodiments are applicable to this embodiment. Accordingly, this embodiment can achieve the same beneficial effects as any of the above embodiments.
[0158] This invention provides an intelligent medical guidance system for the elderly based on multimodal interaction, such as... Figure 1 As shown, the system includes:
[0159] The intelligent medical guidance terminal is worn on the wrist of elderly users to obtain the user's destination for medical treatment and to monitor the user's status data in real time.
[0160] The hospital server is used to monitor hospital environmental data in real time.
[0161] A multimodal interaction engine is used to run a path planning algorithm with the triage destination as the endpoint, and performs the following operations at each step of the path planning:
[0162] S1. Real-time updates of user status data, voice commands, and environmental data. User status data includes the user's location, posture, and physiological status indicators, while environmental data includes crowd density, environmental noise, and waiting time in the examination room.
[0163] S2. The user's basic information, user location, posture, voice commands and environmental data from multiple recent time steps are fused to obtain contextual features that comprehensively reflect the current medical environment.
[0164] S3. Update the user's triage scenario based on the aforementioned contextual features;
[0165] S4. Based on the triage scenario, the medical rules triggered by the physiological state indicators, and environmental risks, adjust the cost function of each candidate node in the path planning, and select the next node of the triage path according to the adjusted cost function.
[0166] S5. Based on the pre-constructed registration transformation matrix, the real-time images captured by multiple cameras closest to the user's future entry position at the next node are registered to the user's future entry view. Then, the registered images are registered a second time to obtain a final image consistent with the entry view, which is used by the user to identify the next node.
[0167] The system in this embodiment is based on the same inventive concept as the method in any of the above embodiments, and the limitations of any of the above embodiments are applicable to this embodiment. Accordingly, this embodiment can achieve the same beneficial effects as any of the above embodiments.
[0168] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 6 As shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more. Figure 6 Taking a processor 60 as an example; the processor 60, memory 61, input device 62, and output device 63 in the device can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.
[0169] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the multimodal interaction-based intelligent medical guidance method for the elderly in this embodiment of the invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, thereby realizing the aforementioned multimodal interaction-based intelligent medical guidance method for the elderly.
[0170] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0171] Input device 62 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 63 may include display devices such as a display screen.
[0172] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the intelligent medical guidance method for the elderly based on multimodal interaction according to any embodiment.
[0173] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0174] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0175] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0176] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. An old-age intelligent doctor guiding method based on multi-modal interaction, characterized in that, The method comprises the following steps: acquiring a guide destination of an elderly user; running a path planning algorithm with the guide destination as the end point, and performing the following operations in each step of the path planning: S1, updating user state data, voice instructions and environment data in real time, wherein the user state data includes the user's position, posture and physiological state indicators, and the environment data includes crowd density, environmental noise and clinic queuing progress; S2, fusing the user's basic information, the user's position, posture, voice instructions and environment data at the last multiple time steps to obtain context features that comprehensively reflect the current medical environment; S3, updating the user's guide scene according to the context features, including: inputting the context features into a classifier to obtain the probability that the user's guide scene belongs to a normal scene, an emergency scene, an elderly patient scene or a noise-sensitive scene; S4, adjusting the cost function of each candidate node in the path planning according to the guide scene, the physiological state indicators triggered medical rules and environmental risks, and selecting the next node of the guide path according to the adjusted cost function; S5, according to the pre-constructed registration transformation matrix, the real-time images of multiple cameras closest to the user's future entering position at the next node are respectively registered to the user's future entering view angle, and then the registered images are secondarily registered to obtain the final image consistent with the entering view angle, which is used for the user to identify the next node; specifically, the user's future entering position refers to the intersection position of the channel between the current node and the next node and the next node, and the user's future entering view angle refers to the view angle pointing to the next node along the channel; the secondary registration of the registered images comprises: determining the intersection of the registered images; if the intersection exists outside the field of view of the entering view angle, the intersection outside the field of view of the entering view angle is deleted from the registered images, and the deleted images are secondarily registered; Before S5, it further comprises: the images taken at the same time at the entering position along the entering view angle, and the images taken by a single camera closest to the entering position from multiple cameras constitute an image pair; fine-tune the neural network model for outputting the registration transformation matrix between the two images in the image pair by using the image pairs from the multiple cameras, so that the registration matrices corresponding to each image pair from the same camera tend to be consistent, and the registration transformation matrices corresponding to each image pair from different cameras are maximized; the fine-tuned neural network model outputs the registration transformation matrix from each camera to the entering view angle to obtain the pre-constructed registration transformation matrix. 2.The multi-modal interaction based intelligent guide doctor method for the elderly according to claim 1, wherein, S1 comprises: collecting the user's position and posture through a millimeter wave radar; collecting the user's physiological state indicators through a physiological sensor; collecting the user's voice instructions through a voice device; collecting the crowd density, environmental noise and clinic queuing progress of the hospital through an environment sensor. 3.The multi-modal interaction based intelligent guide doctor method for the elderly according to claim 1, wherein S2 It comprises: generating a total feature vector at the same time step according to the user's basic information and the position, posture, voice instructions and environment data at the same time step; The total feature vector of the recent multiple time steps is fused through an adaptive attention mechanism to obtain a current context feature. 4.The multi-modal interaction based intelligent guide doctor method for the elderly according to claim 1, wherein, The cost function is: , , , , wherein, denotes the total cost of a candidate node , denotes the actual cost of the current node to the candidate node , denotes the heuristic cost of the candidate node to the end point, denotes the cost of medical rules and environmental risks, denotes a weighted weight of , , and respectively denote the path length, the number of steps, the slope and the noise exposure risk of the current node to the candidate node , , , and respectively denote the weight coefficients of , , and ; and respectively denote the spatial distance of the candidate node to the end point and the environmental parameter on the spatial distance, and respectively denote the weighted weights of and ; denotes the number of medical rules triggered by the physiological state of the user, denotes the adjustment intensity of the th medical rule to the node cost, denotes an indicator function whether the th medical rule is triggered, and if triggered, otherwise; denotes the environmental risk of a candidate node , denotes a weighted weight of the environmental risk.
5. The multi-modal interaction based intelligent guide doctor method for the elderly according to claim 4, characterized in that S4 Comprise: In the case where the triage scenario includes an emergency scenario, increase the weight of the path length ; In the case where the triage scene includes an elderly patient scene, increase the weight of the path gentleness and ; In the case that the guidance scene comprises a noise sensitive scene, increasing the weight of the noise exposure risk ; In case the user's heart rate > first threshold: if the candidate node is a staircase node, set to 1 and assign a value greater than 0; if the candidate node is an elevator node, set to 1 and assign a value less than 0; In case the user blood oxygen is < the second threshold value: if the candidate node is a climbing node, set 1, and assign a value greater than 0; In the case that the corridor flow density is greater than the third threshold value, if the candidate node is a corridor node, the set proportion is increased; and the candidate node is selected as the target node. In case the elevator waiting time is greater than a fourth threshold value: if the candidate node is an elevator node, increase the set proportion; if the candidate node is a walking staircase node on the same floor as the elevator, decrease the set proportion. In case the elevator waiting time is greater than a fourth threshold value: if the candidate node is an elevator node, increase the set proportion; if the candidate node is a walking staircase node on the same floor as the elevator, decrease the set proportion.
6. The multi-modal interaction-based intelligent medical guide method for the elderly according to claim 3, characterized in that, The multi-modal interaction-based intelligent medical guide method for the elderly further comprises: The parameters of the attention mechanism and the classifier and the respective weighting weights are trained and updated according to the guide path accuracy and the user interaction satisfaction; The rationality of the parameters of the adaptive attention mechanism is verified according to the uniformity of the weight distribution between different modal features in a general scenario and the emphasis of the weight of each modal feature in different guide scenarios.
7. An elderly intelligent doctor guiding system based on multi-modal interaction, characterized in that, Comprise: The intelligent medical guide terminal is worn on the wrist of the elderly user and is used to obtain the guide destination of the user and detect the user state data in real time; The hospital server is used to detect the hospital environment data in real time; The multi-modal interaction engine is used to run a path planning algorithm with the guide destination as the terminal point and perform the following operations in each step of the path planning: S1, real-time update of user state data, voice instructions and environment data, wherein the user state data includes the user's position, posture and physiological state indicators, and the environment data includes crowd density, environmental noise and clinic queuing progress; S2, fusion of the user's basic information, user position, posture, voice instructions and environment data of the recent multiple time steps to obtain context features for comprehensively reflecting the current treatment environment; S3, updating the guide scenario of the user according to the context features; comprising: inputting the context features into a classifier to obtain the probability that the guide scenario of the user belongs to a general scenario, an emergency scenario, an elderly patient scenario or a noise-sensitive scenario; S4, according to the guide scenario, the medical rules triggered by the physiological state indicators, and the environmental risks, adjusting the cost function of each candidate node in the path planning, and selecting the next node of the guide path according to the adjusted cost function; S5, according to the pre-constructed registration transformation matrix, the real-time images of the multiple cameras closest to the user's future entry position at the next node are respectively registered to the user's future entry perspective, and then the registered images are secondarily registered to obtain the final image consistent with the entry perspective for the user to identify the next node; specifically, the user's future entry position refers to the intersection position of the channel between the current node and the next node and the next node, and the user's future entry perspective refers to the perspective along the channel pointing to the next node; the second registration of the registered images comprises: determining the intersection of the registered images; if the intersection exists part exceeding the field of view of the entry perspective, the part exceeding the field of view of the entry perspective in the intersection is deleted from the registered images, and the deleted images are secondarily registered; Before S5, further comprising: an image pair composed of an image taken by a same camera at a same time at the entering position along the entering view angle, and an image taken by a single camera closest to the entering position; fine-tuning a neural network model for outputting a registration transformation matrix between two images in an image pair using the image pairs from the plurality of cameras, so that registration matrices corresponding to each image pair from a same camera tend to be consistent, and registration transformation matrices corresponding to each image pair from different cameras are maximized; outputting registration transformation matrices from each camera to the entering view angle using the fine-tuned neural network model, to obtain the pre-constructed registration transformation matrix.
Citation Information
Patent Citations
Indoor positioning system based on registration of live-action picture and building three-dimensional model
CN116823906A
Intelligent hospital guide and tracking return visit recording integrated system for medical examination
CN119446454A