An internet of things medical device intelligent guiding system and method based on multi-modal interaction
By acquiring and processing medical device operation information through multimodal interaction technology and combining it with biometric verification, a personalized intelligent guidance scheme is generated, which solves the problems of accuracy, safety and adaptability of operation guidance for IoT medical devices, and realizes efficient and safe operation guidance in network-limited environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-14
AI Technical Summary
Existing IoT medical device operation guidance solutions cannot meet the needs for accurate, safe, and convenient operation guidance in medical scenarios. They cannot adapt to users with different cognitive levels, suffer from response delays in network-limited scenarios, lack identity verification and privacy protection, cannot dynamically adjust the level of detail in guidance, and lack closed-loop optimization.
By acquiring multimodal input information and biometric data, and combining ASR speech recognition and voiceprint recognition verification, personalized intelligent guidance is generated. It integrates text, diagram annotations, voice guidance and dynamic projection, and achieves offline operation on local area network and privacy and security protection. Personalized guidance is generated through local knowledge base and large model optimization.
It improves the accuracy, safety, and adaptability of medical device operation, ensures continuous availability in network-constrained environments, adapts to different terminals and user needs, and enhances the accuracy of guidance and user adaptability.
Smart Images

Figure CN121364915B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an intelligent guidance system and method for IoT medical devices based on multimodal interaction. Background Technology
[0002] Existing IoT medical device operation guidance solutions have significant technical flaws, making it difficult to meet the needs for accurate, safe, and convenient operation guidance in medical scenarios. Specific problems are as follows:
[0003] Traditional guidance solutions often rely on single text or voice output, failing to offer diverse interactive options tailored to the characteristics of medical scenarios and user habits. This results in a disconnect between the guidance and the actual user experience, hindering the development of a comprehensive system. Furthermore, it fails to meet the needs of users with varying cognitive levels, such as healthcare professionals and elderly patients, and struggles to match the output characteristics of multiple terminals, including AR glasses and device screens. Consequently, the guidance becomes detached from the user's actual usage scenario, reducing its efficiency.
[0004] Existing solutions largely rely on cloud-based knowledge bases and model inference, lacking a localized data processing and storage system. In core medical scenarios with limited network access, such as ICUs and operating rooms, issues like delayed guidance responses and functional malfunctions can easily arise, compromising the continuity and timeliness of medical equipment operation and hindering the progress of the diagnosis and treatment process.
[0005] The lack of operator authentication mechanisms specific to medical scenarios and the absence of biometric recognition and other technologies for controlling access permissions pose a risk of unauthorized personnel misoperating the equipment. Furthermore, the absence of a robust mechanism for de-identifying sensitive information and implementing tiered access control makes it susceptible to issues such as leakage of patient privacy and confidential device parameters during data transmission and storage, failing to comply with medical data security standards.
[0006] The system lacks integration of local structured operation graphs and user operation history data, resulting in mostly fixed templates for guidance content that cannot be dynamically adjusted in terms of detail based on scenario requirements and user proficiency. Furthermore, the absence of a closed-loop evaluation and optimization mechanism prevents continuous iteration and updates to the knowledge base and model parameters based on guidance performance, making it difficult to improve guidance accuracy and user adaptability.
[0007] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore includes information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0008] According to one aspect of this application, a method for intelligent guidance of IoT medical devices based on multimodal interaction is provided, comprising: acquiring multimodal input information related to medical device operation and operator biometric data, including voice commands, device operation trigger signals, voiceprint features, and medical scenario requirement information; processing the multimodal input information and biometric data, parsing the command intent through an ASR speech recognition engine, completing operator compliance verification with the help of voiceprint recognition, and extracting core keywords of device operation; processing the verified command intent and core keywords, retrieving and matching a structured operation graph based on a local knowledge base engine, and generating multimodal guidance materials adapted to the scenario through a target large language model combined with the operation history trajectory recorded by a context-aware module. The system processes multimodal guidance materials, analyzes user cognitive levels through a dynamic guidance generation system, adaptively adjusts the level of detail in the guidance, imposes offline operation constraints and privacy protection requirements in a local area network environment, and generates a multi-channel output adaptation scheme. It then processes the multi-channel output adaptation scheme, integrating text prompts, diagram annotations, voice guidance, and dynamic projections based on different terminal characteristics to generate personalized intelligent guidance execution instructions. Finally, it processes operator compliance verification results, multimodal guidance materials, multi-channel output adaptation schemes, and personalized intelligent guidance execution instructions, using a local knowledge base OTA silent update mechanism and large-model dynamic reasoning optimization to generate comprehensive evaluation information on medical device operation guidance, including guidance accuracy and user adaptability.
[0009] Another aspect of this application discloses an intelligent guidance device for IoT medical devices based on multimodal interaction, comprising: an acquisition module for acquiring multimodal input information related to medical device operation and operator biometric data, including voice commands, device operation trigger signals, voiceprint features, and medical scenario requirement information; a processing module for processing the multimodal input information and biometric data, parsing the command intent through an ASR speech recognition engine, verifying operator compliance with voiceprint recognition, and extracting core keywords for device operation; processing the verified command intent and core keywords, retrieving and matching a structured operation graph based on a local knowledge base engine, and generating multimodal guidance for the appropriate scenario through a target large language model combined with the operation history trajectory recorded by the context awareness module. Modal guidance materials; The system processes multimodal guidance materials, analyzes user cognitive levels through a dynamic guidance generation system, adaptively adjusts the level of detail in guidance, imposes offline operation constraints in a local area network environment and privacy and security protection requirements, and generates a multi-channel output adaptation scheme; The multi-channel output adaptation scheme is further processed, combining text prompts, diagram annotations, voice guidance, and dynamic projections to generate personalized intelligent guidance execution instructions, taking into account the characteristics of different terminals; The system processes operator compliance verification results, multimodal guidance materials, multi-channel output adaptation schemes, and personalized intelligent guidance execution instructions, and through a local knowledge base OTA silent update mechanism and large-model dynamic reasoning optimization, generates comprehensive evaluation information on medical device operation guidance, including guidance accuracy and user adaptability.
[0010] According to another aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a second processor, implements the above-described intelligent guidance method for IoT medical devices based on multimodal interaction.
[0011] This application presents an intelligent guidance system and method for IoT medical devices based on multimodal interaction. The server focuses on the pain points of operation guidance for IoT medical devices, proposing a multimodal interaction scheme that integrates a local knowledge base and a lightweight large model. By collecting multimodal data such as voice commands and voiceprint features, the system uses ASR speech recognition to analyze intent and voiceprint recognition to verify compliance, extracting core keywords. It retrieves operation graphs based on the local knowledge base, combines them with user operation history to generate suitable materials, and adaptively adjusts the level of guidance detail according to cognitive level, balancing offline operation on local area networks and privacy protection. Through multi-terminal integration of text, diagrams, voice, and dynamic projections, it outputs personalized guidance commands, which are then updated via OTA and the model optimized to generate comprehensive evaluation information including guidance accuracy and user adaptability. This technology effectively solves the problems of network dependence, single interaction, and insufficient compliance in traditional solutions, improving the accuracy, security, and adaptability of medical device operation guidance.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0013] Figure 1 A flowchart illustrating an embodiment of the present application provides a method for intelligent guidance of IoT medical devices based on multimodal interaction;
[0014] Figure 2 This illustration shows a structural schematic diagram of an IoT medical device intelligent guidance device based on multimodal interaction, provided in one embodiment of this application. Detailed Implementation
[0015] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0016] The following is combined with Figure 1 This application describes an intelligent guidance method for IoT medical devices based on multimodal interaction, according to exemplary embodiments thereof. In one embodiment, this application also proposes an intelligent guidance system and method for IoT medical devices based on multimodal interaction. Figure 1 As shown, this method is applied to a server and includes:
[0017] S101, acquire multimodal input information related to the operation of medical equipment and operator biometric data.
[0018] In one implementation, multimodal input information focuses on the core needs of medical device operation scenarios. A dedicated acquisition module simultaneously captures diverse input data to ensure coverage of different user operating habits and scenario requirements. An integrated ASR speech recognition engine, supporting targeted optimization for medical terminology, captures users' verbal commands in real time via the device's built-in microphone or an external speech acquisition device. For example, when a medical professional says "Start venous pressure monitoring" next to a hemodialysis machine, the system accurately recognizes this medical-specific command; when an elderly patient states "I need to register; my symptoms are chest tightness and shortness of breath," the engine accurately extracts the core needs and symptom keywords, avoiding misunderstandings of terminology.
[0019] By interfacing with medical device hardware, the system captures physical operation signals in real time, such as user pressing physical buttons, touching the control screen, and adjusting device knobs, and records the operation time and location simultaneously. When a user presses the "mode switch" button on a ventilator, the system immediately captures the trigger signal and associates it with the subsequent mode adjustment guidance process; touching the "parameter settings" option on a hemodialysis machine screen triggers the corresponding operation guidance preparation process. Through built-in sensors, an in-hospital scene positioning module, and user-initiated input, the system collects key information about the operation scenario, including scenario type, environmental parameters, and special requirements. In the ICU, the system identifies environmental attributes through the scene positioning module and automatically adapts to a highly reliable, low-interference guidance mode; during operations in the operating room, the system simultaneously captures the "aseptic operation" scenario requirement, and the guidance content avoids prompts requiring contact with non-sterile areas; when the user actively selects the "emergency priority" scenario tag, the system prioritizes responding to emergency operation-related guidance.
[0020] Biometric data acquisition focuses on operator compliance verification, collecting only voiceprint features as the core biometric identifier to ensure the security and necessity of data collection. The collection process is simple and convenient and does not affect the operation flow: Voiceprint samples are entered when the user first uses the system via a built-in or external voiceprint acquisition module. In subsequent operations, voiceprint features in the operator's speech are captured in real time for matching and verification. When a nurse uses the hemodialysis machine guidance system for the first time, she says "device operation verification" as prompted. The system records her voiceprint features and associates them with operation permissions. When the nurse subsequently says "replace blood pump tubing," the system simultaneously extracts voiceprint features and compares them with the stored sample. Only after successful verification is the subsequent guidance triggered. If other personnel without operation permissions attempt to issue the "adjust treatment parameters" command, voiceprint verification fails, the system refuses to execute the guidance and prompts "permission verification failed, please contact authorized personnel."
[0021] S102 processes multimodal input information and biometric data, parses the intent of commands through the ASR speech recognition engine, completes operator compliance verification with the help of voiceprint recognition, and extracts core keywords for device operation.
[0022] In one implementation, multimodal input information and biometric data are classified and processed to generate an adaptation association result between data types and processing mechanisms. The multimodal input information includes voice commands optimized for medical terminology, device operation trigger signals, and medical scenario requirement information; the biometric data consists of the operator's voiceprint feature data. First, the collected multimodal input information and biometric data are clearly classified to ensure accurate matching between data types and corresponding processing mechanisms, laying the foundation for subsequent processing steps. Multimodal input information is divided into three categories based on source and purpose: voice commands optimized for medical terminology, device operation trigger signals, and medical scenario requirement information; biometric data focuses solely on the operator's voiceprint feature data to avoid redundant collection. Based on the classification results, adaptation association rules are established to clarify the specific processing mechanisms corresponding to different data types: voice commands are adapted to ASR speech recognition processing optimized for medical terminology; device operation trigger signals are adapted to hardware interface linkage parsing processing; medical scenario requirement information is adapted to scenario attribute recognition and requirement extraction processing; and voiceprint feature data is adapted to voiceprint comparison verification processing. Finally, a one-to-one correspondence adaptation association result between data types and processing mechanisms is generated.
[0023] After collecting the voice command "Start the hemodialysis machine venous pressure monitoring", the trigger signal of pressing the device "start" button, ICU scene positioning information and operator voiceprint data, it is classified into four categories: voice command optimized with medical terminology, device operation trigger signal, medical scene requirement information and voiceprint feature data. These are then associated with ASR voice recognition processing, hardware interface parsing processing, scene attribute recognition processing and voiceprint comparison and verification processing to form a complete adaptation and association result.
[0024] The adaptation and association results are analyzed in a targeted manner to generate a list of core processing variables, including medical terminology recognition variables from the ASR speech recognition engine, compliance verification variables from voiceprint recognition, and device operation keyword extraction variables. The adaptation and association results are then broken down to extract key variables that need to be controlled during the processing of each type of data, constructing a list of core processing variables to ensure that the processing process is quantifiable and controllable. These core processing variables are set around key aspects of data processing, covering core dimensions such as recognition accuracy, verification standards, and extraction rules. Specifically, the list of core processing variables includes: medical terminology recognition variables from the ASR speech recognition engine (including medical terminology recognition accuracy threshold, professional vocabulary error tolerance rate, and dialect adaptation coefficient); compliance verification variables from voiceprint recognition (including voiceprint matching similarity threshold, number of sample comparisons, and permission association level); and device operation keyword extraction variables (including keyword extraction accuracy, core terminology screening rules, and multi-command keyword priority ranking standards).
[0025] For the voice command "Start venous pressure monitoring" in the ICU scenario, the medical terminology recognition variables of the ASR speech recognition engine were analyzed in a targeted manner. The recognition accuracy threshold was set at 95%, the professional vocabulary error tolerance rate at 5%, and the dialect adaptation coefficient at 0.8. For the compliance verification variable of voiceprint recognition, the matching similarity threshold was set at 90%, the number of sample comparisons was 3, and the permission association level was set to "medical operation level". For the equipment operation keyword extraction variable, the extraction accuracy was set at 98%, the core terminology selection rule was "equipment name + operation action + monitoring object", and the priority ranking standard for multi-command keywords was "emergency operation > routine operation > auxiliary operation".
[0026] The system performs simultaneous verification and compliance screening of the adaptation and association results, the list of core processing variables, and the original multimodal inputs and biometric data. Combined with a context-aware preprocessing mechanism, it generates instruction intent parsing results, operator compliance verification conclusions, and a core keyword set for device operation. This comprehensive simultaneous verification ensures data validity, the rationality of the processing mechanism, and the scientific nature of variable settings. Simultaneously, strict compliance screening eliminates non-compliant data and invalid processing logic. During the simultaneous verification process, the context-aware preprocessing mechanism correlates the user's historical operation trajectory, current medical scenario characteristics, and device operating status to comprehensively assess data integrity, consistency, and compliance. Compliance screening, based on medical data security regulations and device operation permission management requirements, filters out data from unauthorized operations, sensitive information, and invalid input data, ultimately generating three target results: instruction intent parsing results, operator compliance verification conclusions, and a core keyword set for device operation.
[0027] When a user triggers the "Start Venous Pressure Monitoring" voice command and the device start button signal in an ICU setting, voiceprint data is collected simultaneously. The context-aware preprocessing mechanism associates this with the user's historical operation records (all routine operations of the hemodialysis machine), the current ICU setting's "aseptic operation + high-priority response" characteristics, and the device's standby status to verify the data: confirming that the voice command is complete, the trigger signal is valid, the voiceprint data is clear, and the core processing variable settings meet the ICU setting operation requirements; the compliance screening process verifies that the operator's voiceprint matches the authorized medical staff sample, with no risk of sensitive information leakage. The final command intent parsing result is "Start the hemodialysis machine venous pressure monitoring function", the operator compliance verification conclusion is "verification passed, operation authority is granted", and the core keyword set for device operation is "hemodialysis machine, start, venous pressure monitoring".
[0028] S103 processes the verified instruction intent and core keywords, retrieves and matches the structured operation graph based on the local knowledge base engine, and generates multimodal guidance materials adapted to the scenario by combining the target large language model with the operation history trajectory recorded by the context-aware module.
[0029] In one implementation, verified instruction intents are categorized by scenario, broken down by requirement dimensions, and refined by core demands to generate basic parameters for intent parsing. First, verified instruction intents are broken down into multiple dimensions. Scenario categorization clarifies applicable medical scenarios, requirement dimension breakdown refines functional demands, and core demand refinement focuses on core objectives. Finally, standardized basic parameters for intent parsing are generated, providing clear direction for subsequent searches. Scenario categorization is based on medical scenario attributes, including outpatient registration, equipment operation, and emergency response; requirement dimension breakdown breaks down requirements into functional needs, operational needs, and timeliness needs; and core demand refinement extracts the user's most critical operational goals or service requirements. The verified instruction intent is "to activate the venous pressure monitoring function of the hemodialysis machine in the ICU setting," and the scenario classification is "equipment operation - ICU-specific equipment." The requirement dimension is broken down into functional requirement "activate venous pressure monitoring," operational requirement "hemodialysis machine-specific operation," and timeliness requirement "real-time monitoring." The core demand is extracted as "quickly activate the venous pressure monitoring of the hemodialysis machine." The basic parameters of the final generated intent parsing include the scenario tag "ICU + hemodialysis machine," the requirement type "equipment function activation," and the core objective "venous pressure monitoring."
[0030] The core keywords undergo medical terminology matching, semantic expansion, and operational scenario mapping to generate keyword retrieval adaptation parameters. These parameters include a terminology accuracy matching threshold, semantic relevance calculation rules, and a mapping table between keywords and equipment operation graphs. Core keywords related to equipment operation are professionally processed. Medical terminology matching ensures accuracy, semantic expansion enriches keyword coverage, and operational scenario mapping clarifies the corresponding equipment operation scenarios. This generates keyword retrieval adaptation parameters containing quantitative standards and rules to guarantee retrieval accuracy. Medical terminology matching is validated against authoritative medical terminology dictionaries; semantic expansion expands related terms based on the semantic network of the medical field; operational scenario mapping establishes the correspondence between keywords and specific equipment operation scenarios; and keyword retrieval adaptation parameters include a terminology accuracy matching threshold, semantic relevance calculation rules, and a mapping table between keywords and equipment operation graphs.
[0031] The core keyword set is "hemodialysis machine, start-up, venous pressure monitoring". Medical terminology matching verification shows that "venous pressure monitoring" is a standardized medical term without ambiguity. Semantic extension generates related keywords such as "hemodialysis machine, start-up, venous pressure monitoring". Operation scenario mapping determines the corresponding "hemodialysis machine - monitoring function start-up" scenario. In the generated keyword retrieval adaptation parameters, the terminology accuracy matching threshold is set to 90%, and the semantic relevance calculation rule is "core terminology matching weight 60% + scenario relevance weight 40%". The mapping relationship table between keywords and equipment operation graph clearly shows that "hemodialysis machine" corresponds to the "equipment body - hemodialysis machine" node in the graph, "start-up" corresponds to the "operation action - function start-up" node, and "venous pressure monitoring" corresponds to the "functional module - monitoring function - venous pressure" node.
[0032] Based on a local knowledge base engine, the structured operation graph undergoes hierarchical retrieval, multimodal content association, and precise fragment extraction. Through operation step matching, risk point annotation association, and binding of various material formats, graph retrieval result parameters are generated. The local knowledge base engine performs precise retrieval and content extraction on the structured operation graph. Hierarchical retrieval locates relevant content for corresponding devices and functions, multimodal content association integrates different forms of guidance materials, and precise fragment extraction filters core operations and risk warning content, ultimately generating graph retrieval result parameters to provide data support for guidance material generation. Hierarchical retrieval searches the graph according to a hierarchical structure of "device type - functional module - operation step"; multimodal content association binds different forms of materials such as text descriptions, 3D animations, and emergency videos; precise fragment extraction filters key content based on operation logic and risk level; and graph retrieval result parameters include operation step sequences, risk point annotation information, and a multimodal material association list.
[0033] Based on keyword retrieval and parameter adaptation, a hierarchical search is performed within the structured operation graph of the local knowledge base. First, the "hemodialysis machine" device node is located, then the "monitoring function" module is entered, ultimately pinpointing the "venous pressure monitoring activation" operation branch. Multimodal content is used to extract the corresponding textual step descriptions, 3D animations of tubing connections, and emergency response videos. Precise segment extraction identifies three core operation steps, two high-risk points (sensor interface leakage and tubing blockage), and corresponding warning labels. The generated graph retrieval results clearly define the operation steps as "check sensor → connect tubing → activate monitoring function," and the risk points are marked with "red for sensor interface, yellow for tubing direction." The multimodal material association list includes textual documentation, 3D animation file paths, and emergency video segment numbers.
[0034] By combining the target large language model with the operation history recorded by the context-aware module, user operation habits are analyzed, cognitive level is determined, and historical guidance preferences are adapted to generate personalized inference adjustment parameters. The target large language model (a lightweight open-source DeepseekR1 model) is used in conjunction with the user operation history recorded by the context-aware module to conduct in-depth analysis from three dimensions: operation habits, cognitive level, and guidance preferences. This generates personalized inference adjustment parameters to adapt guidance content to different user characteristics. User operation habit analysis statistically analyzes characteristics such as the frequency and sequence of user historical operations; cognitive level determination is based on historical interaction feedback and operation proficiency to classify cognitive levels; historical guidance preference adaptation analyzes the user's past preference for text, animation, voice, and other guidance formats.
[0035] The context-aware module recorded that the user was an ICU nurse. Historical operation trajectory showed that she operated the hemodialysis machine an average of 20 times per month, with standardized operation procedures and a preference for concise instructions. She had skipped basic operation instructions and directly viewed risk warnings on several occasions. The user's operation habit analysis result was "skilled operator with standardized operation procedures"; the cognitive level was determined to be "professional level (medical staff)"; the historical guidance preference adaptation result was "prioritize risk warnings and simplify basic steps"; the generated personalized inference adjustment parameters included an operation proficiency coefficient of "0.9", a cognitive level of "Level 3 (professional level)" and a guidance preference label of "risk priority + concise steps".
[0036] By imposing offline operation constraints on the local area network and multimodal content adaptation requirements, the above parameters are integrated and optimized. The validity of search results is verified through offline availability checks, and the content presentation format is adjusted based on user adaptability to generate multimodal guidance materials adapted to specific scenarios. Taking into account both offline operation constraints on the local area network and multimodal content adaptation requirements, the above-mentioned basic parameters for intent parsing, keyword search adaptation parameters, graph search result parameters, and personalized inference adjustment parameters are integrated and optimized. Offline availability checks ensure normal use in offline environments, and the presentation format of guidance content is adjusted based on user adaptability to ultimately generate multimodal guidance materials adapted to specific scenarios and users. Integration and optimization merge parameters according to the logic of "scenario adaptation - user adaptation - offline adaptation"; offline availability checks detect the offline call response speed and storage usage of materials; user adaptability adjustments adjust the level of detail and presentation format of guidance based on personalized parameters; multimodal guidance materials include various forms of guidance content such as text, 3D animation, voice, and diagrams.
[0037] After integrating all parameters, offline operation constraints on the local area network are applied, and the multimodal materials in the map retrieval results are lightweighted to ensure that the offline call response time is ≤1 second and the storage size of a single material is ≤50MB. Combined with personalized inference to adjust parameters, the text descriptions of basic operation steps are simplified, risk points are highlighted in red, and 3D animations and emergency videos are prioritized for association. The generated multimodal guidance materials adapted to the scenario include: concise text operation steps (only 3 core steps are retained), 3D demonstration animation of pipeline connection (focusing on sensor interface connection), voice prompt "Please confirm that there is no leakage at the sensor interface to avoid monitoring abnormalities", and pipeline connection diagram with red markings of common mistakes.
[0038] S104 processes the multimodal guidance materials, analyzes the user's cognitive level through the dynamic guidance generation system, adaptively adjusts the level of detail in the guidance, imposes offline operation constraints and privacy and security protection requirements in the local area network environment, and generates a multi-channel output adaptation scheme.
[0039] In one implementation, multimodal guidance materials are categorized and their content features extracted to generate features for text-based materials, 3D animation materials, emergency video materials, and voice-guided materials. The multimodal guidance materials for the applicable scenarios are clearly classified into four categories based on content format: text-based, 3D animation, emergency video, and voice-guided. Specific content features are then extracted for each category to form standardized feature descriptions, providing a basis for subsequent adaptation processing. Text-based material features focus on conciseness of expression, logical steps, and clarity of risk warnings; 3D animation material features focus on image clarity, accuracy of operational process reproduction, and labeling of key components; emergency video material features emphasize scene realism, timeliness of emergency response, and standardized operation demonstration; and voice-guided material features include voice clarity, instruction recognition, and appropriate speech rate.
[0040] For multimodal guidance materials on starting venous pressure monitoring in hemodialysis machines, the characteristics of text-based materials are: "concise steps (3 core operations), logical coherence, and two high-risk points highlighted in red"; the characteristics of 3D animation materials are: "1080P resolution, 1:1 reproduction of the operation process, and highlighted sensor interfaces"; the characteristics of emergency video materials are: "simulating the real ICU environment, emergency response time ≤30 seconds, and operation demonstration conforming to medical standards"; and the characteristics of voice-guided materials are: "clear voice without background noise, prominent key words in the instructions, and moderate speaking speed (150 words / minute)".
[0041] Based on the goal of user cognitive adaptation, a dynamic guidance generation system analyzes users' operational proficiency, knowledge background, and historical interaction feedback to quantify the differences in the level of detail required by different user groups, generating cognitive level adaptation parameters. With user cognitive adaptation as the core objective, the dynamic guidance generation system deeply analyzes user characteristics, quantifying the differences in users' needs for guidance detail from three dimensions: operational proficiency, knowledge background, and historical interaction feedback. This ultimately generates cognitive level adaptation parameters that can be directly used for adaptation adjustments. Operational proficiency is quantified through historical operation counts and error rates; knowledge background is categorized into levels based on groups such as medical staff, elderly patients, and general users; historical interaction feedback analyzes data such as skip rates and dwell times for past guidance content; and cognitive level adaptation parameters include cognitive level labels, guidance detail coefficients, and preferred key content directions.
[0042] For ICU nurses, the system analyzes their historical operation data: they operate the hemodialysis machine 20 times per month, with an operation error rate of 0.5%, and their knowledge background is that of professional medical staff. In historical interactions, the probability of skipping basic steps is as high as 80%, and the time spent on risk warnings accounts for 60%. The generated cognitive level adaptation parameters are: cognitive level label "professional-medical staff", guidance detail coefficient "0.3 (simplify basic content, strengthen core and risk)", and key content preference direction "risk warning > operation steps > auxiliary instructions".
[0043] Based on the constraints of offline operation on the local area network (LAN), the transmission rate of materials, storage usage, and offline call response time are detected and optimized to generate offline adaptation guarantee parameters, including lightweight material processing standards, offline caching priority rules, and emergency plans for material calls in offline environments. Strictly adhering to the requirements of offline operation on the LAN, key indicators such as the transmission, storage, and call of multimodal guide materials are detected and optimized, clarifying quantitative standards and emergency rules, and generating offline adaptation guarantee parameters to ensure normal use in offline environments. Material transmission rate detection and optimization ensure fast transmission within the LAN; storage usage is optimized and controlled within the device's local storage threshold; and offline call response time optimization ensures immediate response. Specific offline adaptation guarantee parameters include lightweight material processing standards, offline caching priority rules, and emergency plans for material calls in offline environments.
[0044] For guidance materials related to hemodialysis machines, the original 3D animation materials were found to occupy 150MB of storage, with a transmission rate of 2MB / s and an offline call response time of 1.5 seconds. In the optimized offline adaptation guarantee parameters, the material lightweighting standard is "3D animation compressed to within 50MB, transmission rate increased to above 5MB / s, and offline call response time ≤ 1 second"; the offline caching priority rule is "voice guidance > text description > 3D animation > emergency video"; the material call emergency plan in the offline environment is "prioritize calling local cached materials, and if the cache is corrupted, automatically use simplified text + voice guidance as a substitute".
[0045] To ensure privacy and security, sensitive information in the materials is anonymized, a tiered access control mechanism is established, and privacy protection and management parameters are generated. To further ensure medical data privacy and security, multimodal guidance materials undergo security processing. Sensitive information anonymization eliminates the risk of privacy leaks, a tiered access control mechanism ensures compliant use, and privacy protection and management parameters are generated. Sensitive information anonymization involves hiding or replacing patient information and confidential equipment parameters in the materials; the tiered access control mechanism divides access permissions according to user identity; and the privacy protection and management parameters include anonymization rules, permission level classification standards, and access verification requirements.
[0046] If the guidance materials contain confidential information such as equipment calibration parameters, the desensitization rules in the privacy protection and control parameters are "hide the core equipment calibration parameters and only retain operation-related prompts"; the permission level classification standards are "medical staff (can access the complete materials), equipment administrators (can access calibration-related materials), and ordinary users (only access basic operation materials)"; the access verification requirements are "secondary voiceprint verification is required to access confidential materials".
[0047] By combining the characteristics of multi-terminal output, cognitive level adaptation parameters, offline adaptation guarantee parameters, privacy protection control parameters, and terminal performance indicators are correlated and matched to dynamically adjust the presentation format and output priority of guidance content. Taking into account the characteristics of different terminals such as AR glasses, device screens, and voice broadcasts, cognitive level adaptation parameters, offline adaptation guarantee parameters, privacy protection control parameters, and terminal performance indicators are precisely correlated and matched to dynamically adjust the presentation format and output priority of guidance content to ensure adaptation to the usage needs of different terminals. Terminal performance indicators include display resolution, audio output capabilities, and storage capacity; presentation format adjustments optimize material formats based on terminal compatibility; output priority is sorted according to user needs and terminal characteristics; the correlation and matching follows the principle of "terminal capability adaptation + user needs priority + security and compliance as a safety net".
[0048] For AR glasses terminals, the performance specifications are "720P projection resolution, support for voice interaction, and 10GB local storage capacity"; after association and matching, the presentation format is adjusted to "3D animation adapted to 720P resolution, voice guidance synchronized with the projected image, and text descriptions simplified to floating prompts"; the output priority order is "voice guidance > 3D animation (highlighted) > risk warning text"; for device screen terminals, the presentation format is adjusted to "complete display of text descriptions, high-definition playback of 3D animations, and clickable emergency videos", and the output priority order is "operation steps text > 3D animation > voice guidance > emergency videos".
[0049] The above processing results are integrated and verified. Through offline usability testing, privacy compliance review, and user compatibility simulation assessment, a multi-channel output adaptation solution is generated. A comprehensive integration and verification of the above-mentioned cognitive level adaptation parameters, offline adaptation guarantee parameters, privacy protection control parameters, and terminal adaptation adjustment results are conducted. Three types of core tests ensure the solution's compliance, usability, and compatibility, ultimately generating a multi-channel output adaptation solution covering different terminals. Offline usability testing verifies the success rate of material retrieval and response time in offline environments; privacy compliance review verifies the completeness of sensitive information anonymization and the effectiveness of access permissions; user compatibility simulation assessment verifies the matching degree between the guidance content and user needs; and the multi-channel output adaptation solution clearly defines the material format, output sequence, and presentation priority for each terminal.
[0050] After integrating all parameters and performing verification, the offline availability test results showed a 100% success rate in calling materials and a response time of 0.8 seconds in a network-off environment; the privacy compliance audit confirmed that confidential parameters had been de-identified and access permissions were effective; the user compatibility simulation evaluation showed that user satisfaction with risk warnings reached 95%; the generated multi-channel output adaptation scheme is as follows: AR glasses terminal "720P 3D animation (highlighted) + synchronous voice guidance + floating risk text prompts, output sequence: voice first, animation plays synchronously, text is overlaid in real time"; device screen terminal "complete text description + 1080P 3D animation + clickable emergency video, output sequence: text is displayed first, animation plays with a 2-second delay, video is called on demand"; voice broadcast terminal "simplified voice commands + risk warning voice, output sequence: core operation commands are broadcast in sequence, risk warning is repeated once".
[0051] S105 processes multi-channel output adaptation solutions, combines the characteristics of different terminals, integrates multimodal content such as text prompts, diagram annotations, voice guidance, and dynamic projection, and generates personalized intelligent guidance execution instructions.
[0052] In one implementation, the terminal adaptation type, content presentation priority, and offline operation requirements in the multi-channel output adaptation scheme are classified and quantified to generate basic terminal adaptation parameters. Each terminal type corresponds to a unique output format adaptation coefficient, and each content type corresponds to a presentation weight coefficient. The core adaptation elements in the multi-channel output adaptation scheme are classified and quantified, clarifying the specific quantitative standards for terminal adaptation type, content presentation priority, and offline operation requirements. Basic terminal adaptation parameters, including output format adaptation coefficients and presentation weight coefficients, are generated to provide a quantitative basis for subsequent integration and matching. Terminal adaptation types are divided into three mainstream terminal categories: AR glasses, device screens, and voice broadcasts. Each terminal corresponds to a unique output format adaptation coefficient, quantifying the terminal's compatibility with different content formats. Content presentation priority is divided into four content types: text descriptions, 3D animations, voice guidance, and emergency videos. Each type corresponds to a presentation weight coefficient, quantifying the importance of the content in the terminal output. Offline operation requirements are quantified as the minimum local cache capacity of the terminal and the offline call response time threshold.
[0053] For the multi-channel output adaptation scheme for venous pressure monitoring in hemodialysis machines, the terminal adaptation basic parameters are as follows: terminal output format adaptation coefficient is set to AR glasses (0.9, adapting to 3D animation + dynamic projection), device screen (0.8, adapting to all types of content), and voice broadcast (0.6, adapting to voice guidance + text-to-speech); content presentation weight coefficient is set to risk warning text (0.4), 3D animation (0.3), voice guidance (0.2), and emergency video (0.1); offline operation requirements are quantified as a minimum local cache capacity of 1GB and an offline call response time of ≤1 second.
[0054] The system integrates and matches basic terminal adaptation parameters, different terminal characteristics, and multimodal content features to generate a multi-dimensional adaptation matrix that includes terminal adaptation, content adaptation, and experience optimization dimensions. The row vectors of the adaptation matrix represent the adaptation dimensions, and the column vectors represent specific quantitative indicators. The system deeply integrates and matches basic terminal adaptation parameters with different terminal characteristics and multimodal content features, breaking down quantitative indicators from the three dimensions of terminal adaptation, content adaptation, and experience optimization to construct a structured multi-dimensional adaptation matrix that clearly presents the quantitative requirements for each dimension. The terminal adaptation dimension focuses on the matching degree between terminal hardware performance and output format, with quantitative indicators including terminal resolution adaptation value, audio output compatibility, and storage capacity utilization. The content adaptation dimension focuses on the fit between content features and terminal characteristics, with quantitative indicators including content format adaptation rate, content size compression ratio, and multimodal content synchronization accuracy. The experience optimization dimension focuses on the potential for improvement in user experience, with quantitative indicators including operation guidance response speed, content comprehension difficulty coefficient, and user interaction convenience. The row vectors of the multi-dimensional adaptation matrix represent the three adaptation dimensions, and the column vectors represent the specific quantitative indicator values corresponding to each dimension.
[0055] Taking AR glasses as an example, the system integrates basic terminal adaptation parameters (output format adaptation coefficient 0.9, content presentation weight coefficient and risk warning text 0.4, etc.), terminal characteristics (projection resolution 720P, support for voice interaction, local storage 10GB), and multimodal content features (3D animation 1080P / 50MB, voice guidance 150 words / minute) to generate a multi-dimensional adaptation matrix: the terminal adaptation dimension quantitative indicators are resolution adaptation value 720P (1080P compression adaptation), audio output compatibility 100%, and storage capacity utilization rate 0.5%; the content adaptation dimension quantitative indicators are content format adaptation rate 90%, content size compression ratio 80%, and multimodal content synchronization accuracy ±0.5 seconds; the experience optimization dimension quantitative indicators are operation guidance response speed 0.8 seconds, content understanding difficulty coefficient 0.2 (low difficulty), and user interaction convenience 95%.
[0056] The multi-dimensional adaptation matrix is input into the intelligent guidance instruction generation model and subjected to weighted fusion calculations to generate multimodal content integration rules adapted to different terminals. These rules include the logic for simultaneous output of text and illustrations, the linkage mechanism between voice guidance and operation steps, and the matching strategy between dynamic projection and real-time operation. The multi-dimensional adaptation matrix is input into the intelligent guidance instruction generation model, and a weighted fusion algorithm is used to calculate the quantitative indicators of each dimension. Combined with terminal characteristics and user needs, multimodal content integration rules adapted to different terminals are generated, clarifying the logic, mechanism, and strategy of content output. In the weighted fusion calculation, the terminal adaptation dimension accounts for 40% of the weight, the content adaptation dimension accounts for 35%, and the experience optimization dimension accounts for 25%. The multimodal content integration rules specifically include the logic for simultaneous output of text and illustrations, the linkage mechanism between voice guidance and operation steps, and the matching strategy between dynamic projection and real-time operation, ensuring collaborative output of multimodal content and adaptation to terminal usage scenarios.
[0057] The multi-dimensional adaptation matrix corresponding to the AR glasses is input into the model. After weighted fusion calculation, the multimodal content integration rules are as follows: Synchronous output logic of text and diagrams: "The diagram is displayed floating above the operation area, and the text of key steps pops up synchronously in the form of a highlighted floating window"; Linkage mechanism of voice guidance and operation steps: "After each operation step is completed, the voice guidance delays by 0.3 seconds to broadcast the next step instruction. If an operation error is detected, the voice is immediately paused and an error correction prompt is broadcast"; Matching strategy of dynamic projection and real-time operation: "The dynamic projection screen adjusts its angle according to the operator's hand movements. The operation progress and the projection content are synchronized in real time with an error of ≤0.5 seconds."
[0058] Based on a multimodal content integration rule-triggered instruction generation mechanism, personalized intelligent guidance execution instructions are generated. These instructions include terminal startup parameters, content output timing, and interactive feedback response rules. Based on the generated multimodal content integration rules, the instruction generation mechanism is triggered, combining specific operational scenarios and user characteristics to generate personalized intelligent guidance execution instructions containing core elements such as terminal startup, content output, and interactive feedback. This ensures the instructions are directly executable and highly adaptable. Terminal startup parameters specify the terminal startup method and initialization configuration; content output timing specifies the output order and time intervals for various types of content; interactive feedback response rules specify the recognition method and response strategy for user operation feedback; personalized intelligent guidance execution instructions must be consistent with the aforementioned scenarios and user characteristics to ensure effective execution.
[0059] For ICU nurses using AR glasses to operate a hemodialysis machine and initiate venous pressure monitoring, the generated personalized intelligent guidance execution instructions are as follows: Terminal startup parameters: "Start AR glasses through voiceprint verification (match success threshold 90%), initialize projection mode to 'medical equipment operation exclusive mode', resolution adapted to 720P"; Content output sequence: "First project pipeline connection diagram (0 seconds) → synchronously pop up 'Check sensor' text prompt (0.2 seconds) → voice broadcast 'Please check sensor interface for leaks' (0.5 seconds) → after detecting sensor connection operation, project dynamic monitoring interface (0.3 seconds after operation completion)"; Interactive feedback response rules: "Recognize operation actions through AR glasses camera. If the operation is correct, voice feedback 'Operation successful' will be given and the next step will be advanced. If the operation is wrong, an error correction diagram will be projected immediately + voice broadcast 'Please reconnect the sensor to ensure no leakage'. If there are 3 consecutive errors, a manual intervention request will be triggered."
[0060] S106 processes the operator compliance verification results, multimodal guidance materials, multi-channel output adaptation solutions, and personalized intelligent guidance execution instructions. Through the local knowledge base OTA silent update mechanism and large model dynamic reasoning optimization, it generates comprehensive evaluation information of medical device operation guidance, including guidance accuracy and user adaptability.
[0061] In one implementation, the system collects data on operator compliance verification pass rate, multimodal guidance material call success rate, multi-channel output adaptation compatibility, and personalized intelligent guidance execution instruction completion rate. Core performance indicators for each stage are extracted to generate a basic evaluation data vector. The system comprehensively collects performance data from all key stages, focusing on four core stages: operator compliance verification, multimodal guidance material call, multi-channel output adaptation, and personalized intelligent guidance execution. Quantitative indicators are extracted and a basic evaluation data vector is constructed to provide raw data support for subsequent evaluations. The operator compliance verification pass rate is calculated as the ratio of the number of successful compliance verifications to the total number of verifications within a certain period; the multimodal guidance material call success rate is calculated as the ratio of the number of successful material calls to the total number of call requests; the multi-channel output adaptation compatibility is calculated as the ratio of the number of scenarios successfully adapted to different terminals to the total number of scenarios; and the personalized intelligent guidance execution instruction completion rate is calculated as the ratio of the number of complete instruction executions to the total number of issuances. The basic evaluation data vector is presented in the format of "[compliance verification pass rate, material call success rate, output adaptation compatibility, instruction execution completion rate]".
[0062] For the venous pressure monitoring guidance scenario of hemodialysis machines, relevant operational data within one month were collected: a total of 100 compliance verifications were conducted, with 98 passing; a total of 80 material retrieval requests were made, with 78 successful; 28 scenarios were successfully adapted to 30 scenarios in total, involving 3 types of terminals: AR glasses, device screens, and voice broadcasts; 75 instructions were issued, with 72 being fully executed; the generated evaluation base data vector is "[98%, 97.5%, 93.3%, 96%]".
[0063] The evaluation process comprehensively verifies and quantifies the compliance verification effectiveness, material adaptation accuracy, output channel stability, and instruction execution integrity within the basic data vector, generating a multi-dimensional preliminary evaluation result. The core indicators in the basic data vector are comprehensively verified, and quantitative scoring is conducted across four dimensions: compliance verification effectiveness, material adaptation accuracy, output channel stability, and instruction execution integrity. This results in a multi-dimensional preliminary evaluation result, clarifying the performance level of each stage. Compliance verification effectiveness is scored based on pass rate and analysis of verification error causes (maximum score 100 points); material adaptation accuracy is scored based on call success rate and material matching deviation (maximum score 100 points); output channel stability is scored based on adaptation compatibility and terminal adaptation failure frequency (maximum score 100 points); and instruction execution integrity is scored based on execution completion rate and analysis of instruction interruption causes (maximum score 100 points). The multi-dimensional preliminary evaluation result is based on the scores of each dimension and the overall average score.
[0064] Based on the aforementioned basic data vector, a preliminary evaluation result was generated after comprehensive verification: compliance verification effectiveness score: 95 points (pass rate 98%, 2 failures were due to user voiceprint input deviation, not system issues); material adaptation accuracy score: 94 points (success rate 97.5%, 2 failures were due to temporary damage to material cache); output channel stability score: 88 points (adaptation compatibility 93.3%, 2 scene adaptation failures were due to AR glasses resolution adaptation deviation); instruction execution integrity score: 93 points (execution completion rate 96%, 3 interruptions were due to user operation exiting midway); the overall average score is (95+94+88+93) / 4=92.5 points.
[0065] Based on the preliminary assessment results, the latest operating standards and optimized cases are synchronized through the local knowledge base OTA silent update mechanism. Combined with the dynamic reasoning of the target large language model, the assessment weights are adjusted to correct the guidance accuracy calculation results. Based on the multi-dimensional preliminary assessment results, the latest standards are synchronized and assessment weights are adjusted using the local knowledge base OTA silent update mechanism and the dynamic reasoning capabilities of the target large language model, thus correcting the guidance accuracy calculation results and improving assessment accuracy. The local knowledge base OTA silent update mechanism automatically synchronizes the latest operating standards for medical equipment, fault handling cases, and other content to ensure the timeliness of the assessment basis; the target large language model (a lightweight open-source deepseekR1 model) combines historical assessment data and new scenario data to dynamically reason and adjust the assessment weights of each dimension; the guidance accuracy is calculated as the sum of "(corrected accuracy of each step × corresponding weight)".
[0066] After initial assessment, the local knowledge base silently updated the latest operating procedures for venous pressure monitoring of hemodialysis machines and added one new fault handling case via OTA. After dynamic reasoning of the large model, the evaluation weights were adjusted: compliance verification effectiveness weight 20%, material adaptation accuracy weight 30%, output channel stability weight 30%, and instruction execution integrity weight 20%. Before the correction, the adaptation accuracy rates of each link were 98%, 97.5%, 93.3%, and 96%, respectively. After the correction, combined with the newly synchronized content and weight adjustments, the final guidance accuracy rate was calculated as 98%×20%+97.5%×30%+93.3%×30%+96%×20%=96.04%.
[0067] By utilizing user interaction history recorded by the context-aware module, we analyze the differences in feedback from users with different cognitive levels to the guided service, quantify user fit levels, and formulate personalized fit assessment rules. We also use the context-aware module to deeply analyze the differences in feedback from users with different cognitive levels to the guided service, quantify user fit levels, and formulate standardized personalized fit assessment rules to supplement assessment dimensions. User cognitive levels are divided into three categories: professional medical staff, general users, and elderly / special users. Fit is quantified by analyzing data such as operation time, error rate, and feedback scores for different groups. User fit levels are divided into four levels: excellent (above 90 points), good (80-89 points), satisfactory (70-79 points), and needs improvement (below 70 points). The personalized fit assessment rules clearly define the fit standards and optimization directions for users at each cognitive level.
[0068] Based on data from the context-aware module, feedback from three user groups was analyzed: professional medical staff (such as ICU nurses) had an average operation time of 2 minutes per operation, an error rate of 0.5%, and a feedback score of 95; ordinary users had an average operation time of 3.5 minutes per operation, an error rate of 3%, and a feedback score of 85; and elderly users had an average operation time of 5 minutes per operation, an error rate of 6%, and a feedback score of 78. The user fit level was quantified as good (average score of 86). The resulting personalized fit assessment guidelines clearly state that "professional medical staff need to strengthen the simplification of risk prompts, ordinary users need to optimize the intuitiveness of operation steps, and elderly users need to increase the frequency of voice guidance repetition."
[0069] The revised guidance accuracy rate, user suitability level, and performance scores of each stage are integrated and weighted using a multi-dimensional evaluation model to generate a comprehensive evaluation of medical device operation guidance, including data support, level determination, and optimization directions. The multi-dimensional evaluation model assigns 60% weight to guidance accuracy and 40% weight to user suitability. Level determination is based on the overall score, categorized as Excellent (above 90 points), Good (80-89 points), Satisfactory (70-79 points), and Needs Optimization (below 70 points). Optimization directions propose specific improvement measures for shortcomings in each stage.
[0070] The integrated and revised guidance achieved an accuracy rate of 96.04%, a user suitability score of 86, and performance scores for each stage (compliance 95, materials 94, channels 88, instructions 93). Through a multi-dimensional evaluation model with weighted calculation, the overall score was 96.04% × 60 + 86 × 40% = 92.02, resulting in an "Excellent" rating. The generated comprehensive evaluation information for medical device operation guidance is as follows: Data supports "guidance accuracy rate of 96.04%, good user suitability, and an average score of 92.5 for each stage"; the rating is "Excellent"; optimization directions include: "1. Optimize the AR glasses resolution adaptation algorithm to improve output channel stability; 2. Add a voice guidance repetition function for elderly users to improve suitability for special groups; 3. Optimize the material caching mechanism to reduce material retrieval failure rate."
[0071] In one implementation, such as Figure 2 As shown, this application also provides an intelligent guidance device for IoT medical devices based on multimodal interaction, comprising:
[0072] The acquisition module 201 is used to acquire multimodal input information related to the operation of medical equipment and operator biometric data, including voice commands, equipment operation trigger signals, voiceprint features and medical scenario requirement information;
[0073] Processing module 202 is used to process multimodal input information and biometric data. It parses the command intent using an ASR speech recognition engine, verifies operator compliance with voiceprint recognition, and extracts core keywords for device operation. It then processes the verified command intent and core keywords, retrieves and matches structured operation graphs based on a local knowledge base engine, and generates multimodal guidance materials adapted to the specific scenario using a target large language model combined with the operation history recorded by the context-aware module. Finally, it processes the multimodal guidance materials, analyzes the user's cognitive level through a dynamic guidance generation system, and adaptively adjusts the level of detail in the guidance. Imposing offline operation constraints and privacy protection requirements in a local area network environment, a multi-channel output adaptation scheme is generated. This scheme is then processed, and multimodal content including text prompts, diagram annotations, voice guidance, and dynamic projections is integrated to generate personalized intelligent guidance execution instructions, taking into account the characteristics of different terminals. The operator compliance verification results, multimodal guidance materials, multi-channel output adaptation scheme, and personalized intelligent guidance execution instructions are processed, and through a local knowledge base OTA silent update mechanism and large-model dynamic reasoning optimization, a comprehensive evaluation of medical device operation guidance, including guidance accuracy and user adaptability, is generated.
[0074] The various embodiments in this application are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for evaluating the intelligent guidance method, electronic device, electronic device, and readable storage medium for IoT medical devices based on multimodal interaction are basically similar to the above-described embodiments of the intelligent guidance method for IoT medical devices based on multimodal interaction, and therefore are described relatively simply. Relevant parts can be referred to in the descriptions of the above-described embodiments of the intelligent guidance method for IoT medical devices based on multimodal interaction.
Claims
1. A method for intelligent guidance of IoT medical devices based on multimodal interaction, characterized in that, include: Acquire multimodal input information related to the operation of medical devices and operator biometric data, including voice commands, device operation trigger signals, voiceprint features and medical scenario requirements; The system processes multimodal input information and biometric data, analyzes the intent of commands through the ASR speech recognition engine, verifies operator compliance with voiceprint recognition, and extracts core keywords for device operation. The verified instruction intent and core keywords are processed, and the structured operation graph is retrieved and matched based on the local knowledge base engine. The operation history trajectory recorded by the target large language model and the context awareness module is combined to generate multimodal guidance materials adapted to the scenario. The multimodal guidance materials are processed, and the user's cognitive level is analyzed through the dynamic guidance generation system. The level of detail of the guidance is adaptively adjusted, and offline operation constraints and privacy and security protection requirements in the local area network environment are imposed to generate a multi-channel output adaptation scheme. The multi-channel output adaptation scheme is processed, and multimodal content such as text prompts, diagram annotations, voice guidance, and dynamic projection are integrated to generate personalized intelligent guidance execution instructions, taking into account the characteristics of different terminals. The system processes operator compliance verification results, multimodal guidance materials, multi-channel output adaptation solutions, and personalized intelligent guidance execution instructions. Through a local knowledge base OTA silent update mechanism and dynamic reasoning optimization using a large model, it generates comprehensive evaluation information on medical device operation guidance, including guidance accuracy and user adaptability. This includes collecting operator compliance verification pass rate, multimodal guidance material retrieval success rate, multi-channel output adaptation compatibility, and personalized intelligent guidance execution instruction completion rate. Core performance indicators for each stage are extracted to generate a basic evaluation data vector. The system then comprehensively verifies and quantifies the compliance verification effectiveness, material adaptation accuracy, output channel stability, and instruction execution completeness within the basic evaluation data vector, generating multi-dimensional preliminary evaluation results. Based on these preliminary evaluation results, the system synchronizes the latest operating standards and optimized cases through the local knowledge base OTA silent update mechanism, and adjusts the evaluation weights using dynamic reasoning combined with the target large language model to correct the guidance accuracy calculation results. By utilizing the user interaction history recorded by the context-aware module, we can analyze the differences in feedback from users with different cognitive levels to the guidance service, quantify the user fit level, and form personalized fit assessment rules. By integrating the revised guidance accuracy, user fit level, and performance scores of each stage, we can generate comprehensive evaluation information for medical device operation guidance that includes data support, level determination, and optimization direction through weighted calculation by a multi-dimensional evaluation model.
2. The method as described in claim 1, characterized in that, The system processes multimodal input information and biometric data, analyzes the command intent using an ASR speech recognition engine, verifies operator compliance with voiceprint recognition, and extracts core keywords for device operation, including: The multimodal input information and biometric data are classified and processed to generate the matching and association results of data types and processing mechanisms. The multimodal input information includes voice commands optimized for medical terminology, device operation trigger signals, and medical scenario requirement information, while the biometric data is the operator's voiceprint feature data. The adaptation and association results are analyzed in a targeted manner to generate a list of core processing variables, including medical terminology recognition variables of the ASR speech recognition engine, compliance verification variables of voiceprint recognition, and device operation keyword extraction variables. The system performs synchronous verification and compliance screening on the adaptation and association results, the list of core processing variables, and the original multimodal inputs and biometric data. Combined with the context-aware preprocessing mechanism, it generates instruction intent parsing results, operator compliance verification conclusions, and a set of core keywords for equipment operation.
3. The method as described in claim 1, characterized in that, The verified instruction intent and core keywords are processed, and a structured operation graph is retrieved and matched based on the local knowledge base engine. Then, using a target large language model combined with the operation history recorded by the context-aware module, multimodal guidance materials adapted to the scenario are generated, including: The verified instruction intents are classified by scenario, broken down by requirement dimensions, and extracted by core demands to generate basic parameters for intent parsing. The core keywords are matched with medical terms, expanded with related semantics, and mapped to operational scenarios to generate keyword retrieval adaptation parameters, including the terminology accuracy matching threshold, semantic relevance calculation rules, and a mapping relationship table between keywords and equipment operation graphs. Based on the local knowledge base engine, the structured operation graph is subjected to hierarchical retrieval, multimodal content association and precise fragment extraction. Through operation step matching, risk point annotation association and multi-form material binding, graph retrieval result parameters are generated. By combining the target large language model with the operation history trajectory recorded by the context awareness module, user operation habits are analyzed, cognitive level is determined, and historical guidance preferences are adapted to generate personalized reasoning adjustment parameters. By imposing constraints on offline operation on the local area network and requirements for multimodal content adaptation, the above parameters are integrated and optimized. The validity of the search results is verified through offline availability, and the content presentation format is adjusted according to user adaptability to generate multimodal guidance materials adapted to the scenario.
4. The method as described in claim 1, characterized in that, The multimodal guidance materials are processed, and the user's cognitive level is analyzed through a dynamic guidance generation system. The level of detail in the guidance is adaptively adjusted, and offline operation constraints and privacy and security protection requirements in a local area network environment are imposed. A multi-channel output adaptation scheme is generated, including: Multimodal guidance materials are categorized and their content features are extracted to generate features for text description materials, 3D animation materials, emergency video materials, and voice guidance materials. Based on the user cognitive adaptation target, the dynamic guidance generation system analyzes the user's operational proficiency, knowledge background and historical interaction feedback, quantifies the differences in the needs of different user groups for the level of detail of guidance, and generates cognitive level adaptation parameters. Based on the constraints of offline operation on the local area network, the transmission rate of materials, storage usage, and offline call response time are detected and optimized to generate offline adaptation guarantee parameters, including material lightweight processing standards, offline caching priority rules, and emergency plans for material calls in the absence of network. In accordance with privacy and security protection requirements, sensitive information in the materials is anonymized, a hierarchical access mechanism for operation permissions is set up, and privacy protection and control parameters are generated; By combining the characteristics of multi-terminal output, the cognitive level adaptation parameters, offline adaptation guarantee parameters, privacy protection and control parameters and terminal performance indicators are correlated and matched to dynamically adjust the presentation format and output priority of the guidance content; The above processing results are integrated and verified. Through offline usability testing, privacy compliance auditing, and user adaptability simulation evaluation, a multi-channel output adaptation solution is generated.
5. The method as described in claim 4, characterized in that, The multi-channel output adaptation scheme is processed, and multimodal content including text prompts, diagram annotations, voice guidance, and dynamic projections is integrated to generate personalized intelligent guidance execution instructions, including: The terminal adaptation type, content presentation priority, and offline operation requirements in the multi-channel output adaptation scheme are classified and quantified to generate basic terminal adaptation parameters. Among them, each terminal type corresponds to a unique output format adaptation coefficient, and each content type corresponds to a presentation weight coefficient. The basic parameters of terminal adaptation, the characteristics of different terminals, and the features of multimodal content are integrated and matched to generate a multi-dimensional adaptation matrix that includes terminal adaptation dimension, content adaptation dimension, and experience optimization dimension. The row vectors of the adaptation matrix are the adaptation dimensions, and the column vectors are the specific quantitative indicators. The multi-dimensional adaptation matrix is input into the intelligent guidance instruction generation model and weighted and fused to generate multimodal content integration rules adapted to different terminals, including the logic of simultaneous output of text and diagrams, the linkage mechanism between voice guidance and operation steps, and the matching strategy of dynamic projection and real-time operation. Based on the multimodal content integration rule-triggered instruction generation mechanism, personalized intelligent guidance execution instructions are generated. These instructions include terminal startup parameters, content output timing, and interactive feedback response rules.
6. A smart guidance device for IoT medical devices based on multimodal interaction, characterized in that, The device includes: The acquisition module is used to acquire multimodal input information related to the operation of medical devices and operator biometric data, including voice commands, device operation trigger signals, voiceprint features and medical scenario requirements information; The processing module processes multimodal input information and biometric data. It parses command intent using an ASR speech recognition engine, verifies operator compliance with voiceprint recognition, and extracts core keywords for device operation. It then processes the verified command intent and core keywords, retrieves and matches structured operation graphs based on a local knowledge base engine, and generates multimodal guidance materials adapted to specific scenarios using a target large language model combined with operation history records from a context-aware module. The module further processes these multimodal guidance materials, analyzes user cognitive levels through a dynamic guidance generation system, adaptively adjusts the level of detail, imposes offline operation constraints and privacy protection requirements in a local area network environment, and generates a multi-channel output adaptation scheme. This multi-channel output adaptation scheme is then processed, integrating text prompts, diagram annotations, voice guidance, and dynamic projections to generate personalized intelligent guidance execution instructions, taking into account different terminal characteristics. Finally, it processes the operator compliance verification results, multimodal guidance materials, multi-channel output adaptation schemes, and personalized intelligent guidance execution instructions, optimizing them through a local knowledge base OTA silent update mechanism and dynamic reasoning of a large model to generate a package. This system provides a comprehensive evaluation of medical device operation guidance, including accuracy and user fit. The evaluation process involves collecting data on operator compliance verification pass rate, multimodal guidance material retrieval success rate, multi-channel output compatibility, and personalized intelligent guidance instruction execution completion rate. Core performance indicators for each stage are extracted to generate a basic evaluation data vector. The system then comprehensively verifies and quantifies the compliance verification effectiveness, material fit accuracy, output channel stability, and instruction execution completeness within this data vector, generating a multi-dimensional preliminary evaluation result. Based on this preliminary evaluation result, the system synchronizes the latest operation standards and optimization cases through a local knowledge base OTA silent update mechanism. Combined with dynamic reasoning using a target large language model, the evaluation weights are adjusted to correct the guidance accuracy calculation result. Utilizing user interaction history recorded by the context-aware module, the system analyzes the feedback differences of users with different cognitive levels regarding the guidance service, quantifies user fit levels, and forms personalized fit evaluation rules. Finally, the system integrates the corrected guidance accuracy, user fit levels, and performance scores for each stage, and performs weighted calculations using a multi-dimensional evaluation model to generate a comprehensive evaluation of medical device operation guidance, including data support, level determination, and optimization directions.
7. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the IoT medical device intelligent guidance method based on multimodal interaction as described in any one of claims 1 to 5 by executing the executable instructions.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the second processor, it implements the intelligent guidance method for IoT medical devices based on multimodal interaction as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Remote guidance system and method, computer equipment and readable storage medium
CN112216376A
Method for improving man-machine interaction efficiency by sensing environmental information through AI
CN120561360A