A medical device respiratory parameter intelligent regulation system and method based on voice interaction
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-16
- Publication Date
- 2026-08-11
AI Technical Summary
本发明专利通过采用语音采集与增强模块及指向性麦克风阵列结合自适应波束成形算法,本发明系统实现了在充斥监护仪警报、电刀噪声、人员对话等复杂声学环境的医疗场景中,对操作者语音指令的清晰、稳定拾取,该模块能够动态追踪声源方向,智能增强目标语音,同时强力抑制非目标方向的背景噪声,从根本上解决了传统语音控制系统在医疗环境下因噪声干扰导致识别率急剧下降的难题,为后续高可靠性的语音交互奠定了坚实的基础;通过构建操作者身份与指令安全验证模块,并创新性地采用声纹生物识别与动态口令双重认证的融合决策模型(C = α * S_norm + β * F),系统构筑了指令准入的第一道坚实防线,声纹识别确保了“谁在说”的身份合法性,动态口令验证则增加了“在正确场景下说”的上下文安全性。这种双重因子认证机制,有效解决了无线遥控或简单语音控制方案身份验证薄弱、易被仿冒或误触发的高风险问题,确保了只有经过严格授权且意图明确的操作者才能启动控制流程;通过开发医疗语义精准解析与意图理解模块及其后台的医疗知识图谱,系统实现了对专业化、口语化医疗指令的深度理解。模块内置的医疗术语优先匹配算法(W_medical >> W_general)确保了解析的准确性,能够正确理解“潮气量”、“PEEP”、“SIMV”等专业术语以及“提高百分之二十”、“同步降低”等相对性和关联性指令,解决了通用语音助手无法理解医疗专业语境、需要用户使用固定格式化命令的痛点,使人机交互自然、高效,显著降低了医护人员的认知负担与学习成本;通过引入患者状态自适应的参数安全边界决策模块及基于动态风险指标R = f(ΔP, S, Φ)的评估体系,系统将安全控制从被动的“事后报警”提升至主动的“事前预防”和“事中干预”,该模块能够结合患者的实时生命体征(如血氧饱和度、气道压力),智能判断拟调整参数是否安全、有效,并能自动协调多个关联参数的联动设置,避免参数间冲突,解决了手动操作模式下,依赖个人经验进行复杂安全校验效率低、易出错的重大缺陷,为患者提供了基于实时数据的、个性化的智能安全守护;通过设计多模态交互式安全确认与执行模块及防误触物理确认装置(触发条件:压力>P_th且时间>T_min),系统建立了“听觉复核、视觉聚焦、触觉确认”的三重安全冗余机制,语音复述让操作者进行首次听觉核对;屏幕高亮提供直观的视觉确认点;最终的物理确认动作则创造了一个必须有意为之的操作中断点,彻底杜绝了因语音识别偶然错误或操作者口误而导致的误执行风险,完美解决了简单语音控制系统缺乏有效二次确认、可靠性不足的核心安全问题;通过集成分级应急响应处理子模块,系统实现了常态安全流程与应急处理效率的智能平衡,在常规操作中,完备的安全流程得以严格执行;而当系统识别到包含“抢救”、“纯氧”等危机关键词且身份验证通过的指令时,自动进入红色应急通道,可简化流程、快速调用预设方案。这解决了在紧急情况下,冗长的安全确认流程可能延误抢救时机的矛盾,使系统既能保障日常操作的安全严谨,又能满足急救场景下的快速响应需求;通过采用遵循国际/行业标准的控制指令标准化封装与通信模块,实现了与不同品牌、型号医疗设备的标准化对接,基于通用协议的设计,使得本系统可以作为一个开放的智能控制平台,广泛适配于手术室、ICU内各类符合标准的麻醉机、呼吸机,解决了私有协议导致的兼容性差、集成成本高的行业痛点,提升了方案的普适性与推广价值;通过部署非操作者语音过滤子模块及基于概率P_operator的判定机制,系统显著增强了在特殊场景下的适用性与安全性。在儿科病房或患者可能发声的床边,该模块能有效区分医护人员与患者/家属的语音特征,自动过滤掉非授权声源发出的、可能包含控制词汇的指令,防止意外触发。综上,本发明通过整合高鲁棒性语音处理、多重安全认证、医疗人工智能、多模态交互及标准化设备通信等关键技术,构建了一个安全等级高、响应速度快、交互体验自然的医疗设备智能调控系统,实现了“非接触、动口不动手”的便捷操作,大幅提升工作效率并降低感染风险,更重要的是通过从前端到后端的多层次、闭环式安全设计,为高风险医疗操作构筑了坚实的可靠性屏障。
Smart Images

Figure CN122551786A_ABST
Abstract
Description
Technical Field
[0001] This invention patent relates to the field of intelligent control technology for medical devices, specifically to an intelligent control system and method for respiratory parameters of medical devices based on voice interaction. Background Technology
[0002] In modern clinical medicine, anesthesia machines and ventilators are core equipment for maintaining patients' respiratory function, ensuring surgical safety, and providing critical care. The accuracy of their parameter settings and the timeliness of their adjustments are directly related to the patient's life and safety.
[0003] Currently, the mainstream method of parameter adjustment relies entirely on medical staff manually operating physical knobs or touchscreen interfaces on the equipment. This traditional interaction mode has a series of inherent drawbacks that urgently need to be addressed in complex medical environments: The process is inefficient and carries the risk of delayed response. When a patient's vital signs suddenly change and require urgent adjustments to respiratory parameters, medical staff must interrupt their other work and turn to the equipment for multi-step menu navigation and numerical settings. This process, from the generation of the intention to adjust to the completion of the physical operation, involves a significant time delay, which, in a race against time in a rescue scenario, may delay the optimal intervention time.
[0004] The procedure disrupts the sterile environment and increases the risk of infection. In operating rooms where strict aseptic techniques are required, adjusting equipment parameters often necessitates medical staff removing their gloves or requesting assistance from others. This not only compromises the sterile barrier and increases the risk of surgical site infection for patients, but also affects the smoothness of collaboration within the surgical team.
[0005] The complexity and potential errors of multi-parameter coordinated adjustment. Advanced respiratory support modes often require the coordinated setting of multiple interrelated parameters (such as tidal volume, respiratory rate, inspiratory-expiratory ratio, positive end-expiratory pressure, etc.). When operating manually, medical staff need to frequently switch between different menus and perform mental calculations or estimations, which is highly complex and prone to causing parameter mismatches or unreasonable settings due to distraction or calculation errors, posing safety hazards.
[0006] Existing technological solutions have significant shortcomings. To improve convenience, some solutions have introduced wireless remote control based on common Bluetooth or Wi-Fi, but this introduces network latency, signal interference, and data security risks. Other solutions that attempt to apply general voice recognition technology lack deep adaptation to specific medical scenarios: their recognition engines struggle to effectively filter out complex background noise such as monitor alarms and electrosurgical unit noise; they cannot accurately understand medical instructions containing numerous technical terms, abbreviations, and idioms; and most importantly, they fail to build a multi-layered safety confirmation and anti-misoperation mechanism that matches the high-risk operation of medical equipment, resulting in reliability that does not meet clinical requirements.
[0007] In view of this, we propose a voice-interactive intelligent control system and method for respiratory parameters of medical devices.
[0008] Invention Patent Content The purpose of this invention is to provide a voice-interactive intelligent control system and method for respiratory parameters of medical devices, in order to solve the problems existing in the background art.
[0009] To achieve the above objectives, this invention provides the following technical solution: A voice-interactive intelligent control system for respiratory parameters in medical devices, the system comprising: The voice acquisition and enhancement module is used to acquire the operator's original voice commands through a directional microphone array, and to process the original voice signal using an adaptive beamforming algorithm and an environmental noise suppression model to enhance the target voice and suppress background noise interference in the medical environment. The operator identity and command security verification module is used to extract voiceprint features from the enhanced voice signal and match and verify them with the pre-stored authorized operator voiceprint model. At the same time, it verifies whether the voice command contains a security dynamic password that conforms to preset rules. By fusing the voiceprint matching degree and the dynamic password verification results, a comprehensive credibility score is generated. Only when the score exceeds the security threshold is the source of the command deemed legitimate and allowed to enter the subsequent processing flow. The Medical Semantic Precision Analysis and Intent Understanding module is used to convert verified voice commands into text and call the built-in medical respiratory parameter knowledge graph and natural language processing model to perform deep analysis of the text, identify the target medical device identifier, the set of respiratory parameters to be adjusted, the specific parameter values or relative adjustment amounts contained in the command, and understand the implicit relationships and collaborative adjustment intentions between parameters. The patient-state adaptive parameter safety boundary decision module is used to obtain the current operating parameters of the target medical device and the patient's real-time physiological monitoring data in real time according to the parsed parameter adjustment intention. Based on the pre-set respiratory mechanics model, pharmacokinetics model and patient-specific safety parameter library, it dynamically evaluates the safety, effectiveness and internal logical consistency of the parameter values or parameter combinations to be adjusted in the current patient state, and generates safety assessment conclusions and executable parameter adjustment schemes. The multimodal interactive safety confirmation and execution module is used to initiate a confirmation process with at least two independent feedback channels after the parameter adjustment scheme passes the safety assessment. This process includes at least repeating the complete instructions to be executed to the operator through a voice synthesis unit, and visually highlighting the parameter items to be modified and their new settings on the human-machine interface of the target medical device. Finally, the module generates the final control command only after receiving a trigger signal from an independent physical confirmation device. The control command standardization encapsulation and communication module is used to encapsulate the final control command according to the standardized communication protocol supported by the target medical device, and send it to the main controller of the target medical device through a wired or wireless medical dedicated communication link to drive it to perform corresponding parameter adjustment operations. The end-to-end audit traceability and execution feedback module is used to record the entire chain of events from the start of voice command acquisition to the completion of device execution, including but not limited to command timestamps, voiceprint feature summaries, parsed command text, security assessment process logs, physical confirmation action records, issued control command content, and device execution result feedback, generating encrypted and tamper-proof audit logs; at the same time, it provides feedback on the device execution status to the operator through voice and interface.
[0010] Preferably, in the operator identity and instruction security verification module, the calculation model for the comprehensive credibility score C is: ; where S_norm is the normalized voiceprint matching score, F is the Boolean value mapping of the dynamic password verification result, α and β are weight coefficients, and satisfy α+β= 1. This model quantifies the credibility of the instruction source by weighted fusion of voiceprint biometrics and dynamic knowledge factors.
[0011] Preferably, in the patient state adaptive parameter safety boundary decision module, the safety and effectiveness assessment of the proposed parameter combination is based on a dynamic risk index R. This index is calculated by integrating the degree of parameter deviation from the baseline, the patient's physiological state sensitivity, and the coupling effect between parameters: where ΔP represents the deviation of the parameter adjustment vector from the current baseline value, S represents the state sensitivity coefficient calculated based on the patient's real-time physiological data, and Φ represents the parameter association constraint function defined by the medical knowledge graph. This module determines the feasibility of the adjustment scheme by evaluating whether R exceeds the dynamic adjustment safety threshold θ.
[0012] Preferably, the medical semantic accurate parsing and intent understanding module has a built-in medical terminology priority weighting algorithm. When the speech recognition result has multiple meanings of a word, the system prioritizes matching the entries in the medical professional terminology database. Its priority weight W_medical is significantly higher than the weight W_general of general words, that is, W_medical>W_general, to ensure the accuracy of parsing in the medical context.
[0013] Preferably, in the multimodal interactive safety confirmation and execution module, the independent physical confirmation device is configured as an anti-accidental touch foot switch or a dedicated confirmation button. Its effective trigger signal must meet the condition that the pressure value exceeds the set threshold P_th and the duration is greater than T_min, so as to prevent false confirmation caused by unintentional contact.
[0014] Preferably, the system further includes a tiered emergency response processing submodule. This submodule monitors the acoustic characteristics and text content of voice commands in real time. When a preset set of crisis keywords is identified and voiceprint verification is quickly passed, the processing priority of the command is automatically increased to the highest level. Under this priority, the system can simplify some non-critical confirmation steps according to a preset plan and prioritize calling the set of device parameter plans corresponding to the crisis scenario for quick setting.
[0015] Preferably, the adaptive beamforming algorithm in the speech acquisition and enhancement module is based on the real-time calculation and updating of the weighting vector w(θ,t) of the microphone array, so that the main lobe of the array receiving mode is continuously aligned with the estimated sound source direction θ_s(t), while suppressing noise from other interference directions θ_i(t) as much as possible. The objective function of the weighting vector optimization aims to maximize the gain of the target direction signal and the environmental noise suppression ratio.
[0016] Preferably, the standardized encapsulation of control commands and the communication module follow specific international or industry standard protocols for medical device information exchange to ensure reliable and secure data interaction with life support devices from different manufacturers that are compatible with the standard.
[0017] Preferably, the system further includes a non-operator voice filtering submodule. This submodule analyzes the acoustic feature spectrum of the input audio signal and compares it with a pre-stored typical authorized operator voiceprint model to calculate the probability P_operator that the voice source is an authorized operator. When P_operator is lower than a preset filtering threshold λ, it is determined that the voice command is likely from a patient, family member or other unauthorized person, and the system will automatically ignore the command or only log it without triggering any control process.
[0018] A method for intelligent control of respiratory parameters in medical devices based on voice interaction, the method comprising: S1. The operator's original voice commands are acquired through a directional microphone array, and the original voice signal is processed using an adaptive beamforming algorithm and an environmental noise suppression model to enhance the target voice and suppress background noise interference in the medical environment. S2. Extract voiceprint features from the enhanced voice signal and match and verify them with the pre-stored authorized operator voiceprint model. At the same time, verify whether the voice command contains a security dynamic password that conforms to the preset rules. Generate a comprehensive credibility score by fusing the voiceprint matching degree and the dynamic password verification results. Only when the score exceeds the security threshold is the source of the command deemed legitimate and allowed to enter the subsequent processing flow. S3. Convert the verified voice commands into text, and call the built-in medical respiratory parameter knowledge graph and natural language processing model to perform deep analysis of the text, identify the target medical device identifier, the set of respiratory parameters to be adjusted, the specific parameter values or relative adjustment amounts contained in the command, and understand the implicit relationship and collaborative adjustment intention between parameters. S4. Based on the parsed parameter adjustment intention, obtain the current operating parameters of the target medical device and the real-time physiological monitoring data of the patient in real time. Based on the pre-set respiratory mechanics model, pharmacokinetics model and patient individualized safety parameter library, dynamically evaluate the safety, effectiveness and internal logical consistency of the parameter value or parameter combination to be adjusted in the current patient state, and generate safety assessment conclusions and executable parameter adjustment plans. S5. After the parameter adjustment scheme passes the safety assessment, a confirmation process with at least two independent feedback channels is initiated. This process includes at least repeating the complete instructions to be executed to the operator through a voice synthesis unit, and visually highlighting the parameter items to be modified and their new settings on the human-machine interface of the target medical device. Finally, the module generates the final control command only after receiving a trigger signal from an independent physical confirmation device. S6. The final control command is encapsulated according to the standardized communication protocol supported by the target medical device, and sent to the main controller of the target medical device through a wired or wireless medical dedicated communication link to drive it to perform the corresponding parameter adjustment operation. S7. Record the entire chain of events from the start of voice command acquisition to the completion of device execution, including but not limited to command timestamps, voiceprint feature summaries, parsed command text, security assessment process logs, physical confirmation action records, issued control command content, and device execution result feedback, generating encrypted and tamper-proof audit logs; at the same time, provide feedback on the device execution status to the operator through voice and interface.
[0019] By employing the above technical solution, this invention patent provides a voice-interactive intelligent control system and method for respiratory parameters in medical devices. It possesses at least the following beneficial effects: This invention, by employing a voice acquisition and enhancement module and a directional microphone array combined with an adaptive beamforming algorithm, enables the system to clearly and stably pick up operator voice commands in complex acoustic environments such as medical settings filled with monitor alarms, electrosurgical noise, and conversations. This module can dynamically track the direction of the sound source, intelligently enhance the target voice, and strongly suppress background noise from non-target directions. This fundamentally solves the problem of traditional voice control systems experiencing a sharp drop in recognition rate due to noise interference in medical environments, laying a solid foundation for subsequent highly reliable voice interaction. Furthermore, by constructing an operator identity and command security verification module and innovatively adopting a fusion decision model (C = α * S_norm + β * F) combining voiceprint biometrics and dynamic password dual authentication, the system builds a solid first line of defense for command access. Voiceprint recognition ensures the legitimacy of "who is speaking," while dynamic password verification increases the contextual security of "speaking in the correct context." This dual-factor authentication mechanism effectively solves the high-risk problems of weak identity verification, easy imitation, or accidental triggering in wireless remote control or simple voice control schemes, ensuring that only strictly authorized operators with clear intentions can initiate the control process. By developing a medical semantic precision parsing and intent understanding module and its backend medical knowledge graph, the system achieves a deep understanding of professional and colloquial medical instructions.The module's built-in medical terminology priority matching algorithm (W_medical >> W_general) ensures accurate parsing, correctly understanding professional terms such as "tidal volume," "PEEP," and "SIMV," as well as relative and related commands such as "increase by 20%" and "synchronously decrease." This addresses the pain point of general voice assistants failing to understand medical terminology and requiring users to use fixed, formatted commands, making human-computer interaction natural and efficient, and significantly reducing the cognitive burden and learning cost for medical staff. Furthermore, by introducing a patient-state adaptive parameter safety boundary decision module and a dynamic risk indicator based on R = f(ΔP, S, The system's assessment system (Φ) elevates safety control from passive "post-event alarms" to proactive "pre-event prevention" and "in-event intervention." This module can intelligently determine the safety and effectiveness of proposed parameter adjustments based on the patient's real-time vital signs (such as blood oxygen saturation and airway pressure). It can also automatically coordinate the linkage settings of multiple related parameters to avoid conflicts between parameters. This solves the major shortcomings of manual operation mode, which relies on personal experience for complex safety verification, resulting in low efficiency and high error rates. It provides patients with personalized intelligent safety protection based on real-time data. By designing a multimodal interactive safety confirmation and execution module and an anti-accidental touch physical confirmation device (trigger condition: pressure > P_th and time > T_min), the system establishes "auditory verification, visual focusing, and tactile confirmation." The system employs a triple safety redundancy mechanism: voice repetition allows operators to perform initial auditory verification; a brightly lit screen provides intuitive visual confirmation; and a final physical confirmation action creates a necessary point of interruption, completely eliminating the risk of erroneous execution due to accidental errors in voice recognition or operator slips of the tongue. This perfectly solves the core safety problem of simple voice control systems lacking effective secondary confirmation and having insufficient reliability. By integrating a tiered emergency response processing submodule, the system achieves an intelligent balance between routine safety procedures and emergency response efficiency. In routine operations, complete safety procedures are strictly enforced; and when the system recognizes an instruction containing crisis keywords such as "rescue" or "pure oxygen" and the identity verification is successful, it automatically enters the red emergency channel, simplifying procedures and quickly recalling preset plans.This resolves the contradiction that lengthy safety confirmation processes can delay rescue opportunities in emergency situations, enabling the system to ensure both the safety and rigor of daily operations and the rapid response requirements of emergency scenarios. By adopting standardized encapsulation of control commands and communication modules that comply with international / industry standards, it achieves standardized interoperability with different brands and models of medical equipment. Based on a universal protocol design, this system can serve as an open intelligent control platform, widely adaptable to various standard-compliant anesthesia machines and ventilators in operating rooms and ICUs. This addresses the industry pain points of poor compatibility and high integration costs caused by proprietary protocols, enhancing the universality and promotional value of the solution. By deploying a non-operator voice filtering submodule and a probability-based P_operator-based decision mechanism, the system significantly enhances its applicability and safety in special scenarios. In pediatric wards or at the bedside where patients may make sounds, this module can effectively distinguish the voice characteristics of medical staff from those of patients / family members, automatically filtering out commands from unauthorized sources that may contain control words, preventing accidental triggering. In summary, this invention integrates key technologies such as highly robust voice processing, multiple security authentication, medical artificial intelligence, multimodal interaction, and standardized device communication to construct a medical device intelligent control system with high security level, fast response speed, and natural interactive experience. It realizes convenient operation of "non-contact, hands-free" operation, greatly improves work efficiency and reduces infection risk. More importantly, through multi-level, closed-loop security design from front end to back end, it builds a solid reliability barrier for high-risk medical operations. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of the invention, form part of this application: Figure 1 This is a schematic diagram of the overall structure of this invention. Detailed Implementation
[0021] The technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, those skilled in the art, without making innovative embodiments, are all within the scope of protection of this invention.
[0022] Please see Figure 1 The present invention discloses a voice-interactive-based intelligent control system for respiratory parameters in medical devices, comprising: The voice acquisition and enhancement module 101 is used to acquire the operator's original voice commands through a directional microphone array, and to process the original voice signal using an adaptive beamforming algorithm and an environmental noise suppression model to enhance the target voice and suppress background noise interference in the medical environment. It should be noted that the microphone array deployed in this module typically employs a linear or ring structure, with its spatial directivity pre-configured or aligned via a self-learning algorithm to target the area where medical personnel routinely operate in front of the device. The adaptive beamforming algorithm can estimate and track the direction of the sound source in real time, dynamically adjusting the weighting coefficients of each microphone unit to form a spatial filter. Its core objective is to maximize the gain of the speech signal from the operator's direction while simultaneously aligning the beam's nulls with persistent, strong noise sources (such as monitor alarms or high-frequency electrosurgical units). This significantly improves the signal-to-noise ratio of the input speech in complex, multi-noise operating room or ICU environments, providing a high-quality audio front-end for subsequent speech recognition.
[0023] The operator identity and instruction security verification module 102 is used to extract voiceprint features from the enhanced voice signal and match and verify them with the pre-stored authorized operator voiceprint model. At the same time, it verifies whether the voice instruction contains a security dynamic password that conforms to preset rules. By fusing the voiceprint matching degree and the dynamic password verification results, a comprehensive credibility score is generated. Only when the score exceeds the security threshold is the instruction source deemed legitimate and allowed to enter the subsequent processing flow. It should be noted that this module integrates biometric recognition and knowledge factor authentication. Voiceprint feature extraction typically employs algorithms such as Mel-frequency cepstral coefficients and linear predictive coding, while modeling can utilize Gaussian mixture models. The dynamic password design is flexible and scenario-dependent; for example, it could be the last few digits of the surgery number for the day, the patient's bed number, or a dynamic token that changes over time. This two-factor authentication mechanism—"Who are you?" (voiceprint) + "What do you know?" (dynamic password)—signally enhances system security, effectively preventing recording playback attacks, voiceprint imitation attacks, and accidental triggering of commands by unauthorized personnel, establishing a reliable identity access barrier for the entire high-risk medical device control process.
[0024] The medical semantic precision parsing and intent understanding module 103 is used to convert verified voice commands into text and call the built-in medical respiratory parameter knowledge graph and natural language processing model to perform deep parsing of the text, identify the target medical device identifier, the set of respiratory parameters to be adjusted, the specific parameter values or relative adjustment amounts contained in the command, and understand the implicit relationship and collaborative adjustment intent between parameters. It's important to note that the speech recognition engine in this module has been trained on a large amount of medical scenario speech data and loaded with a specialized medical pronunciation dictionary. The subsequent natural language processing (NLP) process is crucial; it goes beyond simply converting speech to text. It also needs to understand common informal expressions, technical abbreviations (such as "PEEP," "FiO2," and "SIMV"), and contextual omissions in clinical instructions. For example, when the instruction is "Increase tidal volume to 500 ml, simultaneously decrease respiratory rate to 12," the parsing module needs to accurately extract the parameters "tidal volume = 500 ml" and "respiratory rate = 12 breaths / min," and understand the implied meaning of "simultaneously" in the word "simultaneously." The knowledge graph stores the logical relationships between parameters, providing a rule-based foundation for subsequent safety decisions.
[0025] The patient-state adaptive parameter safety boundary decision module 104 is used to obtain the current operating parameters of the target medical device and the real-time physiological monitoring data of the patient in real time according to the parsed parameter adjustment intention. Based on the pre-set respiratory mechanics model, pharmacokinetics model and patient-specific safety parameter library, it dynamically evaluates the safety, effectiveness and internal logical consistency of the parameter value or parameter combination to be adjusted in the current patient state, and generates safety assessment conclusions and executable parameter adjustment schemes. It's important to note that this module is the core of the system's intelligent and preventative safety control. It doesn't simply compare thresholds; instead, it performs dynamic calculations based on the patient's real-time condition. For example, when assessing whether to increase positive end-expiratory pressure (PEEP), the module combines real-time monitoring of airway plateau pressure, blood pressure, and estimated lung compliance, using a respiratory mechanics model to predict changes in adjusted plateau pressure and determine whether it will cause barotrauma or affect circulatory stability. For linked adjustments (such as simultaneously changing tidal volume and respiratory rate), it calculates derived indicators such as adjusted minute ventilation to ensure they remain within safe ranges. This process transforms traditional judgments relying on healthcare professionals' experience into objective and rapid safety checks based on data and models.
[0026] The multimodal interactive safety confirmation and execution module 105 is used to initiate a confirmation process containing at least two independent feedback channels after the parameter adjustment scheme passes the safety assessment. This process includes at least repeating the complete instructions to be executed to the operator through a voice synthesis unit, and visually highlighting the parameter items to be modified and their new settings on the human-machine interface of the target medical device. Finally, the module generates the final control command only after receiving a trigger signal from an independent physical confirmation device. It's important to note that the multi-level confirmation process designed for this module aims to build redundant safety barriers and minimize the risk of misoperation. Voice repetition provides an "auditory verification" opportunity, allowing operators to immediately identify voice recognition or parsing errors. Visual highlighting on the interface (such as flashing or color changing) provides a "visual verification" focus, ensuring the operator notices the correct parameters. The final "physical confirmation" (such as pressing the foot switch) is a crucial, conscious point of interruption, requiring explicit final authorization from the operator. This effectively prevents misoperation due to verbal misanswers or distraction. This multimodal confirmation mechanism combining "hearing, seeing, and moving" is a key design feature for high-risk medical procedures.
[0027] The control command standardization encapsulation and communication module 106 is used to encapsulate the final control command according to the standardized communication protocol supported by the target medical device, and send it to the main controller of the target medical device through a wired or wireless medical dedicated communication link to drive it to perform corresponding parameter adjustment operations. It should be noted that, in order to achieve interoperability with medical devices (anesthesia machines, ventilators) from different manufacturers and models, this module follows specific international or industry standard protocols for medical device information exchange, such as the IEEE 11073 Service-Oriented Device Connectivity Standard or HL7 FHIR. The module encapsulates the structured control commands generated internally by the system (e.g., {“device”:“ventilator_01”, “command”: “set”, “parameter”: “FiO2”, “value”: 60}) into specific data packet formats according to standard protocol specifications, and transmits them via serial port, Ethernet, or wireless channels that meet medical electromagnetic compatibility requirements. This standardized design improves the system's compatibility and integrability, avoiding the closed nature problems caused by proprietary protocols.
[0028] The full-process audit traceability and execution feedback module 107 is used to record the entire chain of events from the start of voice command acquisition to the completion of device execution, including but not limited to command timestamps, voiceprint feature summaries, parsed command text, security assessment process logs, physical confirmation action records, issued control command content, and device execution result feedback, generating encrypted and tamper-proof audit logs; at the same time, it provides feedback on the device execution status to the operator through voice and interface. It's important to note that this module is crucial for medical quality control and post-event traceability. All log entries are stored in a timestamped chain structure, and technologies such as hash algorithms ensure their integrity and immutability. The audit logs meticulously record "who, when, what instructions were given, how the system understood and verified them, how they were ultimately executed, and the results," forming a complete closed-loop evidence chain. This not only meets stringent medical compliance and review requirements but also provides comprehensive data support for analyzing operational habits, optimizing system performance, or investigating anomalies. Simultaneously, real-time execution feedback (such as "oxygen concentration has been set to 60%)" closes the human-machine interaction loop, enhancing the operator's sense of control and system transparency.
[0029] In the operator identity and command security verification module, the calculation model for the comprehensive credibility score C is: ; where S_norm is the normalized voiceprint matching score, F is the Boolean value mapping of the dynamic password verification result, α and β are weight coefficients, and satisfy α+β= 1. This model quantifies the credibility of the command source by weighted fusion of voiceprint biometrics and dynamic knowledge factors. It is worth noting that this model demonstrates the flexibility and configurability of security strategies, with weighting coefficients α and β adjustable according to the security level requirements of different application scenarios. For example, in routine wards, the weight of the dynamic password factor β can be appropriately increased; while in environments with relatively fixed personnel, such as operating rooms, the voiceprint factor α can be emphasized more. Normalization of S_norm ensures the comparability of scores between different voiceprint models. This quantitative scoring model provides a clear and objective decision-making basis for determining "legitimate commands," rather than a simple binary pass / fail, allowing system administrators to finely adjust the balance between security and convenience.
[0030] In the patient-state adaptive parameter safety boundary decision module, the safety and effectiveness assessment of the proposed parameter combination is based on a dynamic risk index R. This index is calculated by integrating the degree of parameter deviation from the baseline, the patient's physiological state sensitivity, and the coupling effect between parameters: where ΔP represents the deviation of the parameter adjustment vector from the current baseline value, S represents the state sensitivity coefficient calculated based on real-time patient physiological data, and Φ represents the parameter association constraint function defined by the medical knowledge graph. This module determines the feasibility of the adjustment scheme by evaluating whether R exceeds the dynamic adjustment safety threshold θ. It is worth noting that ΔP (parameter deviation) reflects the magnitude of the adjustment; S (state sensitivity coefficient) quantifies the patient's current physiological state's vulnerability to changes in specific parameters, such as hypotensive patients being more sensitive to changes in PEEP; and Φ (association constraint function) encodes complex rules in the medical knowledge graph, such as the relationship between tidal volume and plateau pressure. By integrating these three into a comprehensive risk indicator R and comparing it with a dynamic threshold θ, the system can make more personalized safety judgments that align with the thinking of clinical experts, rather than mechanically applying fixed rules.
[0031] The medical semantic accurate parsing and intent understanding module has a built-in medical terminology priority weighting algorithm. When the speech recognition result has multiple meanings of a word, the system prioritizes matching the entries in the medical professional terminology database. Its priority weight W_medical is significantly higher than the weight W_general of general words, that is, W_medical>W_general, to ensure the accuracy of parsing in the medical context. It is worth noting that the medical terminology priority weighting algorithm (W_medical >> W_general) built into the medical semantic accurate parsing and intent understanding module is a key strategy for resolving speech recognition ambiguity in medical scenarios. For example, the syllable "feilv" might be preferentially identified as "fee rate" in a general context, but with the support of the system's medical knowledge graph, it will be preferentially matched as the respiratory parameter "frequency" due to its extremely high W_medical weight. This domain knowledge-driven disambiguation mechanism significantly improves the accuracy of professional terminology recognition, enabling the system to truly "understand" the professional conversations of medical staff.
[0032] In the multimodal interactive safety confirmation and execution module, the independent physical confirmation device is configured as an anti-accidental touch foot switch or a dedicated confirmation button. Its effective trigger signal must meet the condition that the pressure value exceeds the set threshold P_th and the duration is greater than T_min, so as to prevent false confirmation caused by unintentional contact. It is worth noting that in the multimodal interactive safety confirmation and execution module, the design to prevent accidental touches of independent physical confirmation devices (such as foot switches) (which must meet the requirements of pressure > P_th and time > T_min) is crucial. The P_th threshold excludes minor touches such as accidental drops of objects; the T_min duration requirement prevents accidental and brief contact by the operator's foot. This "intentional" trigger condition design is the last hardware guarantee to ensure that the physical confirmation action represents the operator's clear intention, eliminating hasty or unintentional misconfirmations from the interaction logic.
[0033] The system also includes a tiered emergency response processing submodule, which monitors the acoustic features and text content of voice commands in real time. When a preset set of crisis keywords is identified and the voiceprint verification is passed quickly, the processing priority of the command is automatically raised to the highest level. Under this priority, the system can simplify some non-critical confirmation steps according to a preset plan and prioritize calling the set of device parameter plans corresponding to the crisis scenario for quick setting. It is worth noting that the system includes a tiered emergency response submodule, demonstrating the intelligent resilience of the security mechanism in emergency situations. Upon recognizing critical keywords such as "rescue" and "cardiac arrest" and quickly completing core identity verification, the system automatically enters the "emergency channel." This may bypass some non-critical security calculations or simplify confirmation steps to gain valuable rescue time. However, it still performs core security checks (such as avoiding setting obviously fatal parameters) and mandates the simplest physical confirmation (such as quickly pressing the confirmation button), ensuring that basic security controls are not lost while speeding up the process, achieving a balance between prioritizing safety in normal times and efficiency in emergencies.
[0034] The adaptive beamforming algorithm in the speech acquisition and enhancement module is based on the real-time calculation and updating of the weighting vector w(θ,t) of the microphone array, so that the main lobe of the array receiving mode is continuously aligned with the estimated sound source direction θ_s(t), while suppressing noise from other interference directions θ_i(t) as much as possible. The objective function of the weighting vector optimization aims to maximize the gain of the target direction signal and the environmental noise suppression ratio. It is worth noting that the adaptive beamforming algorithm in the speech acquisition and enhancement module involves a continuous real-time optimization of its weight vector w(θ, t). The algorithm not only tracks the sound source direction θ_s(t) but also dynamically identifies and suppresses changing or moving interference source directions θ_i(t). Its optimization objective function, while maximizing the signal-to-noise ratio, typically also considers constraints on speech distortion to preserve speech features for subsequent speaker recognition and semantic analysis. This enables the front-end processing module to proactively adapt to complex and ever-changing medical acoustic environments.
[0035] The standardized encapsulation of control commands and the communication module follow specific international or industry standard protocols for medical device information exchange to ensure reliable and secure data interaction with life support devices from different manufacturers that are compatible with the standard. It is worth noting that the standardized encapsulation of control commands and the specific international / industry standard protocols (such as IEEE 11073 SDC) followed by the communication module are the foundation for the system's open interconnection and future expansion. Adopting such standards means that this system can be seamlessly integrated into broader medical IoT or smart hospital platforms, interacting with central monitoring stations, electronic medical record systems, etc., providing possible data interfaces and control channels for advanced applications such as perioperative big data analysis and AI-assisted decision-making.
[0036] The system also includes a non-operator voice filtering submodule. This submodule analyzes the acoustic feature spectrum of the input audio signal and compares it with a pre-stored typical authorized operator voiceprint model to calculate the probability P_operator that the voice source is an authorized operator. When P_operator is lower than a preset filtering threshold λ, it is determined that the voice command is likely to come from a patient, family member or other unauthorized person. The system will automatically ignore the command or only log it without triggering any control process. It is worth noting that the system includes a non-operator voice filtering submodule, which employs a statistically based intelligent filtering mechanism by calculating the probability P_operator and comparing it with a threshold λ. This module effectively distinguishes the voice characteristics of adult medical staff from those of children, some elderly patients, or emotionally agitated family members by analyzing acoustic features such as pitch, timbre, and formants. This is particularly important for deploying the system at the bedside of pediatric, geriatric, or confused patients, effectively preventing accidental device triggering due to the patient's groans, crying, or delirium, greatly expanding the system's applicable scenarios and overall safety.
[0037] A method for intelligent control of respiratory parameters in medical devices based on voice interaction, the method comprising: S1. The operator's original voice commands are acquired through a directional microphone array, and the original voice signal is processed using an adaptive beamforming algorithm and an environmental noise suppression model to enhance the target voice and suppress background noise interference in the medical environment. S2. Extract voiceprint features from the enhanced voice signal and match and verify them with the pre-stored authorized operator voiceprint model. At the same time, verify whether the voice command contains a security dynamic password that conforms to the preset rules. Generate a comprehensive credibility score by fusing the voiceprint matching degree and the dynamic password verification results. Only when the score exceeds the security threshold is the source of the command deemed legitimate and allowed to enter the subsequent processing flow. S3. Convert the verified voice commands into text, and call the built-in medical respiratory parameter knowledge graph and natural language processing model to perform deep analysis of the text, identify the target medical device identifier, the set of respiratory parameters to be adjusted, the specific parameter values or relative adjustment amounts contained in the command, and understand the implicit relationship and collaborative adjustment intention between parameters. S4. Based on the parsed parameter adjustment intention, obtain the current operating parameters of the target medical device and the real-time physiological monitoring data of the patient in real time. Based on the pre-set respiratory mechanics model, pharmacokinetics model and patient individualized safety parameter library, dynamically evaluate the safety, effectiveness and internal logical consistency of the parameter value or parameter combination to be adjusted in the current patient state, and generate safety assessment conclusions and executable parameter adjustment plans. S5. After the parameter adjustment scheme passes the safety assessment, a confirmation process with at least two independent feedback channels is initiated. This process includes at least repeating the complete instructions to be executed to the operator through a voice synthesis unit, and visually highlighting the parameter items to be modified and their new settings on the human-machine interface of the target medical device. Finally, the module generates the final control command only after receiving a trigger signal from an independent physical confirmation device. S6. The final control command is encapsulated according to the standardized communication protocol supported by the target medical device, and sent to the main controller of the target medical device through a wired or wireless medical dedicated communication link to drive it to perform the corresponding parameter adjustment operation. S7. Record the entire chain of events from the start of voice command acquisition to the completion of device execution, including but not limited to command timestamps, voiceprint feature summaries, parsed command text, security assessment process logs, physical confirmation action records, issued control command content, and device execution result feedback, generating encrypted and tamper-proof audit logs; at the same time, provide feedback on the device execution status to the operator through voice and interface.
[0038] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0039] Although embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the present invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A voice-interactive-based intelligent control system for respiratory parameters in medical devices, characterized in that, The system includes: The voice acquisition and enhancement module is used to acquire the operator's original voice commands through a directional microphone array, and to process the original voice signal using an adaptive beamforming algorithm and an environmental noise suppression model to enhance the target voice and suppress background noise interference in the medical environment. The operator identity and command security verification module is used to extract voiceprint features from the enhanced voice signal and match and verify them with the pre-stored authorized operator voiceprint model. At the same time, it verifies whether the voice command contains a security dynamic password that conforms to preset rules. By fusing the voiceprint matching degree and the dynamic password verification results, a comprehensive credibility score is generated. Only when the score exceeds the security threshold is the source of the command deemed legitimate and allowed to enter the subsequent processing flow. The Medical Semantic Precision Analysis and Intent Understanding module is used to convert verified voice commands into text and call the built-in medical respiratory parameter knowledge graph and natural language processing model to perform deep analysis of the text, identify the target medical device identifier, the set of respiratory parameters to be adjusted, the specific parameter values or relative adjustment amounts contained in the command, and understand the implicit relationships and collaborative adjustment intentions between parameters. The patient-state adaptive parameter safety boundary decision module is used to obtain the current operating parameters of the target medical device and the patient's real-time physiological monitoring data in real time according to the parsed parameter adjustment intention. Based on the pre-set respiratory mechanics model, pharmacokinetics model and patient-specific safety parameter library, it dynamically evaluates the safety, effectiveness and internal logical consistency of the parameter values or parameter combinations to be adjusted in the current patient state, and generates safety assessment conclusions and executable parameter adjustment schemes. The multimodal interactive safety confirmation and execution module is used to initiate a confirmation process with at least two independent feedback channels after the parameter adjustment scheme passes the safety assessment. This process includes at least repeating the complete instructions to be executed to the operator through a voice synthesis unit, and visually highlighting the parameter items to be modified and their new settings on the human-machine interface of the target medical device. Finally, the module generates the final control command only after receiving a trigger signal from an independent physical confirmation device. The control command standardization encapsulation and communication module is used to encapsulate the final control command according to the standardized communication protocol supported by the target medical device, and send it to the main controller of the target medical device through a wired or wireless medical dedicated communication link to drive it to perform corresponding parameter adjustment operations. The end-to-end audit traceability and execution feedback module is used to record the entire chain of events from the start of voice command acquisition to the completion of device execution, including but not limited to command timestamps, voiceprint feature summaries, parsed command text, security assessment process logs, physical confirmation action records, issued control command content, and device execution result feedback, generating encrypted and tamper-proof audit logs; at the same time, it provides feedback on the device execution status to the operator through voice and interface.
2. The intelligent control system for respiratory parameters of medical devices based on voice interaction according to claim 1, characterized in that, In the operator identity and command security verification module, the calculation model for the comprehensive credibility score C is: ; where S_norm is the normalized voiceprint matching score, F is the Boolean value mapping of the dynamic password verification result, α and β are weight coefficients, and satisfy α+β= 1. This model quantifies the credibility of the command source by weighted fusion of voiceprint biometrics and dynamic knowledge factors.
3. The intelligent control system for respiratory parameters of medical devices based on voice interaction according to claim 1, characterized in that, In the patient-state adaptive parameter safety boundary decision-making module, the safety and effectiveness assessment of the proposed parameter combination is based on a dynamic risk index R. This index is calculated by integrating the degree of parameter deviation from the baseline, the patient's physiological state sensitivity, and the coupling effect between parameters: where ΔP represents the deviation of the parameter adjustment vector from the current baseline value, S represents the state sensitivity coefficient calculated based on the patient's real-time physiological data, and Φ represents the parameter association constraint function defined by the medical knowledge graph. This module determines the feasibility of the adjustment scheme by evaluating whether R exceeds the dynamic adjustment safety threshold θ.
4. The intelligent control system for respiratory parameters of medical devices based on voice interaction according to claim 1, characterized in that, The medical semantic accurate parsing and intent understanding module incorporates a medical terminology priority weighting algorithm. When the speech recognition result has multiple meanings, the system prioritizes matching entries in the medical professional terminology database. Its priority weight W_medical is significantly higher than the weight W_general of general terms, i.e., W_medical>W_general, to ensure the accuracy of parsing in the medical context.
5. The intelligent control system for respiratory parameters of medical devices based on voice interaction according to claim 1, characterized in that, In the multimodal interactive safety confirmation and execution module, the independent physical confirmation device is configured as an anti-accidental touch foot switch or a dedicated confirmation button. Its effective trigger signal must meet the condition that the pressure value exceeds the set threshold P_th and the duration is greater than T_min, so as to prevent false confirmation caused by unintentional contact.
6. The intelligent control system for respiratory parameters of medical devices based on voice interaction according to claim 1, characterized in that, The system also includes a tiered emergency response processing submodule, which monitors the acoustic features and text content of voice commands in real time. When a preset set of crisis keywords is identified and voiceprint verification is passed quickly, the processing priority of the command is automatically raised to the highest level. Under this priority, the system can simplify some non-critical confirmation steps according to a preset plan and prioritize calling the set of device parameter plans corresponding to the crisis scenario for quick setting.
7. The intelligent control system for respiratory parameters of medical devices based on voice interaction according to claim 1, characterized in that, The adaptive beamforming algorithm in the speech acquisition and enhancement module is based on the real-time calculation and updating of the weighting vector w(θ,t) of the microphone array, so that the main lobe of the array receiving mode is continuously aligned with the estimated sound source direction θ_s(t), while suppressing noise from other interference directions θ_i(t) as much as possible. The objective function of the weighting vector optimization aims to maximize the gain of the target direction signal and the environmental noise suppression ratio.
8. The intelligent control system for respiratory parameters of medical devices based on voice interaction according to claim 1, characterized in that, The standardized encapsulation of control commands and the communication module follow specific international or industry standard protocols for medical device information exchange to ensure reliable and secure data interaction with life support devices from different manufacturers that are compatible with the standard.
9. The intelligent control system for respiratory parameters of medical devices based on voice interaction according to claim 1, characterized in that, The system also includes a non-operator voice filtering submodule. This submodule analyzes the acoustic feature spectrum of the input audio signal and compares it with a pre-stored typical authorized operator voiceprint model to calculate the probability P_operator that the voice source is an authorized operator. When P_operator is lower than a preset filtering threshold λ, it is determined that the voice command is likely from a patient, family member or other unauthorized person, and the system will automatically ignore the command or only log it without triggering any control process.
10. A method for intelligent control of respiratory parameters in medical devices based on voice interaction, characterized in that, The method includes: S1. The operator's original voice commands are acquired through a directional microphone array, and the original voice signal is processed using an adaptive beamforming algorithm and an environmental noise suppression model to enhance the target voice and suppress background noise interference in the medical environment. S2. Extract voiceprint features from the enhanced voice signal and match and verify them with the pre-stored authorized operator voiceprint model. At the same time, verify whether the voice command contains a security dynamic password that conforms to the preset rules. Generate a comprehensive credibility score by fusing the voiceprint matching degree and the dynamic password verification results. Only when the score exceeds the security threshold is the source of the command deemed legitimate and allowed to enter the subsequent processing flow. S3. Convert the verified voice commands into text, and call the built-in medical respiratory parameter knowledge graph and natural language processing model to perform deep analysis of the text, identify the target medical device identifier, the set of respiratory parameters to be adjusted, the specific parameter values or relative adjustment amounts contained in the command, and understand the implicit relationship and collaborative adjustment intention between parameters. S4. Based on the parsed parameter adjustment intention, obtain the current operating parameters of the target medical device and the real-time physiological monitoring data of the patient in real time. Based on the pre-set respiratory mechanics model, pharmacokinetics model and patient individualized safety parameter library, dynamically evaluate the safety, effectiveness and internal logical consistency of the parameter value or parameter combination to be adjusted in the current patient state, and generate safety assessment conclusions and executable parameter adjustment plans. S5. After the parameter adjustment scheme passes the safety assessment, a confirmation process with at least two independent feedback channels is initiated. This process includes at least repeating the complete instructions to be executed to the operator through a voice synthesis unit, and visually highlighting the parameter items to be modified and their new settings on the human-machine interface of the target medical device. Finally, the module generates the final control command only after receiving a trigger signal from an independent physical confirmation device. S6. The final control command is encapsulated according to the standardized communication protocol supported by the target medical device, and sent to the main controller of the target medical device through a wired or wireless medical dedicated communication link to drive it to perform the corresponding parameter adjustment operation. S7. Record the entire chain of events from the start of voice command acquisition to the completion of device execution, including but not limited to command timestamps, voiceprint feature summaries, parsed command text, security assessment process logs, physical confirmation action records, issued control command content, and device execution result feedback, generating encrypted and tamper-proof audit logs; at the same time, provide feedback on the device execution status to the operator through voice and interface.