A somatic intelligent platform based on a full-modal large model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-11
AI Technical Summary
其中,单一语音控制轮椅仅能识别固定标准语音指令,极易受环境噪声干扰,完全无法适配老年用户、脑卒中康复患者、口齿不清、口音偏重、音量微弱的特殊群体,且无任何环境感知联动能力,仅能被动执行运动指令,无法预判行驶风险,存在极大安全隐患;单一视觉避障轮椅依托基础视觉算法完成近距离障碍物识别,但仅能实现简单避障动作,无自主交互能力,且环境感知维度单一,无法适配复杂路况场景;少量多模态交互轮椅虽同时集成了语音交互与视觉感知功能,但本质属于硬件与功能的简单叠加,未实现多模态信息的深度融合推理,无专属的医疗级安全约束机制,模型决策可靠性差,完全无法适配医疗场景的高安全标准
[0016] The embodied intelligence platform based on a multimodal large model of the present invention achieves joint perception of user intent information, environmental visual information, and wheelchair operating status information through the voice intent perception submodule, visual perception submodule, and wheelchair posture perception submodule in the multimodal perception module. This multimodal perception information is then input into the multimodal large model, and a fusion reasoning module generates decision results and calculates confidence levels. A safety decision module classifies safety levels according to confidence levels and outputs corresponding control commands for the intelligent wheelchair. Because of the use of these modules, the embodied intelligence platform of the present invention achieves an integrated human-machine-environment proactive perception and decision-making closed loop. It can avoid the risk of collisions with conventional obstacles and predict hidden risks of loss of control such as slopes, tilts, and bumps. Simultaneously, by calling the multimodal large model to calculate and classify confidence levels, it can eliminate dangerous operations caused by model misjudgments, thereby comprehensively ensuring the travel safety of medical users with weak self-rescue capabilities. This completely solves the problems of numerous safety hazards and weak protection capabilities in existing technologies, and ultimately meets the high safety operation standards of medical equipment.
Smart Images

Figure CN122546744A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of embodied intelligence and robot control technology, specifically to an embodied intelligence platform based on a full-modal large model. Background Technology
[0002] In recent years, the rapid iterative breakthroughs in artificial intelligence technology, especially multimodal large-scale model technology, have enabled smart devices to possess comprehensive cognitive abilities that simultaneously understand voice, images, and environmental information, providing core algorithmic support for the deep intelligent upgrade of smart wheelchairs. Simultaneously, the mature implementation of edge computing technology allows high-performance AI models to be lightweightly deployed locally on terminal devices, completely eliminating reliance on cloud servers and effectively addressing industry pain points such as high latency in cloud inference, strong network dependence, and significant data privacy risks. Furthermore, the rise of embodied intelligence technology has fundamentally changed the traditional operating logic of smart devices—"passively executing instructions"—driving them towards intelligent upgrades that "actively perceive the environment, autonomously predict risks, and adaptively adjust their state," providing a new technological approach for the embodied, safe, and personalized upgrades of medical wheelchairs.
[0003] Currently, most publicly available smart wheelchair technologies focus on optimizing single functions, mainly falling into three categories: single voice-controlled wheelchairs, single visual obstacle avoidance wheelchairs, and simple multimodal wheelchairs. Single voice-controlled wheelchairs can only recognize fixed standard voice commands, are highly susceptible to environmental noise interference, and are completely unsuitable for elderly users, stroke rehabilitation patients, and special groups with unclear speech, heavy accents, or weak voices. They also lack any environmental perception and linkage capabilities, passively executing movement commands without anticipating driving risks, posing significant safety hazards. Single visual obstacle avoidance wheelchairs rely on basic visual algorithms to identify nearby obstacles, but can only perform simple obstacle avoidance actions, lacking autonomous interaction capabilities. Their environmental perception is limited, making them unsuitable for complex road conditions. While a few multimodal interactive wheelchairs integrate both voice interaction and visual perception functions, they are essentially a simple combination of hardware and functions, failing to achieve deep fusion and reasoning of multimodal information, lacking dedicated medical-grade safety constraints, and exhibiting poor model decision reliability, making them completely unsuitable for the high safety standards of medical scenarios.
[0004] Existing technologies generally suffer from low technical barriers, limited functionality, and insufficient adaptability. Most solutions only achieve basic functions such as speech recognition, basic visual obstacle avoidance, and simple path planning, failing to form a complete embodied intelligent decision-making loop. Existing technologies generally exhibit four core shortcomings: First, the fusion of multimodal information is shallow, merely at the data surface level, unable to achieve a unified understanding of user intent, environmental status, and device status; second, there is a lack of security reasoning mechanisms specific to medical scenarios, resulting in AI model output decisions without risk verification or fault tolerance, making model misjudgments highly susceptible to security incidents; third, there is a high dependence on cloud computing power, leading to poor real-time control stability and serious privacy risks related to medical users' voice, travel, and environmental data; fourth, the models are fixed pre-trained models, lacking autonomous iterative optimization capabilities, unable to adapt to the personalized needs of different users, resulting in poor long-term user experience and low adaptability.
[0005] In summary, existing intelligent wheelchair technologies suffer from drawbacks such as poor interaction tolerance, incomplete perception, lack of security safeguards, lack of privacy protection, weak adaptability, and low level of intelligence. They lack true embodied intelligent perception and decision-making capabilities and cannot simultaneously meet the high security, high adaptability, high privacy, and high intelligence requirements of medical-grade devices. Therefore, there is an urgent need for a new, targeted, full-modal embodied intelligent wheelchair control solution that addresses the pain points of medical scenarios. Summary of the Invention
[0006] This invention is made to solve the above-mentioned problems, and aims to provide an embodied intelligence platform based on a full-modal large model.
[0007] This invention provides an embodied intelligence platform based on a multimodal large model for controlling the motion posture and operational status of an intelligent wheelchair during use. It features the following components: a multimodal perception module for acquiring multimodal perception information including user intent information, intelligent wheelchair operational status information, and environmental visual information of the surrounding environment; a fusion reasoning module containing a multimodal large model, which, based on the multimodal perception information, infers and generates decision results using the multimodal large model and calculates the confidence level of the decision results; a safety decision module, which, based on the numerical range of the confidence level, classifies the safety level of the intelligent wheelchair's operation using the multimodal large model and outputs corresponding control commands for the intelligent wheelchair according to the safety level; and an execution control module, which drives the intelligent wheelchair to perform real-time state adjustments according to the control commands and feeds back the real-time state adjustment information of the intelligent wheelchair to the multimodal large model. The multimodal perception module includes: a voice intent perception submodule, which recognizes the user's voice to obtain user intent information; a visual perception submodule, which collects environmental visual information; and a wheelchair posture perception submodule, which acquires the motion posture and operational status of the intelligent wheelchair as operational status information.
[0008] The embodied intelligence platform based on a multimodal large model provided by this invention may also have the following features: the multimodal large model extracts multiple operation-related features from multimodal perception information, performs cross-modal fusion processing on the multiple operation-related features, infers and generates decision results, and calculates the confidence level of the decision results. The operation-related features include semantic intent features, visual environment features, and ontological posture features.
[0009] The embodied intelligence platform based on a full-modal large model provided by this invention may also have the following features: In the safety decision module, the threshold of the confidence level numerical range is preset in the full-modal large model. The numerical range is divided into a high confidence range, a medium confidence range, and a low confidence range. The corresponding safety levels of the intelligent wheelchair are Level 1 normal operation, Level 2 degraded operation, and Level 3 safety lock. When the safety level is Level 1 normal operation, the control command output by the safety decision module is: grant full permission for real-time status adjustment of the intelligent wheelchair; when the safety level is Level 2 degraded operation, the control command is: limit the driving speed and steering range of the intelligent wheelchair; when the safety level is Level 3 safety lock, the control command is: retain only the safety protection functions of the intelligent wheelchair, including obstacle avoidance and emergency braking, control the intelligent wheelchair to run at the minimum safe speed, and simultaneously output risk warnings.
[0010] The embodied intelligence platform based on a full-modal large model provided by this invention may also have the following feature: the execution control module includes a manual emergency submodule, which is used to switch the operation state of the intelligent wheelchair to manual control in extreme abnormal scenarios.
[0011] The embodied intelligence platform based on a multimodal large model provided by this invention may also include the following features: an offline optimization module, used to encrypt and store multimodal data during the use of the intelligent wheelchair, and to iteratively optimize the multimodal large model based on the multimodal data when the intelligent wheelchair is idle, so as to adapt to the user's accent characteristics, operating habits and common scenarios. The multimodal data includes information, decision results, control commands of the intelligent wheelchair and real-time status adjustment information.
[0012] The embodied intelligence platform based on a full-modal large model provided by this invention may also have the following features: the operation process of the offline optimization module includes: multimodal data acquisition, local encrypted storage, idle state monitoring, small sample screening, model fine-tuning and updating, and model adaptation optimization.
[0013] The embodied intelligence platform based on a full-modal large model provided by this invention may also have the following features: the visual perception submodule includes a camera and a depth camera, which are used to collect color images and depth information, respectively, thereby realizing three-dimensional environmental perception, obstacle recognition and spatial distance judgment.
[0014] The embodied intelligence platform based on a full-modal large model provided by this invention may also have the following features: the voice intent perception submodule has a built-in fuzzy voice fault-tolerant interaction logic algorithm, which is used to identify fuzzy commands, non-standard pronunciations, and low-speed speech of special medical users, including elderly users, stroke users, and postoperative rehabilitation users.
[0015] Compared with the prior art, the functions and effects of the present invention include:
[0016] The embodied intelligence platform based on a multimodal large model of the present invention achieves joint perception of user intent information, environmental visual information, and wheelchair operating status information through the voice intent perception submodule, visual perception submodule, and wheelchair posture perception submodule in the multimodal perception module. This multimodal perception information is then input into the multimodal large model, and a fusion reasoning module generates decision results and calculates confidence levels. A safety decision module classifies safety levels according to confidence levels and outputs corresponding control commands for the intelligent wheelchair. Because of the use of these modules, the embodied intelligence platform of the present invention achieves an integrated human-machine-environment proactive perception and decision-making closed loop. It can avoid the risk of collisions with conventional obstacles and predict hidden risks of loss of control such as slopes, tilts, and bumps. Simultaneously, by calling the multimodal large model to calculate and classify confidence levels, it can eliminate dangerous operations caused by model misjudgments, thereby comprehensively ensuring the travel safety of medical users with weak self-rescue capabilities. This completely solves the problems of numerous safety hazards and weak protection capabilities in existing technologies, and ultimately meets the high safety operation standards of medical equipment.
[0017] In the embodied intelligence platform based on a full-modal large model of this invention, the offline optimization module and the full-modal large model perform local calculations entirely offline, without relying on the cloud network. This eliminates control lag and failure issues caused by network fluctuations and outages, ensuring stable and reliable operation. It can achieve millisecond-level real-time decision response, adapting to emergency protection needs such as sudden obstacles and unexpected road conditions. Furthermore, it ensures that all user data is encrypted and stored locally, fully complying with medical privacy and security standards, thus solving the dual pain points of poor real-time performance and high privacy risks of existing cloud-based solutions. Through this offline optimization module, the embodied intelligence platform also possesses the ability to autonomously iterate and evolve, continuously adapting to users' personalized usage habits. As usage time increases, the accuracy of speech recognition, environmental adaptation precision, and decision rationality continuously improve, achieving a personalized adaptation effect of "becoming smarter with use, one model per person," completely overcoming the shortcomings of poor adaptability of general-purpose models.
[0018] In the embodied intelligence platform based on a multimodal large model of this invention, the fuzzy speech error-tolerant interaction logic algorithm built into the voice intent perception submodule can perfectly adapt to various special groups such as the elderly, those with unclear speech, stroke rehabilitation patients, and those with postoperative mobility difficulties. It eliminates the need for standard pronunciation and precise operation by the user, greatly reducing the barrier to entry for using smart wheelchairs. Furthermore, because the multimodal perception module of this invention possesses full-scene adaptive perception and control capabilities, it can adapt to various complex indoor and outdoor scenarios such as homes, hospitals, nursing homes, and communities. Its scene compatibility and user adaptability far exceed those of existing general-purpose smart wheelchairs.
[0019] The embodied intelligent platform system based on the full-modal large model of the present invention has a simple architecture, high stability, controllable transformation cost, strong implementation, and is easy to industrialize and popularize. In addition, the system has strong scalability. On the basis of the existing functions of intelligent wheelchairs such as assisted driving, safety protection, and personalized adaptation, it can be expanded to include medical adaptation functions such as rehabilitation guidance, health monitoring, emergency assistance, intelligent voice companionship, and route memory navigation. There is no need to reconstruct the system architecture, which has the potential for long-term iterative upgrades and a longer product life cycle. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the structure of the embodied intelligent platform in an embodiment of the present invention;
[0021] Figure 2 This is a diagram of the fusion reasoning logic architecture of the fusion reasoning module in an embodiment of the present invention.
[0022] Figure 3 This is a flowchart of the offline optimization module in an embodiment of the present invention.
[0023] In the diagram, 1. Multimodal perception module; 11. Voice intent perception submodule; 12. Visual perception submodule; 13. Wheelchair posture perception submodule; 2. Fusion reasoning module; 21. Full-modal large model; 3. Safety decision module; 4. Execution control module; 41. Manual emergency submodule; 5. Offline optimization module; 100. Embodied intelligence platform. Detailed Implementation
[0024] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0025] To make the technical means, creative features, objectives and effects of this invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, provide a detailed description of the embodied intelligence platform based on a full-modal large model of this invention.
[0026] Figure 1 This is a schematic diagram of the embodied intelligence platform in this embodiment.
[0027] like Figure 1 As shown, this embodiment provides an embodied intelligence platform 100 based on a multimodal large model, which is used in the intelligent control system of a wheelchair to control the motion posture and operating status of the intelligent wheelchair during use. It includes a multimodal perception module 1, a fusion reasoning module 2, a safety decision module 3, an execution control module 4, and an offline optimization module 5.
[0028] The multimodal perception module 1 includes a voice intent perception submodule 11, a visual perception submodule 12, and a wheelchair posture perception submodule 13.
[0029] The voice intent perception submodule 11 uses a microphone array to recognize the user's voice to obtain user intent information. Through its built-in fuzzy speech error-tolerant interaction logic algorithm, it identifies ambiguous commands, non-standard pronunciations, and slow speech from special medical users, including elderly users, stroke patients, and post-operative rehabilitation users. The core function of this voice intent perception submodule 11 is intent semantic understanding, filtering environmental noise interference, accurately capturing the user's actual operational needs, and addressing the pain points of difficult voice interaction and low error tolerance for special populations.
[0030] The visual perception submodule 12 is used to collect environmental visual information, including cameras and depth cameras, to simultaneously collect environmental color images and depth information, thereby three-dimensionally reconstructing the surrounding environmental structure, accurately identifying environmental elements such as obstacles, traversable paths, and spatial boundaries, and providing accurate environmental data for obstacle avoidance and path decision-making.
[0031] The wheelchair posture perception submodule 13 includes an IMU posture sensor to acquire the motion posture and operating status of the smart wheelchair as operating status information. Specifically, the wheelchair posture perception submodule 13 collects motion posture and operating status data such as pitch, tilt, vibration, and balance during the wheelchair's movement in real time as operating status information. This accurately identifies hidden road condition hazards such as slopes, bumps, road inclination, and rollover risks, making up for the technical blind spots of pure visual perception and achieving comprehensive perception of both visible obstacles and hidden risks.
[0032] The multimodal perception module 1 is the information collection foundation of the embodied intelligence platform 100. It breaks the traditional single perception mode and constructs a three-dimensional integrated perception system of "voice intention perception + visual environment perception + body posture perception". It collects multimodal perception information in all aspects, including user intention information, environmental visual information, and wheelchair operation status information.
[0033] The fusion reasoning module 2 includes a full-modal big model 21, which extracts multiple operation-related features, including semantic intent features, visual environment features and ontological posture features, from the above multimodal perception information, performs cross-modal fusion processing on these operation-related features, infers and generates decision results, and calculates the confidence level of the decision results.
[0034] Figure 2 This is a diagram of the fusion reasoning logic architecture of the fusion reasoning module 2 in this embodiment.
[0035] Specifically, such as Figure 2 As shown, the fusion inference module 2 comprises an input layer, a fusion layer, an inference layer, and an output layer. The full-modal large model 21 extracts various operation-related features at the input layer, including semantic intent features from user intent information, visual environment features from environmental visual information, and ontological pose features from operational state information. The fusion layer performs feature alignment and projection, cross-modal attention fusion, and fusion feature representation. This inference layer contains a full-modal encoder and an inference decoder for cross-modal joint understanding. At the output layer, the full-modal large model 21 outputs the decision result and a risk assessment report including the confidence level of the decision result, as well as the natural language understood text obtained after cross-modal joint understanding. The fusion reasoning module 2 relies on a local medical-grade lightweight reasoning terminal installed on the smart wheelchair to achieve fully offline autonomous computation. The multimodal large model 21 performs all data processing, model reasoning, and decision result generation locally on the wheelchair. Unlike general large models, the multimodal large model 21 in this embodiment is deeply optimized for the characteristics of medical wheelchairs, such as low-speed operation, safety priority, fixed scene, and special user. It can simultaneously receive, fuse, and parse user intent information, visual environment information, and operating status information to achieve deep fusion reasoning of multimodal information, rather than simple data splicing. It can comprehensively understand "what the user wants to do, what the environment allows to do, and whether the current state of the device can execute", providing accurate intelligent reasoning results for subsequent safety decisions.
[0036] The safety decision module 3 classifies the safety level of the intelligent wheelchair based on the numerical range of the confidence level using the full-modal large model 21, and outputs the corresponding control commands for the intelligent wheelchair according to the safety level. The threshold of the numerical range of the confidence level of the decision result is preset in the full-modal large model 21.
[0037] Specifically, the security decision module 3 adopts a three-level security degradation inference mechanism, which divides the numerical range of the confidence level into a high-confidence range, a medium-confidence range, and a low-confidence range. The corresponding security levels for the intelligent wheelchair operation are divided into Level 1 normal operation, Level 2 degraded operation, and Level 3 security lockout. The control permissions and operating parameters of the wheelchair are tightened at each level to achieve full-scene security protection. The complete hierarchical logic, access control, and scene adaptation relationship of this three-level security degradation inference mechanism are shown in Table 1 below.
[0038] Table 1
[0039]
[0040] According to Table 1 above: When the confidence level is in the high confidence range (confidence level > 70%), the safety level is Level 1 normal operation. The control command output by the safety decision module 3 is: grant full permission for real-time status adjustment of the intelligent wheelchair. At this time, the embodied intelligent platform 100 fully responds to the user's voice commands, combines environmental perception information and the operating status information of the intelligent wheelchair, and outputs a refined and adaptive assisted driving control strategy to ensure ease of use. When the confidence level is in the medium range (30% < confidence level ≤ 70%), and there is slight environmental interference or ambiguous commands, the safety level is Level 2 downgraded operation. The full-modal large model 21 adjusts the weight of multimodal perception information, takes visual environmental information and the operating status information of the intelligent wheelchair as the core decision basis, weakens the weight of user voice commands, and limits the driving speed and steering range of the intelligent wheelchair. When the confidence level is in the low confidence range (confidence level ≤ 70%), the safety level is low. When the speed reaches 30%, the safety level is locked at Level 3. Safety Decision Module 3 forcibly blocks all unnecessary movement commands, retaining only the intelligent wheelchair's safety protection functions, including obstacle avoidance and emergency braking. It controls the intelligent wheelchair to operate at the minimum safe speed and simultaneously outputs risk warnings. This three-level safety degradation inference mechanism employed by Safety Decision Module 3 proactively protects against various high-risk scenarios such as ambiguous commands, complex environments, abnormal postures, and low model confidence. It eliminates safety incidents caused by model misjudgments at the root of algorithmic inference, filling the industry gap in AI safety decision-making mechanisms for medical-grade intelligent wheelchairs and achieving comprehensive, all-time safety assurance.
[0041] The execution control module 4 serves as the action implementation carrier of the embodied intelligent platform 100. Based on the control commands output by the safety decision module 3, it drives the intelligent wheelchair to make real-time state adjustments, complete assisted driving actions such as forward movement, backward movement, steering, speed adjustment, braking, path fine-tuning, and posture adaptive correction, and feeds back the real-time state adjustment information of the intelligent wheelchair, such as real-time driving speed, posture changes, operating status, and execution results, to the full-modal large model 21.
[0042] Specifically, the execution control module 4 enables the full-modal large model 21 to monitor the equipment's operational feedback in real time, dynamically adjust subsequent decision-making strategies, avoid the problem of rigid execution of single commands, and ensure the smoothness, continuity, and safety of the assisted driving process. At the same time, the execution control module 4 includes a manual emergency submodule 41, which is used to switch the intelligent wheelchair's operating state to manual control in extreme and abnormal scenarios, forming a dual safety guarantee.
[0043] The offline optimization module 5 is used to encrypt and store multimodal data during the use of the smart wheelchair, and to iteratively optimize the full-modal large model 21 based on the multimodal data when the smart wheelchair is idle, so as to adapt to the user's accent characteristics, operating habits and common scenarios. The multimodal data includes information, decision results, control commands of the smart wheelchair and real-time status adjustment information.
[0044] Figure 3 This is a flowchart of the offline optimization module 5 in this embodiment.
[0045] Specifically, such as Figure 3 As shown, the operation of the offline optimization module 5 includes: multimodal data acquisition, local encrypted storage, idle state monitoring, small sample screening, model fine-tuning and updating, and model adaptation optimization. Multimodal data acquisition refers to inputting the multimodal perception information acquired by the multimodal perception module 1 into the offline optimization module 5 in data form. Local encrypted storage refers to storing the multimodal perception information in an encrypted database by category and time in an offline state, and performing data anonymization processing, allowing access only to local devices to protect user privacy and security. Idle state monitoring refers to determining whether the smart wheelchair is in an idle state when it is charging, idle, during low-frequency nighttime use, or when the battery is fully charged. When the smart wheelchair is in an idle state, small sample screening is performed, and the screened data includes: recent high-frequency usage data, high-confidence interaction data, representative scenario data, and user error correction data. After screening, model fine-tuning and updating are performed, including efficient parameter fine-tuning of the full-modal large model 21, updating model parameters, maintaining the original capabilities of the model, local secure replacement, and retaining historical versions. Model adaptation optimization refers to improving speech recognition capabilities and the ability to understand user commands, making the use of smart wheelchairs more in line with user habits and more suitable for various scenarios. In short, the offline optimization module 5 can continuously adapt to the accent characteristics, operating habits, and common travel routes of different users, achieving a personalized upgrade effect of "becoming more adaptable and intelligent with use," thereby solving the industry pain points of poor adaptability and inability to iterate on general models.
[0046] The complete workflow of the embodied intelligence platform 100 based on a full-modal large model in this embodiment includes the following steps S1-S5.
[0047] S1, the user's voice wakes up the embodied intelligent platform 100, starts the multimodal perception module 1, and collects multimodal perception information in real time;
[0048] S2, transmit all the above multimodal perception information to the full-modal large model 21, and the fusion reasoning module 2 completes the full-modal fusion reasoning and outputs the decision result and the confidence level of the decision result;
[0049] S3, the safety decision module 3 determines the corresponding safety level based on the value range of the confidence level and outputs the corresponding intelligent wheelchair control command;
[0050] S4, the execution control module 4 executes the control command (i.e., hierarchical safety execution) and provides real-time feedback of the running status information to the full-modal large model 21 (i.e., running safety feedback).
[0051] S5, along with the offline optimization module 5, accumulates user-specific local data and completes model fine-tuning and upgrades during idle time, continuously improving the adaptability and intelligence level of the embodied intelligence platform 100.
[0052] Users can turn off the embodied intelligent platform 100 at any time via voice command. The embodied intelligent platform 100 will then automatically enter a low-power sleep state. The entire process requires no manual intervention, thus achieving fully automatic closed-loop operation.
[0053] The role and effects of the embodiments:
[0054] The embodied intelligence platform 100 based on a multimodal large model in this embodiment achieves joint perception of user intent information, environmental visual information, and wheelchair operating status information through the voice intent perception submodule 11, visual perception submodule 12, and wheelchair posture perception submodule 13 in the multimodal perception module 1. It breaks through the single limitation of pure visual perception in traditional intelligent wheelchairs, and deeply integrates and reasones the operating status information of the wheelchair itself with the environmental visual information. It is no longer limited to recognizing conventional obstacles, but can accurately identify hidden road conditions that traditional vision cannot capture, such as slopes, road inclination, bumps, step edges, and rollover risks, to achieve two-way embodied perception of "environment + body". Based on the above multimodal perception information, the fusion reasoning module 2 calls the multimodal large model 21 to generate decision results and calculate confidence, and the safety decision module divides the safety level and outputs control commands, thereby achieving comprehensive recognition of explicit obstacles and hidden risks, greatly improving driving safety in complex indoor and outdoor scenarios, and solving the core pain points of incomplete risk recognition and easy loss of control in existing technologies.
[0055] In this embodiment of the embodied intelligence platform 100 based on a full-modal large model, the offline optimization module 5 and the full-modal large model 21 perform local calculations entirely offline, without relying on the cloud network. This eliminates control lag and failure issues caused by network fluctuations and outages, ensuring stable and reliable operation. It can achieve millisecond-level real-time decision response and also ensures that all user data is encrypted and stored locally, fully complying with medical privacy and security standards. This solves the dual pain points of poor real-time performance and high privacy risks of existing cloud solutions. Through the offline optimization module 5, the embodied intelligence platform 100 also has the ability to autonomously iterate and evolve, continuously adapting to users' personalized usage habits. As usage time increases, the accuracy of speech recognition, environmental adaptation precision, and decision rationality continuously improve, achieving a personalized adaptation effect of "becoming smarter with use and one model per person," completely overcoming the shortcomings of poor adaptability of general models.
[0056] In this embodiment of the embodied intelligence platform 100 based on a multimodal large model, the fuzzy speech fault-tolerant interaction logic algorithm built into the voice intent perception submodule 11 can perfectly adapt to various special groups such as the elderly, those with unclear speech, stroke rehabilitation patients, and those with postoperative mobility difficulties. It eliminates the need for standard pronunciation and precise operation by the user, greatly reducing the barrier to entry for using smart wheelchairs. Furthermore, because the multimodal perception module 1 of this invention possesses full-scene adaptive perception and control capabilities, it can adapt to various complex indoor and outdoor scenarios such as homes, hospitals, nursing homes, and communities. Its scene compatibility and user adaptability far exceed those of existing general-purpose smart wheelchairs.
[0057] The embodied intelligence platform 100 based on the full-modal large model in this embodiment has a simple system architecture, high stability, controllable transformation costs, strong feasibility, and is easy to industrialize and popularize. In addition, the system has strong scalability. On the basis of the existing functions of intelligent wheelchairs such as assisted driving, safety protection, and personalized adaptation, it can be expanded to include medical adaptation functions such as rehabilitation guidance, health monitoring, emergency assistance, intelligent voice companionship, and route memory navigation without reconstructing the system architecture. It has the potential for long-term iterative upgrades and a longer product life cycle.
[0058] Those skilled in the art should understand that this invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to this invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. An embodied intelligence platform based on a full-modal large model, used to control the motion posture and operating status of an intelligent wheelchair during use, characterized in that, include: The multimodal perception module is used to acquire multimodal perception information, including user intent information, the operating status information of the smart wheelchair, and environmental visual information of the surrounding environment. The fusion reasoning module includes a full-modal large model. Based on the multimodal perception information, it generates decision results through reasoning using the full-modal large model and calculates the confidence level of the decision results. The safety decision module classifies the safety level of the intelligent wheelchair based on the numerical range of the confidence level using the full-modal large model, and outputs the corresponding control commands for the intelligent wheelchair based on the safety level. The execution control module drives the intelligent wheelchair to perform real-time status adjustments according to the control commands, and feeds back the real-time status adjustment information of the intelligent wheelchair to the full-modal large model; The multimodal sensing module includes: The voice intent perception submodule obtains the user's intent information by recognizing the user's voice; The visual perception submodule is used to collect the visual information of the environment; The wheelchair posture perception submodule is used to acquire the motion posture and operating status of the smart wheelchair as the operating status information.
2. The embodied intelligence platform based on a full-modal large model according to claim 1, characterized in that: in, The full-modal large model extracts multiple operation-related features from the multimodal perception information, performs cross-modal fusion processing on these features, infers and generates the decision result, and calculates the confidence level of the decision result. The operation-related features include semantic intent features, visual environment features, and ontological posture features.
3. The embodied intelligence platform based on a full-modal large model according to claim 1, characterized in that: in, In the safety decision module, the threshold of the confidence level's numerical range is preset in the full-modal large model. The numerical range is divided into a high-confidence range, a medium-confidence range, and a low-confidence range. The corresponding safety levels for the intelligent wheelchair's operation are Level 1 normal operation, Level 2 degraded operation, and Level 3 safety lockout. When the security level is Level 1 and the system is operating normally, the control command output by the security decision module is: grant full permission to adjust the real-time status of the intelligent wheelchair; When the safety level is downgraded to level two, the control command is to limit the driving speed and steering range of the intelligent wheelchair. When the safety level is three-level safety lock, the control command is: retain only the safety protection functions of the smart wheelchair, including obstacle avoidance and emergency braking, control the smart wheelchair to run at the minimum safe speed and output risk warnings simultaneously.
4. The embodied intelligence platform based on a full-modal large model according to claim 1, characterized in that: in, The execution control module includes a manual emergency submodule, which is used to switch the operation of the intelligent wheelchair to manual control in extreme abnormal scenarios.
5. The embodied intelligence platform based on a full-modal large model according to claim 1, characterized in that, Also includes: An offline optimization module is used to encrypt and store multimodal data during the use of the smart wheelchair, and to iteratively optimize the full-modal model based on the multimodal data when the smart wheelchair is not in use, in order to adapt to the user's accent characteristics, operating habits and common scenarios. The multimodal data includes the information, the decision results, the control commands of the smart wheelchair, and real-time status adjustment information.
6. The embodied intelligence platform based on a full-modal large model according to claim 5, characterized in that: in, The operation of the offline optimization module includes: multimodal data acquisition, local encrypted storage, idle state monitoring, small sample screening, model fine-tuning and updating, and model adaptation optimization.
7. The embodied intelligence platform based on a full-modal large model according to claim 1, characterized in that: in, The visual perception submodule includes a camera and a depth camera, which are used to acquire color images and depth information, respectively, thereby enabling three-dimensional environmental perception, obstacle recognition, and spatial distance judgment.
8. The embodied intelligence platform based on a full-modal large model according to claim 1, characterized in that: in, The voice intent perception submodule has a built-in fuzzy voice fault-tolerant interaction logic algorithm, which is used to identify fuzzy commands, non-standard pronunciations, and low-speed speech of special medical users, including elderly users, stroke patients, and postoperative rehabilitation users.