Language barrier intention analysis system based on multi-modal perception and dynamic interaction

Through the multimodal perception matrix and dynamic reliability fusion algorithm, non-invasive sensors are used to fuse multiple data sources, and the problems of single interaction, limited scenarios and misjudgment of intentions in the prior art are solved, and high-precision intention analysis and low-cost accurate analysis of those with language impairments are achieved.

CN120197029APending Publication Date: 2025-06-24GUANGZHOU ROBOTZERO SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510334787.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing assistive technologies have pain points in terms of single interaction, limited scenarios and misjudgment of intentions, and it is difficult to meet the needs of people with severe disorders, especially those with movement disorders.

Method used

Using multimodal perception matrix and dynamic reliability fusion algorithm, a variety of data sources are fused to achieve high-precision demand inference through non-invasive sensors such as millimeter wave radar, thermal imaging camera, throat vibration sensor and smart carpet.

Benefits of technology

It realizes high-precision intention analysis for those with language impairments, reduces the misjudgment rate, is suitable for a variety of scenarios, and uses non-invasive equipment, and is low in cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197029A_ABST
    Figure CN120197029A_ABST
Patent Text Reader

Abstract

The invention discloses a language barrier person intention analysis system based on multi-modal perception and dynamic interaction, which realizes high-precision demand inference through a non-intrusive multi-modal perception matrix and dynamic credibility fusion algorithm. The system supports two deployment forms: 1) an independent unit (spherical equipment comprises a mechanical arm); (2) an embedded module (a standard interface is adaptive to various devices); the innovativeness comprises throat vibration spectrum analysis, thermal imaging emotion recognition and federal learning personalized modeling. According to the embodiment of the invention, the intention recognition accuracy is high and the response speed is higher than that of a traditional scheme in a family scene-autistic child assistance, a public place-apoplexy patient navigation, disaster rescue-aphasia trapped person communication and other scenes. The system supports independent deployment of minimum units or is used as a module to be integrated to equipment such as robots, household appliances and public facilities, and is suitable for families, medical treatment and public rescue scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of artificial intelligence and assistive technology, and specifically relates to a multimodal intention analysis system for people with language disorders. Through non - invasive sensor fusion, dynamic interaction strategies, and an adaptive decision - making model, it realizes demand understanding and behavior prediction. The system supports independent deployment of the smallest unit or integration as a module into devices such as robots, household appliances, and public facilities, and is applicable to scenarios such as home, medical, and public rescue. Background Art

[0002] Existing assistive technologies have the following pain points: First, the interaction is single: Traditional eye - tracking devices / button devices rely on active operations and cannot meet the needs of severely disabled people; Second, the scenario is limited: Fixed devices are difficult to adapt to dynamic environments (such as rescue sites); Third, the intention is misjudged: Single - modality data (such as gestures) is prone to misreading (misjudgment rate > 35%).

[0003] CN118672410A "SSVEP - eye movement Chinese association input method based on large language model" proposes a brain - computer interface solution to provide an effective communication tool for people with motor disabilities, but it requires invasive devices and is costly. The present invention realizes low - cost and accurate analysis through a multimodal fusion architecture and a dynamic credibility evaluation algorithm. Summary of the Invention

[0004] The present invention provides an intention analysis system for people with language disorders based on multimodal perception and dynamic interaction. Through a non - invasive multimodal perception matrix and a dynamic credibility fusion algorithm, it realizes high - precision demand inference. The system architecture is shown in Figure 1 , and it consists of a perception layer, a data processing layer, and an application layer. The core innovative designs include a non - invasive perception matrix and a dynamic credibility fusion algorithm.

[0005] The non - invasive perception matrix is composed of multi - form perception devices such as millimeter - wave radar, thermal imaging cameras, laryngeal vibration sensors, and intelligent carpets ( Figure 2 ), and its detection dimensions and technical indicators are shown in Table 1: Table 1 Detection Dimensions and Technical Indicators of Each Sensor Sensor Detection dimension Technical index 60GHz millimeter-wave radar Chest fluctuation (accuracy 0.1mm) Respiratory rate error < 0.2 times / minute Thermal imaging camera Temperatures of 42 facial feature points Emotion recognition accuracy > 89% Laryngeal vibration sensor Vocal cord micro-vibration frequency (50 - 200Hz) Detection delay of vocalization intention < 0.3 seconds Intelligent carpet Foot pressure distribution (resolution 1cm²) Recognition rate of abnormal gait > 93% .

[0006] The present invention adopts a dynamic credibility fusion algorithm for intention analysis, and the formula is: Where: wi: The weight of the i - th sensor, reflecting its reliability; Si: The signal strength of the i - th sensor (such as breathing amplitude, gesture amplitude); λ: Time decay factor, controlling data timeliness; t: The time difference from data acquisition to the current time (unit: second).

[0007] The core idea of this algorithm: 1) Time decay: The weight of sensor data decays as the "freshness" decreases, and the influence of old data weakens; 2) Multimodal weighting: Different sensors have different importance, which is reflected by the weight wi; 3) Dynamics: The weight wi and the decay factor λ can be dynamically adjusted according to the scenario.

[0008] The following explains the calculation process of judging the user's intention through gestures (radar) and expressions (thermal imaging) in combination with a scenario example.

[0009] 1) Time decay: For the gesture signal, λ1 = 0.1, and it decays to e−0.1 ≈ 0.905 after 1 second; for the expression signal, λ2 = 0.05, and it decays to e−0.05 ≈ 0.951 after 1 second.

[0010] 2) Dynamic weight: The recent accuracy rate of the gesture sensor is 80%, and the basic weight is 0.6 → w1 = 0.7 ⋅ 0.6 + 0.3 ⋅ 0.8 = 0.66; the recent accuracy rate of the thermal imaging sensor is 90%, and the basic weight is 0.4 → w2 = 0.7 ⋅ 0.4 + 0.3 ⋅ 0.9 = 0.55.

[0011] 3) Comprehensive calculation: The intensity of the gesture signal S1 = 0.8, and the time difference t = 2 seconds → S1′ = 0.8 ⋅ e−0.1 ⋅ 2 ≈ 0.654; the intensity of the expression signal S2 = 0.9, and the time difference t = 2 seconds → S2′ = 0.9 ⋅ e−0.05 ⋅ 2 ≈ 0.812; the final credibility Cfinal = 0.66 ⋅ 0.654 + 0.55 ⋅ 0.812 ≈ 0.879.

[0012] The interactive decision-making process is shown in Figure 4 , first, the user's behavior triggers multimodal signals. After the perception layer perceives these signals, it sends timestamped data to the aggregation engine, calculates the required probability through the formula, then generates an intention model. The system combines with the interactive device to generate an interactive plan and provides it to the decision-making engine. Under the voice query and action guidance, the execution layer will execute these interactive plans. On the other hand, the user's feedback will also confirm or correct the system, and finally update the personal behavior pattern to the database.

[0013] The present invention has two deployment forms: independent unit ( Figure 2 ), embedded module ( Figure 3 ).

[0014] Independent unit ( Figure 2 ): A spherical device with a diameter of 15 cm, containing an annular radar array + fixed by magnetic adsorption / folding robotic arm bracket, supporting dual-mode power supply of solar energy + lithium battery. Typical scenarios: Placed on the desktop to assist in eating, installed on the wall to monitor falls.

[0015] Embedded module ( Figure 3 ): Standard interface encapsulation (Type-C + PoE). 1) It can be retrofitted to the wheelchair armrest: Perceive hand movements through vibration; 2) It can be retrofitted to the refrigerator door: Monitor the food intake frequency and conduct nutritional analysis; 3) It can be retrofitted to the rescue drone: Recognize the gestures of trapped persons in the air. Description of the Drawings

[0016] Figure 1 : System architecture diagram.

[0017] Figure 2 : Exploded view of the independent unit structure. Outer layer in the figure: Transparent wave composite material spherical shell; Middle layer: Ring-shaped radar array (8 groups of 60GHz modules); Core: AI processing chip + Foldable robotic arm joint.

[0018] Figure 3 : Definition diagram of the embedded module interface. In the figure, data interface: USB4 + Fiber optic hybrid port; Power supply interface: PoE++ (90W) + Wireless charging coil; Mechanical interface: Universal magnetic suction base + 3 types of screw holes.

[0019] Figure 4 : Interactive decision-making flowchart.

[0020] Figure 5 : Rescue application scenario diagram. Detailed Implementation Modes

[0021] Example 1: Home scenario - Autism children assistance. Deployment plan: The independent unit is placed in the children's room, and the embedded module is installed in the refrigerator / medicine cabinet.

[0022] Operation process: 1) The radar detects that the child waves repeatedly (frequency 2Hz) + The thermal imaging shows anxiety (the temperature between the eyebrows rises by 1.2°C); 2) The system plays a voice: "Do you want to open the snack cabinet? Please nod or shake your head"; 3) The thermal imaging captures the nodding action → Unlock the snack cabinet and give a voice prompt: "No more than three times a day". Technical indicators: Intention recognition accuracy: Higher than the traditional solution; Response delay: <1.5 seconds.

[0023] Example 2: Public place - Stroke patient navigation. Deployment plan: The embedded module is integrated into the mall guide robot, and the pressure sensor is embedded in the armrest. Operation process: 1) The patient holds the armrest for 3 seconds (pressure > 2kg), and the thermal imaging shows the gazing direction; 2) The system asks a voice question: "Do you need to go to the restroom? (The light arrow points to the east)"; 3) Detect that the exhalation frequency increases (positive intention) → The robot leads the way and clears the crowd along the way. Innovation point: Distinguish urgent needs through the holding time of the grip strength (3 seconds = restroom, 5 seconds = medical help).

[0024] Example 3: Disaster rescue - Communication with aphasic trapped persons (Figure 5 ). Deployment plan: An independent unit is equipped with a rescue drone, and a module is embedded with a life detector. Operation process: 1) The drone hovers to identify the knocking rhythm (3 short, 3 long, 3 short = SOS); 2) Project a laser keyboard for the trapped person to select the injury condition (red light on the finger = fracture, green light = bleeding); 3) After the selection is confirmed by thermal imaging, an emergency kit is automatically dropped and the GPS coordinates are marked. Technical advantages: The recognition distance through the ruins is > 5 meters; multi-modal verification reduces the false alarm rate to < 5%.

Claims

1. A system for analyzing the intention of people with speech impairments based on multimodal perception and dynamic interaction, characterized in that Includes: non-invasive multimodal perception matrix (millimeter wave radar + thermal imaging + bone conduction microphone + pressure sensor); dynamic credibility fusion algorithm based on time attenuation factor; supports two physical forms of independent unit and embedded module.

2. The system according to claim 1, characterized in that Independent unit design: spherical shell with built-in circular radar array (beam coverage 360°x120°); foldable robotic arm (unfolded length 30cm, load 500g) for item delivery; solar panels integrated into the shell surface (conversion efficiency >22%).

3. The system according to claim 1, characterized in that Embedded module interface: power supply and data interface comply with USB4+Power over Ethernet standard; adaptive mounting bracket (compatible with M3-M8 screw holes); environmental shielding layer (reduces electromagnetic interference of main equipment> 20dB).

4. A method for inferring intention, characterized by the steps of: Active vocalization and unconscious sounds are distinguished through laryngeal vibration frequency; positive / negative intentions are judged by combining thermal imaging of zygomatic muscle temperature changes (±0.3°C); when multimodal conflicts occur, a three-level confirmation protocol (voice + light + vibration) is initiated.

5. A personalized learning mechanism, characterized by: Establish user-specific behavior maps (including 500+ micro-action features); update group models through federated learning, and retain local data for no more than 24 hours.