Real object pattern and Chinese character element fused cross-modal environment interaction method and system
By integrating physical patterns with Chinese character elements and combining them with multimodal input technology, the problems of cultural barriers in graphical interfaces and fragmented multi-terminal experiences are resolved, achieving cross-cultural intuitive interaction and efficient operation.
Patent Information
- Application Number
- CN202510999796.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies have problems such as cultural barriers to graphical interfaces, fragmented multi-terminal experiences, and separation of multi-modal inputs, which lead to high user learning costs and low operational efficiency.
A cross-modal environment interaction method that integrates physical patterns with Chinese character elements is adopted. The physical image of the device and the Chinese character control instructions are dynamically synthesized into an interactive interface through a semantic association engine. Combined with multimodal input decoding technology, natural intent recognition is achieved and it is adapted to different terminals under a layered rendering architecture.
It achieves intuitive interaction of cross-cultural cognition, reduces the user's cognitive load, improves operational efficiency, enables different users to complete intuitive operations within 0.5 seconds, and adapts to efficient control in multiple scenarios.
Smart Images

Figure CN120803272A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to intelligent environment control technology, especially suitable for smart home, medical care for the aged, vehicle-mounted system and other scenes requiring efficient human-computer interaction. BACKGROUND
[0002] The prior art has three defects: (1) Cultural barriers of graphical interface: abstract icons need to be learned by users (such as European users do not understand that the "fan icon" represents a fan); (2) Fragmentation of multi-terminal experience: the layout of the mobile terminal App and the wall panel control is inconsistent; (3) Separation of multi-modal input: voice, touch, and gesture need to be operated independently.
[0003] The patents of the same applicant, CN2025106599753 Lighting control panel interaction method and system based on Chinese Hanzi elements, CN2025210350749 Lighting air conditioning sound three-in-one environment adjustment panel based on Chinese Hanzi elements, CN2025214036074 Villa mansion Hanzi element full-scene interaction system, 2025109205838 Environment state interaction method and system based on anti-sense Hanzi combination and dynamic mask, and 2025210801508 Hanzi element interaction large screen device with decoration art and environment adjustment function, although they realize the interaction of Hanzi elements and environment control, they do not solve the problem of real object cognition fusion. SUMMARY
[0004] The present application discloses a cross-modal environment interaction method and system integrating real object patterns and Hanzi elements. Through a semantic association engine, device real object images and Hanzi control instructions are dynamically synthesized into an interaction interface, combined with multi-modal input decoding technology (touch / voice / EEG, etc.) to realize natural intent recognition, and adapted to wall panels, robot screens and other terminals under a layered rendering architecture. Figure 1 As shown in the accompanying drawings, the present application adopts a three-layer modular architecture, including an input layer, a processing layer, and an output layer.
[0005] Input layer: receives multi-modal signals such as touch, voice, eye movement, and EEG, and captures user intent through physical sensors (capacitive screen, microphone, camera, EEG electrode).
[0006] The processing layer is composed of a semantic association engine and a cross-modal decoder: (1) Semantic association engine: dynamically binds device real object images (such as air conditioner line drawing) with Hanzi instructions (such as "clear wind"), and establishes a mapping database; (2) Cross-modal decoder: uniformly converts input signals into control instructions (such as voice→Hanzi, eye movement→coordinate positioning).
[0007] The output layer is composed of a terminal rendering engine and an environment control bus: (1) Terminal rendering engine: adaptively generate mixed interfaces according to device types (wall panel / AR glasses); (2) Environment control bus: drive lighting, air conditioning and other devices through MQTT / CoAP protocol.
[0008] The core innovation of the present application: physical patterns provide intuitive cognitive anchors, Chinese character elements inject cultural semantics, and double channels reduce interactive cognitive load.
[0009] The core interactive interface of the present application ( Figure 2 ) adopts a three-layer dynamic fusion architecture: the bottom layer is an artistic line drawing of air conditioning and other devices (precisely retaining key structural features such as air outlets and blades), the middle layer superimposes the semi-transparent regular script Chinese character "Snow Mountain Breathing" (patented white ink printing process with a light transmission rate of 92%), and the top layer maps physical parameters through mask transparency: when the wind speed increases from level 1 to level 5, the "Snow" character mask gradually changes from 0% to 100%, triggering the ice crystal particle special effect to simulate the snow melting dynamic. This design converts abstract parameters such as temperature and wind speed into visual intuitive cognition (e.g., 80% mask coverage is perceived as a strong wind state). The cross-terminal rendering engine ( Figure 3 ) realizes scene adaptive layout: in the wall panel, it adopts a minimalist design of "large physical drawing + central floating Chinese character" (70% area displaying device line drawing), ensuring instant recognition for the elderly; the vehicle-mounted screen dynamically partitions based on driving safety logic, displaying only key instructions such as "stable" in the main driver area to ensure 0.5 seconds of operation during driving; and the AR glasses suspend the "cold" character at the real air conditioner outlet through spatial anchoring technology, allowing users to adjust by focusing on the gesture halo - the three terminals share the same semantic mapping rules (e.g., the "cold" character always corresponds to a temperature range below 24°C), maintaining the consistency of interactive logic in multiple scenarios.
[0010] The workflow closed-loop control mechanism of the present application is as follows: the system runs from multi-modal signal input, when the user's finger touches the air conditioner line drawing on the wall panel (or the eye gaze focuses on the virtual marker in the AR glasses), the voice command "increase wind power" is simultaneously transmitted into the processing layer. The semantic engine first locates the device type, retrieves the mapping library to generate the Chinese character instruction "strong wind", and the cross-modal decoder fuses voice and touch coordinates into a four-level wind speed control instruction. The output layer simultaneously drives the environmental devices (air conditioner wind speed increases) and updates the interface (the "Snow Mountain Breathing" mask transparency rises to 80% + ice crystal special effect is triggered), and returns the real-time wind speed data to the processing layer through the control bus. When the sensor detects that the actual wind speed meets the standard, the new interface state (e.g., the mask transparency stabilizes at 80%) is immediately fed back to the terminal screen, forming a closed loop of "intention input → instruction generation → environmental response → state visualization" ( Figure 4). The process is faster than the traditional interface in completing complex environmental adjustment, indicating that the physical-Hanzi dual-channel significantly compresses the cognitive path.
[0011] Application value and field empowerment: In the field of medical rehabilitation, patients with frozen shoulder can trigger the call system independently by staring at the "an" word (covered on the line drawing of the medicine bottle) on the bed screen within 3.5 seconds, breaking through the body dependence of traditional interaction; In the cultural export scene, Dubai hotels use the combination interface of "fan physical drawing + wind + al-riyah (Arabic 'wind')", which makes the understanding of Middle East users reach 100%; The "an / danger" opposite Hanzi mask system in the industrial control room significantly improves the accident response speed. These cases confirm the core value of the invention: establishing a cross-cultural cognitive benchmark with physical patterns, injecting precise semantics with Hanzi elements, and restructuring the intuitive paradigm of human-computer interaction under the condition of zero additional hardware. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 : System architecture diagram (showing the coupling relationship between the semantic engine and the multi-modal gateway).
[0013] Figure 2 : Mixed interface example (air conditioner physical line drawing + "snow mountain breathing" gradient mask Hanzi).
[0014] Figure 3 : Cross-terminal rendering comparison (layout scheme of wall panel / vehicle-mounted screen / AR glasses).
[0015] Figure 4 : Workflow diagram.
[0016] Figure 5 : Interaction interface diagram for the elderly voice control scene.
[0017] Figure 6 : Interaction interface diagram for children's room scene.
[0018] Figure 7 : Interaction interface diagram for hospital bed scene.
[0019] Figure 8 : Interaction interface diagram for hotel desk scene.
[0020] Figure 9 : Interaction interface diagram for smart conference room scene. DETAILED DESCRIPTION
[0021] Example 1: Elderly voice control scene ( Figure 5 ). Interaction process: (1) The user says "a little stuffy" → the system matches the "stuffy" word associated with the air conditioner and fresh air; (2) The wall panel displays the fan physical drawing + "clear wind" Hanzi sentence; (3) The user nods to confirm → the air conditioner starts with level 2 wind speed + fresh air oxygenation.
[0022] Example 2: Children's room scene (voice + touch interaction, Figure 6 ). Problem: Children have difficulty understanding the abstract control interface, and traditional buttons pose a risk of accidental touches. Solution: (1) Interface design: The wall panel displays a cartoon-like physical pattern (cloud-shaped lamp + moon-shaped air-conditioning outlet), with superimposed luminous Chinese characters ("Dream" covers the cloud lamp, and "Cool" is embedded in the moon-shaped air-conditioning outlet); (2) Interaction process. The child says: "Tell me a story from Journey to the West" → Voice recognition triggers the appearance of a real-life picture of the story machine and the word "悟" (enlightenment); Touch the word "悟" → The light changes to a warm orange light (Flame Mountain effect), and the air conditioner sends out a breeze (Somersault Cloud wind speed); Gaze at the word "梦" (dream) for 2 seconds → The light gradually dims to night reading mode (3000K color temperature); (3) Industrial design highlights: Chinese characters are inlaid with food-grade silicone, which is soft to the touch and impact-resistant; the surface of the physical pattern is covered with a luminous coating, which shows the outline in a dark environment.
[0023] Example 3: Hospital bed scene (eye control + EEG interaction, Figure 7 ). Problem: Severely ill patients have limited limb movement and require silent control of the environment. Solution: (1) Interface design: The bed support screen displays simple drawings of medical equipment (IV stand icon + heart rate monitor line drawing); superimposed medical Chinese characters (the character "安" is semi-transparently covering the IV stand, and the character "宁" is suspended on the electrocardiogram); (2) Interaction process. EEG detects anxiety signals → the interface highlights the medicine bottle pattern + the word "Shu" (blue-green breathing light effect); eyes focus on the word "Shu" → lavender fragrance is released, and the mattress starts the massage waveform; blink twice to trigger the word "An" → call the nurse station and turn on the bedside lamp; (3) Industrial design highlights: The screen is integrated with millimeter-wave radar to identify eye movements without the need for wearing equipment; Chinese characters use low-saturation Morandi colors to avoid visual stimulation in the medical environment.
[0024] Example 4: Hotel desk scene (gesture + environmental perception interaction, Figure 8 ). Problem: Business users need efficient multi-device collaborative control. Solution: (1) Interface design: The embedded screen on the desk displays a 3D device model (rotatable desk lamp + coffee machine cross-section); Chinese characters are integrated into the model in the form of calligraphy watermarks (the word "ink" appears through the desk lampshade, and the word "chun" is imprinted on the coffee cup); (2) Interaction process. Gesture air grab "ink" word → desk lamp brightness with gesture height infinitely adjustable; Coffee machine water temperature anomaly → "alcohol" word red flashing + mobile phone push "boil spring guest" prompt sentence; Say "conference mode" → curtains automatically closed, projector down, Chinese character switch to "meeting" word (square lishu body); (3) Industrial design highlights: screen surface etching micro-texture light guide groove, making Chinese characters produce relief effect under side light; Coffee machine pattern built-in temperature sensing color change layer, "alcohol" word automatically red above 65℃.
[0025] Example 5: Intelligent conference room scene (multi-user collaborative interaction, Figure 9 ). Question: When there are many people in the meeting, the environment parameters are difficult to adjust. Solution: (1) Interface design: conference table central projection dynamic heat map (region division physical icon: projector / air conditioner / sound); Each partition superimposes the opposite Chinese character combination (cold / hot, light / dark, high / low); (2) Interaction process. Participant A eye gaze "cold" word → its seat area air conditioner cooling 2℃ (map "cold" word mask expansion); Participant B gesture rotation "light" word → projector area brightness increase ("light" word transparency to 30%); System detects argument → automatically trigger tea cup pattern + "and" word, start background guqin music; (3) Industrial design highlights: use capacitive ink screen table mat, Chinese character touch area supports multi-user operation at the same time; Laser indication channel is set between physical icon and Chinese character, which clearly defines the control object ownership.
[0026] These five examples prove that through the combination of universal recognition of physical patterns and cultural accuracy of Chinese characters, combined with multi-modal interaction technology, a truly "zero learning cost" environment control experience can be created. From children to patients, from travelers to business elites, users in different scenarios can complete intuitive operation within 0.5 seconds, which is the key leap for Chinese character interaction from "technological innovation" to "humanistic care".
Claims
1. An interactive method for integrating physical patterns with Chinese character elements, characterized in that including: a. Generation of mixed visual elements: Dynamically synthesize an interface by using the physical image / 3D model of environmental devices and Chinese character elements (single characters / words / short sentences) that map control instructions through a semantic association engine; b. Cross-modal intention recognition: Receive at least one input of touch screen, voice, motion, eye tracking, and electroencephalogram signals, and convert it into a control instruction through a multi-modal fusion decoder; c. Dynamic response feedback: Adjust environmental parameters according to the instruction, and real-time update the transparency / color / morphology of Chinese characters and the status of physical patterns on the interaction interface.
2. The method according to claim 1, wherein the operations of the semantic association engine include: (1) Construct a physical object-Chinese character mapping database (such as "physical fan image + 'gentle breeze' corresponding to wind speed adjustment"); (2) Automatically match calligraphy fonts (regular script for home scenes / running script for art spaces) by using an aesthetic weight algorithm; (3) Generate a dynamic mask layer to make Chinese character strokes gradually change with parameters (such as the coverage ratio of the character "cold" reflecting the actual temperature).
3. The method according to claim 1, wherein the cross-modal intention recognition is specifically: (1) Parse voice input into Chinese character instructions through a dialect-adapted ASR module; (2) Lock the control target through a fixation point-Chinese character hot zone association algorithm for eye tracking; (3) Generate control words such as "relaxed / focused" from electroencephalogram signals by an EEG-Chinese character semantic conversion model.
4. An interactive system implementing claims 1-3, characterized in that including: (1) Mixed interface rendering module: Adaptively layout physical patterns and Chinese characters based on the terminal type (wall panel / robot screen / AR glasses); (2) Multi-modal input gateway: Integrate touch control, microphone array, millimeter wave radar, and EEG sensors; (3) Environmental control actuator: Drive lighting, air conditioning, and audio equipment through MQTT / CoAP protocols.
5. The system according to claim 4, wherein the mixed interface rendering module supports: (1) 3D perspective deformation of physical patterns (such as the air conditioner outlet model rotates with the wind direction change); (2) Scene-based reorganization of Chinese character elements (display "energy-saving mode" + light bulb icon on the mobile terminal, and display "follow me" + footprint icon on the robot screen).
Citation Information
Cited By
Screen control method and system
CN121171004A