A mixed reality-based adhd supporter emotional healing method and system
Patent Information
- Application Number
- CN202610797529.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]然而,在现有的实施方式中,这类干预存在一个深层技术矛盾,即管理导向或脱离情境的治疗路径难以在支持者日常承受情感压力的真实物理空间中建立安全的情感锚点,从而无法实质性缓解因长期照护引发的情感耗竭
[0014]与现有技术相比,本发明具有如下有益效果:本发明通过构建一个由实体装置和混合现实数字层深度融合的治疗空间,形成了物理锚定和数字引导的混合框架,解决了现有管理导向的干预技术难以在支持者日常承受情感压力的真实物理空间中建立安全情感锚点的问题;触觉真实的实体装置为支持者提供了持续、可触及的物理安全感基底,避免了纯虚拟现实体验中因脱离现实环境而产生的定向障碍或失序感,支持者在进行内在情感探索时能够始终感知到现实世界的存在,确保了心理层面的稳定与安全;同时,混合现实数字层将接纳与承诺疗法的核心过程转化为交互式隐喻叙事,使支持者能够通过与虚实融合元素的具身交互完成情感外化,克服了传统纯语言干预难以触及深层次情感创伤的局限;系统中匿名化的集体叙事单元设计,允许支持者在其熟悉的物理空间中接触其他支持者的情感自述,在保持隐私的前提下有效消解了支持者的孤立感和羞耻感,建立了基于共同人性认知的情感支持网络;此外,纯语音引导和价值导向的沉浸式视听组合降低了社交压力和表现焦虑,赋予了支持者高度的自主权和对体验进程的控制力,支持其完成从被动接收信息到主动进行自我探索和意义重构的转变,从而在真实生活场景中更有效地实现内在情感的修复和价值重建。
Smart Images

Figure CN122822233A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer technology, specifically relating to an emotional healing method and system for ADHD supporters based on mixed reality. Background Technology
[0002] Mixed reality technology, by overlaying digital information in real-time onto a real physical environment, constructs an interactive space that blends the virtual and real worlds, demonstrating unique potential in the field of mental health intervention in recent years. Particularly for special groups undertaking long-term caregiving tasks, how to transform the framework of psychological therapy into a perceptible and participatory immersive experience has become a cutting-edge direction in the interdisciplinary research of human-computer interaction and digital health. This direction attempts to break through the traditional psychological support's singular reliance on language and cognition, deeply integrating bodily perception, physical environment, and emotional processes to address the emotional exhaustion and loss of self-worth experienced by caregivers under continuous stress. Among these, emotional support technologies for supporters of ADHD—parents, partners, or close friends who provide long-term companionship and care for ADHD patients—often employ an intervention logic centered on teaching management skills. Specifically, existing solutions typically use structured courses or virtual reality scenarios to teach supporters behavioral management strategies or guide them in cognitive restructuring to improve their ability to cope with external behavioral problems. Some virtual reality psychological support systems construct virtual scenarios detached from the real environment within a head-mounted display, allowing users to complete emotional expression or cognitive exercises in an isolated space before returning to their daily routines.
[0003] However, existing implementations of this type of intervention suffer from a deep-seated technical contradiction: management-oriented or detached therapeutic pathways struggle to establish safe emotional anchors within the real physical spaces where supporters experience daily emotional stress, thus failing to substantially alleviate emotional exhaustion caused by long-term care. This is because supporters' emotional stress does not originate from artificially designed training scenarios but is deeply rooted in their daily environments, such as family and close relationships. When psychological intervention occurs in virtual spaces or purely verbal conversations disconnected from these environments, while users acquire strategies at the cognitive level, they cannot establish a tangible, safe connection with the real space where stress arises at the level of bodily perception and emotional memory. This absence of physical context makes it difficult for the emotional regulation experience formed during therapy to automatically transfer to everyday stressful situations. As a result, supporters often quickly revert to a state of isolation, self-blame, and exhaustion upon returning to reality, ultimately causing the intervention to remain at the "skill acquisition" level rather than achieving sustained inner emotional repair. This makes the construction of a healing space that can provide a sense of security through physical anchoring in a real environment while simultaneously using digital guidance to externalize emotions and reconstruct values a critical technical bottleneck that urgently needs to be overcome in this field. Summary of the Invention
[0004] The purpose of this invention is to provide an emotional healing method and system for supporters of attention deficit hyperactivity disorder based on mixed reality, which can effectively solve the problems in the background art.
[0005] To achieve the above objectives, the technical solution adopted by this invention is as follows: an emotional healing method for supporters of attention deficit hyperactivity disorder (ADHD) based on mixed reality, comprising the following specific steps: S1: Deploying a physical device with a preset curved surface shape and tactile feedback characteristics in a real physical space, and superimposing virtual visual elements onto the physical device through a mixed reality head-mounted display device to construct a therapeutic space that integrates physical anchoring and digital guidance; S2: Guiding supporters to interact tactilely with the physical device under the visual content and audio guidance presented by the mixed reality head-mounted display device, performing a safety establishment and emotional activation phase, wherein the system renders a soft visual scene without specific spatial reference cues in the mixed reality environment, and evokes memories related to caregiving stress in supporters only through audio commands, while prompting supporters to adjust their body posture to a comfortable state; S3: Entering the narrative exploration and group connection phase, the system in the mixed reality... A central virtual object is presented in the field of view. When a supporter triggers this central virtual object through physical contact, the central virtual object decomposes into multiple anonymous narrative units. Each anonymous narrative unit carries anonymized audio narration fragments from other supporters. Supporters can freely choose and listen to these audio narration fragments. S4: Entering the value-oriented guidance and immersive integration stage, the system plays a customized audiovisual combination that transforms abstract emotional states into dynamic visual metaphors within the supporter's mixed reality field of view. At the end of this audiovisual combination, a resilient voice layer from the supporter group is seamlessly integrated, followed by a quiet period for supporters to engage in introspection and value clarification. S5: Executing the procedural convergence and return to reality stage, the system empowers supporters with the autonomy to end the experience through progressive voice prompts. After the supporter removes the mixed reality head-mounted display, a guided reflection manual is provided to consolidate the emotional integration results.
[0006] Preferably, in the security establishment and emotional activation stage of step S2, the mixed reality head-mounted display device utilizes its depth sensor and real-time positioning and mapping algorithm to perform 3D reconstruction of the physical interaction space and generate a persistent spatial map during the initialization stage. Based on this persistent spatial map, the virtual and real fusion rendering engine accurately registers and anchors the pre-configured virtual scene and physical device, so that the virtual interactive elements and the physical surface of the physical device achieve a consistent correspondence in spatial position. When the supporter is in this stage, the stable contact between the supporter and the surface of the physical device is detected in real time. The system adjusts the softness of the visual feedback and the speed of the audio guidance in the virtual environment according to the contact state, so that the supporter's psychological sense of security is gradually enhanced.
[0007] Preferably, the visual content within the mixed reality field of view contains only a visual metaphor system based on natural elements and abstract geometric forms, the themes of which are chosen autonomously by the supporter at the beginning of the experience; the audio guidance is pure voice guidance based entirely on pre-recorded audio material, without any artificially synthesized voice or real-time instructor interruptions.
[0008] Preferably, the process of the supporter triggering the central virtual object through physical contact in step S3 is specifically implemented using an embodied interaction technology module based on physical object anchoring. This module includes two processing layers: the first processing layer is a precise hand-to-surface contact detection layer, which determines whether the supporter intends to make contact with a specific surface area by fusing inside-out tracking data from the mixed reality head-mounted display device with sensor feedback signals sparsely distributed inside the physical device; when the spatial distance between the hand position and the surface area is less than a preset distance threshold and the sensor feedback signal exceeds a preset signal threshold, the contact state is determined to be valid; the second processing layer is a synchronous multimodal feedback coupling layer, which immediately triggers visual changes and audio responses that are closely coupled with the tactile properties of the physical device material when the first processing layer determines that the contact is valid. The visual changes are manifested as the decomposition animation of the central virtual object into multiple narrative units, and the audio response is the anonymous voice narration playback that starts synchronously with the visual decomposition.
[0009] Preferably, in step S4, the abstract emotional state is transformed into a customized audiovisual combination of dynamic visual metaphors. The generation process is controlled by a dynamic content management and scheduling engine. This engine orchestrates the delivery of media assets based on the supporter's experience process and real-time behavioral signals. Its functions include asynchronous playback of anonymous audio narratives, real-time streaming of abstract visual content, and coordinated control and smooth transition of ambient sound, guiding narration, and hierarchical audio. The engine also includes a lightweight adaptive modulation mechanism that reads state variables representing the supporter's state, which consist of the supporter's spatial location, interaction speed, and interaction intensity. The audiovisual content is modulated based on the state variables. The modulation process is constrained by a content change threshold, ensuring that the deviation of the content from the baseline value remains within the preset threshold range, thereby achieving adaptive fine-tuning while maintaining narrative coherence and emotional continuity.
[0010] Preferably, in step S5, the system supports multiple emotional safety protection mechanisms, including supporters being able to immediately exit by directly removing the mixed reality head-mounted display device, a progressive audio guidance degradation initiated by the system when it detects that the supporter is continuously unresponsive, and the option for supporters to accept the intervention of assistants; the guided reflection manual includes writing pages to guide supporters to transform emotional fragments in the experience into visual text or graphic symbols.
[0011] Preferably, the physical device is a single curved surface with an enveloping feel, and its surface material has a preset thermal conductivity and surface roughness to provide a warm and smooth tactile feedback; the physical device serves as the only continuously tactile object throughout the treatment process, providing a non-judgmental physical safety boundary for the supporter.
[0012] Preferably, the soft visual experience in step S2 specifically refers to a spacious visual experience in a mixed reality environment that does not superimpose reference space markers or virtual boundaries unrelated to physical devices, in order to reduce unnecessary visual competition for the supporter's attention.
[0013] The present invention also provides an emotional healing system for attention deficit hyperactivity disorder supporters based on mixed reality, utilizing the method described in any of the above.
[0014] Compared with existing technologies, this invention has the following beneficial effects: By constructing a therapeutic space deeply integrated with physical devices and a mixed reality digital layer, this invention forms a hybrid framework of physical anchoring and digital guidance, solving the problem that existing management-oriented intervention techniques struggle to establish safe emotional anchors in the real physical space where supporters experience daily emotional stress. The tactilely realistic physical devices provide supporters with a continuous and tangible foundation of physical security, avoiding the disorientation or sense of disorder that arises from detachment from the real environment in pure virtual reality experiences. Supporters can always perceive the existence of the real world while exploring their inner emotions, ensuring psychological stability and security. Simultaneously, the mixed reality digital layer transforms the core process of acceptance and commitment therapy into an interactive metaphorical narrative. This system enables supporters to externalize their emotions through embodied interaction with elements that blend the virtual and real worlds, overcoming the limitations of traditional purely verbal interventions that struggle to reach deep-seated emotional trauma. The anonymized collective narrative units within the system allow supporters to access the emotional narratives of other supporters in their familiar physical spaces, effectively mitigating feelings of isolation and shame while maintaining privacy, and establishing an emotional support network based on shared humanistic cognition. Furthermore, the combination of purely voice-guided and value-oriented immersive audiovisual experiences reduces social pressure and performance anxiety, granting supporters a high degree of autonomy and control over the experience, supporting their transition from passively receiving information to actively exploring themselves and reconstructing meaning. This allows them to more effectively repair their inner emotions and rebuild their values in real-life scenarios. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the overall technical solution architecture of an ADHD supporter emotional healing method based on mixed reality proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework for the transformation of the core process of acceptance and commitment therapy into interactive metaphorical narrative in this invention; Figure 3 This is a logical flowchart of the security establishment and emotional activation phases in this invention; Figure 4 This is a logical flowchart of the narrative exploration and group connection phases in this invention; Figure 5 This is a schematic diagram of the multi-level interaction relationship and data flow between the supporter and the treatment system in this invention; Figure 6 This is a schematic diagram comparing the technical effects of the mixed reality therapy space in this invention with purely physical or purely virtual spaces in establishing a safe emotional anchor point. Detailed Implementation
[0016] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below. Example 1
[0017] This embodiment provides a mixed reality-based emotional healing method and system for supporters of Attention Deficit Hyperactivity Disorder (ADHD). The system and method are specifically designed for supporters of long-term caregivers of ADHD patients, including the patient's parents, partners, or close friends. The aim is to construct a healing environment that integrates physical anchoring and digital guidance within the real physical spaces where these supporters experience daily emotional stress, such as a quiet corner of a family living room, private office, or community counseling room. This environment transforms the core therapeutic processes of acceptance and commitment therapy into a perceptible and participatory mixed reality interactive metaphorical narrative. This allows supporters to externalize emotions, reconstruct stress, and clarify values through embodied interaction with elements blending virtual and real elements within a familiar environment. This addresses the deep-seated technical contradiction that existing management-oriented or detached intervention techniques struggle to establish safe emotional anchors in real-world stressful scenarios and fail to substantially alleviate emotional exhaustion caused by long-term care.
[0018] Before detailing the treatment process of this embodiment, the hardware architecture, module composition, and their physical and logical connections on which the system depends will be clearly described first. The system in this embodiment is deployed in an indoor space with clearly defined physical boundaries. Its core hardware components include: a central processing and rendering host serving as the computing and control hub of the entire system; a mixed reality head-mounted display device for presenting mixed reality visuals and capturing user behavior data; a specially designed physical interaction device for providing continuous physical tactile reference and an emotional safety foundation; a spatial sensing subsystem consisting of depth sensors and environmental perception modules; and a remote or local content server for storing, scheduling, and managing all audiovisual media assets. All hardware entities are interconnected via high-speed wired and wireless communication links, forming a networked system with real-time perception, dynamic rendering, and closed-loop feedback capabilities.
[0019] The central processing and rendering host is a workstation-class computer equipped with a high-performance multi-core CPU, an independent graphics processing unit (GPU), and at least 32 gigabytes of dynamic random access memory (DRAM). Internally, the host carries a software stack that implements all the logical steps of the method described in this embodiment, including a spatial mapping and positioning module, a virtual-real fusion rendering engine, an interaction state determination and multimodal feedback coupling module, a dynamic content management and scheduling engine, and a lightweight AI state evaluation and content modulation module running in the background. The GPU of this host has at least 8 gigabytes of dedicated video memory and supports hardware-level real-time ray tracing and depth texture sampling to ensure that the high frame rate, low latency visual content rendered in the mixed reality head-mounted display device can achieve precise spatial alignment with the physical surface of the physical device. The host connects to the mixed reality head-mounted display device via a USB-C or dedicated fiber optic hybrid cable with a bandwidth of at least 10 gigabits per second, and is responsible for receiving real-time sensor data streams collected by the head-mounted device and transmitting back rendered stereoscopic visual images for the left and right eyes and spatialized audio signals to the head-mounted device.
[0020] A mixed reality head-mounted display is a wearable device with inward-to-outward spatial tracking capabilities, integrating a depth sensor, a high-definition see-through camera array, and an inertial measurement unit (IMU). Specific models can include the Microsoft HoloLens 2 or devices with equivalent technical specifications. The core functional components of this device include: a depth sensor, employing time-of-flight or stereo vision principles, capable of generating a dense depth point cloud of the physical environment within its field of view at a refresh rate of at least 30 frames per second; four or more visible light cameras forming an inward-to-outward tracking array for real-time calculation of six-DOF head pose and 3D reconstruction of hand keypoint positions; an IMU containing a three-axis accelerometer and a three-axis gyroscope, providing high-frequency pose interpolation and prediction when camera data experiences brief failures; a waveguide-type see-through holographic lens for projecting virtual images into the user's eyes while allowing the user to clearly see the real physical environment; a spatialized speaker array capable of accurately reproducing audio guidance and narrative content in 3D space based on the user's head pose and the position of virtual sound sources; and a built-in near-infrared illuminator for assisting hand tracking in low-light conditions. The operating system and runtime environment of the head-mounted device are responsible for managing the acquisition of all sensor data, initial timestamp synchronization and spatial coordinate system alignment, and transmitting the processed raw data stream, including the three-dimensional coordinates of the 21 joints of the hand, the environmental point cloud captured by the depth camera, and the six-degree-of-freedom pose of the head, to the central processing and rendering host via the aforementioned high-speed cable.
[0021] The physical interactive device is a key hardware element that distinguishes this system from pure virtual reality therapy solutions, providing supporters with a non-judgmental physical safety boundary. This physical device is designed as a single, enveloping curved surface, its physical contour approximating the inner surface of a large, longitudinally cut cylinder, or a deeply recessed, smooth, shell-like structure, ensuring that the user's torso and outstretched arms are naturally accommodated by its curved surface when seated or semi-reclining. The device's physical dimensions can be set as follows: approximately 1.2 meters in height, 1.5 meters in width, and approximately 0.8 meters in curved surface depth. Its main structural frame is assembled from lightweight, high-strength 6061 aluminum alloy profiles or carbon fiber composite tubing using 3D-printed joint connectors, covered by a continuous and seamless polymer composite shell. The outermost surface material of the shell undergoes rigorous tactile engineering selection, consisting of silicone-modified polyurethane or medical-grade thermoplastic elastomers with preset thermal conductivity and surface roughness. The material's thermal conductivity was chosen to be between 0.2 and 0.4 Kelvin per meter, and its surface roughness Ra value was controlled within the range of 1.0 to 2.5 micrometers. This provides a warm, smooth, and slightly soft non-cold, hard tactile feedback, avoiding the aloofness and medical connotations associated with metals or hard plastics. Between the inner surface of the housing and the outer cladding, a sparsely embedded array of 12 to 20 distributed force-sensitive resistors or capacitive proximity sensors is incorporated into the interlayer. These sensors are connected via thin-film cables to a low-power microcontroller unit (MCU) equipped with Bluetooth 5.0 or Zigbee wireless communication capabilities. This MCU is responsible for acquiring the pressure or capacitance changes of each sensor at a frequency of 100 Hz and wirelessly transmitting the packaged sensor data packets to the central processing and rendering host. The MCU is powered by an embedded lithium polymer battery, ensuring the entire device is free of external cables, maintaining a visually clean and user-friendly appearance.
[0022] The spatial sensing and environmental awareness subsystem, in addition to the depth sensor built into the head-mounted device, optionally includes one or two auxiliary RGB-D cameras installed in the corners of the room to expand the field of view and robustness for tracking the user's full-body posture. These auxiliary cameras are also connected to the central processing unit via USB 3.0 interfaces. The content server is a standard rack-mount server or a high-performance network-attached storage device, with an internal solid-state drive array storing all media assets, including pre-recorded anonymous voice narration clips, abstract animated video streams with dynamic visual metaphors, ambient background audio tracks, guiding narration audio files, and various visual theme asset packages. The content server connects to the central processing and rendering unit via a gigabit Ethernet or Wi-Fi 6 wireless network, with data transmission latency controlled to within 5 milliseconds.
[0023] The above is a detailed explanation of the static architecture of the system in this embodiment. Next, in conjunction with the S steps S1 to S5 of the invention as described above, we will describe in detail how the components of the system work together in the running state to execute the complete dynamic workflow from environment construction to emotional convergence.
[0024] In the initial phase of system startup, corresponding to step S1, physical devices are deployed in the real physical space and a mixed reality therapy space is constructed. First, staff guide supporters into a quiet room selected as the therapy space, where the aforementioned physical interaction devices are already in place. Supporters are assisted in wearing the mixed reality head-mounted display and adjusting it to a comfortable position. Once powered on, the device's depth sensor and inside-out tracking camera array immediately begin operation. The spatial mapping and localization module inside the central processing and rendering host receives continuous image frames and depth point cloud streams from the head-mounted device and initiates an instantaneous localization and mapping algorithm. This algorithm first extracts and matches robust visual feature points such as ORB or SuperPoint from the continuous image frames, while simultaneously using inertial measurement unit data for inter-frame pose prediction. Through iterative nearest point or global optimization methods based on sparse bundle adjustment, the depth point cloud of each frame is fused into a unified world coordinate system, generating a persistent 3D spatial map composed of voxels or triangular meshes with centimeter-level or even millimeter-level resolution. During this process, the 3D contours of the physical devices are accurately reconstructed within this spatial map. Next, the virtual-to-real rendering engine loads a pre-configured virtual scene containing natural elements and abstract geometric shapes. The engine parses this persistent spatial map, identifies the mesh regions corresponding to the surface of the physical device's shell, and performs precise registration and anchoring between the virtual scene and the physical environment. This anchoring process essentially involves solving for and continuously optimizing a rigid body transformation matrix, ensuring that the reprojection error term between the local coordinate system of specific interactive meshes representing visual elements such as "water surface," "light flow," and "particle fog" in the virtual scene and the coordinate system of the physical device's physical surface continuously approaches zero. When the supporter looks around, what they see through the holographic lens of their head-mounted device is a mixed reality therapy space within a familiar room, where virtual aurora or morning mist gently flows across the curved surface of the physical device, while the rest of the room remains visible.
[0025] Subsequently, the system seamlessly enters the security establishment and emotional activation phase described in step S2. The system's visual rendering direction within the mixed reality environment is limited to the local space anchored to the aforementioned physical device, and the view presented to the supporter by the head-mounted display is configured as a soft view. Specifically, within the mixed reality field of view, the rendering engine does not render any reference space indicators, virtual menus, bounding boxes, or arrows unrelated to the device, except for the virtual visual metaphors superimposed on the surface of the physical device and the device itself. The head-mounted display utilizes its inside-out tracking array and the image segmentation pipeline of the central processing unit to identify and mask the virtual avatar model of the user's hands in real time, minimizing visual contention and maintaining a pure, open visual experience. The audio guidance in this phase is entirely based on pre-recorded pure speech audio material stored on the content server, without containing any artificially synthesized speech. The system uses a spatialized speaker array on both sides of the head-mounted device to guide supporters to close their eyes or, while maintaining visual contact with the physical device, gradually adjust their sitting or semi-reclining posture, allowing their backs and arms to make extensive contact with the smooth curves of the device, until they find the most relaxed position. Simultaneously, the audio guidance begins to gently prompt open-ended questions, evoking specific memories related to caregiving stress, such as, "Try to recall an afternoon when you felt powerless or tired; what was the lighting like in the room then?"
[0026] During this psychological arousal process, the system performs real-time safety checks and closed-loop adjustments. A distributed sensor array within the physical device's casing captures multi-point pressure signals generated by the supporter's hand or arm contacting the curved surface at a sampling rate of 100 Hz. The microcontroller unit wirelessly transmits these signals to the central processing unit. The interaction state determination module continuously analyzes these pressure signals and hand position data transmitted from the head-mounted device. When a stable pressure signal is detected—that is, at least three sensor points return pressure values greater than 0.5 Newtons for a duration exceeding two seconds—the supporter is determined to be in a stable contact state. Simultaneously, the AI state assessment module within the main unit begins operation. This module estimates the supporter's real-time anxiety or tension level by analyzing features such as motion jitter frequency and the standard deviation of minute body sway amplitude extracted from the head-mounted device's inertial measurement unit and hand tracking data. Once stable contact is detected, the dynamic content management and scheduling engine adjusts the playback speed of the audio accordingly, fine-tuning its rate between 0.9x and 1.0x. Simultaneously, it instructs the graphics processing unit to adjust the transparency and color saturation of the virtual visual elements superimposed on the physical device's surface, gradually transitioning them from a slightly bright initial brightness to a softer, paler tone. For example, the particle speed of the virtual light flow is reduced by 15%, and the color changes from cool white to warm yellow. This entire process forms a closed loop: the supporter's stable physical contact is perceived by the system, which then provides softer, slower audiovisual feedback. This feedback, in turn, gradually enhances the supporter's sense of psychological security, creating a positive cycle.
[0027] After emotional activation, the system enters the narrative exploration and group connection phase described in step S3. The dynamic content management and scheduling engine of the central processing and rendering host retrieves an extremely detailed central virtual object model from the content server and precisely anchors it in the center of the curved surface of the physical device, approximately 0.6 meters directly in front of the supporter's line of sight, through the rendering engine. This central virtual object appears as a translucent cocoon-shaped luminous body surrounded by countless tiny light particles, with a warm, breathing-like pulse inside. The system invites the supporter through audio guidance to "reach out and touch the warm light in front of you."
[0028] At this point, the core computation of the system enters the processing logic of the embodied interaction technology module. The first processing layer of this module—the precise hand-to-surface contact detection layer—is fully activated. This layer executes a fusion judgment algorithm 60 times per second. The algorithm reads the three-dimensional position coordinates h of the key point at the center of the user's index finger or palm, calculated by the head-mounted device's inside-out tracking system, and calculates the Euclidean distance d(h,s) between it and the pre-anchored collision boundary mesh element s of the central virtual object. Simultaneously, the algorithm continuously polls the feedback signals fs of several specific force-sensitive resistor sensors located directly behind the virtual object within the physical device. The system sets two concurrent thresholds: a spatial distance threshold δ of 0.04 meters and a sensing signal threshold τ, which is the digital quantity converted from analog to digital corresponding to a pressure of 0.3 Newtons. The contact state is determined to be valid if and only if the conditions (d(h,s) < 0.04 meters) and (fs > τ) are met, i.e., the supporter's intentional limb contact is confirmed, rather than an unintentional wave.
[0029] Once the effective contact state is set to logical true by the first processing layer, the second processing layer—the synchronous multimodal feedback coupling layer—is immediately triggered. This layer is responsible for generating a set of highly coupled audiovisual tactile feedback. First, the central processing unit sends a brief tactile feedback command to the microcontroller unit within the physical device via a serial bus. The microcontroller unit then drives a linear resonant actuator or voice coil motor embedded in the back of the device housing to generate a gentle vibration lasting 100 milliseconds at a frequency of 150 Hz, simulating the physical sensation of "poking" or "touching." Completely synchronized with this physical vibration, the graphics processing unit immediately executes a pre-baked shader effect: the light particles of the central virtual object rapidly expand and split in all directions, decomposing into 5 to 7 smaller, floating, and differently shaped abstract anonymous narrative units, each resembling a translucent origami or leaf, slowly suspending at a spatial coordinate different from the curved surface of the physical device. Simultaneously, without frame-by-frame interval with the visual decomposition animation, a spatialized speaker array begins playing voice-over segments from different directions, processed with anonymized voiceprints. Each anonymous narrative unit is a spatialized audio source. Supporters can instinctively turn their heads or reach out to different narrative units. The head-tracking system of the mixed reality head-mounted display dynamically adjusts the audio's location and volume, making it sound as if each story is floating in different corners of space. Supporters freely choose to hover their hands over a narrative unit for more than 1.5 seconds; the system recognizes this as a selection and then plays a 20-40 second segment of life experiences and emotional struggles narrated by other anonymous supporters, clearly and without background interference. Through this precise coupling of physical touch, visual decomposition, and spatialized anonymous audio, the system practices the core processes of acceptance and commitment therapy, namely "cognitive dissociation" and "shared humanity."
[0030] The experience continues to the value-oriented guidance and immersive integration stage described in step S4. This stage is initiated by the dynamic content management and scheduling engine after the supporter has independently listened to several narrative segments, or after the system determines that the interaction activity has decreased. The engine first sends an instruction to the rendering host to reduce the transparency of all narrative units in the virtual space until they completely disappear, and then unfolds a grand, continuously evolving, customized audiovisual composition within the curved space of the entire physical device—a "digital audiovisual poem" that transforms abstract emotional states into dynamic visual metaphors. This process is based on a strict choreography timeline. The visual content is pre-rendered and stored in the content server as a high dynamic range video stream at 60 frames per second, and loaded into the texture buffer of the graphics processing unit via real-time streaming technology. The video content is presented as a series of immaterial visual metaphors: such as golden vines slowly growing on the surface of the physical device, symbolizing "patience"; ink-colored droplets dripping from the edge of the device and evaporating into mist, symbolizing "exhaustion"; and a stream of light constantly surging upwards from the bottom but never overflowing, symbolizing "tenacity." The audio tracks construct a complex, layered sound field within spatialized speakers, including ambient sounds at the bottom layer such as the wind in distant mountains or the sound of rain pattering on leaves, guiding narration in the middle layer, and rhythmic elements at the top layer simulating heartbeats or tides. The dynamic content management and scheduling engine's coordination control module precisely manages the volume envelope and crossfade-in / fade-out of these layered audio tracks, achieving smooth transitions.
[0031] At the end of this audiovisual ensemble, the engine seamlessly integrates resilient voice layers from the supporters. These voice layers are not independent audio files, but rather short phrases expressing "hope," "persistence," or "acceptance" extracted from multiple anonymous interviews by a lightweight AI-generated module. These phrases are then processed by a vocoder to generate a chorus that blends multiple vocal characteristics but makes individual voices unidentifiable, rendering it as a warm, spatialized sound emanating from all directions. Subsequently, the entire audiovisual ensemble fades, and the system proactively introduces a 120-second period of complete silence. During this time, only a faint, nebulous, slowly breathing monochromatic halo remains on the surface of the physical device within the mixed reality field of view. This period of silence aims to provide supporters with a cognitive space for introspection and value clarification, practicing the process of "being oneself in the scene." During the silent period, a lightweight adaptive modulation mechanism continues to run in the background, but its modulation behavior is constrained by extremely strict content change thresholds to ensure that narrative coherence and emotional continuity are not disrupted.
[0032] Finally, the system enters the procedural convergence and return to reality phase described in step S5. Thirty seconds before the end of the quiet period, the central processing unit triggers a progressive audio prompt sequence. A gentle voice softly speaks through the headset's speakers: "When you feel ready, take a slow, deep breath and bring your attention back to this space. When you are ready to leave, you can remove the headset." This prompt grants the supporter absolute autonomy to end the experience. Simultaneously, the system runs a background daemon for the emotional safety protection mechanism. This process continuously monitors the headset's status; if the device is detected as being removed directly by the user, all processes immediately terminate, and the safety protection mechanism responds normally. If the system detects that the supporter remains unresponsive for more than 300 seconds after the audio prompt, a progressive audio guidance degradation scheme is activated, changing the guidance from an interrogative tone to a more explicit, non-intrusive prompt. Ultimately, if necessary, a low-priority notification is sent to a professional in a nearby room, who then uses a gentle, manual guidance protocol to assist the supporter in removing the device and providing brief, non-judgmental reassurance.
[0033] After supporters remove their mixed reality headsets and complete a brief status check, staff provide them with a carefully designed guided reflection manual. Unlike traditional questionnaires, this manual contains several pages of illustrated writing prompts. For example, a large blank geometric shape or a softly colored area is printed in the center of each page, accompanied by handwritten prompts such as, "Write down the most touching moment of your experience in this space, using a sentence or a simple symbol." This allows supporters to transform the emotional fragments and inner clarifications gained from the experience into visual text or graphic symbols in a low-cognitive-load, non-verbal way, thus consolidating the initial results of emotional integration in the physical world after leaving the mixed reality environment.
[0034] Through the system architecture and workflow described in this embodiment, a therapeutic space deeply integrated with "physical anchoring and digital guidance" was successfully constructed. The physical device, through its preset thermal conductivity, surface roughness, and enveloping form, provides supporters with a continuous and tangible foundation of security at the tactile and proprioceptive levels, acting as a "safety anchor" in the emotional flow. This avoids the orientational obstacles and sense of disconnection from reality commonly found in pure virtual reality experiences, solving the problem of missing emotional anchors of security described in the background art. Distributed sensors and high-speed processing links establish a closed-loop channel with extremely low latency between physical touch and digital feedback. Simultaneously, the mixed reality digital layer, through spatial mapping, virtual-real fusion rendering, and a dynamic content engine, transforms the core processes of acceptance and commitment therapy—especially cognitive dissociation, shared humanity, and value clarification—into a narrative experience completed through embodied interaction with virtual-real fusion elements. In this embodiment, isolated and unspeakable emotional pressure is externalized into tangible and audible anonymous narrative units, which are then integrated into dynamic visual metaphors. Ultimately, guided by quiet reflection and the manual, it leads to value cognition defined by the supporter, thereby realizing a shift from passively receiving information to actively reconstructing meaning. Example 2
[0035] Based on the system architecture and interaction logic of Embodiment 1, this embodiment provides another preferred specific design scheme for aspects such as the specific implementation of physical interaction devices, the audiovisual guidance mode in the security establishment phase, and the narrative interaction granularity in the group connection phase. It aims to further enhance the physical immersion, emotional security boundaries, and depth of narrative exploration of the system.
[0036] Firstly, at the level of the physical structure of the system hardware, this embodiment deepens and modifies the design of the physical interaction device. Unlike the lightweight skeleton covered by a single-layer composite shell used in Embodiment 1, the physical device in this embodiment adopts a thick shell structure with multi-layer composite materials. The physical form of the device is also a single curved surface with a sense of enclosure, but to achieve better tactile transition and immersive experience, its shell is composed of three layers of functional materials from the outside to the inside. The outermost layer is a self-healing coating, about 0.5 mm thick, made of polyurethane-based material, which can self-repair within a few hours after minor scratches occur through the movement of polymer chain segments at room temperature, maintaining a smooth and flawless tactile feel. The middle layer is the main structural layer, about 15 mm thick, made of rigid polyurethane foam sculpted by a CNC machine tool into a curved shell. The inner wall of this layer has pre-fabricated grooves and wiring channels for embedding sensors and actuators, providing rigid support. The innermost layer is the interactive surface layer that the user directly contacts, about 5 mm thick, made of a silicone material filled with phase change microcapsules. The phase change material's phase change temperature is designed to be between 31 and 33 degrees Celsius, close to the surface temperature of human skin. When a supporter's arm or body comes into contact with this surface, the contact area rapidly rises from ambient temperature to around 32 degrees Celsius. Because the phase change material absorbs heat, it remains stable at this temperature plateau, providing a warm, tactile sensation that feels like body heat, surpassing the simple warmth of traditional materials. Simultaneously, distributed between the middle and innermost layers is a high-density sensor array consisting of 24 piezoelectric dynamic force sensors. Unlike simple force-sensitive resistors, these sensors can simultaneously capture the static amplitude and dynamic micro-vibrations of contact pressure at a sampling rate of 1 kHz, more precisely reflecting the supporter's emotional fluctuations during contact, such as subtle hand tremors caused by anxiety.
[0037] Secondly, in the security establishment and emotional activation phase described in step S2, this embodiment has implemented specific technical processing for the generation logic of "soft visuals" and "audio guidance" to enhance the gradual construction of psychological security. Building upon the open visuals of Embodiment 1, the rendering engine of this embodiment adds active optical flow suppression functionality for visual elements unrelated to the physical device. The texture details of the walls, floors, and furniture in the physical room environment seen by the supporter through the mixed reality head-mounted display are "soft-focused" through a real-time post-processing pipeline. Specifically, the rendering pipeline applies a variable Gaussian blur and hue / saturation reduction composite filter to all pixel areas in each frame of perspective view that do not overlap with the spatial grid of the physical device using a computational shader. The blur radius starts at an initial value of 3.0 pixels when the supporter puts on the device and smoothly interpolates to a maximum value of 8.0 pixels within 5 minutes of the emotional activation phase, based on the detected stable contact state. This causes the physical environment to visually fade into a blurry, relaxing patch of color, thereby minimizing the potential cognitive burden and visual competition caused by ambient light and object edges. Meanwhile, the audio guidance system in this embodiment introduces a spatialized sound anchoring technology. In addition to the central spatialized audio in Embodiment 1, the rendering engine virtually creates two static environmental sound sources that generate extremely low-frequency (40-60 Hz) sine waves in a specific location within the virtual space—approximately 1 meter behind and to the left of the physical device. These two low-frequency sound sources do not contain any melody or rhythm, but their presence creates a subconscious "acoustic boundary" or "acoustic embrace" for the entire sound field. When the supporter feels pressured during recollection, the system slightly increases the gain of these two low-frequency sound sources disproportionately based on the anxiety score assessed in real-time, creating a psychological suggestion that the environment is physically "moving closer" to provide support. This is a non-verbal, non-invasive emotional soothing technique.
[0038] Furthermore, in the narrative exploration and group connection phase of step S3, this embodiment introduces a multi-granularity interaction and dynamic narrative branching mechanism. In Embodiment 1, supporters can only freely choose which anonymous narrative units to listen to. In this embodiment, each anonymous narrative unit is designed as a "narrative sphere" carrying hierarchical content. When a supporter first selects a narrative sphere by hovering their hand, the mixed reality head-mounted display plays a "core emotional summary" of the narrative, a concise sentence of about 8 to 12 seconds containing the most core emotional impact, such as "I often cry alone at night because I feel like a failed mother." The system uses high-precision hand tracking to identify the supporter's subconscious reaction after hearing this summary, such as whether the hand quickly retracts, pauses, or moves closer to touch it. The built-in micro-decision tree model of the dynamic content management and scheduling engine judges the supporter's willingness to accept based on this behavior: if hesitation or avoidance is detected, the narrative sphere gently expands the spatial trigger area for playing the next complete segment, and soft, reassuring text guidance appears on the sphere's surface, such as "You don't have to bear all this alone"; if active touching is detected, the narrative sphere immediately decomposes and plays a more detailed 60-second, more comprehensive, in-depth voice narration that includes the narrator's specific situation, psychological struggle, and a moment of subtle change. This dynamic narrative depth adjustment based on real-time behavioral signals endows the system with higher interactive intelligence, enabling it to sensitively respond to the supporter's psychological defense boundaries and maximize the emotional healing and dissociation effects of collective narratives while avoiding secondary trauma.
[0039] Finally, in the value-oriented guidance phase of step S4, this embodiment enhances the generation mechanism of dynamic visual metaphors. In addition to playing a pre-made abstract video stream, the graphics processing unit utilizes a lightweight generative adversarial network inference model to sample a 2D noise field synchronized with the ambient background rhythm in real time. This noise field is then mapped onto the emissivity, lifetime, and color gradient of the particle system, thereby generating layered, unique, and never-repeating light and shadow brushstroke effects on specific areas of the physical device's surface. For example, when expressing the metaphor of "resilience," the model generates a cluster of digital silver vines that continuously grow upwards from the bottom of the device, with its branching forms spontaneously changing but its overall upward trend remaining constant. This non-repetitive, slightly random generative visual enhances the immersiveness and uniqueness of the experience, making each therapy session an unrepeatable artistic existence that resonates with the supporter's current state, thus more profoundly supporting its value reconstruction. Example 3
[0040] This embodiment further provides specific technical selections and operation procedures for the system in terms of communication architecture, user status assessment algorithm, and collective narrative content storage and distribution. It also describes an extended application scenario for long-distance home support to demonstrate the extended capabilities of the technical solution of the present invention in terms of networking and intelligence.
[0041] In terms of communication architecture, this embodiment decouples the central processing host function from Embodiment 1. The mixed reality head-mounted display device no longer connects directly to the local workstation, but instead connects to an edge computing node located in the venue via its built-in multi-mode wireless communication module supporting Wi-Fi 6E and Bluetooth 5.2. This edge computing node itself is a miniaturized edge server with a passive cooling design, integrating a dedicated AI inference chip, such as NVIDIA Jetson Orin or a similar embedded artificial intelligence computing platform. This node is responsible for handling all real-time tasks that are extremely sensitive to latency, including spatial mapping updates, fusion of hand tracking data, and determination of interaction states. Tasks with higher computing power requirements and slightly higher latency tolerance, such as complex generative dynamic visual metaphor rendering, large-scale high-fidelity audio streaming decoding, and deep learning training of novel narrative content, are offloaded to a central computing cluster deployed in a regional data center or cloud via a secure socket-encrypted gigabit fiber optic link. This "end-edge-cloud" three-level collaborative architecture not only reduces the equipment cost and maintenance complexity of individual supporter locations, but also enables the system to aggregate anonymized behavioral and status data from multiple user endpoints. It then performs federated learning-style continuous optimization of the lightweight AI adaptation and modulation model in the cloud, thereby allowing the overall system's emotional regulation intelligence level to increase in tandem with the number of users served without leaking individual privacy.
[0042] Regarding the real-time user state assessment algorithm, this embodiment provides a more specific engineering definition and calculation of the state quantity St mentioned in step S4. The state quantity St is no longer merely a primitive physical quantity such as spatial location, velocity, and intensity, but a higher-dimensional emotional state feature vector calculated by the feature extraction pipeline on the edge computing nodes. When the supporter enters the immersion integration phase, the feature extraction pipeline calculates the following quantitative indicators within a continuous 30-second time window: spatial location stability, i.e., the 95% confidence ellipsoidal volume of the supporter's head trajectory; interaction fluidity, i.e., the LZ complexity of the supporter's unconscious hand movements in space; and the physiological synchronization index, which obtains heart rate variability data through a reflective photoelectric pulse wave sensor integrated near the nose pad inside the head-mounted device and calculates its normalized high-frequency power. These three indicators are then input into a lightweight gradient boosting decision tree model pre-trained with the supporter's physiological data, outputting a continuous value, namely the "emotional immersion depth index." The content modulation mechanism reads this index. When the index is low, the system fine-tunes the dynamic contrast of the visual metaphor; when the index is within the ideal range, the system remains stable; when the index is too high and drifts towards anxiety, the system immediately shifts the visual tone to a more soothing warm tone and reduces particle speed. The entire modulation process strictly follows the formula. The constraints ensure that the overall perceived difference between the newly generated content Ct and the baseline content Cbase is always less than an acceptable threshold, which is preset to 0.15 in this embodiment.
[0043] Regarding the storage and distribution of collective narrative content, this embodiment provides a specific anonymization review and content library construction process. When the system wants to include new supporter narratives, it can collect their spoken words through an independent recording module with their informed consent. The audio file first passes through a speech activity detection and signal-to-noise ratio filtering module to filter out silence and low-quality segments. Then, it undergoes voiceprint separation and resynthesis processing via a parameterized vocoder. This processing decouples the biometric features of the original speech, such as timbre, intonation, and rhythm, and replaces them with a general voiceprint base randomly selected from a preset voiceprint library, introducing 0.5% random pitch jitter to fundamentally eliminate the possibility of identifying the speaker through speech. The processed audio is automatically transcribed into text, and a sensitive information filtering model detects and removes any information that may reveal the speaker's identity, such as place names, personal names, or specific organization names, while retaining the core content of emotional expression. The anonymized audio that passes the review is stored along with corresponding text tags in the solid-state drive array of the content server, and is tagged with the date, preset topic tags (such as "tired," "blame," "hope"), and voiceprint type. When subsequent supporters enter the group connection phase, the dynamic content management and scheduling engine intelligently recalls and arranges a set of anonymous stories with the greatest potential for resonance from the content library based on the visual theme selected by the supporters in the early stages of the experience and their current state. This achieves a continuously growing, self-enriching, and absolutely anonymous collective narrative ecosystem network. Through the extended design of this embodiment, the system of the present invention is not limited to single-point healing scenarios, but also possesses the technical foundation to evolve into a distributed, adaptive, and continuously evolving daily emotional repair platform.
Claims
1. An emotional healing method for supporters of attention deficit hyperactivity disorder based on mixed reality, characterized in that, Includes the following steps: S1: Deploy a physical device with a preset curved shape and tactile feedback characteristics in a real physical space, and overlay virtual visual elements onto the physical device using a mixed reality head-mounted display to construct a therapeutic space that integrates physical anchoring and digital guidance; S2: Guide the supporter to interact tactilely with the physical device under the visual content and audio guidance presented by the mixed reality head-mounted display, performing the safety establishment and emotional activation phase. The system renders a soft visual scene without specific spatial reference cues in the mixed reality environment, and evokes memories related to caregiving stress in the supporter solely through audio commands, while simultaneously prompting the supporter to adjust their body posture to a comfortable state; S3: Enter narrative exploration and group... In the body connection phase, the system presents a central virtual object in the mixed reality field of view. When the supporter triggers the central virtual object through physical contact, the central virtual object decomposes into multiple anonymous narrative units. Each anonymous narrative unit carries a voice self-narration fragment anonymized by other supporters. Supporters can freely choose and listen to the voice self-narration fragment. S4: Entering the value-oriented guidance and immersive integration phase, the system plays a customized audiovisual combination that transforms abstract emotional states into dynamic visual metaphors in the supporter's mixed reality field of view. At the end of the customized audiovisual combination, a resilient voice layer from the supporter group is seamlessly integrated. Then, a quiet period is introduced for supporters to conduct introspective reflection and value clarification. S5: Execute the procedural convergence and return to reality phase. The system empowers the supporter with the autonomy to end the experience through progressive voice prompts. After the supporter removes the mixed reality head-mounted display, a guided reflection manual is provided to consolidate the emotional integration results.
2. The method according to claim 1, characterized in that, Step S2 includes: the mixed reality head-mounted display device uses its depth sensor and real-time localization and mapping algorithm to perform 3D reconstruction of the physical interaction space and generate a persistent spatial map during the initialization phase; the virtual and real fusion rendering engine, based on the persistent spatial map, precisely registers and anchors the pre-configured virtual scene with the physical device, so that the virtual interactive elements and the physical surface of the physical device achieve a consistent spatial correspondence; the stable contact between the supporter and the surface of the physical device is detected in real time, and the system adjusts the softness of the visual feedback and the speech rate of the audio guidance in the virtual environment according to the contact state.
3. The method according to claim 1, characterized in that, The visual content within the mixed reality field of view consists only of a visual metaphor system based on natural elements and abstract geometric forms. The themes of this visual metaphor system are chosen autonomously by the supporter at the beginning of the experience. The audio guidance is pure voice guidance based entirely on pre-recorded audio material, without any artificially synthesized voice or real-time instructor interjections.
4. The method according to claim 1, characterized in that, In step S3, the process of a supporter triggering the central virtual object through physical contact relies on an embodied interaction technology module based on physical object anchoring. This module includes a first processing layer and a second processing layer. The first processing layer is a precise hand-to-surface contact detection layer. By fusing inside-out tracking data from the mixed reality head-mounted display device with sensor feedback signals sparsely distributed within the physical device, it determines whether the supporter has made intentional contact with a specific surface area. When the spatial distance between the hand position and the surface area is less than a preset distance threshold and the sensor feedback signal exceeds a preset signal threshold, the contact state is determined to be valid. The second processing layer is a synchronous multimodal feedback coupling layer. When the first processing layer determines that the contact is valid, the second processing layer immediately triggers visual changes and audio responses that are closely coupled with the tactile properties of the material of the physical device. The visual changes are manifested as the decomposition animation of the central virtual object into multiple narrative units, and the audio response is the anonymous voice narration playback that starts synchronously with the visual decomposition.
5. The method according to claim 1, characterized in that, In step S4, the abstract emotional state is transformed into a customized audiovisual combination of dynamic visual metaphors, which is generated under the control of a dynamic content management and scheduling engine. The engine orchestrates the delivery of media assets based on the supporter's experience process and real-time behavioral signals, including asynchronous playback of anonymous audio narratives, real-time streaming of abstract visual content, and coordinated control and smooth transition of ambient sound, guiding narration, and hierarchical audio. The engine includes a lightweight adaptive modulation mechanism that reads state variables representing the supporter's state, which consist of the supporter's spatial location, interaction speed, and interaction intensity, and modulates the audiovisual content according to the state variables. The modulation process is constrained by a content change threshold, ensuring that the deviation of the content from the baseline value remains within the preset threshold range.
6. The method according to claim 1, characterized in that, In step S5, the system supports multiple emotional safety protection mechanisms, including supporters being able to exit immediately by directly removing the mixed reality head-mounted display device, the system initiating a progressive audio-guided degradation when it detects that the supporter is continuously unresponsive, and the supporter can choose to accept the intervention of assistants; the guided reflection manual includes writing pages to guide supporters to transform emotional fragments in the experience into visual text or graphic symbols.
7. The method according to claim 1, characterized in that, The physical device is a single curved surface with an enveloping feel, and its surface material has a preset thermal conductivity and surface roughness; the physical device serves as the only continuously tactile object throughout the entire treatment process.
8. The method according to claim 1, characterized in that, The soft visual experience in step S2 specifically refers to the spacious visual experience in a mixed reality environment where no reference space markers or virtual boundaries unrelated to the physical device are superimposed.
9. The method according to claim 2, characterized in that, A distributed sensor array is embedded in the interlayer between the inner surface of the physical device housing and the outer covering material. The sensor array transmits sensor data packets to the central processing host via wireless communication. The real-time detection of stable contact with the surface of the physical device in step S2 is achieved by capturing multi-point pressure signals from the distributed sensor array. When at least a preset number of sensor points return pressure values greater than a preset pressure threshold and the duration exceeds a preset time, it is determined that the supporter is in a stable contact state.
10. An emotional healing system for supporters of attention deficit hyperactivity disorder based on mixed reality, characterized in that, The system utilizes the method described in any one of claims 1 to 9.