An embedded data query and multi-modal interaction-based childcare meta-universe simulation system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本发明的目的在于提供一种基于嵌入式数据查询与多模态交互的托育元宇宙仿真系统,以解决上述背景技术中提出的单一接触阈值参数无法捕捉非标动作的连续物理特征变化,造成系统输出的实训诊断评价与真实规范发生严重偏移的问题
1、该基于嵌入式数据查询与多模态交互的托育元宇宙仿真系统中,针对现有系统依赖单一接触阈值导致无法评价非标动作的缺陷,本发明通过提取高频三轴加速度物理流中的瞬态动能衰减梯度,对目标虚拟对象的初始碰撞盒执行法向内陷的弹性映射,生成柔性物理边界张量;能够高保真地量化受训者在拍嗝等操作中的“卸力缓冲”与“按压深度”连续物理特征;使得系统能精准判别动作是否满足托育医疗规范中的柔和要求,从根源上消除了实训诊断评价与真实岗位标准之间的严重偏移。
Smart Images

Figure CN122548969A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of virtual reality simulation technology, and more specifically, to a childcare metaverse simulation system based on embedded data query and multimodal interaction. Background Technology
[0002] With the development of virtual reality and digital twin technologies, metaverse simulation systems have been introduced into the professional training system for childcare services. They are used to conduct practical exercises in high-risk or rare scenarios (such as sudden illness of infants and young children, loss of control of group behavior, etc.). A typical metaverse simulation system collects the limb movement and voice information of trainees through wearable sensors, and compares the above physical interaction quantities (the limb movement and voice information of trainees) with the built-in childcare standard database to drive the three-dimensional virtual infant entity to perform corresponding physiological state feedback and evaluation.
[0003] The effectiveness determination of existing multimodal interactions generally relies on discrete, single, rigid physical parameters (such as triggering "contact / non-contact" state transitions via spatial collision boxes). However, in childcare tasks such as emotional soothing, feeding, and burping, operational compliance is built upon the coordination of continuous temporal actions (such as patting frequency and force attenuation) and associated modalities (the acoustic envelope of soothing sounds). Traditional single contact threshold parameters cannot capture the continuous physical characteristic changes of non-standard actions, causing the system to fail to accurately map and correct the multimodal temporal sequences extracted from the physical space to the industry standard job baseline requirements. This results in a significant deviation between the system's output training diagnostic evaluation and actual standards. Therefore, there is an urgent need for a childcare metaverse simulation system based on embedded data querying and multimodal interaction to overcome the aforementioned physical communication bottlenecks and parameter evaluation errors. Summary of the Invention
[0004] The purpose of this invention is to provide a childcare metaverse simulation system based on embedded data query and multimodal interaction, in order to solve the problem mentioned in the background art that the single contact threshold parameter cannot capture the continuous physical characteristic changes of non-standard actions, resulting in a serious deviation between the system output training diagnosis evaluation and the real standard.
[0005] To achieve the above objectives, the present invention aims to provide a childcare metaverse simulation system based on embedded data query and multimodal interaction. This childcare metaverse simulation system is deployed on a simulation terminal equipped with a local graph database and a graphics rendering engine, and is communicatively connected to a dual-track reference data base. The dual-track reference data base pre-loads a multimodal interaction reference dataset and initial collision box volume parameters of the target virtual object. The childcare metaverse simulation system includes: The multimodal acquisition unit is used to acquire the absolute spatial coordinates and view frustum deflection vector of the trainee, and simultaneously acquire the three-dimensional coordinate sequence, the three-axis acceleration physical flow, and the concurrent audio Mel-Cepstral feature sequence. An embedded pre-addressing unit generates a pre-addressing index vector based on the absolute spatial coordinates and the view frustum deflection angle vector. It then performs addressing in the local map database according to the pre-addressing index vector and loads the matched multimodal interaction benchmark dataset and the initial collision box volume parameters into the local process memory stack. The kinetic energy correction and feature fusion unit is used to perform time-domain difference decomposition on the triaxial acceleration physical flow to calculate the transient kinetic energy decay gradient, and to perform elastic mapping on the initial collision box volume parameters based on the transient kinetic energy decay gradient to generate a flexible physical boundary tensor. Furthermore, when it is determined that the three-dimensional coordinate sequence intrudes into the flexible physical boundary tensor surface and reaches the extreme intrusion depth, the transient kinetic energy decay gradient, extreme intrusion depth and audio Mel-frequency cepstral feature sequence within the same underlying clock cycle are extracted and spliced together to generate an operation behavior feature matrix. The deviation comparison rendering unit is used to compare the multimodal interaction benchmark dataset with the operation behavior feature matrix and generate a physical operation deviation index. The physical operation deviation index is mapped to the rendering pipeline timing compensation frame number, and the graphics rendering engine is rewritten at the hardware level to drive the mesh deformation and audio synchronization update of the target virtual object.
[0006] As a further improvement to this technical solution, the multimodal acquisition unit includes a spatial dynamics capture module and a visual cone acoustic synchronous tracking module; The spatial dynamics capture module is used to collect the absolute spatial coordinates of the trainee and extract the three-dimensional coordinate sequence of the joints of the upper limb bones and the three-axis acceleration physical flow at the end of the force application. The visual cone acoustic synchronous tracking module acquires concurrent audio Mel-Cepstral feature sequences through a hardware pickup array and extracts the visual cone deflection angle vector based on the head-mounted display device.
[0007] As a further improvement to this technical solution, the triaxial acceleration physical flow is a discrete time series matrix with the physical reference timestamp as the independent variable. The data elements of the discrete time series matrix include transient acceleration component parameters and resultant acceleration dynamic amplitude. The frustum deflection angle vector includes the origin coordinates of the viewpoint calculated from the trainee's head-mounted display device, the central optical axis direction vector mapped by the attitude Euler angles, and the effective field of view boundary parameters.
[0008] As a further improvement to this technical solution, the embedded front-end addressing unit includes a line-of-sight priority correction module and a memory front-end loading module; The view distance priority correction module is used to obtain the initial Euclidean distance scalar between the absolute spatial coordinates and the three-dimensional anchor point of the target virtual object, and introduces the view dwell time integral extracted based on the view frustum deflection angle vector. The view dwell time integral is used as a nonlinear weight to correct the priority of the initial Euclidean distance scalar and generate the pre-addressing index vector. The memory preloading module addresses the local graph database based on the pre-addressing index vector and preloads the matched multimodal interaction benchmark dataset and initial collision box volume parameters into the local process memory stack of the simulation terminal.
[0009] As a further improvement to this technical solution, the line-of-sight priority correction module performs priority correction on the initial Euclidean distance scalar to generate a pre-addressing index vector. The specific steps involved are as follows: A ray bounding box is constructed based on the absolute spatial coordinates and the view frustum deflection angle vector. The spatial geometry falling within the ray bounding box is resolved into the currently observed view frustum, and the three-dimensional anchor point of the target virtual object inside the currently observed view frustum is locked. Calculate the physical space straight-line geometric distance between the absolute spatial coordinates and the three-dimensional anchor point, and generate an initial Euclidean distance scalar that is negatively correlated with the distance value; Within the current observation frustum, the angle between the central optical axis of the frustum deflection angle vector and the three-dimensional anchor point is within a preset effective field of view deflection threshold for a continuous hardware clock cycle. A time-domain integration operation is performed on the continuous hardware clock cycle to generate the line-of-sight dwell time integral. Construct a priority correction function with the initial Euclidean distance scalar as the base and the line-of-sight dwell time integral as the correction exponent, perform exponential weighted reconstruction on the initial Euclidean distance scalar, and output a pre-addressing index vector representing the urgency of data addressing.
[0010] As a further improvement to this technical solution, the kinetic energy correction and feature fusion unit includes a flexible boundary mapping module and a feature matrix splicing module; The flexible boundary mapping module is used to perform first-order time-domain difference on the triaxial acceleration physical flow to extract the transient kinetic energy decay gradient characterizing the contact buffer features, and to perform indentation mapping along the surface normal on the initial collision box volume parameters using the transient kinetic energy decay gradient as the deformation excitation variable to generate a flexible physical boundary tensor. The feature matrix splicing module is used to monitor the spatial interference state between the three-dimensional coordinate sequence acquired by the multimodal acquisition unit and the flexible physical boundary tensor, determine the extreme point trigger of the physical intrusion of the three-dimensional coordinate sequence and lock the hardware interrupt timestamp; call the hardware interrupt timestamp to extract the transient kinetic energy decay gradient within the same underlying clock cycle, and perform memory splicing with the audio Mel-frequency cepstral feature sequence to output the operation behavior feature matrix.
[0011] As a further improvement to this technical solution, the flexible boundary mapping module generates a flexible physical boundary tensor, involving the following specific steps: The transient acceleration component parameters are extracted from the triaxial acceleration physical flow, and a first-order time-domain difference operation is performed on the transient acceleration component parameters to solve the rate of change of acceleration at the moment of contact. The negative extremum in the scalar of the rate of change of acceleration is analyzed as the transient kinetic energy decay gradient. Access the local process memory stack, extract the initial collision box volume parameters corresponding to the interaction area of the target virtual object, and parse the three-dimensional mesh vertex clusters and corresponding face normal vectors that constitute the initial collision box volume parameters; A mesh elastic deformation function is constructed with the transient kinetic energy decay gradient as the core independent variable. The mesh elastic deformation function is called to calculate the mesh vertex offset. The three-dimensional mesh vertex cluster is driven to perform a depth translation in the mesh along its corresponding surface normal vector. The translated and reconstructed three-dimensional mesh dataset is encapsulated and output as a flexible physical boundary tensor representing the surface physical pressure rebound depth.
[0012] As a further improvement to this technical solution, the feature matrix splicing module performs the following steps to output the operation behavior feature matrix: The three-dimensional coordinate sequence is imported into the bounding box collision detection pipeline of the simulation terminal, and the relative spatial position relationship between the physical force application end represented by the three-dimensional coordinate sequence and the outer surface of the flexible physical boundary tensor is calculated in real time. When it is determined that the spatial coordinates of the physical force application end penetrate the flexible physical boundary tensor surface and reach the extreme penetration depth, a high-level transition signal is sent to the interrupt request pin of the central processing unit to forcibly suspend the non-concurrent polling thread, and the underlying clock scheduler forcibly generates an absolute timing hardware interrupt timestamp. Create an extremely narrow aligned window in physical memory with the hardware interrupt timestamp as the addressing index key; Access the direct memory access cache configured at the bottom layer, extract the audio Mel-Cepstral feature sequence and the corresponding transient kinetic energy decay gradient that fall within the extremely narrow alignment window, perform multidimensional tensor fusion calculation, and concatenate the audio Mel-Cepstral feature sequence, the transient kinetic energy decay gradient and the extreme value intrusion depth on continuous physical memory addresses to generate and output an operational behavior feature matrix that characterizes the concurrent features of acoustic expression and dynamic pressure under the current spatiotemporal node.
[0013] As a further improvement to this technical solution, the deviation comparison rendering unit includes a dual-track benchmark comparison module and a timing compensation direct drive module. The dual-track benchmark comparison module is used to parse the multimodal interaction benchmark dataset from the local process memory stack and decompose it into explicit dynamic physical benchmark track and implicit cognitive and acoustic interaction benchmark track. The operation behavior feature matrix is then imported into the explicit dynamic physical benchmark track and the implicit cognitive and acoustic interaction benchmark track respectively to perform heterogeneous distance calculation and generate a quantified physical operation deviation index. The timing compensation direct drive module is used to receive the physical operation deviation index, map it to the rendering pipeline timing compensation frame number through a preset nonlinear compensation function; call the rendering pipeline timing compensation frame number to perform hardware-level interruption and state rewriting of the physics engine and audio mixer in the graphics rendering engine, and drive the execution of the mesh deformation and audio synchronization update of the target virtual object.
[0014] As a further improvement to this technical solution, the dual-track benchmark comparison module performs the following steps to generate the physical operation deviation index: The preset standard buffer gradient scalar and standard safety intrusion limit are analyzed and extracted from the explicit dynamic physical reference track, and the preset standard soothing voiceprint reference tensor is analyzed and extracted from the implicit cognitive and acoustic interaction reference track. The operational behavior feature matrix is decomposed into its lowest-level dimensions to extract the transient kinetic energy decay gradient and extreme intrusion depth actually triggered by the trainee. The first Euclidean distance between the transient kinetic energy decay gradient and the standard buffer gradient scalar, and the second Euclidean distance between the extreme intrusion depth and the standard safe intrusion limit are calculated. A weighted summation operation is performed on the first Euclidean distance and the second Euclidean distance to generate a dynamic deviation sub-index; The audio Mel-Cepstral feature sequence actually triggered by the trainee is extracted from the operational behavior feature matrix in a synchronous manner. The cosine similarity algorithm is used to calculate the angle between the vectors of the audio Mel-Cepstral feature sequence and the standard soothing voiceprint reference tensor in the multidimensional acoustic space, and to generate an acoustic deviation feature sub-index. A preset global penalty leveling coefficient is introduced, and a cross-dimensional linear fusion calculation is performed on the dynamic deviation sub-index and the acoustic deviation feature sub-index. The fused result is then encapsulated into a single scalar physical operation deviation index.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. In this childcare metaverse simulation system based on embedded data query and multimodal interaction, addressing the shortcomings of existing systems that rely on a single contact threshold and thus cannot evaluate non-standard movements, this invention extracts the transient kinetic energy decay gradient in the high-frequency triaxial acceleration physical flow and performs a normal indentation elastic mapping on the initial collision box of the target virtual object to generate a flexible physical boundary tensor. This enables high-fidelity quantification of the continuous physical characteristics of "force relief buffering" and "pressure depth" of trainees in operations such as burping. This allows the system to accurately determine whether the movements meet the gentle requirements of childcare medical standards, fundamentally eliminating the serious deviation between practical training diagnosis and evaluation and real job standards.
[0016] 2. In this childcare metaverse simulation system based on embedded data query and multimodal interaction, the present invention addresses the pain point of audio-visual asynchrony caused by the heterogeneity of the underlying clock of multimodal devices (such as motion capture and voice pickup). The invention triggers the underlying hardware-level interrupt timestamp at the moment of physical extreme value intrusion, and forces the underlying memory splicing of dynamic, geometric and acoustic matrices within an extremely narrow physical alignment window. Furthermore, it maps the multimodal comparison deviation to the rendering pipeline timing compensation frame number, and directly rewrites the GPU graphics shader and DAC sound card ring buffer state through the hardware bus. This achieves absolute spatiotemporal synchronization feedback from "trainee deviation behavior" to "body deformation vision" and "crying and soothing hearing", completely eliminating sensory phase distortion in VR medical training.
[0017] 3. In this childcare metaverse simulation system based on embedded data query and multimodal interaction, to address the network addressing lag and frame drops caused by inconsistent hand-eye displacement in emergency scenarios (such as handling spit-up emergency care), this invention constructs a nonlinear priority correction function by introducing the trainee's visual cone deflection angle vector and the integral of gaze dwell time. Within the millisecond-level time window when the trainee's gaze is focused on the sudden lesion but the hand has not yet reached it, the dual-track reference data is forcibly copied to the local process memory stack in advance via a direct memory access channel in block transfer mode. This achieves zero-network-dependent local data retrieval within seconds, ensuring the high smoothness of multimodal computing power allocation and rendering pipeline during high-risk emergency training. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the overall process of the present invention.
[0019] The meanings of the labels in the diagram are as follows: 1. Multimodal acquisition unit; 11. Spatial dynamics capture module; 12. View cone acoustic synchronous tracking module; 2. Embedded front-end addressing unit; 21. Line-of-sight priority correction module; 22. Memory front-end loading module; 3. Kinetic energy correction and feature fusion unit; 31. Flexible boundary mapping module; 32. Feature matrix splicing module; 4. Deviation comparison rendering unit; 41. Dual-track benchmark comparison module; 42. Timing compensation direct drive module. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0021] Please see Figure 1 As shown, a childcare metaverse simulation system based on embedded data query and multimodal interaction is provided. The childcare metaverse simulation system is deployed on a simulation terminal equipped with a local graph database and a graphics rendering engine, and is connected to a dual-track reference data base. The dual-track reference data base is pre-set with a multimodal interactive reference dataset for characterizing standardized childcare operation behavior, and initial collision box volume parameters for characterizing the original physical boundary of the target virtual object. The dual-track benchmark data base is a multimodal reference system storage architecture built on a graph database for non-standardized childcare scenarios. It includes two sets of underlying benchmark architectures that are stored in parallel and strictly aligned in timestamps: The first track is the explicit dynamic physical reference track, which encapsulates the dynamic characteristic tensor collected and standardized by authoritative medical and childcare entities, including but not limited to the extreme boundary of the compliant slapping impulse and the standard curve of the kinetic energy decay gradient of the gentle touch. The second track is the implicit cognition and acoustic interaction baseline track, which encapsulates continuous baseline data representing the trainee's attention focus and soothing emotions, including effective visual cone coverage and a standard array of soothing speech acoustic features. The two sets of data are cross-mapped within the nodes of the graph database, together forming a multimodal interactive benchmark dataset. Simultaneously, the static storage area of this base also stores the initial collision box volume parameters representing the original three-dimensional geometric boundaries of various virtual objects (such as different physiological parts of virtual infants at different ages). During the system's runtime, this dual-track benchmark data base serves as the sole benchmark, providing zero-latency data reference for the dynamic correction of pre-addressing and physical collision boundaries through underlying scheduling mechanisms such as Direct Memory Access (DMA).
[0022] Furthermore, the childcare metaverse simulation system includes a multimodal acquisition unit 1, which is used to acquire the absolute spatial coordinates and visual cone deflection angle vector of the trainee, and simultaneously acquire three-dimensional coordinate sequences, three-axis acceleration physical flow, and concurrent audio Mel-Cepstral feature sequences. The multimodal acquisition unit 1 includes a spatial dynamics capture module 11 and a visual cone acoustic synchronous tracking module 12. The spatial dynamics capture module 11 is used to collect the absolute spatial coordinates of the trainee and extract the three-dimensional coordinate sequence of the joints of the upper limb bones and the three-axis acceleration physical flow at the end of the force application. The visual cone acoustic synchronous tracking module 12 acquires concurrent audio Mel-Cepstral feature sequences through a hardware pickup array and extracts the visual cone deflection angle vector based on the head-mounted display device.
[0023] In this embodiment, the triaxial acceleration physical flow is a discrete time series matrix with the physical reference timestamp as the independent variable. The data elements of the discrete time series matrix include spatially orthogonal transient acceleration component parameters and the resultant acceleration dynamic amplitude generated by the real-time superposition of the transient acceleration component parameters. The frustum deflection vector includes the origin coordinates of the viewpoint calculated from the trainee's head-mounted display device, the central optical axis direction vector mapped by the attitude Euler angles, and the effective field-of-view boundary parameters constructed based on the central optical axis direction vector and the lens optical field angle.
[0024] In this embodiment, when performing burping and emergency treatment for sudden spitting up of infants, trainees are required to observe the mouth and nose area of the virtual infant in real time while performing continuous burping to determine whether there are signs of spitting up, and at the same time issue a calm and soothing voice.
[0025] Specifically, in the burping training, in order to capture the kinetic energy characteristics of the trainee's palm touching the back of the virtual infant, the spatial dynamics capture module 11 tracks the infrared reflective markers worn by the trainee through an optical motion capture base station (sampling rate of 120Hz) deployed in the training site; and establishes a global world coordinate system with the center of the site as the origin to obtain the absolute spatial coordinates of the trainee's center of gravity. Furthermore, a global world coordinate system is established with the center of the training ground as the origin, and multiple sets of optical marker points attached to the skin surface of the shoulder, elbow and wrist joints of the trainees are collected in real time. Through high-frequency optical motion capture base stations deployed in the training ground, multiple infrared cameras perform multi-view visual spatial triangulation to directly calculate the absolute physical displacement trajectory of the center points of the three physical joints of the shoulder, elbow and wrist in the global world coordinate system, and perform time-series sampling at a frame rate of 120Hz to generate a three-dimensional coordinate sequence for determining the spatial posture of the limbs. To address the limitations of traditional optical cameras, which are restricted by sampling frame rate and physical obstruction and can only determine whether a hand is in contact with the back but cannot quantify the transient impact force, a hardware acquisition mechanism based on underlying inertial dynamics is adopted. The spatial dynamics capture module 11 calls the inertial measurement unit (IMU) sewn into the palm of the trainee's training glove. When the trainee performs the "burping" action, the IMU continuously collects the three-axis instantaneous acceleration values in the local coordinate system of the palm at a high-frequency hardware clock of 200Hz. When the trainee's palm swings towards the back of the virtual infant and actively decelerates or rebounds at the moment of simulated contact, the IMU will inevitably output a significant reverse acceleration impulse (i.e., physical extreme peak) on the Z-axis (the normal axis perpendicular to the plane of the palm). The system directly extracts the absolute value of this acceleration peak (i.e., the resultant acceleration dynamic amplitude) and the numerical decay slope around the peak (i.e., the transient kinetic energy decay gradient). The spatial dynamics capture module 11 then collects the original acceleration along the three axes at a frequency of 200Hz. Considering that human muscles will produce natural vibrations of 8Hz to 12Hz when exerting force, the module is equipped with a low-pass filter with a cutoff frequency of 5Hz to smooth the original signal and output a three-axis acceleration physical flow.
[0026] In the emergency milk spillage handling process, the visual cone acoustic synchronous tracking module 12 reads the quaternion data of the trainee's VR headset and maps it to a visual cone deflection angle vector. The visual cone deflection angle vector includes the coordinates of the viewpoint origin and the direction vector of the central optical axis. Combined with the inherent 110-degree effective field of view boundary parameter of the headset, a solidified visual cone (prism field of view) geometric boundary moving in three-dimensional space is constructed. This is used to perform the underlying spatial bounding box intersection detection to objectively define and eliminate the trainee's physical observation blind spots, and to provide a rigorous three-dimensional geometric constraint basis for accurately determining whether the milk spillage lesion actually falls into the effective attention field of view. Meanwhile, to capture the trainee's soothing voice, a hardware pickup array is activated. To prevent environmental noise interference and ensure multimodal data alignment, the module employs a Voice Activity Detection (VAD) hardware circuit. When the trainee emits a soothing voice and the sound pressure level exceeds 60dB, the VAD circuit sends a microsecond-level hardware interrupt timestamp to the system. Based on this timestamp, a 39-dimensional audio Mel-frequency cepstral feature sequence (MFCC) is extracted. The clock scheduler then forcibly takes over the data bus, aligning the three-axis acceleration physical flow at the end of the force application with the audio Mel-frequency cepstral feature sequence to the same reference hardware clock line.
[0027] The childcare metaverse simulation system also includes an embedded pre-addressing unit 2. The embedded pre-addressing unit 2 generates a pre-addressing index vector based on absolute spatial coordinates and view frustum deflection angle vector. It performs addressing in the local map database according to the pre-addressing index vector and loads the matched multimodal interaction benchmark dataset and initial collision box volume parameters into the local process memory stack. Specifically, the embedded front-end addressing unit 2 includes a line-of-sight priority correction module 21 and a memory front-end loading module 22; Among them, the view distance priority correction module 21 is used to obtain the initial Euclidean distance scalar between the absolute spatial coordinates and the three-dimensional anchor point of the target virtual object, and introduce the view dwell time integral extracted based on the view frustum deflection angle vector, and use the view dwell time integral as a nonlinear weight to perform priority correction on the initial Euclidean distance scalar to generate the pre-addressing index vector. The memory preloading module 22 performs addressing in the local map database based on the pre-addressing index vector, and preloads the matched multimodal interaction benchmark dataset and initial collision box volume parameters into the local process memory stack of the simulation terminal.
[0028] Furthermore, to address the high-frequency data scheduling conflicts that occur when "spitting up" suddenly happens during the "infant burping" process, and because the trainee's hand is performing periodic patting on the back (at which point it is very close to the 3D anchor point on the back), while the spitting up incident occurs in the mouth and nose area (at which point the hand is far from the mouth and nose anchor point), if a spitting up incident occurs at this time, the traditional collision detection engine that relies on spatial distance will give the spitting up emergency processing data a very low prefetch priority. This will cause the local rendering pipeline to experience frame drops and clipping when the trainee's hand turns to the mouth and nose area, as it is waiting for network data to be sent.
[0029] In this embodiment, the line-of-sight priority correction module 21 receives the spatial absolute coordinates and the view frustum deflection angle vector from the multimodal acquisition unit 1, and the specific steps for generating the pre-addressing index vector are as follows: The viewing distance priority correction module 21 first acquires the absolute spatial coordinates and viewing cone deflection angle vector of the trainee in real time, and constructs a ray bounding box based on the above parameters (the absolute spatial coordinates and viewing cone deflection angle vector of the trainee) and resolves it as the currently observed viewing cone that moves with the trainee's head. When the virtual infant's nasal and oral region's overflow lesion enclosing box (i.e., the specific target virtual object in this embodiment) intrudes into the currently observed frustum, the 3D anchor points of the specific target virtual object within the frustum are located and extracted using a ray intersection algorithm. ; Locking 3D anchor points Then, the absolute spatial coordinates of the trainee's force application end are extracted simultaneously. By calculating the physical spatial straight-line geometric distance between the two (absolute spatial coordinates and three-dimensional anchor point). And generate an initial Euclidean distance scalar that is negatively correlated with it. : ; In the formula, The absolute physical distance scalar from the end of the trainee's hand exerting force to the anchor point of the milk regurgitation lesion, in meters; The system space scaling factor is a normalized mapping parameter based on the standard arm length of the trainee (approximately 0.6m to 0.8m) and the operating radius of the training platform in ergonomics. The preferred value range is 0.5 to 1.5. To prevent overflow, the range of values for this minimal constant is preferably [value range missing]. to This is used to ensure that the hand coordinates and the anchor point coordinates completely coincide on the space grid (i.e., During the burping phase, division will not throw a serious system interrupt. The initial Euclidean distance scalar generated at this time is relatively large. The value is in the extremely low level range; While locking the anchor point, activate the currently observed frustum. Within the currently observed frustum, extract the central optical axis direction vector of the frustum deflection angle vector. , And construct a three-dimensional anchor point from the trainee's viewpoint origin. Spatial ray vector Real-time calculation of the angle parameter between two vectors : ; In the formula, This represents the dot product of the central optical axis direction vector and the spatial ray vector in the global coordinate system. Represents the direction vector of the central optical axis With spatial ray vector The length of the three-dimensional Euclidean model; Set effective field of view deflection threshold (Preferably 10° to 15°), when the included angle parameter At this time, it indicates that the trainee's physical line of sight is directly focused on the lesion of milk regurgitation; for a continuous hardware clock cycle that meets this condition (from the trigger start time) Up to the current moment Perform time-domain integration to generate the line-of-sight dwell time integral. : ; In the formula, Represents the step indicator function, when the condition is satisfied ( Output 1 if the line of sight deviates from the system tolerance period (preferably 100ms to 200ms), otherwise output 0. If the line of sight deviates from the system tolerance period, the integral is cleared to zero. Furthermore, to forcibly reverse the low distance priority in the "hand not yet seen but eye already aware" state, the initial Euclidean distance scalar is used. Using the baseline coefficient, integral over line-of-sight dwell time For the natural index compensation term, construct a nonlinear priority correction function: ; in, This is the pre-address index vector, representing the urgency level weight of the data addressing sent to the underlying bus controller; The attention compensation gain constant is preferably in the range of 2.0 to 5.0. The attention compensation gain constant determines the "aggressiveness" of the system in preempting memory bus bandwidth when a sudden emergency occurs. If the value is too small, it will not be able to effectively increase the prefetch priority in the short time when the trainee moves his hand. If the value is too large, it is very easy to cause frequent useless memory erase and write due to brief misperception.
[0030] When the trainee focuses on the lesion after noticing spitting up milk ( (linear growth), even though its hands are still on its back ( (Very small), the pre-address index vector output by the system. It can also exhibit exponential jumps.
[0031] Furthermore, in obtaining the exponentially amplified pre-address index vector... Then, the memory preloading module 22 directly uses it as a trigger signal to block all regular HTTP / TCP network addressing requests sent to the cloud and initiate local hardware-level takeover; The pre-addressing index vector Input the address mapping table of the local graph database configured on the edge side of the simulation terminal to accurately match the physical starting address of the data storage block that is strongly associated with the emergency milk spillage operation; The specific data structure of the address mapping table is as follows: Key: A unique graph node hardware identifier (Node_ID) that represents each target virtual object in the scene (such as "respiratory tract regurgitation lesion" or "back burping force surface"). Mapping Value: Contains the physical starting address of the corresponding interactive data in a non-volatile storage medium (such as solid-state flash memory), the length of the contiguous data block in bytes, and the matching direct memory access channel number (DMA_CH_ID). Real-time monitoring of the front addressing index vector The current addressing index vector When the gaze dwell time increases exponentially and exceeds the preset emergency pre-fetching threshold, the regular network addressing request sent to the cloud is immediately blocked, and the graph node hardware identifier (Node_ID) of the locked target virtual object (i.e., the milk overflow lesion) is extracted. Subsequently, the hardware identifier of the graph node is entered into the address mapping table for execution. The complex hash collision addressing method is used to resolve the physical starting address and data block byte length that are strongly associated with the dual-track baseline data of the milk overflow emergency operation. Finally, a handshake signal is sent to the motherboard's underlying bus controller to establish a direct memory access channel (DMA channel) between the aforementioned peripheral data storage block and the local process memory stack (RAM). While maintaining the DMA channel open, without occupying the CPU clock cycle and without instruction decoding polling, in block transfer mode, the multimodal interaction benchmark dataset for spitting up (specifically including the respiratory tract clearing standard dynamic feature tensor retrieved from the explicit dynamic physical benchmark track, and the soothing voiceprint baseline array retrieved from the implicit cognitive and acoustic interaction benchmark track), as well as the initial collision box volume parameters of this region, are forcibly copied in full to the active page of the local process memory stack; Based on the aforementioned scheduling mechanism of the vision-driven underlying bus, before the trainee's hand physical spatial displacement actually reaches the three-dimensional boundary of the mouth and nose area, all the underlying reference tensors and geometric parameter boundaries required for emergency judgment are already in a physically ready state in the local chip memory. This completely eliminates the rendering delay caused by communication link fluctuations in the time domain, and realizes a zero-latency physical closed loop for multimodal interaction in childcare emergency scenarios.
[0032] The childcare metaverse simulation system also includes a kinetic energy correction and feature fusion unit 3. The kinetic energy correction and feature fusion unit 3 is used to perform time-domain difference decomposition on the triaxial acceleration physical flow to calculate the transient kinetic energy decay gradient. Based on the transient kinetic energy decay gradient, the initial collision box volume parameters are elastically mapped to generate a flexible physical boundary tensor. Furthermore, when it is determined that the three-dimensional coordinate sequence intrudes into the flexible physical boundary tensor surface and reaches the extreme intrusion depth, the transient kinetic energy decay gradient, extreme intrusion depth and audio Mel cepstral feature sequence within the same underlying clock cycle are extracted and spliced together to generate the operation behavior feature matrix. Among them, the kinetic energy correction and feature fusion unit 3 includes a flexible boundary mapping module 31 and a feature matrix splicing module 32; Among them, the flexible boundary mapping module 31 is used to perform first-order time-domain difference on the triaxial acceleration physical flow to extract the transient kinetic energy decay gradient characterizing the contact buffer feature, and to perform indentation mapping along the surface normal on the initial collision box volume parameters with the transient kinetic energy decay gradient as the deformation excitation variable to generate the flexible physical boundary tensor. The feature matrix splicing module 32 is used to monitor the spatial interference state between the three-dimensional coordinate sequence acquired by the multimodal acquisition unit 1 and the flexible physical boundary tensor, determine the extreme point trigger of the physical intrusion of the three-dimensional coordinate sequence and lock the hardware interrupt timestamp; call the hardware interrupt timestamp to extract the transient kinetic energy decay gradient within the same underlying clock cycle, and perform memory splicing with the audio Mel cepstral feature sequence to output the operation behavior feature matrix.
[0033] In compliant childcare operations, the trainee must actively decelerate and cushion the impact (i.e., "gentle pat") the moment they touch the virtual infant's back. Traditional simulation systems use absolutely rigid collision boxes, where a collision is considered complete once the physical coordinates touch the mesh surface, thus losing the ability to quantify the "force cushioning" and "pressure depth" after contact. In this embodiment, by extracting the dynamic parameters of the multimodal acquisition unit 1 and elastically modifying the geometric boundary loaded by the embedded front-end addressing unit 2, a flexible boundary determination with physical elasticity is achieved. The specific steps are as follows: During the process of the trainee waving their hand towards the virtual infant interaction surface, such as their back, the flexible boundary mapping module 31 reads the three-axis acceleration physical flow input from the multimodal acquisition unit 1 in real time, and extracts the transient acceleration component parameter perpendicular to the normal axis of the trainee's palm from the three-axis acceleration physical flow. ; When transient acceleration component parameters are monitored When the sign of the instantaneous acceleration component parameter is reversed due to physical contact (i.e., collision deceleration occurs), within the time window of physical displacement contact being monitored, the instantaneous acceleration component parameter is... Perform a first-order time-domain difference operation to solve for the rate of change of acceleration at the instant of contact, and analyze the negative extrema of this rate of change scalar as the transient kinetic energy decay gradient. : ; In the formula, Indicates the current time The instantaneous acceleration scalar along the palm normal, in units of ; Represents the instantaneous acceleration scalar of the previous sampling period; This represents the low-level hardware sampling period of the microelectromechanical system (MEMS) inertial sensor, with a value of 5 milliseconds (corresponding to a sampling rate of 200 Hz). The absolute value of the negative extreme value of the difference result (i.e., the extreme point of the most severe transient deceleration) is taken, and the transient kinetic energy decay gradient is also considered. It maps the "force relief / buffering" effectiveness of the trainee's operation end. The smaller the value, the smoother and more compliant the buffering. The larger the value, the more illegal and heavy the action.
[0034] Furthermore, in order to generate flexible boundaries, the local process memory stack loaded by the embedded front addressing unit 2 extracts the initial collision box volume parameters corresponding to the infant's back in the target virtual object interaction area. Using the built-in geometry shader pipeline, the 3D mesh vertex clusters that constitute the initial collision box volume parameters are resolved. and its corresponding surface normal vector ; Obtaining the transient kinetic energy decay gradient and 3D mesh vertex clusters and its corresponding surface normal vector Then, a mesh elastic deformation function is constructed with the transient kinetic energy decay gradient as the core independent variable; the mesh elastic deformation function is called to calculate the offset of each mesh vertex, driving the three-dimensional mesh vertex cluster to perform a depth translation along its corresponding surface normal vector into the mesh. The mesh elastic deformation function is as follows: ; In the formula: This represents the coordinates of the new vertices of the 3D mesh after deformation reconstruction. Represents the vertex coordinates of the original 3D mesh; Represents the vertex The corresponding unit surface normal vector indicates the orientation of the mesh surface; The value is a mapping coefficient for the elastic modulus of the physiological tissue of virtual infants and young children. The preferred value range is 0.01 to 0.05. The value is set according to the Young's modulus of the soft tissue of the back of the infant's body. It is used to control the deformation stiffness of the mesh. The smaller the value, the harder the surface. This represents a damping constant used to prevent register division-by-zero errors; in this embodiment, it is set to... Used to ensure that in the extreme soft state ( The function still converges; the product term This is the scalar value for the calculated mesh vertex offset.
[0035] Among them, the transient kinetic energy decay gradient The smaller the value (the smoother the physical action), the larger the indentation offset displacement output by the product term; this applies to all 3D mesh datasets after translation reconstruction (i.e., all...). The topological surface formed is encapsulated and output as a flexible physical boundary tensor to characterize the physical press-and-rebound depth of the surface.
[0036] In multimodal childcare training (such as burping and soothing simultaneously), due to the inherent heterogeneity of the underlying clock of motion dynamics sensors (such as 200Hz of IMU) and acoustic microphone arrays (such as 16kHz or 44.1kHz of audio sampling), traditional application-layer software polling is easily affected by system thread blocking, resulting in a random drift of tens to hundreds of milliseconds on the time axis between physical patting actions and soothing voice output. The feature matrix splicing module 32 completely removes the time delay interference at the software level through the underlying hardware-level interrupt mechanism and direct memory access (DMA) calls, realizing microsecond-level physical spatiotemporal forced anchoring of heterogeneous multimodal data.
[0037] Specifically, the three-dimensional coordinate sequence is imported into the bounding box collision detection pipeline at the bottom layer of the simulation terminal; within this pipeline, the physical force end represented by the three-dimensional coordinate sequence is calculated in real time. The relative spatial relationship between the flexible physical boundary tensor and the outer surface: Assume the physical force application end The instantaneous depth scalar that penetrates into the interior of the flexible physical boundary tensor along the surface normal is Then, perform a first-order time-domain derivative on the scalar in real time: ; In the formula, It is a continuous physical time variable at the system's underlying level, generated by mapping the hardware clock cycle of the simulation terminal motherboard; Indicates the current time Below, the absolute distance from which the end of the trainee's physical force penetrates into the grid is measured in millimeters. This represents the first derivative of the intrusion depth with respect to a continuous physical time variable at the current moment. The instantaneous value, i.e., the instantaneous pressing speed scalar; When the instantaneous depth scalar is detected When the sign changes from positive to negative (i.e., the hand is pressed to its deepest point and physical rebound begins), the end of the physical force is determined. Reaching extreme penetration depth And lock its value in the register, specifically, the extreme value intrusion depth. This geometric indicator, used to quantify the maximum indentation when a trainee pats an infant's back, is a core geometrical indicator for determining whether the action causes "organ compression injury." Its value is based on a safety redundancy limit established by the elasticity of the thoracic cartilage and the Young's modulus of the back soft tissue in infants aged 0-1 years. The preferred value range is... to .
[0038] In determining the end of the physical force application Reaching the above extreme penetration depth In a microsecond-level physical instant, the feature matrix splicing module 32 directly sends a high-level transition signal to the interrupt request pin of the central processing unit (CPU). This hardware-level electrical transition forcibly suspends all currently non-concurrent polling threads and non-critical rendering pipelines, and the motherboard's underlying clock scheduler forcibly intercepts the current system bus clock, generating a hardware interrupt timestamp with absolute immutability. This is used to establish the unique absolute coordinates of the contact pole of this burping event on the global time axis.
[0039] In the system's physical memory, directly establish a hardware interrupt timestamp. With the central axis of symmetry, Extremely narrow alignment window for tolerance radius : ; In the formula, This represents a continuous timestamp index retrieval range allocated in physical main memory; in, The time window tolerance radius is preferably between 30 milliseconds and 50 milliseconds. Its value is set based on the minimum neural conduction delay of the vocal cords vibrating when the human vocal organs are driven by tactile feedback. Specifically, the minimum neural reflex and muscle coordination delay between the trainee applying a tactile action (hand slapping) and triggering the vocal organs (mouth uttering soothing words such as "baby, be good") is about 30 to 50 milliseconds. Narrowing the window to this physical limit can completely filter out asynchronous environmental noise, just like a hardware-level bandpass filter, ensuring that the system only extracts acoustic signals that have an absolute causal relationship with the slapping action.
[0040] It is worth noting that narrowing the alignment window to this physical limit can completely filter out environmental noise and asynchronous speech, ensuring that only the soothing speech that occurs absolutely concurrently with the "tapping" action is captured. This window also serves as the addressing index key for subsequent direct memory calls.
[0041] Furthermore, by utilizing extremely narrow aligned windows As the addressing index key, access is configured in the underlying Direct Memory Access (DMA) cache, and the retrieved timestamp must fall strictly within this extremely narrow alignment window. The audio Mel-frequency cepstral feature sequence concurrently input by the multimodal acquisition unit 1 (Characteristics of soothing speech features, and transient kinetic energy decay gradient calculated by flexible boundary mapping module 31) (Characterizing contact buffer characteristics).
[0042] Finally, by performing multidimensional tensor fusion computation, heterogeneous parameter execution matrices on the same microsecond-level physical causal chain are concatenated on contiguously allocated physical memory address blocks to generate and output an operation behavior feature matrix. : ; In the formula, The output is the operational behavior feature matrix, used to characterize the multimodal coupled physical slice at a certain extreme contact instant; The transient kinetic energy decay gradient scalar represents the first and second addresses of the occupied matrix dynamics channel, used to characterize the action unloading and buffering state; This indicates the occupied geometric channel address, used to characterize the physical mesh intrusion limit; This represents the transposed audio Mel-Cepstral feature column vector, which occupies the acoustic feature channels of the matrix.
[0043] Among them, the feature matrix The underlying bus is written in a one-time packet, which eliminates the frame misalignment phenomenon caused by subsequent network transmission from the perspective of system engineering architecture, and realizes the rigid binding of the concurrent characteristics of "acoustic expression" (soothing words spoken) and "dynamic pressure" (force and depth of hand slapping) at the current spatiotemporal node.
[0044] The childcare metaverse simulation system also includes a deviation comparison and rendering unit 4, which compares the multimodal interaction benchmark dataset with the operation behavior feature matrix and generates a physical operation deviation index. The physical operation deviation index is mapped to the rendering pipeline timing compensation frame number, and hardware-level state rewriting is performed on the graphics rendering engine to drive the mesh deformation and audio synchronization update of the target virtual object.
[0045] Among them, the deviation comparison rendering unit 4 includes a dual-track benchmark comparison module 41 and a timing compensation direct drive module 42; Among them, the dual-track benchmark comparison module 41 is used to parse the multimodal interaction benchmark dataset from the local process memory stack and decompose it into explicit dynamic physical benchmark track and implicit cognitive and acoustic interaction benchmark track. The operation behavior feature matrix is respectively imported into the explicit dynamic physical benchmark track and the implicit cognitive and acoustic interaction benchmark track to perform heterogeneous distance calculation and generate a quantified physical operation deviation index. The timing compensation direct drive module 42 is used to receive the physical operation deviation index, map it to the rendering pipeline timing compensation frame number through a preset nonlinear compensation function; call the rendering pipeline timing compensation frame number to perform hardware-level interrupt and state rewriting on the physics engine and audio mixer in the graphics rendering engine, and drive the mesh deformation and audio synchronization update of the target virtual object.
[0046] Specifically, when determining the results of multimodal interactions, the childcare metaverse simulation system generally relies on the application-layer software framework to perform serial polling and single threshold truncation. This is highly susceptible to severe phase distortion due to uneven distribution of CPU computing power, resulting in situations where "physical actions have triggered mesh deformation, but the sound card audio is still queuing for decoding." Therefore, the deviation comparison rendering unit 4 directly extracts the operational behavior feature matrix, which has achieved microsecond-level physical alignment, from the underlying process memory. It then performs heterogeneous mathematical space dimensionality reduction comparison in an explicit / implicit dual-track system independent of the application layer. The deviation index output from the solution is directly converted into underlying clock interrupt instructions for the graphics processing unit (GPU) and digital-to-analog converter (DAC), thereby ensuring absolute spatiotemporal synchronization of multimodal feedback at the physics engine computation level. The specific calculation steps and hardware scheduling are as follows: The dual-track benchmark comparison module 41 directly accesses the local process memory stack loaded and ready by the embedded front addressing unit 2 via the bus, parses and extracts the multimodal interaction benchmark dataset; and physically splits the multimodal interaction benchmark dataset into two independent underlying control groups: the explicit dynamic physical benchmark track and the implicit cognitive and acoustic interaction benchmark track. Furthermore, the dual-track benchmark comparison module 41 resolves and extracts the preset standard buffer gradient scalar from the explicit dynamical physical benchmark track. (i.e., the scalarized expression of the kinetic energy decay gradient standard curve of the aforementioned flexible touch) and the standard safe intrusion limit (i.e., the specific threshold of the extreme boundary of the compliant tapping impulse); simultaneously, the pre-set standard soothing voiceprint reference tensor is parsed and extracted from the implicit cognition and acoustic interaction reference track. (That is, the specific tensor form of the soothing speech acoustic feature standard array); By performing low-level dimension decomposition on the operational behavior feature matrix output by the kinetic energy correction and feature fusion unit 3, the transient kinetic energy decay gradient actually triggered by the trainee is extracted. With extreme penetration depth ; Calculate the actual trigger value (the transient kinetic energy decay gradient of the actual trigger) separately. With extreme penetration depth ) and the standard reference value (preset standard buffer gradient scalar) Compared to standard security intrusion limits The Euclidean distance between the two indices is calculated, and a normalized weighted sum is performed to generate a dynamic deviation sub-index. : ; In the formula, It represents the overall violation rate of the current interactive action at the dynamic physics level, and is a dimensionless normalized scalar. The standard buffer gradient scalar is used as the normalized denominator, and its value is based on the fixed value of the standard hollow palm slapping force unloading index defined in the childcare medical guidelines. This indicates that the range of values for the safe compressive deformation limit of the infant's back bones and muscles (i.e., the standard safe intrusion limit) is preferably fixed as follows: to ; and These are the unloading deviation weight and the pressing depth deviation weight, respectively (and Furthermore, the numerator term The first Euclidean distance between the transient kinetic energy decay gradient and the standard buffer gradient scalar is used to characterize the absolute deviation in algebraic space between the actual buffering force relief of the trainee's action and the standard buffering force relief. The second Euclidean distance, representing the extreme intrusion depth and the standard safe intrusion limit, is used to characterize the absolute deviation between the extreme depth and the safe depth of the actual pressure grid on the trainee's hand.
[0047] Since the risk of internal organ compression caused by "excessive pressure" is far greater than that of "insufficient surface pressure relief" in childcare practice, this embodiment preferably sets... , This is to increase the sensitivity of penalties for deep overstepping.
[0048] Simultaneously, the audio Mel-Cepstral feature sequences of actual concurrent triggers by the trainees are extracted from the operational behavior feature matrix. The cosine similarity algorithm is used to solve the audio Mel-frequency cepstral feature sequence. Compared with the standard soothing voiceprint reference tensor Vector angle in multidimensional acoustic space Generate acoustic deviation feature sub-indices : ; Furthermore, ; In the formula, A column vector representing the Mel frequency cepstral coefficients of the speech produced by the trainee; This represents the baseline column vector generated after standard gentle and soothing voices (such as "Baby, be good" with a specific frequency and tone) collected from professional medical personnel are extracted using the same dimensionality reduction method. Represents the algebraic inner product of two multidimensional eigenvectors; This represents the calculation of the Euclidean magnitude of the corresponding multidimensional eigenvectors; the vector angle. This indicates the degree of deviation between the trainee's actual speech in timbre and intonation and the standard soothing speech, expressed in radians (rad), with a range of values of [missing value]. ,in, Pi is a constant and is used as a parameter in the normalized denominator; specifically, when... When this occurs, it indicates that the actual sound is acoustically completely "opposite" to the soothing request (e.g., a high-frequency, sharp, panicked scream); by dividing by Map the deviation state to The scalar range, the larger the value, the more the trainee's speech deviates from the "gentle and soothing" state (e.g., mixed with high-frequency components of panic and impatience), and the output acoustic deviation feature sub-index. .
[0049] Furthermore, by introducing a preset global penalty leveling coefficient... and (satisfy In concurrent interactive scenarios of "burping and soothing," physical impact errors (dynamics and geometric deformation) directly endanger the vital signs of infants, while verbal errors mainly affect cognitive soothing emotions; therefore, physical deviations are more significant. Assigned higher than acoustic deviation The balancing weight is therefore preferably set to... , ), for dynamic deviation sub-indices Acoustic deviation characteristic sub-index Perform cross-dimensional linear fusion and encapsulate the fusion result into a single scalar physical operation deviation index. : ; Furthermore, to prevent drastic deviations from directly driving the underlying engine and causing screen tearing or audio popping, the timing compensation direct drive module 42 receives physical operation deviation indicators. Then, by constructing a nonlinear compensation function, it is mapped to the number of timing compensation frames in the rendering pipeline. : ; In the formula, This represents the final output global multimodal physics operation deviation rate evaluation scalar; This represents the floor function, used to ensure that the output is a discrete physical frame integer; The base number for the underlying rendering buffer is preferably 3 to 8 frames. Specifically, based on the mainstream refresh rate standard of 60Hz to 90Hz for head-mounted display devices, if a value of 5 frames is taken, it corresponds to a physical buffer time of about 55 milliseconds to 83 milliseconds (this time window is sufficient for the physics engine to smoothly insert the deformation tweening animation of "skin turning from white to red and mesh turning from light to dark" to avoid VR motion sickness in trainees, and it is also lower than the upper limit of 100 milliseconds for the human brain to recognize "audio-visual asynchrony"). Represents the natural logarithm function; The above expression shows that the overall operational deviation... The larger the image size, the more transition compensation frames the system needs to visually smoothly render the "deep indentation of the body and reddening of the skin due to the impact." It increases logarithmically.
[0050] Furthermore, in determining the number of compensation frames... Then, the calculated mesh deformation topology coordinates and audio gain frequency shift instructions are encapsulated respectively; a hardware-level clock interrupt signal is sent directly to the direct memory access controller (DMAC) configured on the motherboard of the simulation terminal through the timing compensation direct drive module 42; the interrupt instruction crosses the operating system kernel and forcibly takes over the vertex shader and fragment shader cache in the graphics rendering engine, as well as the digital-to-analog conversion channel of the audio mixer.
[0051] The system's underlying clock scheduler strictly suspends and processes the compensated frame count. After one micro-compensation clock cycle, at the same microsecond-level physical level transition edge, the vertex shader's screen rewriting and the sound card's ring buffer's audio decoding playback are concurrently triggered.
[0052] Through the aforementioned underlying hardware scheduling mechanism, the trainee's "operational deviation" is transformed into the "mesh deformation feedback" and "audio prompt feedback" of the target virtual object with zero phase difference synchronous update, which is used to avoid the timing misalignment problem caused by multimodal heterogeneous processing.
[0053] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A childcare metaverse simulation system based on embedded data query and multimodal interaction, deployed on a simulation terminal equipped with a local graph database and a graphics rendering engine, and communicatively connected to a dual-track reference data base, wherein the dual-track reference data base is pre-loaded with a multimodal interaction reference dataset and initial collision box volume parameters of the target virtual object; characterized in that, The childcare metaverse simulation system includes: The multimodal acquisition unit (1) is used to acquire the absolute spatial coordinates and the view frustum deflection angle vector of the trainee, and simultaneously acquire the three-dimensional coordinate sequence, the three-axis acceleration physical flow and the concurrent audio Mel-Cepstral feature sequence. The embedded front addressing unit (2) generates a front addressing index vector based on the absolute spatial coordinates and the view frustum deflection angle vector, and addresses in the local map database according to the front addressing index vector, loading the matched multimodal interaction benchmark dataset and the initial collision box volume parameters into the local process memory stack; The kinetic energy correction and feature fusion unit (3) is used to perform time-domain difference decomposition on the triaxial acceleration physical flow to calculate the transient kinetic energy decay gradient, and to perform elastic mapping on the initial collision box volume parameters based on the transient kinetic energy decay gradient to generate a flexible physical boundary tensor. When it is determined that the three-dimensional coordinate sequence invades the flexible physical boundary tensor surface and reaches the extreme invasion depth, the transient kinetic energy decay gradient, extreme invasion depth and audio Mel cepstral feature sequence within the same underlying clock cycle are extracted and spliced to generate an operation behavior feature matrix. The deviation comparison rendering unit (4) is used to compare the multimodal interaction benchmark dataset with the operation behavior feature matrix to generate a physical operation deviation index, which is mapped to the rendering pipeline timing compensation frame number, performs hardware-level state rewriting on the graphics rendering engine, and drives the mesh deformation and audio synchronization update of the target virtual object.
2. The childcare metaverse simulation system based on embedded data query and multimodal interaction as described in claim 1, characterized in that, The multimodal acquisition unit (1) includes a spatial dynamics capture module (11) and a visual cone acoustic synchronous tracking module (12). The spatial dynamics capture module (11) is used to collect the absolute spatial coordinates of the trainee and extract the three-dimensional coordinate sequence of the upper limb skeletal joints and the three-axis acceleration physical flow at the end of the force application; The acoustic synchronous tracking module (12) acquires concurrent audio Mel-Cepstral feature sequences through a hardware pickup array and extracts the cone deflection angle vector based on the head-mounted display device.
3. The childcare metaverse simulation system based on embedded data query and multimodal interaction according to claim 2, characterized in that, The triaxial acceleration physical flow is a discrete time series matrix with a physical reference timestamp as the independent variable. The data elements of the discrete time series matrix include transient acceleration component parameters and resultant acceleration dynamic amplitude. The frustum deflection angle vector includes the origin coordinates of the viewpoint calculated from the trainee's head-mounted display device, the central optical axis direction vector mapped by the attitude Euler angles, and the effective field of view boundary parameters.
4. The childcare metaverse simulation system based on embedded data query and multimodal interaction according to claim 1, characterized in that, The embedded front-end addressing unit (2) includes a line-of-sight priority correction module (21) and a memory front-end loading module (22). Among them, the line-of-sight priority correction module (21) is used to obtain the initial Euclidean distance scalar between the absolute spatial coordinates and the three-dimensional anchor point of the target virtual object, and introduce the line-of-sight dwell time integral, and use the line-of-sight dwell time integral as a nonlinear weight to perform priority correction on the initial Euclidean distance scalar, and generate the pre-addressing index vector. The memory preloading module (22) addresses the local graph database based on the pre-addressing index vector and preloads the matched multimodal interaction benchmark dataset and initial collision box volume parameters into the local process memory stack of the simulation terminal.
5. The childcare metaverse simulation system based on embedded data query and multimodal interaction according to claim 4, characterized in that, In the line-of-sight priority correction module (21), the initial Euclidean distance scalar is corrected for priority to generate a preceding addressing index vector. The specific steps involved are as follows: A ray bounding box is constructed based on the absolute spatial coordinates and the view frustum deflection angle vector. The spatial geometry falling within the ray bounding box is resolved into the currently observed view frustum, and the three-dimensional anchor point of the target virtual object inside the currently observed view frustum is locked. Calculate the physical spatial straight-line geometric distance between the absolute spatial coordinates and the three-dimensional anchor point, and generate an initial Euclidean distance scalar; Within the current observation frustum, the angle between the central optical axis of the frustum deflection angle vector and the three-dimensional anchor point is within a preset effective field of view deflection threshold for a continuous hardware clock cycle. A time-domain integration operation is performed on the continuous hardware clock cycle to generate the line-of-sight dwell time integral. Construct a priority correction function with the initial Euclidean distance scalar as the base and the line-of-sight dwell time integral as the correction exponent, perform exponential weighted reconstruction on the initial Euclidean distance scalar, and output the pre-addressing index vector.
6. The childcare metaverse simulation system based on embedded data query and multimodal interaction according to claim 3, characterized in that, The kinetic energy correction and feature fusion unit (3) includes a flexible boundary mapping module (31) and a feature matrix splicing module (32). Among them, the flexible boundary mapping module (31) is used to perform first-order time-domain difference on the triaxial acceleration physical flow to extract the transient kinetic energy decay gradient, and to perform indentation mapping along the surface normal on the initial collision box volume parameters with the transient kinetic energy decay gradient as the deformation excitation variable to generate a flexible physical boundary tensor. The feature matrix splicing module (32) is used to monitor the spatial interference state between the three-dimensional coordinate sequence and the flexible physical boundary tensor, determine the extreme point trigger of the physical intrusion of the three-dimensional coordinate sequence and lock the hardware interrupt timestamp; call the hardware interrupt timestamp to extract the transient kinetic energy decay gradient within the same underlying clock cycle, and splice it with the audio Mel cepstral feature sequence in memory to output the operation behavior feature matrix.
7. The childcare metaverse simulation system based on embedded data query and multimodal interaction according to claim 6, characterized in that, The flexible boundary mapping module (31) generates the flexible physical boundary tensor, and the specific steps involved are as follows: The transient acceleration component parameters are extracted from the triaxial acceleration physical flow, and a first-order time-domain difference operation is performed on the transient acceleration component parameters to solve the rate of change of acceleration at the moment of contact. The negative extremum in the scalar of the rate of change of acceleration is analyzed as the transient kinetic energy decay gradient. Access the local process memory stack, extract the initial collision box volume parameters corresponding to the interaction area of the target virtual object, and parse the three-dimensional mesh vertex clusters and corresponding face normal vectors that constitute the initial collision box volume parameters; A mesh elastic deformation function is constructed with the transient kinetic energy decay gradient as the core independent variable. The mesh elastic deformation function is called to calculate the mesh vertex offset. The three-dimensional mesh vertex clusters are driven to perform a depth translation in the mesh interior along their respective surface normal vectors. The translated and reconstructed three-dimensional mesh dataset is encapsulated and output as a flexible physical boundary tensor.
8. The childcare metaverse simulation system based on embedded data query and multimodal interaction according to claim 6, characterized in that, The feature matrix splicing module (32) performs the following steps to output the operation behavior feature matrix: The three-dimensional coordinate sequence is imported into the bounding box collision detection pipeline of the simulation terminal, and the relative spatial position relationship between the physical force application end represented by the three-dimensional coordinate sequence and the outer surface of the flexible physical boundary tensor is calculated in real time. When it is determined that the spatial coordinates of the physical force application end penetrate the flexible physical boundary tensor surface and reach the extreme penetration depth, a high-level transition signal is sent to the interrupt request pin of the central processing unit to forcibly suspend the non-concurrent polling thread, and the underlying clock scheduler forcibly generates an absolute timing hardware interrupt timestamp. Create an extremely narrow aligned window in physical memory with the hardware interrupt timestamp as the addressing index key; Extract the audio Mel-Cepstral feature sequence and the corresponding transient kinetic energy decay gradient that fall within the extremely narrow alignment window, perform multidimensional tensor fusion calculation, and concatenate the audio Mel-Cepstral feature sequence, the transient kinetic energy decay gradient, and the extreme value intrusion depth on continuous physical memory addresses to generate and output the operation behavior feature matrix.
9. The childcare metaverse simulation system based on embedded data query and multimodal interaction according to claim 1, characterized in that, The deviation comparison rendering unit (4) includes a dual-track reference comparison module (41) and a timing compensation direct drive module (42). Among them, the dual-track benchmark comparison module (41) is used to parse the multimodal interaction benchmark dataset from the local process memory stack and decompose it into explicit dynamic physical benchmark track and implicit cognitive and acoustic interaction benchmark track. The operation behavior feature matrix is respectively imported into the explicit dynamic physical benchmark track and the implicit cognitive and acoustic interaction benchmark track to perform heterogeneous distance calculation and generate physical operation deviation index. The timing compensation direct drive module (42) is used to receive the physical operation deviation index, map it to the rendering pipeline timing compensation frame number through a preset nonlinear compensation function; call the rendering pipeline timing compensation frame number to perform hardware-level interrupt and state rewriting on the physical engine and audio mixer in the graphics rendering engine, and drive the execution of the mesh deformation and audio synchronization update of the target virtual object.
10. The childcare metaverse simulation system based on embedded data query and multimodal interaction according to claim 9, characterized in that, The dual-track benchmark comparison module (41) performs the following steps to generate the physical operation deviation index: The preset standard buffer gradient scalar and standard safety intrusion limit are analyzed and extracted from the explicit dynamic physical reference track, and the preset standard soothing voiceprint reference tensor is analyzed and extracted from the implicit cognitive and acoustic interaction reference track. The operational behavior feature matrix is decomposed into its lowest-level dimensions to extract the transient kinetic energy decay gradient and extreme intrusion depth actually triggered by the trainee. The first Euclidean distance between the transient kinetic energy decay gradient and the standard buffer gradient scalar, and the second Euclidean distance between the extreme intrusion depth and the standard safe intrusion limit are calculated. A weighted summation operation is performed on the first Euclidean distance and the second Euclidean distance to generate a dynamic deviation sub-index; The audio Mel-Cepstral feature sequence actually triggered by the trainee is extracted from the operational behavior feature matrix in a synchronous manner. The cosine similarity algorithm is used to calculate the angle between the vectors of the audio Mel-Cepstral feature sequence and the standard soothing voiceprint reference tensor in the multidimensional acoustic space, and to generate an acoustic deviation feature sub-index. A preset global penalty leveling coefficient is introduced, and a cross-dimensional linear fusion calculation is performed on the dynamic deviation sub-index and the acoustic deviation feature sub-index. The fused result is then encapsulated into a single scalar physical operation deviation index.