Food safety interactive science popularization method and system based on AI digital human

By constructing an interactive science popularization system for food safety based on AI digital humans, and utilizing multi-channel input and 3D engine-driven digital human posture changes, the system solves the problem of lack of personalized interaction in traditional science popularization methods, and achieves efficient food safety knowledge delivery and increased audience participation.

CN121918702APending Publication Date: 2026-04-24WUHAN FOOD & COSMETIC INSPECTION INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN FOOD & COSMETIC INSPECTION INST
Filing Date
2026-01-13
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Traditional methods of popularizing food safety knowledge lack personalized interaction, resulting in low audience participation and inefficient information delivery. In particular, they struggle to provide dynamic adjustments and personalized experiences when facing audiences of different ages and knowledge backgrounds.

Method used

We will build an interactive science popularization system for food safety based on AI digital humans. Through multi-channel input (voice, posture, touch), we will perceive user behavior in real time, dynamically adjust the pace and depth of the explanation, and combine a 3D engine to drive the digital human's posture changes to achieve immersive interaction.

Benefits of technology

It improved the efficiency of food safety knowledge dissemination and audience participation, and enhanced audience immersion and information delivery effectiveness through personalized interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121918702A_ABST
    Figure CN121918702A_ABST
Patent Text Reader

Abstract

The invention provides a food safety interactive science popularization method and system based on an AI digital human, and the method comprises the steps: constructing a digital human skeleton driving frame in a three-dimensional space, defining a node topology connection relation, and binding a grid vertex to a skeleton influence weight; receiving a voice recognition text stream, a user body posture coordinate sequence and a touch screen event, and generating joint semantic description; driving a 3D engine to update the posture of the digital human frame by frame, and synchronously rendering and outputting; the sight focus and head orientation changes of the user are monitored in real time, and the content rhythm is dynamically adjusted; coordinate system mapping of key positions of the exhibition hall is calibrated, an optimal virtual view angle is calculated, and linkage display of external equipment is triggered; a traditional video playing mode is replaced by 3D engine real-time rendering, and real-time interactive response is realized by fusing multi-channel input, so that a digital person can dynamically adjust an explanation strategy according to audience behaviors, and the immersion, individuation degree and knowledge transmission efficiency of food safety science popularization are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of science popularization technology, specifically to an interactive science popularization method and system for food safety based on AI digital humans. Background Technology

[0002] Currently, food safety education mainly relies on traditional media formats, such as static display boards, brochures, and pre-recorded videos—a one-way communication method. Although some venues have introduced interactive touchscreen systems, the content presentation is still limited to preset paths and lacks the ability to dynamically adjust based on audience feedback.

[0003] In recent years, virtual digital human technology has begun to be applied in the field of science popularization. However, existing solutions mostly use a fixed script to splice video clips, and the digital human's image and movements are pre-recorded, making it impossible to generate natural and coherent responses based on real-time user input. This rigid interaction mode results in low audience participation and low information transmission efficiency, especially when facing audiences of different ages and knowledge backgrounds, making it difficult to provide a personalized science popularization experience. Summary of the Invention

[0004] This invention aims to provide an interactive science popularization method and system for food safety based on AI digital humans, which solves the core problems of one-way indoctrination and lack of personalized interaction in traditional science popularization methods.

[0005] To achieve the above objectives, the technical solution adopted by this invention is: an interactive science popularization method for food safety based on AI digital humans, comprising the following steps: A digital human skeleton driving framework is constructed in three-dimensional space, the topological connection relationship between nodes is defined, the influence weight of the mesh vertices is bound to the corresponding bones, and the zero reference coordinate system under the initial posture is set. Based on the skeleton-driven framework, it receives text stream data output by the speech recognition device, captures the key point spatial coordinate sequence of the user's body posture, and integrates click and swipe gesture events on the touch screen; A joint semantic description is generated based on the integrated multi-channel input, a high-level behavior instruction sequence is obtained by searching a preset behavior template library, and the behavior priority order is adjusted by combining historical interaction states. The high-level behavioral instruction sequence is decomposed into a set of primitive actions, and the spatial parameters of the primitive actions are dynamically adjusted based on environmental parameters to generate a time series of skeletal parameters with interpolation labels. Based on the time series of skeletal parameters, the 3D engine updates the digital human pose frame by frame, applies the skeletal transformation matrix to drive mesh deformation, performs viewport rendering, and outputs the image to the display terminal. During image output, the system continuously monitors changes in the user's gaze focus and head orientation to determine the level of comprehension and acceptance of the current explanatory segment, and dynamically adjusts the pace of subsequent content based on the acceptance results. Based on the explanation content and the layout of the exhibition hall, the three-dimensional coordinate system mapping of key positions is determined, the current explanation target is identified and the optimal virtual view is calculated, and external devices are triggered to synchronously display enhanced information. Record the timestamp sequence of key events during the interaction process, preload content resources of adjacent exhibition areas, maintain the system's operating status, and support seamless switching between multiple scenarios.

[0006] Preferably, the construction of the digital human skeleton driving framework in three-dimensional space further includes: presetting a maximum and minimum rotation angle for each pair of movable joints to form a closed interval, performing a compliance check before each posture update, and automatically correcting to the nearest legal value and transmitting a reaction force signal to the parent node when the real-time angle of the joint exceeds the limit range.

[0007] Preferably, the generation of joint semantic description includes: mapping voice sentences, gesture frames and touch events to the nearest time slot to form a set of triples; if there is no valid data for a certain type of input in the time period, it is filled with null values; traversing the elements in each time slot to extract co-occurrence features; and outputting a structured semantic description.

[0008] Preferably, the step of decomposing the high-level behavior instruction sequence into a set of primitive actions further includes: inserting transitional micro-actions between adjacent high-level instructions, calculating the symmetry difference between the sets of major joints that need to be activated for both, and if the distance metric exceeds a threshold, inserting an intermediate state of a certain duration so that the posture change follows the S-shaped curve change pattern.

[0009] Preferably, the generation of the time series of skeletal parameters with interpolation markers further includes: segmenting the speech waveform into a segment sequence by phonemes, with each segment corresponding to a standard lip shape posture, aligning the start time of the segment with the lip shape activation time, and adjusting the head shaking amplitude according to the speech energy envelope so that the phase of the nodding action is synchronized with the accent.

[0010] Preferably, the method of driving the 3D engine to update the digital human pose frame by frame further includes: setting the global refresh rate to 60Hz, maintaining a high-precision timer to record the cumulative running time, checking whether the next frame time has been reached at the beginning of each loop, locating the corresponding data block from the parameter buffer and writing the control parameters into the current pose buffer.

[0011] Preferably, the step of dynamically adjusting the pace of subsequent content based on the acceptance results further includes: maintaining an understanding status flag for each knowledge node, updating the flag status based on user feedback within a preset waiting time, modifying the speech rate parameter of the speech synthesis unit to be proportional to the average attention value, and extending or shortening the silence period between each sentence.

[0012] Preferably, the triggering of external devices to synchronously display enhanced information further includes: when the angle between the direction in which the digital human arm points and the target normal is less than a threshold and the duration exceeds 1.2 seconds, verifying whether the user's gaze is focused on the same area; if confirmed, sending a control command to activate the corresponding display function, and adding supplementary explanations to the digital human's mouth.

[0013] Preferably, the maintenance system operation status and support for seamless switching between multiple scenarios also includes: setting a watchdog timer to monitor frame update signals; if three consecutive frames fail to be drawn, a recovery process is initiated; the system determines whether to return to the starting point of the previous complete semantic unit or to prioritize the emergency prompt action based on the location of the interruption point; and differentiating operation permissions are set according to the user's identity.

[0014] On the other hand, this invention proposes an interactive science popularization system for food safety based on AI digital humans, comprising: The skeleton construction module is used to build a digital human skeleton driving framework in three-dimensional space, define the topological connection relationship between nodes, bind the mesh vertices to the influence weights of the corresponding bones, and set the zero reference coordinate system under the initial pose. The multi-channel input module is used to receive text stream data output by the speech recognition device based on the skeleton driving framework, capture the key point spatial coordinate sequence of the user's body posture, and integrate click and swipe gesture events on the touch screen; The behavior instruction module is used to generate a joint semantic description based on the integrated multi-channel input, search the preset behavior template library to obtain the high-level behavior instruction sequence, and adjust the behavior priority order based on the historical interaction state. The parameter decomposition module is used to decompose the high-level behavior instruction sequence into a set of primitive actions, dynamically adjust the spatial parameters of the primitive actions based on environmental parameters, and generate a time series of skeletal parameters with interpolation labels. The 3D rendering module is used to drive the 3D engine to update the digital human's pose frame by frame based on the time series of bone parameters, apply the bone transformation matrix to drive mesh deformation, perform viewport rendering, and output the image to the display terminal. The feedback adjustment module is used to continuously monitor changes in the user's gaze focus and head orientation during image output, determine the level of comprehension and acceptance of the current explanatory paragraph, and dynamically adjust the pace of subsequent content based on the acceptance results. The environmental perception module is used to map the three-dimensional coordinate system of key locations based on the content of the explanation and the layout of the exhibition hall, identify the current target of the explanation and calculate the angle, and trigger external devices to synchronously display enhanced information; The system management module is used to record the timestamp sequence of key events during the interaction process, preload the content resources of adjacent exhibition areas, maintain the system's operating status, and support seamless switching between multiple scenarios.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention constructs a digital human motion control system that uses a 3D engine for real-time rendering, integrating voice, posture, and touch input to enable the digital human to instantly perceive and respond to user behavior. The digital human can dynamically adjust the pace and depth of its explanations based on the audience's gaze, questions, and body language, and enhance immersion through spatial linkage demonstrations. This effectively improves the efficiency of food safety knowledge dissemination and audience participation, solving the core problems of one-way instruction and lack of personalized interaction in traditional science popularization methods. Attached Figure Description

[0016] Figure 1 This is a flowchart of the interactive science popularization method for food safety based on AI digital humans, as described in this invention. Figure 2 This is a block diagram of the interactive science popularization system for food safety based on AI digital humans, as described in this invention. Detailed Implementation

[0017] The present invention will be further described below with reference to the accompanying drawings and specific embodiments: Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this application; the terms "comprising" and "having," and any variations thereof, in the description, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the description, claims, or accompanying drawings of this application are used to distinguish different objects, not to describe a particular order.

[0018] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0019] like Figure 1As shown, this invention proposes an interactive science popularization method for food safety based on AI digital humans, applied to interactive science popularization scenarios in the field of food safety. This method abandons the pre-recorded, linear, and irreversible information transmission methods of traditional video playback, and achieves two-way linkage between information expression and audience behavior by constructing a virtual narrator image that can dynamically respond to user input. When the public visits food production processes, learns safety knowledge, or participates in interactive Q&A, the digital human can instantly adjust its posture, expression, language content, and spatial position based on on-site voice, gestures, or touch commands, making the entire explanation process more immersive and participatory; specifically, it includes the following steps: A digital human skeleton driving framework is constructed in three-dimensional space, defining the topological connection relationship between nodes, binding the mesh vertices to the influence weights of the corresponding bones, and setting the zero-position reference coordinate system under the initial posture. The construction of the digital human skeleton driving framework in three-dimensional space also includes: presetting the maximum and minimum rotation angles of each pair of movable joints to form a closed interval, performing compliance checks before each posture update, and automatically correcting to the nearest legal value and transmitting the reaction force signal to the parent node when the real-time angle of the joint exceeds the limit range.

[0020] By setting the range of motion of joints and implementing dynamic compliance checks, the digital human maintains a natural ergonomic posture when responding to commands, avoiding visual anomalies such as limb distortion or clipping. A reaction force signal transmission mechanism is introduced during the movement process to effectively maintain the motion coordination between skeletal levels, improve the realism and stability of the overall movement, and ensure consistent performance over long periods of operation.

[0021] Based on the skeleton-driven framework, it receives text stream data output by the speech recognition device, captures the spatial coordinate sequence of key points of the user's body posture, and integrates click and swipe gesture events on the touch screen; it realizes synchronous perception and fusion understanding of the user's multimodal behavior, enhances the system's accuracy in recognizing interaction intentions; through multi-channel information complementarity, it reduces the risk of misjudgment caused by environmental interference from a single input source, and improves the response reliability in complex scenarios.

[0022] Based on integrated multi-channel input, a joint semantic description is generated. A pre-defined behavior template library is searched to obtain high-level behavior instruction sequences, and the priority order of behaviors is adjusted based on historical interaction states. The generation of the joint semantic description includes mapping voice statements, gesture frames, and touch events to the nearest time slot to form a set of triples. If a certain type of input has no valid data within a time period, it is filled with null values. Co-occurrence features are extracted from the elements in each time slot, and a structured semantic description is output. This improves the depth of understanding of user intent and the coherence of context, enabling the system to recognize complex interactive behaviors and distinguish between primary and secondary information. Through time alignment and co-occurrence feature extraction, the system enhances its tolerance for ambiguous or incomplete input, supporting a more natural human-computer dialogue rhythm.

[0023] The high-level behavioral instruction sequence is decomposed into a set of primitive actions. The spatial parameters of these primitive actions are dynamically adjusted based on environmental parameters to generate a time series of skeletal parameters with interpolation labels. This decomposition also includes inserting transitional micro-movements between adjacent high-level instructions, calculating the symmetry difference between the sets of major joints required for activation, and inserting an intermediate state of duration if the distance exceeds a threshold, ensuring that posture changes follow an S-shaped curve. The speech waveform is segmented by phonemes to obtain a sequence of phonemes, with each segment corresponding to a standard lip-sync posture. The start time of each segment is aligned with the lip-sync activation time, and the head movement amplitude is adjusted based on the speech energy envelope to synchronize the nodding phase with the stress.

[0024] To ensure smooth and natural transitions between actions, eliminate the mechanical feel of sudden changes in posture, and enhance the coherence of visual presentation and the realism of biological movement; through parameterized adaptation of primitive actions, the behavior of digital humans is precisely matched to the on-site spatial conditions, improving the accuracy and credibility of directional actions; the high synchronization of lip movements and speech rhythm, combined with the rhythmic response of slight head swaying, strengthens the emotional transmission ability of language expression, significantly improving the audience's immersion and trust in the virtual narrator.

[0025] The 3D engine is driven to update the digital human's pose frame by frame based on the time series of skeletal parameters. The skeletal transformation matrix is ​​applied to drive mesh deformation, viewport rendering is performed, and the image is output to the display terminal. The process of driving the 3D engine to update the digital human's pose frame by frame also includes: setting the global refresh rate to 60Hz, maintaining a high-precision timer to record the cumulative running time, checking whether the next frame has been reached at the beginning of each loop, locating the corresponding data block from the parameter buffer, and writing the control parameters into the current pose buffer.

[0026] It ensures the real-time and smooth movement of digital humans, avoids screen stuttering or delayed response, and achieves closed-loop feedback from millisecond-level input to output; it ensures the consistency of multimodal output through a precise time synchronization mechanism, and eliminates problems that affect immersion, such as audio-visual asynchrony or motion lag.

[0027] During image output, the system continuously monitors changes in the user's gaze focus and head orientation to determine the comprehension and acceptance level of the current explanatory segment. Based on the acceptance results, it dynamically adjusts the pace of subsequent content. Specifically, this includes maintaining a comprehension status flag for each knowledge node, updating the flag status based on user feedback within a preset waiting time, modifying the speech rate parameters of the speech synthesis unit to be proportional to the average attention value, and extending or shortening the silence period between each sentence.

[0028] It enables personalized adaptation of explanation content, matching the pace of information output with the user's cognitive load and improving the efficiency of knowledge absorption; it continuously assesses the state of understanding through non-intrusive behavioral observation, promptly identifies situations of confusion or distraction, and proactively adjusts the expression method.

[0029] Based on the three-dimensional coordinate system mapping of key positions according to the explanation content and exhibition hall layout, the current explanation target is identified and the optimal virtual perspective is calculated, triggering external devices to synchronously display enhanced information; specifically, when the angle between the direction of the digital human arm pointing and the target normal is less than the threshold and the duration exceeds 1.2 seconds, it verifies whether the user's gaze is focused on the same area. If confirmed, a control command is sent to activate the corresponding display function, and supplementary explanations are added to the digital human's mouth.

[0030] Enhanced interaction and perception between virtual explanations and physical exhibits improves the accuracy and depth of audience's spatial information positioning; a multimodal collaborative verification mechanism ensures the accuracy of external device triggers and avoids misoperation from interfering with the presentation process; the synchronous guidance combining virtual and physical elements strengthens the focus of attention, making it easier for audiences to establish content connections and significantly improving the immersive experience and effectiveness of information delivery in science popularization demonstrations.

[0031] Record the timestamp sequence of key events during the interaction process, preload content resources of adjacent exhibition areas, maintain the system's running status, and support seamless switching between multiple scenarios; specifically, this includes: setting a watchdog timer to monitor frame update signals, starting a recovery process if three consecutive frames fail to be drawn, deciding whether to return to the starting point of the previous complete semantic unit or prioritize the emergency prompt action based on the interruption point location, and setting differentiated operation permissions based on the user's identity.

[0032] To ensure the stability and fault tolerance of the system during long-term operation, effectively respond to sudden interruptions or resource delays, and ensure that critical information is not lost and core functions can be recovered; to eliminate content loading wait during scene switching through a preloading mechanism, providing a continuous and smooth visiting experience; and to maintain system configuration and data security while ensuring the freedom of public interaction through a permission-based management strategy, supporting multi-role collaborative operation and maintenance.

[0033] On the other hand, this invention proposes an interactive science popularization system for food safety based on AI digital humans, such as... Figure 2 As shown, it includes: The skeleton construction module is used to build a digital human skeleton driving framework in three-dimensional space, define the topological connection relationship between nodes, bind the mesh vertices to the influence weights of the corresponding bones, and set the zero reference coordinate system under the initial pose. The multi-channel input module is used to receive text stream data output by the speech recognition device based on the skeleton driving framework, capture the key point spatial coordinate sequence of the user's body posture, and integrate click and swipe gesture events on the touch screen; The behavior instruction module is used to generate a joint semantic description based on the integrated multi-channel input, search the preset behavior template library to obtain the high-level behavior instruction sequence, and adjust the behavior priority order based on the historical interaction state. The parameter decomposition module is used to decompose the high-level behavior instruction sequence into a set of primitive actions, dynamically adjust the spatial parameters of the primitive actions based on environmental parameters, and generate a time series of skeletal parameters with interpolation labels. The 3D rendering module is used to drive the 3D engine to update the digital human's pose frame by frame based on the time series of bone parameters, apply the bone transformation matrix to drive mesh deformation, perform viewport rendering, and output the image to the display terminal. The feedback adjustment module is used to continuously monitor changes in the user's gaze focus and head orientation during image output, determine the level of comprehension and acceptance of the current explanatory paragraph, and dynamically adjust the pace of subsequent content based on the acceptance results. The environmental perception module is used to map the three-dimensional coordinate system of key locations based on the content of the explanation and the layout of the exhibition hall, identify the current explanation target and calculate the optimal virtual viewpoint, and trigger external devices to synchronously display enhanced information; The system management module is used to record the timestamp sequence of key events during the interaction process, preload the content resources of adjacent exhibition areas, maintain the system's operating status, and support seamless switching between multiple scenarios.

[0034] The various modules in the above system, when executed, also implement other steps of the interactive science popularization method for food safety based on AI digital humans, as described above: Step 1: Establish a digital human skeleton driving framework in three-dimensional space To achieve precise control over the virtual narrator's full-body movements, a hierarchical skeletal structure must first be constructed within its 3D model, serving as the foundation for all subsequent dynamic performances. This skeleton not only determines the basic rules of model deformation but also undertakes the task of receiving external signals and converting them into specific joint displacements.

[0035] 1.1: Define the topological connections of the digital human model In the 3D modeling environment, a chain-like structure consisting of multiple nodes is created for the anthropomorphic character. Each node represents an anatomically significant joint, such as the shoulder, elbow, or wrist. These nodes are connected sequentially according to biomechanical principles, forming a tree-like branching structure. The root node is located at the center of the pelvis, and the remaining nodes extend outwards from it to the extremities and head. Nodes are connected by directed edges, representing parent-child dependencies, meaning that the position and rotation of a child node are influenced by its direct parent node. This connection relationship is stored in the configuration file as an adjacency list, guiding subsequent skinning weight allocation and inverse dynamics solution.

[0036] For example, the left upper arm node acts as the parent of the left forearm node; when the forearm rotates, the forearm moves accordingly, but the reverse is not true. This unidirectional dependency ensures the consistency of motion propagation direction, avoiding abnormal postures that violate physical laws. Furthermore, to enhance facial expression expressiveness, a separate micro-skeleton network is attached to the skull region, specifically responsible for controlling subtle changes such as eyelid opening and closing, mouth corner twitching, and eyebrow undulation. This network shares the same time reference as the main skeleton but maintains separation in update frequency and driving source to handle language synchronization and emotional feedback tasks separately.

[0037] 1.2: Binding the influence weights of mesh vertices to the corresponding bones After defining the skeleton topology, the surface mesh of the static 3D model needs to be associated with the aforementioned skeletal system, so that each vertex can deform accordingly based on the movement of its assigned bone. A linear blending skinning technique is used, assigning each vertex a set of non-negative values ​​representing the degree to which it is affected by each bone; the sum of all weights is always equal to 1. Let the vertex... by The influence of the root skeleton is denoted by its weight vector. ,satisfy: ; The principle behind this formula is that when a bone undergoes a transformation, the affected vertices will proportionally inherit the bone's displacement and rotation, thus achieving a smooth transition. For example, a vertex near the shoulder may be primarily affected by the combined effects of the clavicle and upper arm bones. If the former has a weight of 0.7 and the latter 0.3, the final position will be the weighted average of the transformed vertices. The weight allocation process can be initially generated using an automatic weight drawing tool, and then manually corrected and optimized for the deformation quality of key areas (such as the armpit and the knee bend) to prevent wrinkle accumulation or volume loss.

[0038] 1.3: Setting the zero-position reference coordinate system under the initial attitude To accurately measure the degree to which subsequent movements deviate from the norm, a standard standing posture must be established as a baseline reference at the beginning of system startup. In this state, the local transformation matrix of each bone is recorded. This includes its translation vector relative to its parent node. With rotation transformation This constitutes the initial attitude set. ,in This represents the total number of bones. This set will be continuously invoked during runtime to calculate the difference between the current pose and the default state, thereby determining whether a specific response action needs to be triggered.

[0039] For example, when a user asks a question, the digital human should change from a static state to a gesture of raising its hand and pointing to the display board. At this point, the system needs to first read the relevant bones of the arm. The value is then combined with the target angle to generate new transformation parameters. Without this reference frame, it is impossible to determine the specific amplitude and direction of the "lifting" action, which can easily cause action drift or misalignment. At the same time, the setting of the zero coordinate also provides a unified standard for subsequent fatigue compensation and posture normalization, ensuring that the continuity of the action is not affected when switching between scenes.

[0040] 1.4: Apply physical constraints to limit the range of motion of the joint. Although digital humans are not bound by real muscles and ligaments, reasonable movement boundaries still need to be introduced to prevent the generation of distorted postures that do not conform to human physiological characteristics. For each pair of movable joints, the maximum and minimum allowable rotation angles are preset to form closed intervals. And a compliance check is performed before each pose update. For 3D rotations, quaternions are typically used. The orientation is indicated, and after converting it to Euler angles, the corresponding axial components are extracted for comparison.

[0041] When the real-time angle of a joint exceeds the limit, the system automatically corrects it to the nearest valid value and sends a reaction force signal to the parent node, triggering a chain reaction of adjustments. For example, if attempting excessive internal rotation of the wrist leads to... In this way, not only will the wrist itself be pulled back, but the forearm may also slightly abduct to relieve tension. This type of constraint data comes from publicly available ergonomic databases and is appropriately relaxed to meet actual demonstration needs, such as allowing slight over-range head rotation to enhance the viewing experience. In this way, while ensuring a natural visual experience, stability over long-term operation is also improved.

[0042] Step 2: Construct a multi-channel input-response path for semantic parsing After completing the structural construction of the digital human entity, the next step is to equip it with the ability to respond to external stimuli. Traditional single-microphone acquisition can only capture sound waveforms, making it difficult to distinguish between semantic intent and environmental noise; while this invention adopts a multi-source sensing collaborative strategy to simultaneously capture speech, gestures, and spatial location information, which are then uniformly encoded and sent to the decision-making center, thereby more comprehensively understanding the user's intent. The core of this pathway design lies in the spatiotemporal alignment and semantic mapping of heterogeneous signals, ensuring that different types of input can be processed equivalently within the same logical framework.

[0043] 2.1: Synchronously acquire text stream data from the speech recognition device High-sensitivity array microphones are deployed within the exhibition area to continuously monitor the audio signals emitted by visitors and utilize sound source localization technology to pinpoint the approximate location of the speaker. The raw audio, after noise reduction filtering, is sent to the speech-to-text unit, outputting a continuous sequence of characters. Each of them This represents a Chinese character or punctuation mark. For ease of subsequent analysis, the text stream is segmented into independent sentence units based on sentence-ending symbols (such as periods and question marks). , forming an ordered set Each statement is accompanied by a timestamp. With confidence score .

[0044] The key to this process lies in maintaining low latency and high robustness, especially in environments with multiple people talking or noisy backgrounds, to accurately extract relevant content. To this end, the system employs a sliding window mechanism, with windows at fixed intervals... An incremental recognition process is performed, processing only audio segments within the most recent time period to reduce accumulated errors. Simultaneously, contextual keyword prediction technology is used to appropriately complete ambiguous pronunciations or discontinuous sentences; for example, "how to eat this…" is inferred as "how to eat this food". The resulting text stream serves as one of the main bases for subsequent intent classification, participating in a comprehensive judgment along with information from other channels.

[0045] 2.2: Capturing the spatial coordinate sequence of key points in the user's body posture A depth camera is placed directly in front of the digital human to capture real-time 3D pose images of visitors and extract the spatial coordinates of key body parts. Using skeletal tracking algorithms, the coordinates of the head, shoulders, hands, and hips are obtained. Real-time position of each key point This constitutes a point cloud sequence that evolves over time. These coordinates are uniformly calibrated with the coordinate system of the digital human as the origin, ensuring that no additional coordinate transformation is required for subsequent calculations.

[0046] The focus is on monitoring hand movement trajectories, as they are frequently used for pointing, gesturing, or indicating, and contain rich interactive intentions. The active phase of a gesture is defined as continuous. The displacement velocity of at least one hand within the frame exceeds the threshold. The time period, i.e.: ; Once the system enters an active period, the gesture pattern recognition process is immediately initiated; otherwise, it remains in standby mode to conserve resources. The obtained spatial coordinate sequence is used not only to identify specific gesture types (such as pointing, waving, and clenching a fist), but also to assist in verifying the authenticity of voice commands—for example, when hearing "Look here" and observing that the user actually points to a certain place, the credibility is significantly improved. This dual verification mechanism effectively reduces the probability of false triggering.

[0047] 2.3: Integrating click and swipe gesture events on the touchscreen An interactive touchscreen display was placed next to the booth, allowing visitors to actively select topics of interest or answer quiz questions. Each time the screen was touched, the operating system reported a series of events, including the starting point. End point Duration And the sequence of intermediate sampling points along the path. Based on these parameters, the basic operation types are divided: if And the displacement is less than If there is obvious directional movement and the length is greater than a certain value, it is considered a light touch; If so, it is considered as sliding.

[0048] All events are packaged into a standardized message format, including action type, occurrence time, area number involved, and additional parameters (such as sliding direction angle). These messages, along with data from the voice and visual channels, enter a central buffer, are sorted by timestamp, and await processing. Specifically, dedicated hot zones are designed for nodes in the food safety knowledge graph. When a user clicks on the "Additives" section, even without speaking, the system can infer their focus and prepare relevant content in advance. This proactive, guided interaction overcomes the limitations of pure voice input in noisy environments.

[0049] 2.4: Aligning time series from multiple sources and generating joint semantic descriptions Because the sampling frequencies and transmission delays of the three input methods are different, time synchronization is necessary to achieve true fusion and understanding. A global clock is selected as the reference, and the voice statements are... attitude frames Touch events Mapped to the nearest time slot respectively Within, a time-aligned set of triples is formed. If a certain type of input has no valid data within the specified time period, it will be filled with a null value.

[0050] Then, the elements within each time slot are traversed, and a rule engine is used to extract co-occurrence features. For example, if the voice "Can you explain this?" is detected at the same time, and the user's action of looking up at the digital face is also captured, then the co-occurrence features are extracted. If there is no interaction on the touchscreen, it is considered a formal question request. Similarly, three quick clicks of the "back" button combined with a head-shaking gesture may be interpreted as impatience or a desire to skip the current step. The final output is a structured semantic description. This summarizes the user's overall intent at that moment, providing a basis for planning the next action.

[0051] Step 3: Generate a sequence of high-level behavioral instructions that match the semantic intent. After obtaining a comprehensive judgment of the user's intent, the system needs to transform it into a series of abstract behavioral instructions that the digital human can execute. These instructions do not directly manipulate skeletal parameters, but rather describe macroscopic action categories such as "start explaining," "point to exhibits," and "express doubt," forming an intermediate bridge to the underlying drivers. The core task of this stage is to establish a semantic-to-behavior mapping dictionary and dynamically arrange the order of instructions according to the context, so that the overall performance conforms to the principles of logical coherence and emotional adaptation.

[0052] 3.1: Search the preset behavior template library based on semantic description. Pre-create a mapping table Each item in the table corresponds to a common interaction scenario and its recommended response. The items in the table are in the following format: The former is a conditional expression, and the latter is an ordered sequence of actions. For example: Condition: "Detection of first entry into the exhibition area and no one speaking for 10 seconds" - Action sequence: [Wave, greet, display welcome banner]; Conditions: "Recognizes interrogative sentences and user's gaze focusing on the guide" - Action sequence: [Nods in confirmation, adjusts posture, initiates voice response]; Conditions: "Received 'Repeat' instruction and current playback description" - Action sequence: [Raise hand to indicate pause, review the previous sentence, and repeat at a slower pace]; When the first Semantic description of each time slot Each entry is compared against the condition fields in the template library to find the first exact match. If multiple candidates exist, the template with higher confidence or higher historical hit rate is selected first. Once a match is found, the corresponding action sequence is retrieved. This serves as the basis for further refinement. If no match is found, the default response process is activated, typically by maintaining a neutral stance and issuing a prompt sound to guide the user to rephrase their statement.

[0053] 3.2: Adjusting behavior priority order based on historical interaction states Given the contextual continuity of dialogue, each input cannot be viewed in isolation. Therefore, before executing the current instruction sequence, it is necessary to review the interaction records from previous rounds. Analyze whether the current request belongs to a branch of an ongoing topic. If it is found... If it is highly relevant to the preceding context, the original narrative rhythm will be maintained, with only necessary supplementary actions inserted; otherwise, it may be necessary to interrupt the current flow and move to a new theme.

[0054] The specific approach is to assign an activity score to each active topic. The initial value is 1.0, and it increases whenever a new statement relates to the topic. Otherwise, it decays exponentially. ,in .when Below the threshold At that point, it is considered that the topic has ended. If a newly triggered instruction pertains to a low-activity topic, its execution priority can be appropriately lowered to avoid frequent interruptions to the user's train of thought. Furthermore, for consecutive questions, a "continuous response mode" can be automatically activated, omitting intermediate transition actions and improving the fluency of communication.

[0055] 3.3: Insert transitional micro-movements to ensure behavioral continuity Directly splicing two significantly different actions (such as suddenly changing from standing to waving wildly) can easily create a visually jarring effect. Therefore, a transitional action is automatically added between adjacent higher-level commands to make the posture change more natural. The selection of the transitional action is based on the distance measurement between the preceding and following states. , defined as the symmetric difference in the set of major joints that need to be activated for both: ; in Indicates the execution of an action The number of joints that change most significantly. If Then insert a segment with a duration of The intermediate states, such as slight shifts in the center of gravity, a downward and then upward gaze, and rising and falling breathing, are general buffering motions. These micro-movements themselves do not carry semantic information, but they can effectively mask abrupt changes in the underlying skeletal parameters, improving visual comfort. The amplitude and speed of the transitional movements follow an S-shaped curve, that is, the acceleration increases first and then decreases, which is consistent with the characteristics of biological motion inertia.

[0056] 3.4: Output a high-level instruction queue with timing markers After priority adjustment and transitional supplementation, a complete stream of instructions to be executed is finally formed. Each element contains an action identifier and an estimated start time. The scheduling follows a non-preemptive principle, meaning that once an action starts, no new, higher-priority instructions (except for emergency stops) are inserted until it finishes. The estimated duration is estimated based on the historical average duration of the action type, with a 10% margin to account for individual variations.

[0057] The queue is written to the shared memory area for the next stage's underlying parser to read. Meanwhile, to prevent unexpected interruptions from causing state corruption, the system maintains an execution log. The system records the actual start and end times of completed actions for subsequent consistency verification and user experience analysis. At this point, the system has completed the entire process of transforming raw input into high-level intent, laying a solid foundation for the generation of specific actions.

[0058] Step 4: Decompose high-level instructions into executable skeletal animation parameter sequences After acquiring the high-level behavioral instruction queue with time planning, it needs to be further broken down into low-level control parameters that can be directly applied to the digital human skeleton. This process involves mapping abstract actions (such as "explaining" or "pointing") to specific joint angle sequences, facial muscle contraction degrees, and overall displacement trajectories. Since the same high-level instruction may be implemented differently due to environmental variables (such as the position of the target object or the distribution of the audience), a parameterized adjustment mechanism needs to be introduced to ensure that the actions conform to the preset specifications and adapt to the on-site conditions.

[0059] 4.1: Parse the action semantic tags of high-level instructions and match primitive action sets. Each high-level instruction Treated as a named label, it is associated with a set of predefined primitive action combinations through a table lookup method. The primitive action is the smallest reusable action unit, such as "bending the elbow 90 degrees", "turning the head 15 degrees to the left", "forming a circle with the lips", etc. Each primitive action comes with a standard execution time. With typical parameter configurations. For example, “pointing to the display board” can be broken down into: [turning head to face the target, focusing eyes, raising right arm to horizontal, extending index finger, and holding for 2 seconds]; while “expressing surprise” includes: [raising eyebrows, widening eyes, slightly opening mouth, and leaning back slightly].

[0060] The matching process employs a combination of string similarity matching and context filtering. If the instruction tag is "Look this way," the system prioritizes searching for primitives containing the semantic meaning of "guiding attention," rather than entries that are literally identical, to enhance generalization capabilities. For complex instructions (such as "walk over and introduce,"), they are split into two parallel sub-task flows, each generating a corresponding primitive sequence, and their timing is coordinated to ensure synchronized movement and speech. All primitive actions are encapsulated in independent namespaces to avoid parameter cross-contamination.

[0061] 4.2: Dynamically adjust the spatial parameters of the primitive actions based on environmental parameters While primitive actions are generally applicable, their specific values ​​need to be dynamically adjusted based on the current scene. Taking "pointing" actions as an example, the arm extension angle depends on the relative position of the target object in three-dimensional space. Let the digital human's position be... The target point coordinates are Then the ideal pointing direction unit vector is: ; The system uses this direction vector to deduce the target rotation angles of the shoulder, elbow, and wrist joints, making the trajectory of the index fingertip approach a straight line. The solution process employs an iterative approximation method, initializing each joint angle to the value of the previous frame and gradually adjusting it until the end effector error is reached. ,in This is the current estimated fingertip position. This method is superior to analytical solutions because it can flexibly handle joint constraints and obstacle avoidance requirements.

[0062] Similarly, when performing the "facing the audience" action, the system reads the set of multiple people's positions provided by the depth camera. Calculate its centroid Then direct the torso toward the target. Orientation. If the audience is too scattered, a rotating gaze strategy is adopted, with the audience facing different areas in turn at regular intervals to reflect inclusivity. All spatial parameters are adjusted before the command is executed to ensure that the action is performed in one go.

[0063] 4.3: Integrating speech rhythm information to adjust lip movements and head shaking frequency The facial movements of a digital human during speech must be highly synchronized with the spoken content, especially the lip movements, which must match the rhythm of the syllables; otherwise, a sense of disharmony will occur. To address this, the system segments the currently playing speech waveform by phonemes, resulting in a series of audio segments. Each syllable corresponds to a standard mouth shape (such as "A", "O", "E" shape, etc.). These shapes are controlled by specific node groups in the facial bones, such as the "zygomaticus major" node that controls the horizontal stretching of the corners of the mouth, and the "orbicularis oris" node that controls the closure of the upper and lower lips.

[0064] On the timeline, mark the start time of each note segment. Aligned with the corresponding lip-sync activation moment, the duration is equal to the length of the syllable segment. At the same time, to mimic the natural slight head movement of a real person speaking, small nodding movements are superimposed at the points of emphasis in sentences, the intensity of which is... With voice energy envelope Positive correlation: in This is a short-term energy value. This represents the maximum energy of the sentence. This is a magnification factor. The nodding motion is characterized by the periodic oscillation of the neck joint along the pitch axis, with the phase synchronized with the accent, and the amplitude varying accordingly. The resulting lip-sync, voice-head movement creates a unified and engaging experience, greatly enhancing the realism and approachability of the expression.

[0065] 4.4: Generate and cache time series of bone parameters with interpolation labels. After all parameter adjustments are complete, the bone transformations required for each primitive motion are packaged into time-series data blocks. Each data block contains a series of snapshots of control parameters at discrete time points, in the following format: ; in For the first The first element in the... Frame timestamps and The first The rotation and translation parameters of the root skeleton at this moment. To ensure smooth animation, quaternion spherical linear interpolation (SLEP) is used for transitions between adjacent frames: ; in The angle between them. This interpolation method can maintain a uniform rotation speed and avoid the gimbal lock problem caused by Euler angles. All data blocks are concatenated into a complete parameter stream in the execution order, temporarily stored in the high-speed cache, waiting for the rendering engine to read it frame by frame and apply it to the digital human model.

[0066] Step 5: Drive the 3D engine to update the digital human's pose frame by frame and render synchronously. After generating and caching the skeletal parameter sequence, the system enters the real-time visualization stage. This stage is led by the graphics processing unit, which reads the pose data to be applied frame by frame according to a preset frame rate and applies it to the skinned 3D model to achieve dynamic deformation and scene fusion. The entire process must strictly follow the timeline to ensure smooth movement, audio-visual synchronization, and timely response to external input. The rendered output is finally presented to the audience through a high-definition display screen or projection device, forming intuitive interactive feedback.

[0067] 5.1: Initialize the graphics rendering pipeline and load the digital human resource pack During the startup phase, a window context is created by calling the graphics interface, and depth testing, lighting models, and anti-aliasing options are configured to build a complete rendering workflow. Subsequently, the digital human-related asset set is loaded from the storage medium, including the basic mesh model, texture maps, normal maps, material parameters, and parameter cache files generated in step four. All resources are categorized by type and stored in a memory pool for easy on-demand access. Specifically, the skeleton structure is reconstructed in a hierarchical tree format in the runtime environment to ensure that each node can be independently addressed and updated.

[0068] To optimize loading efficiency, an asynchronous pre-fetch strategy is adopted, pre-decompressing resource packages for the next potential use case during system standby to reduce switching latency. A resource reference counting mechanism is also established to prevent memory waste caused by repeatedly loading the same content. Once all necessary components are in place, the first rendering loop is triggered, placing the digital human in a default standing posture and displaying it in the center of the screen, with a light gradient background to highlight the main image. Although no interaction has begun at this point, the character already possesses basic physiological characteristics, such as subtle breathing and natural blinking rhythm, enhancing its presence.

[0069] 5.2: Logic for triggering frame-by-frame updates based on the master clock beat Set global refresh rate That is, each Perform a screen redraw. The system internally maintains a high-precision timer to record the cumulative runtime since startup. And at the beginning of each loop, it checks whether the time of the next frame has been reached. If the conditions are met, the current frame processing flow will begin; otherwise, the system will sleep until a critical point is reached before waking up to avoid excessive consumption of computing resources.

[0070] In each frame, the higher-level instruction queue is queried first. Are there any new actions to be taken? Start. If available, locate the corresponding data block from the parameter cache. And write the control parameters of the first frame into the current attitude buffer. If it is in the process of executing an action, then according to The target state of each skeleton is interpolated based on its relative position within the action time period. This process maintains clock alignment with other subsystems such as voice playback and gesture recognition, with errors controlled within a specified range. Within this range, ensure consistency and coordination across multiple channels.

[0071] 5.3: Applying a skeleton transformation matrix to drive mesh deformation After obtaining the complete set of bone parameters required for the current frame, iterate through each bone in turn and perform its local transformation matrix. This is applied to the corresponding key points. Due to the parent-child relationship, the actual spatial position needs to be calculated recursively. ; in This represents the transformation of the bone relative to its parent node. This is its absolute pose in the global coordinate system. The root node's... The default value is the identity matrix. After all world matrices have been calculated, they are passed to the vertex shader program to participate in mesh deformation calculations.

[0072] For each controlled vertex Its final location Determined according to the linear mixing rule: ; Here The original coordinates of the vertex. For the first The weights of the root bones are summed to cover all bones that contribute to that vertex. This formula embodies the core idea of ​​skinning technology: the position of each vertex is the result of the combined effects of the bones it is connected to, with larger weights having a stronger influence. After this step, the static model is transformed into a representation with dynamic poses, ready to enter the lighting and projection stages.

[0073] 5.4: Perform viewport rendering and output the image to the display terminal. After the mesh update is complete, the standard rendering pipeline begins. First, frustum culling is performed to exclude geometry outside the camera's field of view, reducing unnecessary draw calls. Next, the Z-buffer algorithm is used for depth sorting to ensure correct occlusion relationships between objects. Then, per-pixel lighting calculations, texture sampling mapping, and shadow casting are executed sequentially to give the model realistic lighting and shadow textures. Subsurface scattering materials are additionally used on the facial area to simulate the translucency of skin, enhancing realism.

[0074] The final composite image is sent to the frame buffer and, after coordination via vertical synchronization (V-Sync), is pushed to the external display device. If a dual-screen system is connected, the main screen displays the full-body image, while the secondary screen synchronously displays explanatory text or animated illustrations, both sharing the same timeline. Throughout the process, the system continuously listens for user feedback signals. Upon detecting an interruption request (such as a wave indicating stop), it immediately pauses the current action flow and enters a standby state, awaiting further instructions. At this point, the digital human's visual expression is fully closed-loop and can stably serve subsequent interactive tasks.

[0075] Step Six: Establish a two-way feedback mechanism to dynamically adjust the pace of the explanation. To make the science popularization process more aligned with individual cognitive paces, the system incorporates real-time feedback adjustment capabilities, dynamically adjusting information density and explanation speed based on observed user reactions. Unlike one-way broadcasting, this invention allows the digital human to autonomously decide whether to repeat explanations, skip details, or switch topic directions based on the audience's attention level, comprehension progress, and emotional state, thereby achieving personalized communication path planning.

[0076] 6.1: Continuously monitor changes in the user's gaze focus and head orientation. Using 3D pose data provided by a depth camera, the spatial position of the user's eyes can be estimated in real time. This is combined with pupil direction estimation techniques to infer the gaze target. A gaze vector is defined. The center point of the head. If the vector points to the digital face and the distance is less than a threshold. If the deviation exceeds a certain threshold for multiple consecutive frames, it is considered a state of focused listening; Angle, then, is considered a distraction.

[0077] Simultaneously analyze head tilt angle With yaw angle The frequency of fluctuations. A slight nodding motion ( A short, consistent attention span (approximately 0.8–1.5 seconds) usually indicates agreement or encouragement to continue, while frequent head shaking or looking down at the phone suggests diminished interest. These behavioral indicators are quantified as attention scores. The initial value is 1.0, and it is deducted for each instance of inattentiveness detected. Each effective interactive reply increases Overall, it shows a slow decline trend.

[0078] 6.2: Assess the level of comprehension and acceptance of the current explanatory paragraph. By combining voice interaction history with visual feedback, the system assesses the user's understanding of the content being presented. If the user does not raise any questions and maintains eye contact after an explanation, it is considered that they have a basic understanding and is marked as "passed". If they frown, tilt their head in thought, or repeat keywords in a low voice, it is marked as "questionable". If they directly say "I didn't understand" or "Say it again", it is clearly judged as "not mastered".

[0079] The system maintains an understanding status flag for each knowledge node. . After the digital human completes its explanation of this node, wait... Observe the response. If there is no negative feedback, set it to "confirmed"; if there are follow-up questions but the questions have already been answered, keep it "pending"; if multiple explanations are ineffective, downgrade it to "unknown" and try a different way of expressing it. This mechanism avoids blindly pushing forward and causing information to be missed.

[0080] 6.3: Extend or shorten the duration of subsequent content based on acceptance results. Based on the current level of understanding, the density of subsequent information output is dynamically adjusted. If most nodes are quickly confirmed, "efficiency mode" is activated, automatically omitting redundant examples and increasing the speaking speed. To shorten pauses, if repeated questions occur in multiple stages, switch to "detailed mode," insert more analogies and explanations, and slow down the speaking pace. Increase the time spent in the digestive system.

[0081] In practice, the speech rate parameter of the speech synthesis unit is modified. With the duration of the action Proportional relationship: ; in This represents the average attention level over the past minute. This is the adjustment coefficient (taken as 0.3). When When, the speaking speed slows down; when At the same time, increase the speed appropriately. Simultaneously, lengthen or shorten the silence period between each sentence. This allows the information to be absorbed in a way that matches the listener's pace of thinking.

[0082] 6.4: Generate adaptive narrative paths and re-plan action sequences Based on the adjustment of the content rhythm, the overall explanation route is further restructured. If a user shows strong interest in a certain sub-topic (such as asking questions multiple times or actively approaching the display board), relevant extended knowledge points are temporarily inserted to form a branch path; if a user still cannot understand certain professional terms, they are replaced with more colloquial metaphors and accompanied by gestures to help visualize them.

[0083] The system maintains a dynamic priority graph. Nodes represent knowledge items, and edge weights reflect the strength of association and access frequency. Whenever user behavior indicates a preference in a certain direction, the weight of the corresponding node is adjusted. Increase: ; in Score the event type (Question = 2, Click = 1, Views = 0.5). This is the decay factor. The content to be presented is then reordered based on the updated weights, prioritizing high-priority items. Low-weight sections in the original flow can be postponed or hidden, achieving true personalized instruction.

[0084] Step 7: Integrate the environmental perception module to support spatial linkage demonstration To further enhance the immersive experience, the system expands its cognitive ability regarding physical space, enabling the digital human not only to "speak," but also to "point," "move," and interact with real exhibits. By integrating spatial positioning, object recognition, and path planning functions, the virtual character can simulate behaviors such as explaining different exhibits, guiding eye movement, and even simulating walking without leaving the screen, breaking the limitations of traditional two-dimensional interfaces.

[0085] 7.1: Calibrate the 3D coordinate system mapping of key locations within the exhibition hall In the initial deployment phase, laser rangefinders or visual calibration boards were used to survey the exhibition area and determine the virtual positions of the digital humans. Center point of each major exhibit The precise coordinates are determined. All locations are uniformly transformed to a global coordinate system with the exhibition hall entrance as the origin, and a mapping table is established for runtime lookup. The orientation normal vector of each exhibit is also recorded. This is used to determine the best angle for explanation.

[0086] The calibration results are saved as a configuration file, containing the name, type, size, visible location, and recommended explanation distance range for each target. For example, the ideal explanation point for a "preservative information sign" is 1.5 meters directly in front of it, with a viewing angle not exceeding 30 degrees. This type of data provides a geographical basis for subsequent spatial reasoning, ensuring that the digital human's pointing actions match the actual layout.

[0087] 7.2: Identify the current teaching objective and calculate the optimal virtual perspective. When a higher-level instruction pertains to a specific exhibit, the system automatically retrieves its spatial attributes and calculates the optimal viewing posture the digital human should adopt. Let the target location be... The current virtual location is Then the expected direction of the line of sight is Based on this, adjust the rotation angle of the digital human's head and torso so that its front is roughly aligned. direction.

[0088] If the target is too high or too low, the neck tilt angle needs to be fine-tuned. This draws the eye closer to the center of the exhibit. The calculation formula is as follows: ; in For digital human eye coordinates, The target center height is set at this angle. This angle, after being limited, is transmitted to the skeletal controller to drive the head to rotate up and down. Combined with local eye rotation, this achieves the effect of "looking up" or "looking down," enhancing the realism of spatial orientation.

[0089] 7.3: Simulate virtual displacement effects to represent the intention of position transfer. Although the digital human is fixed within the screen, its movement in space can still be simulated through animation. When it is necessary to switch from explaining exhibit A to exhibit B, the system generates a transition animation: first, the body faces the direction of B, one foot is raised in a stepping posture, and a blurred trailing effect is added to the background to create the illusion of moving forward; then, it returns to standing, completing the perspective switch.

[0090] This process does not change the actual coordinates; it only conveys the intention of displacement through changes in posture. To enhance persuasiveness, the volume and reverberation parameters of the explanation are adjusted simultaneously to simulate the acoustic differences caused by distance. ; in To perceive loudness, This is a virtual distance. For reference distance (1 meter), is the environmental reverberation gain. When approaching the exhibit, the sound is slightly amplified and the echo is reduced, as if getting closer to a microphone; the opposite happens when moving away. This audio-visual synesthesia design strengthens the psychological hint of spatial movement.

[0091] 7.4: Trigger the synchronous display of enhanced information for external devices When the digital human points to a certain intelligent exhibit item, its supporting display function can be activated through a wireless communication protocol. For example, when explaining the "cold chain transportation" link and reaching out to indicate the refrigerated cabinet model, the system sends a signal to make the internal lights turn on, the fan start, and the temperature and humidity curves scroll on the screen. Such linkages are bound through preset rules, and the action tags correspond one-to-one with the device IDs.

[0092] The triggering condition is that the included angle between the direction pointed by the digital human's arm and the target normal is less than and the duration exceeds . To prevent accidental touch, it is also necessary to verify whether the user is also focusing on the same area. Once confirmed, a control instruction is immediately issued, and a supplementary note "Please look here" is added to the digital human's mouth to form multiple guides.

[0093] Step 8: Maintain the operation log and support seamless switching between multiple scenarios To ensure the long-term stable operation of the system and adapt to diverse application scenarios, a complete operation management mechanism needs to be established. This mechanism covers functions such as status recording, exception recovery, scene preloading, and permission control, ensuring smooth transitions under different exhibition themes, visiting crowds, or hardware configurations, and enabling self-adjustment and resource allocation without manual intervention.

[0094] 8.1: Record the timestamp sequence of key events for each round of interaction During operation, continuously collect the occurrence moments of important operations and their context information to form structured log entries. Each record includes the event type (such as "start explanation", "user question", "action completed"), the occurrence time , the objects involved (such as exhibit item numbers, statement fragments), the execution result (success / failure), and additional parameters (such as response time, confidence level). All entries are appended to the local storage file in chronological order.

[0095] The log adopts a segmented archiving strategy. Every 1000 entries or after one hour of continuous recording, they are sealed into independent files, and the naming rule includes a date and a serial number prefix. At the same time, keep the summaries of the 10 most frequent events in memory for the real-time monitoring panel to call. These data can be used for later analysis of user behavior patterns, optimizing the coverage rate of the action library, and troubleshooting the causes of occasional failures.

[0096] 8.2: Detect abnormal interruptions and perform state rollback recovery During the execution of an action, power fluctuations, signal loss, or resource access conflicts may occur, forcing the current process to terminate. The system sets a watchdog timer; if the next frame update signal is not received within the expected time, it is considered a lag; if three consecutive frames fail to be drawn, the recovery process is initiated.

[0097] The recovery strategy depends on the location of the interruption point: if it occurs in the middle of a non-critical action (such as a regular explanation), it will revert to the starting point of the previous complete semantic unit and replay; if it occurs during an emergency alert (such as a safety warning), the action will be completed first before returning to the original position. All incomplete instructions are marked as "interrupted" and will not be counted in subsequent scheduling. If multiple failures occur consecutively, the rendering process will be automatically restarted and a backup resource package will be loaded to maintain service availability to the maximum extent.

[0098] 8.3: Preload content resources from adjacent exhibition areas to achieve smooth transitions. When the system determines that the current topic is about to end (e.g., due to prolonged inactivity or user selection of the next area), it downloads knowledge materials, action templates, and space configurations from neighboring exhibition areas from the server in advance. The download process occurs in the background, limiting bandwidth usage to no more than 30% of the total capacity, and does not affect current operating performance.

[0099] Once resources are in place, integrity checks and dependency resolution are performed to confirm there are no missing textures or link errors. Upon detecting a transition command (such as the voice prompt "Go to the next place" or a touch swipe), the current screen is immediately frozen, a fade-out animation plays, the scene switches to the initial layout of the new scene, and the welcome sequence begins. The entire process is completed within 2 seconds to avoid any waiting time. Resources from older scenes are gradually released after they are confirmed to be no longer needed to prevent memory accumulation.

[0100] 8.4: Enable different levels of access points based on permission levels Different access permissions are set based on different user identities (such as regular visitors, administrators, and maintenance personnel). Visitors can only ask questions or select content via voice and gestures; administrators can use a dedicated terminal to enable presentation mode, adjust volume, or view statistical reports; and maintenance personnel have full permissions to access logs, update resource packages, and calibrate sensors.

[0101] Access verification is implemented through an encrypted token mechanism, requiring scanning a QR code or entering a dynamic password each time you log in. High-risk operations (such as deleting data or formatting storage) require double confirmation, and the operator's ID and time are recorded. All permission changes are synchronized to the cloud for record-keeping, supporting remote auditing and traceability. Through hierarchical management, both the convenience of public use and the security boundaries of system operation and maintenance are ensured.

[0102] This invention utilizes a high-performance graphics processing unit as its computational foundation, combining natural language understanding and human kinematics simulation to drive a virtual character to complete a full closed loop from perception to feedback within a unified timeline. Compared to fixed-path animation playback modes, this invention emphasizes real-time performance, contextual adaptability, and multimodal fusion capabilities, enabling the digital human to autonomously organize its expression logic and enhance information delivery through body language when facing audiences of different ages, cognitive levels, or interests.

[0103] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. An interactive science popularization method for food safety based on AI digital humans, characterized in that, Includes the following steps: A digital human skeleton driving framework is constructed in three-dimensional space, the topological connection relationship between nodes is defined, the influence weight of the mesh vertices is bound to the corresponding bones, and the zero reference coordinate system under the initial posture is set. Based on the skeleton-driven framework, the system receives text stream data output from the speech recognition device, captures the spatial coordinate sequence of key points of the user's body posture, and integrates click and swipe gesture events on the touch screen. It generates a joint semantic description based on the integrated multi-channel input, searches a preset behavior template library to obtain high-level behavior instruction sequences, and adjusts the behavior priority order based on historical interaction states. The high-level behavior instruction sequences are decomposed into primitive action sets, and the spatial parameters of the primitive actions are dynamically adjusted based on environmental parameters to generate a skeleton parameter time series with interpolation labels. Based on the skeleton parameter time series, the 3D engine is driven to update the digital human's posture frame by frame, and a skeleton transformation matrix is ​​applied to drive mesh deformation. Viewport rendering is then performed, and the image is output to the display terminal. During image output, the system continuously monitors changes in the user's gaze focus and head orientation to determine the comprehension and acceptance level of the current explanatory segment. Based on the acceptance results, the system dynamically adjusts the pace of subsequent content. It also maps the three-dimensional coordinate system of key locations according to the explanatory content and exhibition hall layout, identifies the current explanatory target, calculates the optimal virtual perspective, and triggers external devices to synchronously display enhanced information. Record the timestamp sequence of key events during the interaction process, preload content resources of adjacent exhibition areas, maintain the system's operating status, and support seamless switching between multiple scenarios.

2. The interactive science popularization method for food safety based on AI digital humans according to claim 1, characterized in that, The construction of the digital human skeleton driving framework in three-dimensional space also includes: presetting the maximum and minimum rotation angles for each pair of movable joints to form a closed interval, performing compliance checks before each posture update, and automatically correcting the joint's real-time angle to the nearest valid value and transmitting the reaction force signal to the parent node when the joint's real-time angle exceeds the limit range.

3. The interactive science popularization method for food safety based on AI digital humans according to claim 1, characterized in that, The generation of joint semantic description includes: mapping voice sentences, gesture frames and touch events to the nearest time slot to form a set of triples; if there is no valid data for a certain type of input in the time period, it is filled with null values; traversing the elements in each time slot to extract co-occurrence features; and outputting a structured semantic description.

4. The interactive science popularization method for food safety based on AI digital humans according to claim 1, characterized in that, The process of decomposing the high-level behavior instruction sequence into a set of primitive actions further includes: inserting transitional micro-actions between adjacent high-level instructions, calculating the symmetry difference between the sets of major joints that need to be activated, and if the distance metric exceeds a threshold, inserting an intermediate state of a certain duration so that the posture change follows the S-shaped curve change pattern.

5. The interactive science popularization method for food safety based on AI digital humans according to claim 1, characterized in that, The generation of the time series of skeletal parameters with interpolation markers also includes: segmenting the speech waveform into a segment sequence by phonemes, with each segment corresponding to a standard mouth shape posture, aligning the start time of the segment with the mouth shape activation time, and adjusting the head shaking amplitude according to the speech energy envelope to synchronize the nodding phase with the accent.

6. The interactive science popularization method for food safety based on AI digital humans according to claim 1, characterized in that, The method of driving the 3D engine to update the digital human pose frame by frame also includes: setting the global refresh rate to 60Hz, maintaining a high-precision timer to record the cumulative running time, checking whether the next frame time has been reached at the beginning of each loop, locating the corresponding data block from the parameter buffer and writing the control parameters into the current pose buffer.

7. The interactive science popularization method for food safety based on AI digital humans according to claim 1, characterized in that, The method of dynamically adjusting the pace of subsequent content based on acceptance results also includes: maintaining an understanding status flag for each knowledge node, updating the flag status based on user feedback within a preset waiting time, modifying the speech rate parameter of the speech synthesis unit to be proportional to the average attention value, and extending or shortening the silence period between each sentence.

8. The interactive science popularization method for food safety based on AI digital humans according to claim 1, characterized in that, The method of triggering external devices to synchronously display enhanced information also includes: when the angle between the direction in which the digital human arm points and the target normal is less than a threshold and the duration exceeds 1.2 seconds, verifying whether the user's gaze is focused on the same area; if confirmed, sending a control command to activate the corresponding display function, and adding supplementary explanations to the digital human's mouth.

9. The interactive science popularization method for food safety based on AI digital humans according to claim 1, characterized in that, The maintenance system operation status and support for seamless switching between multiple scenarios also include: setting a watchdog timer to monitor frame update signals; if three consecutive frames fail to be drawn, a recovery process is initiated; the system determines whether to return to the starting point of the previous complete semantic unit or prioritize the emergency prompt action based on the interruption point location; and setting differentiated operation permissions based on the user's identity.

10. A food safety interactive science popularization system based on AI digital human for implementing the method as described in any one of claims 1-9, characterized in that, include: The skeleton construction module is used to build a digital human skeleton driving framework in three-dimensional space, define the topological connection relationship between nodes, bind the mesh vertices to the influence weights of the corresponding bones, and set the zero reference coordinate system under the initial pose. The multi-channel input module is used to receive text stream data output by the speech recognition device based on the skeleton driving framework, capture the key point spatial coordinate sequence of the user's body posture, and integrate click and swipe gesture events on the touch screen; The behavior instruction module is used to generate a joint semantic description based on the integrated multi-channel input, search the preset behavior template library to obtain the high-level behavior instruction sequence, and adjust the behavior priority order based on the historical interaction state. The parameter decomposition module is used to decompose the high-level behavior instruction sequence into a set of primitive actions, dynamically adjust the spatial parameters of the primitive actions based on environmental parameters, and generate a time series of skeletal parameters with interpolation labels. The 3D rendering module is used to drive the 3D engine to update the digital human's pose frame by frame based on the time series of bone parameters, apply the bone transformation matrix to drive mesh deformation, perform viewport rendering, and output the image to the display terminal. The feedback adjustment module is used to continuously monitor changes in the user's gaze focus and head orientation during image output, determine the level of comprehension and acceptance of the current explanatory paragraph, and dynamically adjust the pace of subsequent content based on the acceptance results. The environmental perception module is used to map the three-dimensional coordinate system of key locations based on the content of the explanation and the layout of the exhibition hall, identify the current explanation target and calculate the optimal virtual viewpoint, and trigger external devices to synchronously display enhanced information; The system management module is used to record the timestamp sequence of key events during the interaction process, preload the content resources of adjacent exhibition areas, maintain the system's operating status, and support seamless switching between multiple scenarios.