A forest treehouse popularization system fusing multi-modal perception and deep learning feedback
Patent Information
- Application Number
- CN202610998557.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]为克服以上技术中存在的问题,本发明提供了一种融合多模态感知与深度学习反馈的森林树屋科普系统,旨在解决现有技术中儿童行为识别模型冷启动困难、生物识别与儿童行为意图相互割裂、对网络依赖强可能导致场馆内设备延迟、缺乏个性化和沉浸式科普体验等缺陷,所述系统包括:
[0019]本发明的有益效果是:本发明提供了一种融合多模态感知与深度学习反馈的森林树屋科普系统,该系统具有以下优点:1,通过第一预训练模型与第二预训练模型的协同架构,实现了动物/植物精准识别与儿童行为意图的自适应理解:本发明设置了两个独立且协作的预训练模型:第一预训练模型专门对南岭森林场景中的动物物种、植物物种及其行为/物候特征进行分析;第二预训练模型采用“广义预训练+在线增量微调”的两阶段自适应训练策略;这一架构使得系统既能精确识别目标生态对象,又能持续适应现场儿童的行为特点,克服了现有技术中模型静态部署且无法利用隐式反馈进化的缺陷;2,第二预训练模型的两阶段自适应训练解决了“冷启动”问题,实现了开机即用与越用越准:部署前,第二预训练模型通过广义人类多模态数据预训练获得通用行为理解能力,保证设备首次开机即可基于先验映射规则库生成基础反馈,避免了冷启动盲区;部署后,系统自动采集儿童自然互动中的隐式反馈信号对模型进行在线增量微调,逐步进化为独特性第二预训练模型;该过程无需人工标注,识别准确率随使用次数增加而提升,实现了从通用匹配到个性化精细匹配的无缝过渡;3,行为与生态意图映射模块的两阶段匹配策略建立了从儿童行为到生态内容的语义闭环:该模块第一阶段利用广义第二预训练模型输出的通用行为标签和先验规则库进行快速匹配,保证系统即时响应;第二阶段(个性化阶段)利用进化后的独特性第二预训练模型输出的细粒度行为与目的标签,结合第一预训练模型的生物信息进行精细匹配;针对模仿类行为采用多模态相似度计算公式(融合动作、叫声和指向),针对探索类行为根据生物类型分别检索对应知识库;这种两阶段差异化的匹配机制,彻底解决了现有技术中生物识别与儿童行为识别相互割裂的问题,使科普交互自然流畅和教育内容精准触达;4,非接触式多模态感知与本地化边缘计算保障了户外沉浸式体验:通过深度摄像头非接触式捕捉儿童骨骼关键点,无需佩戴任何设备;所有感知、识别、映射和反馈均在本地边缘端完成,结合太阳能混合供电和嵌入式数据库,可在无网络环境中全天候稳定运行,防止网络信号不足导致设备延迟;同时,通过环绕式显示、多声道音响、可编程灯光和环境特效装置实现声、光、画和效的联动反馈,如:动物3D动画、真实叫声、植物生长展示和昼夜模拟,显著增强了体验真实感,并支持对动物和植物的全面科普。
Smart Images

Figure CN122818019A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of forestry science education technology, and more specifically, to a forest treehouse science education system that integrates multimodal perception and deep learning feedback, which is particularly suitable for immersive forest ecology science education scenarios for children. Background Technology
[0002] In recent years, with the development of artificial intelligence, multimodal perception, and immersive interactive technologies, intelligent science education systems for children have gradually attracted attention. Among existing technologies, researchers have developed behavior analysis systems based on multimodal deep learning, using devices such as RGB-D cameras and microphone arrays to collect children's multimodal data for behavioral recognition and psychological monitoring. In the field of ecological science education, immersive eco-house interactive learning systems have emerged, providing ecological image displays and voice Q&A through display modules and image recognition. Treehouse-type amusement facilities with science education functions achieve basic science education displays through static devices such as plant viewing balls and seed display walls. Furthermore, deep learning-based animal recognition and voiceprint monitoring technologies have been applied in nature reserves, capable of identifying specific animal species and their calls. However, the above technologies still have significant shortcomings in the specific scenario of forest treehouse science education for children.
[0003] Current technologies lack the ability to dynamically map children's real-time behavioral intentions to forest ecological content in a closed loop. Treehouse-like facilities are mostly static displays, unable to automatically match animal behavior animations, plant knowledge, or ecological feedback based on children's imitation, pointing, and exploration behaviors. Furthermore, children's behavior recognition models are mostly statically deployed, resulting in a "cold start" blind spot when devices are first turned on due to a lack of data on children in the target location. Existing incremental learning methods rely on manual annotation or explicit datasets, failing to utilize implicit feedback signals from natural interactions to achieve online adaptive evolution without human intervention. In addition, biometric and children's behavior recognition modules are disconnected, lacking a progressive semantic mapping mechanism from general matching to fine-grained matching, and failing to differentiate between imitation and exploratory behaviors. The system's feedback is limited, making it difficult to combine children's historical interests and cognitive levels to achieve personalized learning paths. It also relies on cloud networks, resulting in poor usability in areas with weak signals, such as forests, and lacks effective support for plant science education. Summary of the Invention
[0004] To overcome the problems existing in the above technologies, this invention provides a forest treehouse science popularization system that integrates multimodal perception and deep learning feedback. It aims to solve the shortcomings of existing technologies, such as the difficulty in cold-starting children's behavior recognition models, the disconnect between biometrics and children's behavioral intentions, strong network dependence leading to equipment delays in venues, and a lack of personalized and immersive science popularization experiences. The system includes: The biomimetic treehouse module is used to construct a Nanling forest scene and integrates display, sound, lighting, and environmental effects devices as the physical carrier of an immersive learning space and the terminal execution platform for interactive feedback. Specifically, the biomimetic treehouse module constructs the Nanling forest scene through 3D biomimetic modeling and environmental restoration technology. The biomimetic treehouse module integrates a surround display screen, multi-channel sound system, programmable LED lighting, and environmental effects devices. The Nanling forest scene refers to a virtual ecosystem generated through 3D modeling and multi-channel environmental sound effect restoration technology, which includes unique animal species, typical vegetation distribution, and microclimate characteristics of the Nanling region.
[0005] A multimodal perception and acquisition module is used to simultaneously acquire images, voice, and limb movement information of children through cameras, microphones, and motion sensors. Specifically, the cameras in the multimodal perception and acquisition module include an RGB camera and a depth camera. The RGB camera is used to acquire color facial images and full-body images of children for face recognition and expression recognition. The depth camera, as the motion sensor, is used to achieve non-contact limb movement capture by acquiring the child's three-dimensional skeletal key point sequence for recognizing imitation movements, pointing movements, and sleeping postures. The microphone is a microphone array used to acquire children's voice and environmental sounds.
[0006] The deep learning intelligent processing module, based on a lightweight AI chip, performs edge-end fusion analysis on the collected data. It includes: a first pre-trained model, used to analyze biological and environmental data from the Nanling forest scene, outputting biological species and their behavioral labels, including animal and plant species; and a second pre-trained model, employing a two-stage adaptive training: before deployment, a generalized second pre-trained model is obtained through pre-training with generalized human multimodal data; after deployment, the generalized second pre-trained model is incrementally fine-tuned online using implicit feedback signals generated during children's interactions, gradually forming a unique second pre-trained model used to identify children's specific behaviors and interaction purposes. Specifically, the first pre-trained model adopts an architecture combining convolutional neural networks and recurrent neural networks, using the biological data from the Nanling forest scene as input. The biological data includes: animal image sequences and plant species image sequences collected by a camera, Mel-spectrum maps of animal calls collected by a microphone, animal posture skeleton data collected by a depth camera, and environmental data collected by environmental sensors.
[0007] Specifically, the first pre-trained model is equipped with a multimodal feature fusion layer. The processing flow of the multimodal feature fusion layer includes: first, inputting data from four different modalities in the biological data into the convolutional neural network, and having the convolutional neural network extract their respective spatial feature vectors; then, inputting the four extracted spatial feature vectors into the multimodal feature fusion layer, and merging them into a joint feature vector by concatenation or weighted summation to achieve complementarity and enhancement of cross-modal information; finally, inputting the joint feature vector into the recurrent neural network for temporal dependency modeling, and the final output of the recurrent neural network is a biological species classification label and a biological behavior category label; wherein, the behavioral category label of plant species includes phenological features such as flowering period, fruiting period, and leaf color change.
[0008] Specifically, the generalized human multimodal data in the second pre-trained model before deployment includes: facial images, full-body posture images, voice commands, and body movement data of people of different ages; after deployment, the observable feedback generated during the natural interaction between the child and the system is used as implicit supervision signals to perform online incremental fine-tuning of the generalized second pre-trained model, gradually forming a unique second pre-trained model; the implicit feedback signals include the frequency of the child repeating the same action, the duration of the pause after feedback, and the pleasant and confused voices voluntarily emitted.
[0009] The behavior and ecological intent mapping module employs a two-stage matching strategy: In the first stage, the generalized human behavior labels output by the generalized second pre-trained model and the biological information output by the first pre-trained model are used to perform preliminary semantic alignment based on a priori mapping rule base to generate basic feedback instructions; In the second stage, the fine-grained behavior and purpose labels output by the unique second pre-trained model, which has been formed through online incremental fine-tuning, are combined with the biological information of the first pre-trained model to perform refined semantic matching to generate optimized feedback instructions.
[0010] Specifically, the switching between the first and second stages of the behavior and ecological intent mapping module is gradual and automatic; as the system runs longer and on-site children's interaction data accumulates, the generalized second pre-trained model gradually evolves into a unique second pre-trained model through online incremental fine-tuning; the behavior and ecological intent mapping module automatically switches the input source generalized second pre-trained model to the evolved unique second pre-trained model, achieving a seamless transition from cold-start general matching to personalized fine matching.
[0011] The behavior and ecological intent mapping module executes the following matching rules: When a child's behavior is imitative and the interaction aims to seek an animal response, the module automatically matches the most similar animal species based on the child's imitation characteristics and generates a feedback instruction that displays animations of typical animal behaviors and realistic sounds. When a child's behavior is exploratory and the interaction aims to acquire knowledge, the module dynamically selects appropriate science content from the science database based on the biological species the child is currently interested in, combined with their historical interests and cognitive level, and generates a knowledge push instruction. Specifically, if the child is interested in a plant species, the module pushes plant science knowledge about the plant's name, morphological characteristics, and ecological functions; if the child is interested in an animal species, the module pushes animal ecology knowledge about the animal's habitat, diet, and behavioral habits. If both imitation and exploration objectives exist simultaneously, the module executes them sequentially, prioritizing imitation responses followed by knowledge pushes.
[0012] In the first stage, the behavior and ecological intent mapping module uses a priori mapping rule base for matching. The priori mapping rule base includes multiple preset correspondences, and each rule is in the form of: general human behavior labels are mapped to biological species feature matching rules. For plants, the rule includes: children pointing to green objects are mapped to matching nearby plant species.
[0013] In the second stage, when the behavior and ecological intention mapping module identifies a child's behavior as imitative, it uses a similarity matching formula to calculate a fine-grained matching score between the child's imitative features and the animal species features.
[0014] Where β represents the child's current fine-grained behavioral characteristics output by the second pre-trained model, and α represents the animal species identifier output by the first pre-trained model; and These are the extracted child action feature vectors and the animal feature vectors extracted from the typical behavior database of animal species, respectively, with cos(·) representing the cosine similarity; MFCC(β) sound α sound ) represents the Mel frequency cepstral similarity between the feature vectors of children's vocalizations and those of animal vocalizations, β sound α represents the feature vector of a child's vocalizations. sound For animal vocalization feature vectors; 1 point (β, α) is an indicator function, which takes a value of 1 when the direction of the hand pointing ray calculated based on the child's skeletal key points intersects with the display area of animal α, and otherwise takes a value of 0; ω1, ω2, and ω3 are preset weight coefficients, all of which are positive numbers and satisfy ω1 + ω2 + ω3 = 1; the behavior and ecological intention mapping module selects the animal α with the highest score. * =argmax αScore(β, α) is the matching result.
[0015] In the second stage, when a child's behavior is identified as exploratory, the behavior-ecological intent mapping module does not use a similarity matching formula. Instead, it retrieves corresponding knowledge content from a science database based on the biological species identifiers currently of interest to the child identified by the first pre-trained model, combined with the child's historical interests and cognitive level: if the organism is a plant, it retrieves plant science knowledge such as plant name, morphological characteristics, and ecological functions; if the organism is an animal, it retrieves animal ecological knowledge such as animal habitat, diet, and behavioral habits, and generates corresponding knowledge push instructions.
[0016] The interactive feedback execution module is used to drive the devices in the bionic treehouse module according to the feedback instructions issued by the behavior and ecological intention mapping module, so as to realize multimodal linkage feedback. Specifically, the interactive feedback execution module realizes multimodal linkage feedback in the following ways: when receiving an instruction to display animal behavior, it drives the display screen to play the 3D behavior animation of the animal, drives the sound system to play the corresponding animal's real call synchronously, drives the lights to change according to the preset animal emotion color scheme, and drives the environmental special effects device to release a breeze or mist; when receiving a knowledge push instruction, it drives the display screen to display pictures and / or short videos, drives the sound system to explain in a child-friendly voice, and drives the lights to adjust to a soft and focused mode; when the second pre-trained model recognizes that the child is making a sleeping movement, it automatically triggers the simulation of forest day and night changes, including: the screen brightness gradually dims, the lights switch to a warm yellow moonlight mode, and plays nocturnal animal calls and insect chirps.
[0017] The science popularization database and power supply control module provide a localized Nanling biological knowledge base and employ a hybrid power supply strategy to ensure system operation. Specifically, the science popularization database within the module is stored locally using an embedded method. The database includes tables for plant and animal species, behavioral characteristics, vocalization samples, and ecological knowledge; all tables support millisecond-level access even without a network connection. The power supply control module uses a hybrid power supply system consisting of a maximum power point tracking solar controller and a lithium iron phosphate battery pack. By monitoring the instantaneous power consumption of each module in real time, and combining the battery's state of charge and solar input power, it dynamically adjusts power supply priorities. When solar energy is sufficient, it charges the battery and powers the entire module; when the battery is low, it automatically reduces the display screen brightness and shuts down non-core environmental effects devices, ensuring the continuous operation of the core perception and recognition modules.
[0018] Finally, the deployment of all modules of the system is completed locally at the edge, without relying on the cloud network; the lightweight AI chip adopts a system-on-a-chip with a neural network acceleration unit to perform inference and incremental fine-tuning operations of the first and second pre-trained models locally.
[0019] The beneficial effects of this invention are as follows: This invention provides a forest treehouse science popularization system that integrates multimodal perception and deep learning feedback. This system has the following advantages: 1. Through the collaborative architecture of a first pre-trained model and a second pre-trained model, it achieves accurate animal / plant identification and adaptive understanding of children's behavioral intentions: This invention sets up two independent and collaborative pre-trained models: the first pre-trained model specifically analyzes animal and plant species and their behavioral / phenological characteristics in the Nanling forest scene; the second pre-trained model adopts a two-stage adaptive training strategy of "generalized pre-training + online incremental fine-tuning"; this architecture enables the system to accurately identify target ecological objects and continuously adapt to the behavioral characteristics of children on site, overcoming the limitations of existing technologies. The first model suffers from the drawbacks of static deployment and inability to evolve using implicit feedback. The second pre-trained model addresses the "cold start" problem through two-stage adaptive training, achieving both immediate usability and increasing accuracy with use. Before deployment, the second pre-trained model acquires general behavioral understanding capabilities through pre-training with generalized human multimodal data, ensuring that basic feedback can be generated based on the prior mapping rule base upon device startup, avoiding the cold start blind spot. After deployment, the system automatically collects implicit feedback signals from children's natural interactions to perform online incremental fine-tuning of the model, gradually evolving into a unique second pre-trained model. This process requires no manual annotation, and the recognition accuracy increases with usage, achieving a seamless transition from general matching to personalized fine-grained matching. The third aspect involves behavioral and ecological awareness. The graph mapping module's two-stage matching strategy establishes a semantic closed loop from children's behavior to ecological content: In the first stage, the module uses generalized behavioral labels output by the generalized second pre-trained model and a priori rule base for rapid matching, ensuring immediate system response; in the second stage (personalization stage), it utilizes fine-grained behavioral and purpose labels output by the evolved unique second pre-trained model, combined with biological information from the first pre-trained model, for refined matching; for imitative behaviors, a multimodal similarity calculation formula (integrating actions, vocalizations, and pointing) is used; for exploratory behaviors, corresponding knowledge bases are retrieved based on biological type. This two-stage differentiated matching mechanism completely solves the problem of the separation between biometrics and children's behavior recognition in existing technologies, enabling scientific... The interactive experience is natural and smooth, and educational content is precisely delivered; 4. Non-contact multimodal perception and localized edge computing ensure an immersive outdoor experience: key points of children's skeletons are captured non-contactly through depth cameras without the need for any devices to be worn; all perception, recognition, mapping, and feedback are completed at the local edge, combined with solar hybrid power supply and embedded database, which can operate stably in all weather conditions in environments without network, preventing device delays caused by insufficient network signal; at the same time, the interactive feedback of sound, light, picture, and effect is achieved through surround display, multi-channel sound, programmable lighting, and environmental effects devices, such as: animal 3D animation, realistic sounds, plant growth display, and day and night simulation, which significantly enhances the realism of the experience and supports comprehensive popular science education on animals and plants. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of a forest treehouse science popularization system that integrates multimodal perception and deep learning feedback according to the present invention. Detailed Implementation
[0021] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments given herein are only for illustration and explanation of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions or improvements made by those skilled in the art based on the present invention should be included within the protection scope of the present invention.
[0022] like Figure 1 The diagram shows a forest treehouse science popularization system integrating multimodal perception and deep learning feedback according to the present invention. The diagram includes: a biomimetic treehouse module S10; a multimodal perception acquisition module S20; a deep learning intelligent processing module S30; a behavior and ecological intention mapping module S40; an interactive feedback execution module S50; and a science popularization database and power supply control module S60. The deep learning intelligent processing module S30 includes: a first pre-trained model S301 and a second pre-trained model S302; the behavior and ecological intention mapping module S40 includes: a first stage S401 and a second stage S402. The specific implementation methods of each module are described in detail below with reference to the accompanying drawings.
[0023] The biomimetic treehouse module S10 serves as the physical carrier and immersive learning space of the entire system. Its construction method is as follows: First, a Nanling forest scene is constructed using 3D biomimetic modeling technology. Specifically, 3D scanning or manual modeling methods are used to generate models of the terrain, vegetation, and animals unique to the Nanling region, including animal species such as the Cabot's Tragopan, macaque, and silver pheasant, as well as plant species such as pine, nanmu, and rhododendron. Then, using environmental restoration technology, a multi-channel environmental sound system simulates the natural sounds of the Nanling forest, such as birdsong, stream sounds, and wind sounds, while combining microclimate data, such as temperature, humidity, and light, to generate a realistic virtual ecosystem.
[0024] In terms of hardware integration, the bionic treehouse module S10 has a surround display screen installed on the walls or ceiling, such as multiple high-resolution LCD splicing screens or a projection fusion system; the surround display screen displays dynamic panoramic forest images and 3D biological animations; the multi-channel sound system uses a spatial audio array based on wave field synthesis technology, which can dynamically adjust the sound image positioning according to the child's real-time position in the treehouse to create a realistic surround sound experience; the programmable LED light array uses RGBW four-channel dimmable lamp beads, receives control signals through the DMX512 protocol, and can simulate continuous changes from the color temperature of sunlight 6500K to the color temperature of moonlight 2700K; the environmental effects devices include an ultrasonic atomizer (used to simulate forest morning fog) and a low-frequency airflow generator (used to simulate natural breeze), all of which are driven by the subsequent interactive feedback execution module S50; the entire bionic treehouse module S10, as the terminal execution platform, together with the perception, recognition, and mapping links, forms a complete closed-loop link of "perception-recognition-mapping-presentation".
[0025] The multimodal perception and acquisition module S20 is responsible for simultaneously acquiring image, voice, and body movement information of children inside the bionic treehouse. Its specific implementation is as follows: The module includes an RGB camera, a depth camera, and a microphone array. The RGB camera is a high-resolution color camera that is positioned at multiple angles in the treehouse to capture color facial images and full-body images of children in real time. The captured image data is sent to the deep learning intelligent processing module S30 for face recognition and expression recognition. Face recognition is used to distinguish different children to create personalized user profiles, while expression recognition is used to determine implicit emotional states such as pleasure and confusion.
[0026] The depth camera uses a depth sensor based on the time-of-flight method or structured light principle as a motion sensor. It achieves non-contact limb motion capture by acquiring a three-dimensional skeletal key point sequence of the child's entire body. The specific technical process is as follows: the depth camera outputs a depth image, which is then segmented from the background to extract the human body contour. Then, a human pose estimation algorithm, such as 2D / 3D pose regression based on a convolutional neural network, is used to obtain the three-dimensional coordinates of 28 or more joints, forming a skeletal key point sequence. This sequence can recognize the child's specific movements, including: imitating animal movements (such as: scratching ears and cheeks, flapping wings and jumping), pointing movements (the relationship between the extension direction of the hand joints and the target area), and sleeping postures (curling up and turning the head to the side, etc.). The entire capture process does not require the child to wear any inertial measurement unit or wearable device.
[0027] The microphone array uses a four- or six-microphone circular array, placed on the roof of a tree or on a wall, to enhance children's voice signals and suppress environmental echoes through beamforming algorithms. It collects children's voice commands such as "What animal is this?" "What does it eat?" and imitates animal calls and natural sounds in the environment. After noise reduction and sound source localization, the collected multi-channel audio signals are output as clean voice signals and call characteristics.
[0028] The deep learning intelligent processing module S30 is based on a lightweight AI chip, such as a system-on-a-chip with a neural network acceleration unit, to perform edge-end fusion analysis on the collected multimodal data. This module does not rely on the cloud network, and all calculations are completed locally. Specifically, it includes a first pre-trained model S301 and a second pre-trained model S302, which run independently and cooperate with each other.
[0029] The first pre-trained model S301 adopts an architecture that combines convolutional neural networks and recurrent neural networks to analyze biological and environmental data in the Nanling forest scene and output biological species and their behavioral labels; among them, biological species include animal species and plant species.
[0030] Training data source: The first pre-trained model S301 was pre-trained using a large amount of labeled data from the Nanling forest scene before deployment; this biological data specifically includes: Animal image sequences captured by RGB cameras, such as Cabot's Tragopan and macaques in different postures and lighting conditions; plant image sequences, such as pine needles, camphor tree leaves and rhododendron flowers at different growth stages.
[0031] Mel spectrograms of animal sounds captured by microphones, such as the calls of macaques and silver pheasants.
[0032] Animal posture skeletal data collected through depth cameras, such as three-dimensional keypoint sequences when animals are walking, foraging, and on alert.
[0033] Environmental data collected by environmental sensors includes temperature, humidity, and light intensity.
[0034] During runtime, the processing flow of the first pre-trained model S301 is as follows: First, the data from the different modalities are input into convolutional neural networks. Specifically, animal image sequences and plant image sequences are processed by a shared or independent convolutional neural network, such as ResNet, to extract spatial feature vectors. The Mel spectrogram of animal calls is processed by another convolutional neural network to extract acoustic feature vectors. Animal posture skeleton data is processed by a spatiotemporal graph convolutional network to extract motion feature vectors. Environmental data is processed by a fully connected network to extract environmental feature vectors. Among these, the Mel spectrogram of animal calls is a time-frequency spectrum generated by framing and Fourier transforming the collected animal call audio signals, and mapping the frequency and logarithmic energy according to the Mel auditory scale of the human ear. It is used to characterize the auditory frequency domain features of the calls of different species.
[0035] Then, the first pre-trained model S301 is equipped with a multimodal feature fusion layer; this multimodal feature fusion layer concatenates the above multiple feature vectors in the feature dimension, or fuses them into a joint feature vector by weighted summation; the weights of the weighted summation can be obtained by the network self-learning according to the importance of each modality to the current task; through fusion, the complementarity and enhancement of cross-modal information are realized, for example, making the animal's image appearance and vocal features corroborate each other, and the posture information and environmental information correlated.
[0036] Finally, the fused joint feature vector is input into a recurrent neural network, such as a long short-term memory network or a gated recurrent unit, to model temporal dependencies. Since animal behavior and plant phenological changes are temporally continuous, the recurrent neural network can capture dynamic changes between adjacent frames. The final output of the recurrent neural network is a biological species classification label and a biological behavior category label. For example, the classification labels for two biological species are: Cabot's Tragopan and Pine. For animal species, the behavior labels include: foraging, alertness, running, and calling. For plant species, the behavior category labels include phenological features such as flowering period, fruiting period, and leaf color changes.
[0037] The second pre-trained model S302 is used to identify children’s specific behaviors and interaction purposes; it adopts a two-stage adaptive training strategy, including generalized pre-training before deployment and online incremental fine-tuning after deployment.
[0038] Pre-deployment training: Before the device is first run, it is pre-trained using generalized human multimodal data. This generalized data specifically includes: facial images and full-body posture images of humans of different ages, including children and adults; voice commands, such as ("What is this?" "I want to see it run"); and body movement data of daily actions, such as waving, jumping, pointing, and clapping. Through training with large-scale general data, the second pre-trained model S302 gains the ability to understand general human behaviors, forming a generalized second pre-trained model. This model can output basic behavioral labels, such as "making a sound", "making a stretching movement", and "pointing in a certain direction".
[0039] Post-deployment online incremental fine-tuning: After the system is deployed to a specific forest treehouse science popularization site, the second pre-trained model S302 initiates an online incremental fine-tuning process; the system does not require manual annotation, but automatically collects observable feedback generated during children's natural interaction with the system as implicit supervision signals; these implicit feedback signals include: The frequency with which a child repeats the same action → If a child repeatedly performs an action, it indicates that he is interested in the feedback or is trying to convey some intention.
[0040] The length of time a child stays after a certain feedback → A longer dwell time indicates that the content is attractive or confusing.
[0041] Children's spontaneous vocalizations, whether pleasant or confused, include laughter and exclamations, while confused vocalizations include "Hmm?" and "No, that's not right."
[0042] The system uses these implicit feedback signals as weakly supervised labels to continuously optimize the generalized second pre-trained model. Specifically, an adaptive layer is added to the backend of the model, and model parameters are adjusted based on implicit feedback using online gradient descent or meta-learning methods. As runtime increases and on-site children's interaction data accumulates, the generalized second pre-trained model gradually evolves into a unique second pre-trained model. This evolved model retains the basic ability to understand general behaviors while adding the ability to accurately identify specific behaviors and purposes of children on-site. For example, it can distinguish between "imitating a macaque scratching its ears and cheeks" and "randomly scratching an itch," distinguish between "pointing to a bird on the screen and asking 'What's this?'" and "unintentionally waving," and determine whether a child's deep interaction purpose is "hoping to receive an animal imitation response," "hoping to learn ecological knowledge," or "hoping to trigger environmental changes."
[0043] Through the aforementioned implicit feedback-driven incremental training technique, the system's recognition accuracy increases with the number of uses, achieving continuous adaptive evolution without human intervention.
[0044] The behavior and ecological intent mapping module S40 is responsible for semantic alignment and dynamic matching of the biological information output by the first pre-trained model S301 with the child's behavior and purpose output by the second pre-trained model S302. This module adopts a two-stage progressive matching strategy, including the first stage S401 and the second stage S402. The first stage S401 is called the cold start and general mapping stage. The second stage S402 is called the personalized optimization mapping stage.
[0045] In the first stage S401, after the device is turned on for the first time and before sufficient on-site children's data has been accumulated to effectively fine-tune the generalized second pre-trained model, the behavior and ecological intention mapping module S40 uses the generalized human behavior labels output by the generalized second pre-trained model to perform preliminary semantic alignment with the biological information output by the first pre-trained model S301.
[0046] The specific implementation method is as follows: The system has a pre-set prior mapping rule base; this rule base contains multiple preset correspondences, and each rule is in the form of: "General human behavior label → biological species feature matching rule"; for example: "Children make sounds" → Match the animal species with the highest similarity to the sounds.
[0047] “Children make a fast running motion” → match with animal species that are good at running (such as macaques and serows).
[0048] "Child makes a pointing gesture" → Matches the biological species displayed in the pointed area.
[0049] "Child points to green object" → Match nearby plant species.
[0050] The behavior and ecological intent mapping module S40 looks up the corresponding matching rules in a table based on the currently identified general behavior labels; then, it selects the biological species that meet the conditions from the output of the first pre-trained model S301 and generates basic feedback instructions, such as playing the animal's call or displaying the name of the plant; this stage ensures that the system can provide reasonable and interesting interactive feedback from the first boot, avoiding the "cold start blind spot" of not being able to respond due to lack of personalized data.
[0051] In the second stage S402, as the system running time increases, the generalized second pre-trained model gradually evolves into a unique second pre-trained model through online incremental fine-tuning. At this time, the behavior and ecological intention mapping module S40 automatically switches the input source from the generalized model to the evolved unique model, and adopts differentiated fine matching strategies for different types of children's behaviors, including imitation or exploration.
[0052] (a) Fine-grained matching of imitative behaviors: When the second pre-trained model for uniqueness identifies a child’s behavior as imitative, such as imitating animal sounds and actions, and the interaction purpose is to seek an animal response, the behavior and ecological intention mapping module S40 uses the following similarity matching formula to calculate the fine-grained matching score between the child’s imitative features and the animal species features:
[0053] Where β represents the child's current fine-grained behavioral characteristics output by the second pre-trained model, and α represents the animal species identifier output by the first pre-trained model; and These are the extracted child action feature vectors and the animal feature vectors extracted from the typical behavior database of animal species, respectively, with cos(·) representing the cosine similarity; MFCC(β) sound α sound ) represents the Mel frequency cepstral similarity between the feature vectors of children's vocalizations and those of animal vocalizations, β sound α represents the feature vector of a child's vocalizations. sound For animal vocalization feature vectors; 1 point (β, α) is an indicator function, which takes a value of 1 when the direction of the hand pointing ray calculated based on the child's skeletal key points intersects with the display area of animal α, and otherwise takes a value of 0; ω1, ω2, and ω3 are preset weight coefficients, all of which are positive numbers and satisfy ω1 + ω2 + ω3 = 1; the behavior and ecological intention mapping module selects the animal α with the highest score. * =argmax α Score(β, α) is the matching result.
[0054] (b) Matching of exploratory behaviors: When a child’s behavior is identified as exploratory, including pointing to or observing plants, staring at an image of a living organism for a long time, and asking questions such as “What is this?” or “What does it eat?” and the purpose of the interaction is to acquire knowledge, the behavior and ecological intent mapping module S402 does not use the above similarity matching formula. Instead, it retrieves the corresponding knowledge content from the popular science database based on the biological species identifier that the child is currently interested in, whether it is an animal species or a plant species, combined with the child’s historical interests and cognitive level.
[0055] The specific implementation process is as follows: The system creates a lightweight user profile for each child and stores it in a local database in JSON format. The profile records a list of biological species that the child has actively triggered, the cumulative interaction time for each type of biological species, a set of popular science content tags that the child shows interest in, and statistics on the accuracy of completing interactive Q&A. When the behavior and ecological intent mapping module S402 determines that the behavior is exploratory, it first obtains the biological species identifiers that the child is currently interested in, for example, by tracking the child's gaze or pointing to an area. Then, it retrieves the child's historical profile and uses historical interest preferences as prior weights to sort the search results. At the same time, it dynamically assesses the child's cognitive level based on the accuracy of Q&A and divides popular science content into three levels: beginner (mainly pictures and simple text), intermediate (combination of pictures and text and including simple Q&A), and advanced (including in-depth knowledge such as ecological chain relationships and habitat characteristics). The system selects the difficulty level that matches the child's current cognitive level for push notifications.
[0056] Finally, different push notifications are generated based on the type of organism: If the currently focused organism is a plant, retrieve the plant's name, morphological characteristics (leaf, flower and fruit), ecological functions (such as carbon sequestration and water and soil conservation) and phenological knowledge, and generate a plant popular science push instruction.
[0057] If the currently focused organism is an animal, retrieve animal ecological knowledge including the animal's habitat, feeding habits, behavioral habits and natural enemy relationship, and generate an animal ecological knowledge push instruction.
[0058] If the imitation category and exploration category of intentions exist at the same time, for example: a child imitates an animal's call while pointing at the screen and asks a question; the behavior and ecological intention mapping module S40 executes in series according to the order of "imitation response first, then superimpose knowledge push", first performs imitation category matching, plays animations and calls, and then automatically superimposes knowledge explanation, so as to avoid feedback conflict and improve interaction fluency.
[0059] Wherein, the interactive feedback execution module S50 drives the display screen, audio equipment, lights and environmental special effect devices in the bionic tree house module S10 according to the feedback instruction issued by the behavior and ecological intention mapping module S40, to realize multi-modal linked feedback of sound, light, image and effect; the specific implementation mode is as follows: When receiving an instruction to display animal behavior: the interactive feedback execution module S50 first reads the 3D behavior animation file of the animal species from the popular science database, such as the courtship dance of Cabot's tragopan and the jumping of macaques, drives the surround display screen to play the animation at a frame rate of 60 frames per second, and ensures that the animation is adapted to the child's position and viewing angle; meanwhile, drives the multi-channel audio system to synchronously play the real call recording of the corresponding animal, and the sounds are extracted from the Nanling animal voiceprint database; in terms of lighting, dynamic changes are carried out according to the preset animal emotion color matching scheme; for example: when the animal is excited or alert, the programmable LED light array quickly flickers in warm tones of red and orange; when the animal is foraging quietly, the lights change to soft cool tones of green and blue; for environmental special effect devices: the ultrasonic nebulizer releases mist to enhance the sense of mystery, or the low-frequency airflow generator simulates the breeze driven by the running animal.
[0060] When receiving a plant popular science push instruction: the interactive feedback execution module S50 drives the display screen to display high-resolution images, growth process animations or short videos of the plant; for example: the process of seed germination, flowering and fruiting; the audio system explains the plant characteristics in child-oriented voice; for example: "This is a pine tree, its leaves are like needles", the lights are adjusted to a natural green atmosphere, with a color temperature of about 5000K and moderate brightness.
[0061] When receiving an animal ecological knowledge push instruction: the interactive feedback execution module S50 drives the display screen to display pictures and texts and / or short videos, such as photos of the animal's habitat and schematic diagrams of feeding habits, the audio explains in child-oriented voice, the lights are adjusted to a soft focus mode, with a color temperature of 4000K and slightly lower brightness to help concentrate attention.
[0062] Special Feedback (Day and Night Change Simulation): When the second pre-trained model S302 recognizes that a child is making sleep movements (such as closing their eyes, resting their head on their hand, and curling up) and the interaction aims to trigger environmental changes, the interactive feedback execution module S50 automatically triggers a forest day and night change simulation. The specific steps include: the screen brightness gradually changes from daytime mode to 20% brightness within 3 seconds; the LED light array switches to a warm yellow moonlight mode (color temperature 2700K, illuminance reduced to 50 lux); the sound system switches to playing nocturnal animal calls and insect chirps, while reducing the background volume; the visual, auditory, and lighting changes of the ecological process are directly triggered by the action signal, greatly enhancing the realism of the experience.
[0063] Finally, the S60 science popularization database and power supply control module provide localized Nanling biological knowledge base and energy consumption management support, which are implemented as follows: Science Popularization Database: Utilizing an embedded database, such as SQLite, for local storage, all data tables support millisecond-level access even without a network connection. The database contains the following main data tables: The main table for animal and plant species stores the unique identifier, Chinese name, Latin name, protection level, classification, and typical image path for each species. The behavioral characteristics table stores the behavior name, corresponding 3D animation file name, and action feature vector for animals, and the phenological stage name, corresponding image or animation file name, and phenological feature vector for plants. The vocalization sample table (animals only) stores the audio file paths and Mel-frequency cepstral coefficient feature vectors for animal vocalizations. The ecological knowledge table stores ecological knowledge text, graphic displays, difficulty levels, and associated species identifiers for animals and plants.
[0064] The power supply control module employs a hybrid power supply system consisting of a maximum power point tracking solar controller and a lithium iron phosphate battery pack. In science museums or indoor settings, it can also be connected to mains power as a backup. The power supply control module monitors the instantaneous power consumption of the lightweight AI chip, display screen, speakers, lighting array, and environmental effects devices in real time, and dynamically adjusts the power supply priority based on the battery state of charge and solar input power. The specific strategy is as follows: When solar energy is sufficient or mains power is normal, priority is given to charging the battery and powering the entire module.
[0065] When the battery level is below 20% (and there is no mains power supply), the display screen brightness will be automatically reduced to 60%, non-core environmental effects devices (such as atomizers and airflow generators) will be turned off, and the basic operation of the deep learning intelligent processing module S30 and the interactive feedback execution module S50 will be prioritized to ensure that the core perception and recognition functions are not interrupted.
[0066] It should be noted that all perception, recognition, mapping and feedback of the entire system are completed at the local edge, without relying on the cloud network; the lightweight AI chip adopts a system-on-a-chip with a neural network acceleration unit, which performs inference and incremental fine-tuning operations of the first pre-trained model S301 and the second pre-trained model S302 locally, with inference latency controlled within 100 milliseconds, ensuring the smoothness of real-time interaction.
[0067] In summary, the specific implementation of this invention constructs an immersive scene through the bionic treehouse module S10, captures children's behavior without contact through the multimodal perception and acquisition module S20, achieves biometric recognition and adaptive understanding of children's intentions through the dual models in the deep learning intelligent processing module S30, establishes a semantic closed loop through the two-stage matching of the behavior and ecological intention mapping module S40, drives multimodal linkage through the interactive feedback execution module S50, and ensures a stable local supply of data and energy through the science popularization database and power supply control module S60. This provides children with an all-weather, immersive, and adaptively evolving forest treehouse science popularization experience.
[0068] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention; any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A forest treehouse science popularization system integrating multimodal perception and deep learning feedback, characterized in that, The system includes: The biomimetic treehouse module is used to construct a Nanling forest scene and integrates display, sound, lighting and environmental effects devices as the physical carrier of the immersive learning space and the terminal execution platform for interactive feedback. A multimodal perception and acquisition module is used to simultaneously acquire children's image, voice, and body movement information through a camera, microphone, and motion sensor; The deep learning intelligent processing module, based on a lightweight AI chip, performs edge-end fusion analysis on collected data, including: a first pre-trained model, used to analyze biological and environmental data from the Nanling forest scene, outputting biological species and their behavioral labels, wherein the biological species include animal and plant species; and a second pre-trained model, which employs a two-stage adaptive training: before deployment, a generalized second pre-trained model is obtained through pre-training with generalized human multimodal data; after deployment, the generalized second pre-trained model is incrementally fine-tuned online using implicit feedback signals generated in children's interactions, gradually forming a unique second pre-trained model, used to identify children's specific behaviors and interaction purposes; The behavior and ecological intent mapping module employs a two-stage matching strategy: In the first stage, the generalized human behavior labels output by the generalized second pre-trained model and the biological information output by the first pre-trained model are used to perform preliminary semantic alignment based on a priori mapping rule base to generate basic feedback instructions; In the second stage, the fine-grained behavior and purpose labels output by the unique second pre-trained model, which has been formed through online incremental fine-tuning, are combined with the biological information of the first pre-trained model to perform refined semantic matching to generate optimized feedback instructions. An interactive feedback execution module is used to drive the devices in the bionic treehouse module according to the feedback instructions issued by the behavior and ecological intent mapping module, so as to realize multimodal linkage feedback. The science popularization database and power supply control module are used to provide a localized Nanling biological knowledge base and adopt a hybrid power supply strategy to ensure system operation.
2. The forest treehouse science popularization system integrating multimodal perception and deep learning feedback as described in claim 1, characterized in that: The biomimetic treehouse module constructs a Nanling forest scene using 3D biomimetic modeling and environmental restoration technology; The bionic treehouse module integrates a surround display screen, multi-channel audio, programmable LED lighting, and environmental effects devices. The Nanling Forest Scene refers to a virtual ecosystem generated through 3D modeling and multi-channel environmental sound effect restoration technology, which includes unique animal species, typical vegetation distribution, and microclimate characteristics of the Nanling region.
3. The forest treehouse science popularization system integrating multimodal perception and deep learning feedback according to claim 1, characterized in that: The multimodal perception and acquisition module includes an RGB camera and a depth camera. The RGB camera is used to acquire color facial images and full-body images of children for face recognition and expression recognition. The depth camera serves as the motion sensor and is used to capture non-contact limb movements by acquiring the child's three-dimensional skeletal key point sequence for recognizing imitation movements, pointing movements, and sleeping postures. The microphone is a microphone array used to collect children's voices and ambient sounds.
4. The forest treehouse science popularization system integrating multimodal perception and deep learning feedback according to claim 1, characterized in that: The first pre-trained model adopts an architecture that combines convolutional neural networks and recurrent neural networks, using biological data from the Nanling forest scene as input; The biological data includes: animal image sequences and plant species image sequences acquired by cameras, Mel spectrograms of animal calls acquired by microphones, animal posture skeletal data acquired by depth cameras, and environmental data acquired by environmental sensors. The first pre-trained model is configured with a multimodal feature fusion layer. The processing flow of the multimodal feature fusion layer includes: first, inputting data from four different modalities in the biological data into the convolutional neural network, and having the convolutional neural network extract their respective spatial feature vectors; then, inputting the four extracted spatial feature vectors into the multimodal feature fusion layer, and merging them into a joint feature vector by concatenation or weighted summation to achieve complementarity and enhancement of cross-modal information; finally, inputting the joint feature vector into the recurrent neural network for temporal dependency modeling, and the final output of the recurrent neural network is a biological species classification label and a biological behavior category label; wherein, the behavioral category label of plant species includes phenological features such as flowering period, fruiting period, and leaf color change.
5. A forest treehouse science popularization system integrating multimodal perception and deep learning feedback as described in claim 1, characterized in that: The generalized human multimodal data before deployment in the second pre-trained model specifically includes: facial images, full-body posture images, voice commands, and limb movement data of people of different ages; after deployment, the observable feedback generated during the natural interaction between children and the system is used as an implicit supervision signal to perform online incremental fine-tuning of the generalized second pre-trained model, gradually forming a unique second pre-trained model; The implicit feedback signals include the frequency with which the child repeats the same action, the duration of the pause after feedback, and the spontaneous pleasant and confused sounds.
6. The forest treehouse science popularization system integrating multimodal perception and deep learning feedback according to claim 1, characterized in that, The switching between the first and second phases of the behavior and ecological intent mapping module is gradual and automatic; As the system runs longer and on-site children's interaction data accumulates, the generalized second pre-trained model gradually evolves into a unique second pre-trained model through online incremental fine-tuning. The behavior and ecological intent mapping module automatically switches the input source generalized second pre-trained model to the evolved unique second pre-trained model, achieving a seamless transition from cold start general matching to personalized fine matching. The behavior and ecological intent mapping module executes the following matching rules: When a child's behavior is imitative and the interaction purpose is to seek an animal response, the module automatically matches the most similar animal species based on the child's imitation characteristics and generates a feedback instruction that displays animations of typical animal behaviors and realistic sounds. When a child's behavior is exploratory and the interaction purpose is to acquire knowledge, the module dynamically selects appropriate science content from the science database based on the biological species the child is currently interested in, combined with their historical interests and cognitive level, and generates a knowledge push instruction. Specifically, if the child is interested in a plant species, the module pushes plant science knowledge about the plant's name, morphological characteristics, and ecological functions. If the child is interested in an animal species, the module pushes animal ecological knowledge about the animal's habitat, diet, and behavioral habits. If both imitation and exploration purposes exist simultaneously, the module executes them sequentially, prioritizing imitation responses and then adding knowledge pushes.
7. A forest treehouse science popularization system integrating multimodal perception and deep learning feedback as described in claim 6, characterized in that: In the first stage, the behavior and ecological intent mapping module uses a priori mapping rule base for matching. The priori mapping rule base includes multiple preset correspondences, and each rule is in the form of: general human behavior labels are mapped to biological species feature matching rules. For plants, the rule includes: children pointing to green objects are mapped to matching nearby plant species. In the second stage, when the behavior and ecological intention mapping module identifies that a child's behavior belongs to the imitation category, it uses a similarity matching formula to calculate a fine-grained matching score between the child's imitation features and the animal species features: Where β represents the child's current fine-grained behavioral characteristics output by the second pre-trained model, and α represents the animal species identifier output by the first pre-trained model; and These are the extracted child action feature vectors and the animal feature vectors extracted from the typical behavior database of animal species, respectively, with cos(·) representing the cosine similarity; MFCC(β) sound α sound ) represents the Mel frequency cepstral similarity between the feature vectors of children's vocalizations and those of animal vocalizations, β sound α represents the feature vector of a child's vocalizations. sound For animal vocalization feature vectors; 1 point (β, α) is an indicator function, which takes a value of 1 when the direction of the hand pointing ray calculated based on the child's skeletal key points intersects with the display area of animal α, and otherwise takes a value of 0; ω1, ω2, and ω3 are preset weight coefficients, all of which are positive numbers and satisfy ω1 + ω2 + ω3 = 1; the behavior and ecological intention mapping module selects the animal α with the highest score. * =argmax α Score(β, α) is the matching result; When a child's behavior is identified as exploratory, the behavior-ecological intent mapping module does not use the similarity matching formula. Instead, it retrieves corresponding knowledge content from a science database based on the biological species identifier currently of interest to the child identified by the first pre-trained model, combined with the child's historical interests and cognitive level: if the organism is a plant, it retrieves plant science knowledge including plant name, morphological characteristics, and ecological functions; if the organism is an animal, it retrieves animal ecological knowledge including animal habitat, diet, and behavioral habits, and generates corresponding knowledge push instructions.
8. A forest treehouse science popularization system integrating multimodal perception and deep learning feedback as described in claim 1, characterized in that, The interactive feedback execution module achieves multimodal linkage feedback in the following ways: when receiving an instruction to display animal behavior, it drives the display screen to play a 3D behavioral animation of the animal, drives the speakers to synchronously play the corresponding animal's real calls, drives the lights to change according to a preset animal emotion color scheme, and drives the environmental special effects device to release a breeze or mist; when receiving a knowledge push instruction, it drives the display screen to display images and / or short videos, drives the speakers to provide explanations in a child-friendly voice, and drives the lights to adjust to a soft and focused mode. When the second pre-trained model recognizes a child making sleep movements, it automatically triggers a simulation of day and night changes in the forest, including: the screen brightness gradually dimming, the lights switching to a warm yellow moonlight mode, and playing nocturnal animal calls and insect chirps.
9. A forest treehouse science popularization system integrating multimodal perception and deep learning feedback as described in claim 1, characterized in that, The science database and the science database in the power supply control module are stored locally using an embedded method. The data tables in the science database include a main table of animal and plant species, a behavioral characteristic table, a call sample table, and an ecological knowledge table. All data tables support millisecond-level access in a network-free environment. The power supply control module adopts a hybrid power supply system composed of a maximum power point tracking solar controller and a lithium iron phosphate energy storage battery pack. By monitoring the instantaneous power consumption of each module in real time, and combining the battery state of charge and solar input power, the power supply priority is dynamically adjusted. When solar energy is sufficient, the battery is charged and the entire module is powered. When the battery power is insufficient, the display screen brightness is automatically reduced and non-core environmental effects devices are turned off to ensure the continuous operation of the core perception and recognition modules.
10. A forest treehouse science popularization system integrating multimodal perception and deep learning feedback as described in claim 1, characterized in that, The deployment of all modules of the system is completed locally at the edge, without relying on the cloud network; the lightweight AI chip adopts a system-on-a-chip with a neural network acceleration unit, which performs inference and incremental fine-tuning operations of the first pre-trained model and the second pre-trained model locally.