Self-adaptive explanation method and system and electronic equipment

By acquiring user space and emotional data to make multi-parameter decisions, generating explanation strategies and explanation texts, the system solves the problems of insufficient environmental awareness, low real-time adjustment of strategies, and lack of personalized content adaptation in intelligent explanation, and achieves a rich and personalized explanation experience.

CN121707798APending Publication Date: 2026-03-20MIGU DIGITAL MEDIA CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511739550.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing intelligent explanation technologies suffer from insufficient environmental perception, low real-time and intelligent level of strategy adjustment, and lack of personalized content adaptation capabilities, resulting in an unrich and unpersonalized explanation experience.

Method used

By acquiring user spatial data and sentiment data to make multi-parameter decisions, explanation strategies and explanation texts are generated. 3D spatial point cloud data, multimodal sentiment perception and reinforcement learning-driven dynamic strategy optimization are used to combine knowledge graphs to achieve personalized content generation.

Benefits of technology

It achieves multi-dimensional perception of the environment, real-time strategy adjustment, and personalized content adaptation capabilities, providing a richer and more personalized explanation experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121707798A_ABST
    Figure CN121707798A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a self-adaptive explanation method and system and electronic equipment. The method comprises the steps that environment data in a target scene are acquired, and the environment data comprise user space data and user emotion data; performing multi-parameter decision according to the environmental data to obtain an explanation strategy corresponding to the environmental data; and generating an explanation text corresponding to the environment data, and explaining according to the explanation strategy and the explanation text. By adoption of the scheme, the technical problems of insufficient environment perception dimension, low strategy adjustment real-time performance and intelligent level and lack of personalized content adaptation capability of intelligent explanation in related technologies can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an adaptive explanation method, system, and electronic device. Background Technology

[0002] Intelligent interpretation is a system that uses artificial intelligence technology to provide automated and personalized interpretation services, and is widely used in fields such as guided tours, education, and live streaming. Through 3D adaptive display technology, group behavior analysis technology, and intelligent interactive systems, it replaces the traditional static exhibition model, providing visitors with a unique and rich experience and meeting their needs for immersive experiences. It removes the limitations of time and space from exhibitions and is widely used in museums, science and technology museums, corporate showrooms, and urban planning museums.

[0003] However, among the related technologies, intelligent explanations suffer from insufficient environmental perception, low real-time and intelligent level of strategy adjustment, and lack of personalized content adaptation capabilities. Summary of the Invention

[0004] This disclosure aims to at least partially address one of the technical problems in the related art.

[0005] Therefore, the first objective of this disclosure is to propose an adaptive explanation method to address the technical problems in related technologies, such as insufficient environmental perception dimension, low real-time and intelligent level of strategy adjustment, and lack of personalized content adaptation capabilities.

[0006] The second objective of this disclosure is to propose an adaptive explanation system.

[0007] The third objective of this disclosure is to propose an electronic device.

[0008] The fourth objective of this disclosure is to provide a computer-readable storage medium.

[0009] The fifth objective of this disclosure is to provide a computer program product.

[0010] To achieve the above objectives, the first aspect of this disclosure proposes an adaptive explanation method, comprising: Acquire environmental data in the target scenario, wherein the environmental data includes user space data and user emotional data; Based on the environmental data, a multi-parameter decision is made to obtain the explanation strategy corresponding to the environmental data; Generate explanatory text corresponding to the environmental data, and explain it according to the explanation strategy and the explanatory text.

[0011] Optionally, acquiring environmental data in the target scenario includes: Collect three-dimensional spatial point cloud data of the target scene, and determine user spatial data based on the three-dimensional spatial point cloud data; Collect user voice feature data, user facial feature data, and ambient light data in the target scene, and determine user emotion data based on the user voice feature data, user facial feature data, and ambient light data.

[0012] Optionally, the user space data includes user space distribution data and user behavior trend data, and the step of determining the user space data based on the three-dimensional spatial point cloud data includes: The three-dimensional spatial point cloud data is preprocessed to obtain preprocessed three-dimensional spatial point cloud data. Identify user clustering regions in the preprocessed 3D spatial point cloud data to obtain the user spatial distribution data; Identify the coordinates of user joints in the preprocessed 3D spatial point cloud data, so as to determine the group's line of sight and user limb movement data based on the coordinates of the user joints. User behavior is predicted based on the group's gaze direction and the user's body movement data to obtain the user behavior trend data.

[0013] Optionally, determining user emotion data based on the user's voice feature data, the user's facial feature data, and the ambient light data includes: Determine the weight parameters corresponding to the user voice feature data, the user facial feature data, and the ambient light data; Based on the weight parameters, the user's voice feature data, the user's facial feature data, and the ambient light data are dynamically weighted and fused to obtain user emotion data.

[0014] Optionally, the step of making multi-parameter decisions based on the environmental data to obtain the explanation strategy corresponding to the environmental data includes: Acquire explanation status data, wherein the explanation status data includes explanation progress and explanation resource computing power; A reward function is determined, and based on the reward function, policy parameters are determined according to the explanation state data and the environment data to obtain the explanation policy corresponding to the policy parameters.

[0015] Optionally, generating the explanatory text corresponding to the environmental data includes: Construct a knowledge graph corresponding to the content system of the explanation, wherein the nodes of the knowledge graph are entities, and the edges of the knowledge graph are semantic associations between entities; Acquire historical audience interaction data, and infer interest distribution based on the historical audience interaction data and the environmental data to obtain interest distribution data; Knowledge units corresponding to the interest distribution data are selected from the knowledge graph, and explanatory text is generated based on the knowledge units and their interest matching degree under the interest distribution data.

[0016] Optionally, the explanation according to the explanation strategy and the explanation text includes: The digital narrator is driven to give a presentation in accordance with the presentation strategy and the presentation text.

[0017] Optionally, the method further includes: During the explanation process according to the described explanation strategy and the described explanation text, the text association information corresponding to the explanation text is displayed, wherein the text association information includes images, videos and models.

[0018] To achieve the above objectives, a second aspect of this disclosure provides an adaptive explanation system, comprising: A multi-dimensional environmental perception module is used to acquire environmental data in the target scene, wherein the environmental data includes user space data and user emotion data; The dynamic explanation strategy generation module is used to make multi-parameter decisions based on the environmental data, obtain the explanation strategy corresponding to the environmental data, and generate the explanation text corresponding to the environmental data. The explanation module is used to provide explanations according to the explanation strategy and the explanation text.

[0019] To achieve the above objectives, a third aspect of this disclosure provides an electronic device comprising: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the electronic device to perform the method shown in any of the first aspects above.

[0020] To achieve the above objectives, a fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed, implements the method shown in any of the first aspects above.

[0021] To achieve the above objectives, a fifth aspect of this disclosure provides a computer program product including a computer program that, when executed by a processor, implements the method shown in any of the first aspects above.

[0022] In summary, the method, system, and electronic device provided in this disclosure make multi-parameter decisions based on user spatial data and user emotional data to obtain an explanation strategy corresponding to the environmental data, and generate explanation text corresponding to the environmental data, so as to conduct the explanation according to the explanation strategy and explanation text. Therefore, it can overcome the limitations of related technologies, such as insufficient environmental perception dimension of intelligent explanation, low real-time and intelligent level of strategy adjustment, and lack of personalized content adaptation capability, and provide the audience with a richer and more personalized explanation experience.

[0023] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0024] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which: Figure 1 This is a flowchart illustrating an adaptive explanation method provided in an embodiment of the present disclosure; Figure 2 A flowchart illustrating a presentation strategy generation model provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of the structure of an adaptive explanation system provided in an embodiment of the present disclosure; Figure 4 This is a flowchart illustrating the workflow of an adaptive explanation system provided in an embodiment of this disclosure. Detailed Implementation

[0025] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.

[0026] Among related technologies, intelligent interpretation constructs a dynamic display platform by employing 3D adaptive display technology, crowd behavior analysis technology, and intelligent interactive systems. Specifically, 3D adaptive display technology acquires user position and perspective data through sensors, adjusting the display parameters of the 3D image (such as scaling and rotation) to achieve adaptive viewing angles. Crowd behavior analysis technology uses computer vision and clustering algorithms to model crowd distribution and movement trajectories, applying this to public safety monitoring and path planning. Intelligent interactive systems utilize natural language processing and speech recognition technologies to create interactive guide devices that can respond to user questions according to preset rules (e.g., traditional audio guides in museums).

[0027] Taking one scenario as an example, the first type of intelligent explanation analyzes user performance through a capability assessment module, dynamically adjusts the difficulty of training content, and uses a fixed rule engine to match preset training strategies (such as switching question types based on the accuracy of answers). However, its perception dimension is relatively limited, relying solely on user answer data and failing to consider diverse environmental information such as spatial distribution and emotional state; the strategy adjustment mode is fixed, relying on preset rules rather than real-time learning, which limits its adaptability in complex scenarios. For example, when there are many audience members, this mode cannot flexibly adjust the pace of explanation to adapt to the constantly changing environment; the application scenarios are limited. Although the current focus is mainly on optimizing training content, the potential of multimodal interaction technology goes far beyond this. It involves the dynamic generation of explanation strategies and the implementation of multimodal interaction, such as the coordinated adjustment of gestures and voice volume. These technologies have shown great application potential in many fields such as smart homes, intelligent driving, intelligent education, and intelligent healthcare.

[0028] For example, the second type of intelligent explanation uses a camera to collect user head position data, calculates viewing distance and angle using trigonometric formulas, adjusts the viewpoint and object size of the holographic display, and simultaneously adjusts the sound direction. However, its environmental perception range is limited: it is limited to the user's individual position and perspective in physical space, without delving into group behavior analysis (such as audience density, line of sight, etc.) or accurate perception of emotional states (e.g., subtle fluctuations in group emotions); its content adaptation capability is insufficient: it fails to dynamically reorganize the explanation content according to the audience's interests, and is limited to the passive adjustment of display parameters.

[0029] In summary, the environmental perception dimension of the relevant technologies is insufficient. Although they have made some progress in acquiring audience physical location or single behavioral data (such as answer accuracy), they are still insufficient in integrating multi-dimensional information such as three-dimensional spatial distribution, body movements, and emotional acoustic features. This limits the close integration of explanation strategies with the needs of real scenarios. For example, the first type of intelligent explanation cannot perceive changes in the attention concentration of the audience when they gather and still adopts a fixed explanation rhythm, which easily leads to a decrease in information transmission efficiency. Furthermore, the real-time performance and intelligence of strategy adjustments are low: related technologies mainly rely on preset rules or simple parameter fine-tuning (for example, the second type of intelligent explanation only adjusts the display perspective), lacking the ability to learn and dynamically adapt to complex environmental characteristics in real time. For example, traditional tour guide systems have a delay in responding to group emotional fluctuations and switching the emotional tone of the explanation, usually requiring more than 2 seconds.

[0030] Secondly, the technology lacks personalized content adaptation capabilities: it fails to break down the content into recombinable knowledge units, making it impossible to dynamically generate explanation scripts based on audience interests. For example, the training content adjustment for the first type of intelligent explanation is only based on preset difficulty levels.

[0031] Furthermore, resource utilization efficiency is low: related technologies employ fixed resource configurations across different scale scenarios, leading to resource waste in small-scale scenarios or service lag in large-scale scenarios. For example, traditional digital explanation systems cannot flexibly allocate computing resources when audience density fluctuates.

[0032] The present disclosure will now be described in detail with reference to specific embodiments.

[0033] In the first embodiment, such as Figure 1 As shown, Figure 1 This is a flowchart illustrating an adaptive explanation method provided in an embodiment of the present disclosure. The method can be implemented using a computer program and can run on an adaptive explanation system. This computer program can be integrated into an application or run as a standalone utility application.

[0034] This adaptive explanation method can be executed by an electronic device.

[0035] For example, this adaptive explanation method includes the following steps: S101, acquire environmental data in the target scenario; According to some embodiments, the target scenario refers to the scenario that requires adaptive explanation.

[0036] In some embodiments, environmental data includes user space data and user sentiment data.

[0037] User space data refers to data related to the user in the target scene within a three-dimensional space.

[0038] User sentiment data refers to data related to users' emotions in the target scenario.

[0039] S102, make multi-parameter decisions based on environmental data to obtain the interpretation strategy corresponding to the environmental data; According to some embodiments, multi-parameter decision-making refers to a decision-making process conducted under the influence of multiple parameters or variables. In this embodiment, multi-parameter decision-making is used to make decisions on explanatory strategies under the influence of environmental data.

[0040] In some embodiments, the explanation strategy refers to the strategy employed when giving an explanation. The explanation parameters used in this strategy include, but are not limited to, explanation duration and explanation speed.

[0041] S103, generate explanatory text corresponding to the environmental data, and conduct the explanation according to the explanation strategy and the explanatory text.

[0042] According to some embodiments, explanatory text refers to the text content used when giving an explanation.

[0043] In summary, the method provided in this embodiment makes multi-parameter decisions based on user spatial data and user emotional data to obtain an explanation strategy corresponding to the environmental data, and generates explanation text corresponding to the environmental data, so as to conduct the explanation according to the explanation strategy and explanation text. Therefore, it can overcome the limitations of related technologies, such as insufficient environmental perception dimension of intelligent explanation, low real-time and intelligent level of strategy adjustment, and lack of personalized content adaptation capability, and provide the audience with a richer and more personalized explanation experience.

[0044] Another embodiment of this disclosure provides an adaptive explanation method. This adaptive explanation method can be executed by an electronic device.

[0045] For example, the adaptive explanation method may include the following steps: S201: Collect 3D spatial point cloud data of the target scene and determine user spatial data based on the 3D spatial point cloud data; According to some embodiments, multi-view depth imaging devices (such as LiDAR, binocular vision cameras, etc.) can be used to cover the target scene in an array layout to achieve the acquisition of three-dimensional spatial point cloud data.

[0046] For example, four infrared depth cameras can be deployed in the target scene, covering the top of the display area at 90° intervals to achieve 360° three-dimensional visual coverage. Each camera can be configured with a resolution of no less than 1280×720 and a frame rate of 30fps to ensure high accuracy and real-time performance of the depth information in the three-dimensional spatial point cloud data.

[0047] For example, a laser LiDAR solution can be used to replace a depth camera, improving the accuracy of crowd positioning at long distances (10 meters) (error <3cm), which is suitable for large exhibition scenarios.

[0048] According to some embodiments, after acquiring three-dimensional spatial point cloud data, point cloud preprocessing can be performed on the three-dimensional spatial point cloud data to obtain preprocessed three-dimensional spatial point cloud data; user gathering areas in the preprocessed three-dimensional spatial point cloud data can be identified to obtain user spatial distribution data; user key point coordinates in the preprocessed three-dimensional spatial point cloud data can be identified to determine the group's line of sight direction and user limb movement data based on the user key point coordinates; user behavior prediction can be performed based on the group's line of sight direction and user limb movement data to obtain user behavior trend data.

[0049] In some embodiments, when performing point cloud preprocessing, a point cloud library (PCL) can be used to perform a pass-through filtering operation to effectively remove noise and outliers. At the same time, a voxel downsampling method (e.g., setting the resolution to 0.05m) can be used to reduce the data dimensionality while ensuring that the features of the crowd contours and motion trajectories are preserved.

[0050] In some embodiments, density clustering algorithms can be used to accurately identify user clustering areas in preprocessed 3D spatial point cloud data and calculate key spatial parameters such as group size and density distribution to obtain user spatial distribution data.

[0051] For example, density-based spatial clustering of applications with noise (DBSCAN) or its improved version can be used with the following parameter settings: eps=0.8m (neighborhood radius), min_samples=8 (minimum number of point clusters) to identify user clustering areas and calculate density values ​​(unit: people / m³) to obtain the spatial distribution matrix.

[0052] In some embodiments, the OpenPose human pose estimation model can be used to extract the coordinates of user joints (such as shoulder joints and hip joints) from the preprocessed three-dimensional spatial point cloud data, and then analyze the group's gaze direction (error < 5°) and user limb movement data (such as raising hands and leaning forward) to obtain a behavior trend vector.

[0053] According to some embodiments, edge computing units (such as embedded artificial intelligence (AI) processors) can be configured to support real-time data preprocessing and feature extraction. This configuration can use microprocessor units (MPUs) with considerable computing power as edge computing units to process 3D spatial point cloud data from multiple cameras in real time and efficiently.

[0054] For example, the update frequency of the edge computing unit can be configured to be ≥10Hz, generating a three-dimensional spatial feature vector per second that includes population size, density distribution, and average viewing angle.

[0055] S202, collect user voice feature data, user facial feature data and ambient light data in the target scene, and determine user emotion data based on user voice feature data, user facial feature data and ambient light data; According to some embodiments, a user's voice signal can be acquired through a microphone array and frequency domain analysis can be supported, such as extracting Mel-Frequency Cepstral Coefficients (MFCCs). Then, emotional acoustic features (such as pleasure, focus, and confusion) are classified through a Recurrent Neural Network (RNN) to obtain user voice feature data.

[0056] For example, an array consisting of eight micro-electro-mechanical system (MEMS) microphones can be deployed to collect the user's voice signal at a sampling rate of 44.1 kHz, with a frequency resolution of 128-point FFT. Then, MFCC (13-dimensional), short-time energy, and zero-crossing rate are calculated and input into a bidirectional long short-term memory (LSTM) network for classification to obtain the user's voice feature data (such as laughter corresponding to "pleasure" and faster speech corresponding to "excitement").

[0057] According to some embodiments, for user facial feature data, an image acquisition device can be used to acquire user facial images, and a region of interest (ROI) can be located through a face detection and tracking algorithm; then, micro-expression features (such as the degree of mouth corner raising and changes in the distance between eyebrows and eyes) can be extracted through convolutional neural networks (CNNs) to obtain user facial feature data.

[0058] For example, a depth camera can be used to simultaneously capture user facial images, which are then accurately detected using a multi-task convolutional neural network (MTCNN) algorithm and cropped to 112×112 pixels to obtain the user's facial image. Next, a model combining FaceNet and LSTM can be used. FaceNet (which can be pre-trained based on VGG-Face) is responsible for extracting facial features, and a deep convolutional neural network maps the facial image to a 128-dimensional Euclidean space to generate facial feature vectors. This process achieves an accuracy of 99.63% on the LFW dataset and 95.12% on the YouTube Faces DB dataset. The LSTM network is used to analyze facial expression changes in time series, such as judging the intensity of pleasure by the duration of a smile, thereby obtaining user facial feature data.

[0059] According to some embodiments, wearable device data interfaces such as heart rate monitoring and eye tracking can also be integrated to further improve the accuracy of emotional state perception. Simultaneously, wearable physiological sensors such as the Empatica E4 wristband can be added to collect the audience's heart rate variability (HRV) and electrical conductance per unit skin (EDA) data, thereby enhancing the accuracy of emotion recognition.

[0060] In some embodiments, for ambient light data, sensors such as ambient light and noise sensors can be integrated to obtain physical characteristic parameters of the target scene.

[0061] For example, the BH1750 light sensor can be integrated to monitor the ambient light intensity (unit: lux) in real time and determine the user's concentration level. For example, users tend to be more focused in low-light environments.

[0062] According to some embodiments, weight parameters corresponding to user voice feature data, user facial feature data, and ambient light data can be determined; based on the weight parameters, user voice feature data, user facial feature data, and ambient light data are dynamically weighted and fused to obtain user emotion data.

[0063] For example, attention mechanisms or gating units (such as GRU and LSTM) can be used to fuse multi-source data, including audio, visual, and ambient light features. By dynamically weighting the fusion, the influence of key modalities can be effectively highlighted. For example, in a lecture scenario, the weight of user facial feature data is often higher than that of ambient light data.

[0064] Specifically, for the weight parameters, the fully connected layer can output the probability distribution of eight types of group emotional states. The specific categories include: “focused”, “confused”, “interested”, “pleasant”, “excited”, “calm”, “doubtful”, and “indifferent”. The probability value of each category ranges from [0,1], and the sum is 1.

[0065] It should be noted that by executing steps S201 and S202, and utilizing a multi-view depth camera array and the DBSCAN clustering algorithm, information such as the distribution, density, and gaze direction of the audience in three-dimensional space is captured in real time. Simultaneously, by combining audio sentiment analysis and the Transformer facial micro-expression recognition model, a real-time monitoring mechanism for group emotional states is constructed. Thus, by introducing three-dimensional point cloud analysis and a multimodal sentiment fusion algorithm, the limitations of traditional two-dimensional visual and single-behavioral data are overcome.

[0066] S203, Obtain explanation status data; According to some embodiments, the explanation status data includes explanation progress and explanation resource computing power.

[0067] S204, determine the reward function, and based on the reward function, determine the policy parameters according to the explanation state data and environment data to obtain the explanation policy corresponding to the policy parameters; According to some embodiments, this embodiment constructs a closed-loop system that includes state awareness, action selection, and reward feedback by relying on a reinforcement learning framework, thereby realizing real-time optimization and adjustment of the explanation strategy.

[0068] The parameters in the state space include environmental data and presentation status data. For example, the state space can include 10-dimensional features: group density (people / m²), average gaze angle (relative to the presenter's front), emotion probability vector (top 5 main emotions), current presentation progress, and remaining presentation resource computing power.

[0069] For example, when the group density is high (e.g., 3 people / m²) and the average line of sight is significantly deviated from the front of the presenter (greater than 60°), the state vector will be encoded into the corresponding value to indicate that the audience’s attention may have been distracted.

[0070] The parameters in the Action space are strategy parameters, including but not limited to continuous variables such as the way of explanation and expression (speech rate, volume, gesture amplitude), content organization (knowledge point priority, case selection), and resource allocation (rendering precision, data transmission rate). For example, the Action space can use 20 dimensions and 3 categories of continuous parameters, including: speech rate (0.8-1.5 times the baseline), volume (40-80dB), gesture amplitude (0%~200% of the baseline), content depth (levels 1-5), gaze direction (0-360°), knowledge point priority (levels 1-5), and case complexity (levels 1-3).

[0071] For example, if the state space shows that the audience's emotional bias is "confused" (the probability of confusion in the emotional probability vector is >0.4), the policy network may generate a combination of actions such as "reducing the speaking speed (0.9 times), increasing the content depth (level 3), and increasing the gesture amplitude (150% baseline)" to enhance the clarity of the explanation.

[0072] In some embodiments, the reward function can be optimized with audience attention duration and information absorption rate as the core optimization objectives, and set main rewards, secondary rewards and penalty items in combination with the utilization rate of explanation resources and computing power.

[0073] The main reward includes the audience attention retention time (calculated by the duration of eye contact, +1 point for every 10 seconds). Its main purpose is to encourage the agent to maintain the audience's focus, for example, by flexibly adjusting gestures or speaking speed to attract the audience's attention.

[0074] The secondary rewards include: information absorption rate (based on the accuracy of the questions and answers, +5 points for correct answers and -2 points for incorrect answers), which aims to ensure that the content being explained can be effectively received and understood. For example, this can be achieved by prioritizing key knowledge points and combining them with easy-to-understand examples.

[0075] The penalties include: resource overload (deducting 0.1 points per second when resource computing power usage is >90%), in order to balance computing power performance and user experience, and at the same time avoid excessive resource consumption, for example, by appropriately reducing rendering precision to release more computing power.

[0076] For example, in a museum guided tour scenario test, when the state space shows that visitors are densely lingering in front of a certain artifact and their facial micro-expressions are mainly "focused" (probability > 0.6), the reward function can determine that the current explanation content has aroused high interest, automatically extend the explanation time of that part and add interactive elements.

[0077] It should be noted that a lecturing strategy generation model can be constructed to determine the strategy parameters based on the lecturing state data and environmental data, thereby obtaining the lecturing strategy corresponding to the strategy parameters. This lecturing strategy generation model can employ deep reinforcement learning algorithms (such as PPO, DDPG, etc.) and achieve real-time optimization of the strategy parameters through end-to-end training.

[0078] For example, Figure 2 This is a flowchart illustrating a presentation strategy generation model provided in an embodiment of this disclosure. Figure 2 As shown, the explanation strategy adopts the PPO (Proximal Policy Optimization) algorithm, and the network structure is a 2-layer fully connected layer (256 neurons per layer, ReLU activation). The training data comes from a mixed dataset of simulated scenes and real tour guide records.

[0079] It should be noted that by executing steps S203 and S204, reinforcement learning algorithms are used to continuously optimize the policy parameters within a space containing more than 20 dimensions (such as speech rate, volume, and gesture amplitude). Specifically, a three-dimensional decision space including crowd size, attention index, and emotion intensity is constructed, and deep reinforcement learning is used to achieve real-time iteration of the policy parameters. Furthermore, reinforcement learning techniques enable policy iteration every 0.5 seconds, significantly shortening response time and thus substantially improving the smoothness of the interaction.

[0080] S205, Construct a knowledge graph corresponding to the content system; According to some embodiments, the nodes of a knowledge graph are entities, including but not limited to knowledge points, cases, and interaction methods. The edges of the knowledge graph are semantic relationships between entities, including but not limited to relationships such as "contains," "associates," and "belongs to."

[0081] S206, acquire historical audience interaction data, and infer interest distribution based on historical audience interaction data and environmental data to obtain interest distribution data; For example, interest distribution can be inferred through historical audience interaction data and real-time behavioral characteristics (such as areas of stay and keywords asked in questions).

[0082] S207: Select knowledge units corresponding to interest distribution data from the knowledge graph, and generate explanatory text based on the knowledge units and their interest matching degree under the interest distribution data. According to some embodiments, semantic retrieval technology and path planning algorithms can be used to accurately select knowledge units related to the audience's interests from the knowledge graph, and automatically generate personalized explanatory text based on the principles of interest matching degree and logical coherence.

[0083] In some embodiments, a text generation model can be used to filter out knowledge units corresponding to interest distribution data from a knowledge graph, and generate explanatory text based on the knowledge units and their interest matching degree under the interest distribution data.

[0084] For example, in a knowledge graph construction, a graph database is used to store the explanatory content. Nodes include knowledge points (such as "the history of bronze ware"), cases (such as "the Houmuwu Ding"), and interaction types (such as "3D model rotation"). Edges define semantic associations (such as "belongs to" and "related"). The slicing rules in the text generation model are set as follows: the explanatory text is divided into knowledge units of 50 to 100 words according to semantic content, and each unit is associated with 1 to 3 graph nodes (for example, 'Four-Sheep Square Zun' is associated with nodes such as 'Shang Dynasty Bronze Ware' and 'Sacrificial Ritual Vessels'). The feature inputs of the text generation model can include: the area where the audience stayed (corresponding to the exhibit ID), the keywords of the question (entities can be extracted by BERT), and historical interaction records (such as clicking on the "weapons" tag); The output of the text generation model can include: the probability distribution of the top 5 interest nodes (e.g., "weapon" 0.6, "musical instrument" 0.2); In the process of generating explanatory text, the text generation model can use a greedy algorithm to efficiently retrieve related units in the knowledge graph and intelligently generate explanatory text based on the weight ratio (7:3) of 'interest matching degree' and 'explanation fluency', ensuring that 500 words of content can be accurately reorganized within 5 seconds.

[0085] It should be noted that by executing steps S205 to S207, knowledge unit decomposition technology based on knowledge graphs is used to decompose the content to be explained into semantically related knowledge units. Combined with an interest prediction model, the content can be dynamically reorganized, and the content can be flexibly adjusted according to the audience's interests. In addition, by realizing personalized content reorganization within 5 seconds, the accuracy of knowledge unit association can be greatly improved.

[0086] S208 drives the digital narrator to give a narration according to the narration strategy and narration text, and displays the text association information corresponding to the narration text during the narration process. According to some embodiments, a digital narrator can be obtained by constructing a model skeleton containing 206 bones based on physical skeleton binding and skinning technology.

[0087] In some embodiments, the digital narrator can be driven to retrieve appropriate actions from a library of more than 2,000 action segments (covering various types such as narration and emotion) based on acquired strategy parameters, such as the amplitude of gestures and the direction of gaze. The appropriate actions are then seamlessly connected through an interpolation algorithm to generate a continuous and natural sequence of actions. The naturalness of the actions has also been finely calibrated using inertial motion capture data.

[0088] For example, based on a physical skeletal model (bound by 206 bones), basic movements can be retrieved from a library of over 2000 movements according to the gesture amplitude and gaze direction (e.g., gesture amplitude 150%, gaze direction 30° to the left and front) in the strategy parameters. Then, by using linear interpolation and quaternion interpolation algorithms, the movement segments can be cleverly integrated to generate smooth and natural movements. The relevant movement libraries can be divided into explanatory, emotional, and navigation types according to their corresponding functions. The motion capture data calibration can be performed by using the Perception Neuron inertial motion capture system to collect real-person movements and calibrate the joint angle error (≤5°) and velocity curve of the digital movements.

[0089] In some embodiments, the rendering of facial expressions for digital narrators can employ Blend Shape technology, which presets facial expression fusion coefficients (such as a "smile + nod" combination) and combines them with real-time shadow mapping to enhance realism.

[0090] For example, Blend Shape technology can be used to pre-set 80 sets of facial expression blending coefficients (such as a coefficient of 0.3 for "smile" and 0.2 for "nod"), and then enhance the realism through real-time shadow mapping (PCF algorithm).

[0091] According to some embodiments, the narration output of a digital narrator can adopt multimodal output methods, including speech synthesis and sound field control.

[0092] In some embodiments, the speech synthesis can support a TTS engine with emotion adjustment, synchronize speech rate, tone and strategy (e.g., a 20% increase in tone in the "excitement" mode), and real-time matching of lip movements and speech.

[0093] For example, a streaming deep neural network TTS engine that supports SSML tagging can be used to dynamically adjust the speech rate (0.8-1.5 times the baseline), tone (e.g., increasing tone by 20% in "excited" mode), and emotional tendency (e.g., decreasing speech rate by 15% in "serious" mode) according to policy parameters; at the same time, speech is driven synchronously with actions and expressions (e.g., speech stress is enhanced when gestures are emphasized), and lip-syncing is achieved through Phoneme phoneme matching algorithm (error ≤ 50ms).

[0094] In some embodiments, sound field control can achieve directional sound propagation and dynamic positioning through phased array acoustics or spatial audio algorithms, ensuring clear speech.

[0095] For example, phased array loudspeaker array technology can be used to achieve directional sound propagation and dynamic positioning (e.g., the sound source can move synchronously when the narrator turns around), ensuring that the speech intelligibility (STI) is higher than 0.9 within a 5-meter range.

[0096] According to some embodiments, during the explanation process according to the explanation strategy and explanation text, auxiliary content is displayed by showing the text association information corresponding to the explanation text. This text association information includes, but is not limited to, images, videos, and models, so as to link and display images, videos, or 3D models in the knowledge base, support user interaction (such as touch rotation of models), and enhance the information delivery effect.

[0097] For 2D content display, images and videos can be retrieved from the knowledge base based on the dialogue content and displayed through the UI interface (such as a floating dialog box), and touch zoom or progress control is supported. For 3D model display, 3D models of cultural relics, scenes, etc. can be loaded into a designated area so that users can use a controller or issue voice commands to trigger interactive actions such as rotation and disassembly (e.g., the command 'display the internal structure of bronzeware').

[0098] In summary, the method provided in this embodiment, by integrating core technologies such as three-dimensional spatial crowd analysis, multimodal emotion perception, reinforcement learning-driven dynamic strategy optimization, and knowledge graph-enabled personalized content generation, can construct an intelligent closed loop from environmental perception to behavioral response, achieving precise matching between explanation strategies and scenario requirements. It is applicable to various scenarios such as museum tours, exhibitions, education and training, and virtual conferences.

[0099] To implement the above embodiments, this disclosure also proposes an adaptive explanation system.

[0100] For example, Figure 3 This is a schematic diagram of the structure of an adaptive explanation system provided in an embodiment of this disclosure. Figure 3 As shown, the adaptive explanation system 300 includes: The multi-dimensional environmental perception module 301 is used to acquire environmental data in the target scene, including user space data and user emotional data. The dynamic explanation strategy generation module 302 is used to make multi-parameter decisions based on environmental data, obtain the explanation strategy corresponding to the environmental data, and generate the explanation text corresponding to the environmental data. The explanation module 303 is used to provide explanations according to the explanation strategy and the explanation text.

[0101] To give an example from a scenario, Figure 4 This is a flowchart illustrating the workflow of an adaptive explanation system provided in an embodiment of this disclosure. Figure 4 As shown, this adaptive narration system adopts a three-layer architecture of "perception-decision-execution," consisting of a multi-dimensional environmental perception module, a dynamic narration strategy generation module, a 3D digital narrator execution module, and an edge-cloud collaborative computing platform. The system follows the process below to achieve adaptive narration functionality. Among them, The multi-dimensional environmental perception module 301 can acquire multi-dimensional data such as audience spatial distribution, behavioral characteristics, and emotional state through multiple types of sensors; The dynamic explanation strategy generation module 302 can use a deep learning model to analyze environmental data and then formulate a comprehensive strategy covering explanation content, expression methods and resource allocation. The explanation module 303 can drive the 3D digital model to achieve coordinated display of actions, expressions and voice, and flexibly adjust the strategy based on real-time feedback.

[0102] According to some embodiments, the edge-cloud collaborative architecture adopted in this embodiment can use a hierarchical task scheduling mechanism to intelligently allocate resources based on the computing power characteristics of the edge and the cloud, thereby achieving a balance between efficiency and performance. Through the collaborative complementarity between the edge and the cloud, it can not only meet the low latency requirements of real-time interaction, but also cope with the computing power challenges of large-scale scenarios, and achieve efficient resource utilization across device terminals.

[0103] Regarding task partitioning strategies, in the field of edge computing, edge devices can be responsible for executing lightweight tasks with high real-time requirements, such as environmental data preprocessing (e.g., point cloud filtering, face detection), reinforcement learning inference (generating explanation strategy parameters), local caching and fast retrieval of knowledge graph hotspot data, and real-time rendering of low-precision 3D models. These tasks need to be executed within local computing power constraints to reduce latency and ensure real-time performance. For example, the UHNet lightweight edge detection model achieves efficient edge detection on resource-constrained devices with a minimal number of parameters and fast computation speed. Furthermore, real-time performance analysis of edge computing emphasizes the importance of key performance indicators such as latency and throughput, while the performance parameters of edge computing servers, such as computing power, storage capacity, and network bandwidth, are crucial to ensuring the success of real-time processing tasks at the edge. In the field of cloud computing, the cloud can undertake resource-intensive tasks, such as updating the entire knowledge graph data and performing complex semantic association calculations, as well as remote rendering and streaming of high-precision 3D models, providing computing power support for large-scale scenes or high-end display needs. For example, cloud knowledge graphs (TKG) can support the storage and computation of hundreds of billions of node relationships and can respond to online query operations such as node search and multi-hop queries in near real time, which demonstrates the powerful capabilities of the cloud in handling large-scale knowledge graph tasks.

[0104] Specifically, for dynamic load awareness, an elastic resource scheduling mechanism can be adopted, using the Q-Learning algorithm to optimize resource allocation. Every 5 seconds, the load status of the edge device, including CPU utilization and memory usage, is automatically monitored. If the CPU utilization at the edge device rises above 80%, a cloud collaboration mechanism can be automatically activated to migrate tasks such as high-precision rendering and complex knowledge graph queries to the cloud for processing, ensuring system response latency is within 500 milliseconds.

[0105] Among them, for adaptive rendering accuracy, the LOD (Level of Detail) level can be dynamically switched based on device performance: for high-end devices (such as professional workstations), the full-precision 3D model is rendered; for low-end devices, such as mobile terminals, the model can be automatically downgraded to a simplified version (60% reduction in polygon count) and anti-aliasing algorithms (such as FXAA) can be applied to maintain smooth visual effects and ensure that the frame rate is not lower than 25fps.

[0106] In some embodiments, an LLM service layer can be added to the edge-cloud collaborative architecture, forming a four-layer architecture of "perception-decision-execution-cognition". The edge layer can retain environmental awareness and low-precision rendering, while adding real-time acquisition and preliminary processing of user voice / text questions; the LLM service layer can deploy lightweight LLMs (such as Alpaca, Llama-2-7B, etc.) or call cloud APIs to handle complex semantic understanding, knowledge reasoning, and content generation; the original decision layer can receive LLM outputs and convert them into executable policy parameters (such as speech rate and content depth); the execution layer can drive the 3D digital narrator to achieve multimodal responses and enhance cross-modal interaction, for example, supporting multimodal inputs such as gesture recognition and haptic feedback, enabling proactive interaction between the audience and the digital narrator (such as gesture-based content switching and gamified content presentation).

[0107] In some embodiments, a neural machine translation (NMT) model can be introduced to achieve real-time multilingual conversion and synchronous output of explanatory text, covering multiple languages ​​such as Chinese, English, Japanese, and Korean; by optimizing the NMT model, translation latency can be further reduced and BLEU scores can be improved.

[0108] It should be noted that the foregoing explanation of the adaptive explanation method embodiment also applies to the adaptive explanation system of this embodiment, and will not be repeated here.

[0109] In summary, the system provided in this disclosure, by constructing a fully intelligent digital tour guide system with "environmental awareness - strategy optimization - content adaptation - resource scheduling," can fundamentally improve the scene adaptability and interaction efficiency of digital tour guides, providing a new technological paradigm for fields such as intelligent tour guiding and virtual education. Furthermore, by employing a collaborative rendering architecture of edge computing and central cloud, and incorporating a dynamic resource scheduling strategy based on user behavior prediction, it can ensure that computing resources can be flexibly and efficiently allocated on demand, significantly reducing system energy consumption, ensuring smooth operation in multi-person scenarios, supporting automatic adaptation to multi-person scale scenarios, and solving the problem of inefficient resource allocation in traditional systems.

[0110] To implement the above embodiments, this disclosure also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0111] To implement the above embodiments, this disclosure also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.

[0112] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.

[0113] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this disclosure all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0114] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0115] This disclosure is intended to provide implementation schemes for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.

[0116] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0117] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0118] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.

[0119] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequential list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disks (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, fiber optic devices, and compact disc read-only memory (CDROM). Additionally, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0120] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0121] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium. When executed, the program includes one or a combination of the steps of the method embodiments.

[0122] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0123] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. An adaptive explanation method, characterized in that, include: Acquire environmental data in the target scenario, wherein the environmental data includes user space data and user emotional data; Based on the environmental data, a multi-parameter decision is made to obtain the explanation strategy corresponding to the environmental data; Generate explanatory text corresponding to the environmental data, and explain it according to the explanation strategy and the explanatory text.

2. The method according to claim 1, characterized in that, The acquisition of environmental data in the target scenario includes: Collect three-dimensional spatial point cloud data of the target scene, and determine user spatial data based on the three-dimensional spatial point cloud data; Collect user voice feature data, user facial feature data, and ambient light data in the target scene, and determine user emotion data based on the user voice feature data, user facial feature data, and ambient light data.

3. The method according to claim 2, characterized in that, The user space data includes user space distribution data and user behavior trend data. Determining the user space data based on the three-dimensional spatial point cloud data includes: The three-dimensional spatial point cloud data is preprocessed to obtain preprocessed three-dimensional spatial point cloud data. Identify user clustering regions in the preprocessed 3D spatial point cloud data to obtain the user spatial distribution data; Identify the coordinates of user joints in the preprocessed 3D spatial point cloud data, so as to determine the group's line of sight and user limb movement data based on the coordinates of the user joints. User behavior is predicted based on the group's gaze direction and the user's body movement data to obtain the user behavior trend data.

4. The method according to claim 2, characterized in that, The step of determining user emotion data based on the user's voice feature data, the user's facial feature data, and the ambient light data includes: Determine the weight parameters corresponding to the user voice feature data, the user facial feature data, and the ambient light data; Based on the weight parameters, the user's voice feature data, the user's facial feature data, and the ambient light data are dynamically weighted and fused to obtain user emotion data.

5. The method according to claim 1, characterized in that, The step of making multi-parameter decisions based on the environmental data to obtain the interpretation strategy corresponding to the environmental data includes: Acquire explanation status data, wherein the explanation status data includes explanation progress and explanation resource computing power; A reward function is determined, and based on the reward function, policy parameters are determined according to the explanation state data and the environment data to obtain the explanation policy corresponding to the policy parameters.

6. The method according to claim 1, characterized in that, The explanatory text generated corresponding to the environmental data includes: Construct a knowledge graph corresponding to the content system of the explanation, wherein the nodes of the knowledge graph are entities, and the edges of the knowledge graph are semantic associations between entities; Acquire historical audience interaction data, and infer interest distribution based on the historical audience interaction data and the environmental data to obtain interest distribution data; Knowledge units corresponding to the interest distribution data are selected from the knowledge graph, and explanatory text is generated based on the knowledge units and their interest matching degree under the interest distribution data.

7. The method according to claim 1, characterized in that, The explanation according to the explanation strategy and the explanation text includes: The digital narrator is driven to give a presentation in accordance with the presentation strategy and the presentation text.

8. The method according to claim 1, characterized in that, The method further includes: During the explanation process according to the described explanation strategy and the described explanation text, the text association information corresponding to the explanation text is displayed, wherein the text association information includes images, videos and models.

9. An adaptive explanation system, characterized in that, include: A multi-dimensional environmental perception module is used to acquire environmental data in the target scene, wherein the environmental data includes user space data and user emotion data; The dynamic explanation strategy generation module is used to make multi-parameter decisions based on the environmental data, obtain the explanation strategy corresponding to the environmental data, and generate the explanation text corresponding to the environmental data. The explanation module is used to provide explanations according to the explanation strategy and the explanation text.

10. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the electronic device to perform the method as described in any one of claims 1 to 8.