Accompanying dialogue system and method based on real-time environment perception and knowledge graph enhancement

By combining real-time environmental perception and knowledge graph technology, an intelligent tourism companion robot system has been realized, which solves the problem that existing devices cannot provide real-time interaction and personalized services. It provides accurate scene understanding and in-depth knowledge services, enhancing the immersiveness and fun of the tourism experience.

CN120894187APending Publication Date: 2025-11-04YANGZHOU POLYTECHNIC COLLEGE

Patent Information

Application Number
CN202511001710.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing tourism support equipment and systems cannot achieve real-time interaction with tourists, lack personalized and in-depth information services, and cannot provide personalized and immersive tourism experiences based on users' immediate interests and environmental changes.

Method used

By combining real-time environmental perception, knowledge graph technology, and augmented reality (AR) interaction, the system collects user environmental information through a multimodal perception module, uses knowledge graphs for dynamic association and reasoning, and combines natural language understanding and dialogue management modules to generate personalized recommendations and narratives, thereby achieving multimodal interaction.

Benefits of technology

It achieves accurate scene understanding, in-depth knowledge services, and a highly immersive interactive experience, providing natural and smooth "companion"-style dialogue, dynamically adapting to environmental changes, and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894187A_ABST
    Figure CN120894187A_ABST
Patent Text Reader

Abstract

The invention discloses an accompanying dialogue system and method based on real-time environment perception and knowledge graph enhancement. The system comprises a multi-mode real-time environment perception module, a knowledge graph construction and enhancement module, a natural language understanding and dialogue management module, a personalized recommendation and narration generation module and an immersive multi-mode interaction module. The method comprises the steps of multi-modal real-time environment perception and situation data generation; performing dynamic association, query and enhanced reasoning on the knowledge graph; natural language understanding and dialogue management, personalized recommendation and narrative generation, and immersive multi-modal interaction presentation and feedback reception. According to the method and the system, the concern point of the user, namely scenery or details, can be accurately positioned, and context information required by subsequent service intelligence is provided, so that the problems of poor environment perception ability, weak interaction immersion, dull knowledge service, lack of individuation, insufficient intelligent accompanying experience and the like in the existing tourism auxiliary technology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing and artificial intelligence, and particularly relates to an intelligent travel companion dialogue robot system combining real-time environment perception, knowledge graph technology, natural language processing and augmented reality (AR) interaction. BACKGROUND

[0002] With the rapid development of information technology and mobile Internet, people's demand for tourism mode and tourism experience is also changing. Traditional tourism represented by paper guide, fixed-line group tour or simple and practical information APP cannot meet the demand of people for personalized, intelligent and immersive tourism experience and deep cultural exchange in tourism. At present, various types of tourism auxiliary technologies on the market mainly include the following categories: electronic guide devices or APPs, which are devices or applications that can provide tourists with text, pictures, audio and video introductions of scenic spots and have GPS positioning function, and play the introduction of scenic spots at the specified location. However, in most cases, the content of these devices is pre-recorded and rigid, which informs tourists about the content of the tourist attractions, and does not have the function of real-time interaction with tourists and responding to the immediate interests of tourists, the specific micro-environment in front of them and the questions they raise. The information presentation of such devices is relatively single and lacks a sense of immersion; general dialogue robots / voice assistants have certain natural language understanding ability and dialogue ability, and can answer some simple common sense questions and complete some simple task instructions, but they do not dig deep into professional knowledge in this vertical industry, have weak perception ability for complex tourism scenarios, and cannot combine with the actual environment, so they cannot distinguish the specific scenery in front of them and provide more interesting and deeper information related to the current environment. The knowledge graph-based question and answer system can represent entities such as people, things and events in the real world and their relationships in a structured form, and has higher accuracy and stronger reasoning ability for question and answer. Some knowledge graphs have been applied to the tourism industry, such as building a tourism knowledge graph to recommend scenic spots and plan trips, but they cannot be combined with the real-time physical environment of people, and cannot achieve the effect of intelligent interaction of "what you see is what you get". Augmented reality (AR) tourism uses AR to overlay virtual information in the real scene, such as historical restoration information and cultural relics, to bring people a new tourism experience. However, AR tourism applications have low interactivity, are generally scripted content designed in advance, do not have intelligent dialogue capabilities, cannot adjust the content based on user feedback and changes around them, and do not have the function of deep dialogue and emotional exchange with users like "companion". Therefore, it is necessary to propose a travel companion robot system that can truly understand the needs of tourists, adapt to complex tourism scenarios, provide personalized and deep information services, and bring a new immersive experience. SUMMARY

[0003] The application aims to provide an intelligent travel companion dialogue robot system and method combining real-time environment perception, knowledge graph technology, natural language processing and augmented reality (AR) interaction.

[0004] The companion dialogue system based on real-time environment perception and knowledge graph enhancement comprises a multi-modal real-time environment perception module, a knowledge graph construction and enhancement module, a natural language understanding and dialogue management module, a personalized recommendation and narrative generation module and an immersive multi-modal interaction module.

[0005] Further, the multi-modal real-time environment perception module comprises a visual perception unit, an auditory perception unit and a positioning and attitude perception unit, obtains the geographical position information of the user terminal device, and obtains the attitude information of the user terminal device including the orientation, pitch angle and roll angle through an inertial measurement unit.

[0006] Further, the knowledge graph construction and enhancement module comprises a core tourism knowledge graph, a dynamic knowledge enhancement unit and a knowledge reasoning engine, and the knowledge reasoning engine is configured to perform rule-based reasoning, description logic-based reasoning on the core tourism knowledge graph to respond to complex queries and generate deep knowledge associations.

[0007] Further, the natural language understanding and dialogue management module comprises a natural language understanding unit, a dialogue management unit and a dialogue module; the dialogue management unit tracks and updates the complete information state of the current dialogue by using a dialogue state tracker, infers the next best action according to a model algorithm and outputs the action to the dialogue module.

[0008] The companion dialogue method based on real-time environment perception and knowledge graph enhancement comprises the following steps:

[0009] (1) Collecting physical environment information and user's physiological or behavioral information of the user terminal device through multi-modal sensors, and analyzing and generating structured real-time context data;

[0010] (2) Receiving natural language input of the user or triggering active interaction by the system according to the real-time context data; performing natural language understanding on the user input, identifying user intent and key information; combining the real-time context data and user intent, performing dynamic query, association and reasoning in the pre-constructed tourism knowledge graph to obtain related knowledge;

[0011] (3) Generating information or scene story type content conforming to the current interaction of the dialog parties or dialog decisions to be made by the system according to the user's portrait information, current context data, knowledge obtained from the knowledge graph and current dialog state;

[0012] (4) Presenting the generated recommendation, narrative content or dialog reply to the user in the form of text, voice, augmented reality superposition, virtual reality scene, etc. as an output mode;

[0013] (5) Immersive multi-modal interactive presentation and feedback reception.

[0014] Further, the step (1) is realized by a multi-modal real-time environment perception module, including the following steps:

[0015] (1.1) Data acquisition and preliminary processing, using the sensors integrated in the user terminal to collect visual information including image or video stream V raw , auditory information voice A raw , environmental sound E raw , location information L raw and posture information P raw ;

[0016] (1.2) Visual information processing, applying a Faster RCNN-based model to V raw , the Faster RCNN model including a region proposal network RPN and a Fast RCNN detection network;

[0017] (1.3) Auditory information processing, applying a Conformer-based model for speech recognition ASR to A raw , the Conformer model combining the ability of convolutional neural network to capture local features and the ability of Transformer self-attention mechanism to capture global context, and its training usually adopts CTC loss, Transformer loss or hybrid CTC-Attention loss;

[0018] (1.4) Environmental sound event detection, to E raw perform environmental sound event detection;

[0019] (1.5) Positioning and pose perception, fusing GPS, Wi-Fi fingerprint, iBeacon data to acquire user's precise outdoor or indoor geographic location L fused :

[0020]

[0021] where each weight is dynamically adjusted according to the signal;

[0022] (1.6) Pose estimation, estimating user terminal's real-time pose through IMU data by using extended Kalman filter; EKF's state vector x k may include pose quaternion and gyroscope bias;

[0023] (1.7) Context data fusion and structuring, fusing the information obtained by the above processing into structured real-time context data vector S t :

[0024]

[0025] S t as input of subsequent modules.

[0026] Further, the step (2) is realized by a knowledge graph construction and enhancement module, including the following steps:

[0027] (2.1) Knowledge graph construction: construct a large-scale tourism knowledge graph KG core =(E, R, F), where E is the entity set, R is the relationship set, and F is the set of fact triples (h, r, t),

[0028]

[0029] Each entity is associated with an attribute vector:

[0030] φ(e i )=[loc(e i ), type(e i ), time_peirod(e i ), popularity(e i ),...]∈R d

[0031] where loc(e i )=(lat i ,lon i ) represents the entity geographic location;

[0032] (2.2) Context-driven subgraph activation, according to S t geographical location L fused and visual focus V focus , locate relevant entity nodes in KG core and activate a dynamic subgraph KG active highly relevant to the current context, with an activation degree or node relevance Rel(e i , S t );

[0033] (2.3) Knowledge query and reasoning, the knowledge reasoning engine uses the TransE model based on graph embedding for complex query and relationship reasoning, the TransE model embeds entities and relationships into the same low-dimensional vector space, expecting that for a fact triple (h, r, t), entity embedding: e h , e t ∈R d , relationship embedding: r r ∈R d , the scoring function and loss are:

[0034] f r (h, t) = -||e h +r r -e t ||L2 + b h +b t

[0035] where b h , b t represent entity bias terms;

[0036] Improved loss function:

[0037]

[0038] Negative sampling probability:

[0039]

[0040] where α represents the sampling temperature parameter;

[0041] (2.4) Dynamic knowledge enhancement, dynamically update the information in KG active or KG core according to user interaction feedback and external real-time data sources.

[0042] Further, the step (3) is performed by a natural language understanding and dialogue management module, including the following steps:

[0043] (3.1) Natural language understanding, using a fine-tuned pre-trained language model BERT to process user input Utext and context S t Encoding, the [CLS] output vector of BERT is used for intent recognition, and the output vector at the word level is used for slot filling;

[0044] (3.2) Dialogue state tracking, maintain and update the state DS of the current dialogue t , including historical intent, confirmed slot, system historical action, current environmental focus entity and context information of KG_{active};

[0045] (3.3) Dialogue policy generation, using DQN algorithm based on deep reinforcement learning to optimize dialogue policy, DQN uses neural network Q to approximate state-action value function.

[0046] Further, the step (4) is executed by the personalized recommendation and narrative generation module, including the following steps:

[0047] (4.1) User portrait construction, construct and maintain user portrait UP u , containing user's long-term preference Short-term intent Knowledge level

[0048] (4.2) Personalized recommendation engine, combined with UP u , S t And KG active , using hybrid recommendation algorithm to recommend scenic spots, exhibits, activities and other users. Hybrid recommendation score Score rec (item j |UP u , S t ):

[0049]

[0050] The similarity function is divided into three types: content-based, knowledge graph path-based, and context matching-based.

[0051] (4.3) Dynamic narrative generation, using fine-tuned BART model, BART is a Transformer-based encoder-decoder model, and the pre-training task includes text denoising. For narrative generation, the input is structured knowledge triplets, user portrait features UP u , context S t And style control symbol Style. The model learns to generate coherent, vivid, and personalized narrative text N text :

[0052] Input structured knowledge:

[0053]

[0054] wherein Desc(·) denotes a textual description of an entity or a relation;

[0055] Conditional generation probability:

[0056]

[0057] Encoder input:

[0058]

[0059] wherein Style: style control symbol;

[0060] Loss function:

[0061]

[0062] wherein K input is the information input to the BART encoder.

[0063] Further, the step (5) is executed by an immersive multi-modal interaction module, comprising the following steps:

[0064] (5.1) Natural language generation and speech synthesis, converting the system internal semantic representation A sys , N text into natural language text and synthesizing it into speech output A out through a TTS engine;

[0065] (5.2) Augmented reality content presentation, the 6DoF pose of the user terminal in the world coordinate system is obtained Camera to world by the EKF and VSLAM algorithm provided by the EPM module, VSLAM estimates the camera motion and constructs a sparse or dense environment map by extracting and matching visual feature points ORB in consecutive video frames;

[0066] (5.3) Multi-modal input processing, receiving and processing various inputs of the user, such as voice commands, touch screen gestures, air gestures, head posture, eye movement input, converting them into instructions or feedback that the system can understand, forming a closed-loop interaction.

[0067] Beneficial effects: compared with the prior art, the present application has the following significant advantages:

[0068] (1) Precise scene understanding and context awareness: after multi-modal real-time environment perception, the user's focus point, i.e. the scene or details, can be accurately located, and the context information required by the subsequent service intelligence is provided;

[0069] (2) Deepening and personalizing knowledge services: By combining powerful knowledge graphs and personalized recommendation technology, the system can provide more comprehensive content interpretation and more personalized information recommendations than ordinary tour guides, meeting the needs of users' personalized knowledge-seeking;

[0070] (3) Highly immersive interactive experience: With multi-modal interaction methods such as AR, abstract knowledge or stories can be visualized, making the content of knowledge and stories more vivid and lifelike, and improving the interest and user engagement of tourism;

[0071] (4) Natural and smooth "touring" conversation: Close human-to-human interaction enables intelligent conversation, and the robot is no longer just a machine that provides information for people, but a wise assistant on the journey, better improving the temperature on the journey;

[0072] (5) Dynamic adaptation and active service: The system can update the content of the corresponding service according to environmental changes, user behavior, etc., and can provide relevant recommendation information or suggestions for users without their demands. BRIEF DESCRIPTION OF DRAWINGS

[0073] Figure 1 is the overall framework diagram of the system of the present application;

[0074] Figure 2 is a schematic diagram showing the working principle of the environment perception module;

[0075] Figure 3 is a schematic diagram showing the construction and retrieval of the knowledge graph;

[0076] Figure 4 is a workflow diagram of the method of the present application. DETAILED DESCRIPTION

[0077] The technical solutions of the present application will be further described below in conjunction with the accompanying drawings.

[0078] For example, Figure 1As shown, the companion dialogue system based on real-time environment perception and knowledge graph enhancement described in the present application includes a multi-modal real-time environment perception module 110, a knowledge graph construction and enhancement module 120, a natural language understanding and dialogue management module 130, a personalized recommendation and narrative generation module 140, and an immersive multi-modal interaction module 150. The multi-modal real-time environment perception module 110 collects physical environment information and user's own physiological or behavioral information of the user terminal device, and analyzes the information to generate structured real-time context data. The knowledge graph construction and enhancement module 120 stores and manages a tourism knowledge graph containing tourism domain entities, concepts, relationships and their attributes, and dynamically associates, queries and reasons the knowledge graph according to real-time context data. The natural language understanding and dialogue management module 130 receives user's natural language input, combines real-time context data and knowledge graph query reasoning results, and performs intent recognition, semantic understanding, dialogue state tracking and dialogue strategy generation. The personalized recommendation and narrative generation module 140 generates personalized tourism information recommendation and dynamic narrative content according to user portrait and real-time context data. The immersive multi-modal interaction module 150 displays the reply generated by the natural language understanding and dialogue management module or the content generated by the narrative generation module to the user through one or more output modalities, and accepts multi-modal input of user feedback.

[0079] As shown, Figure 2 The multi-modal real-time environment perception module 110 includes:

[0080] The visual perception unit 111 is configured to acquire real-time images or video streams through image acquisition devices, and analyze the images or video streams using computer vision algorithms to identify objects, text, faces or scene types in the scene, and determine the user's visual focus.

[0081] The auditory perception unit 112 is configured to acquire user speech and environmental sound through audio acquisition devices, convert user speech into text through speech recognition, and optionally perform voiceprint recognition or environmental sound event detection.

[0082] The positioning and attitude perception unit 113 is configured to acquire geographic location information of the user terminal device through at least one or a combination of global positioning system, cellular network positioning, Wi-Fi positioning, Bluetooth positioning, and to acquire attitude information of the user terminal device through an inertial measurement unit, including orientation, pitch angle and roll angle.

[0083] The data received from each of the above perception units are combined together and fused into a configuration for alignment of each of the data, estimation of state, and semantic combination to generate real-time contextualized information data, which contains information including current location information, user's observed direction, current point of interest, and corresponding attributes, and user's original inquiry sentence.

[0084] The target detection algorithm used by the visual perception unit 111 is the Faster R-CNN algorithm, and the image recognition algorithm is an image scene recognition algorithm. The visual perception unit 111 can realize visual SLAM and scene text recognition functions.

[0085] The knowledge graph construction and enhancement module 120 includes:

[0086] A core tourism knowledge graph 121 stored in the form of a graph database, the node types include but are not limited to tourist attractions, cultural exhibits, historical figures, historical events, geographical entities, cultural concepts, catering and accommodation entities, transportation entities, user preference labels, and the edge types include but are not limited to semantic relationships and attributes such as "located in", "belongs to", "author is", "occurred in", "built in", "style is", "recommended dishes", "opening time", "ticket price", etc.

[0087] A dynamic knowledge enhancement unit 122 configured to dynamically activate related subgraphs or nodes in the core tourism knowledge graph according to the real-time context data output by the multi-modal real-time environment perception module 110, in particular the user's visual focus and geographical location, and update the core tourism knowledge graph according to the user's interaction feedback and external data sources.

[0088] A knowledge reasoning engine 123 configured to perform rule-based reasoning, description logic-based reasoning on the core tourism knowledge graph 121, to respond to complex queries, and generate deep knowledge associations.

[0089] The natural language understanding and dialogue management module 130 includes:

[0090] A natural language understanding unit 131 using a pre-trained language model BERT based on deep learning for fine-tuning to achieve high-precision intent recognition, key information slot filling, and context understanding and reference resolution in combination with the real-time context data for user natural language input;

[0091] The dialogue management unit 132 tracks and updates the full information state of the current dialogue, i.e. the user's full history of intents, confirmed slots, the system's full history of actions, the focus entity of the current environment, and all context information currently in effect under the knowledge graph, using a dialogue state tracker. Based on this, the next best action is inferred using a model algorithm, which is a dialogue policy model, and output to the dialogue module 133. The dialogue policy model can be trained based on a deep reinforcement learning algorithm (DQN / PPO, etc.), and optimized and trained based on feedback from real users.

[0092] The personalized recommendation and narrative generation module 140 includes:

[0093] The user portrait unit 141 is configured to build and maintain a portrait of each user, which includes the user's long-term preferences (such as interest levels in history, art, natural scenery, and specific culture), short-term intents (inferred from the current dialogue and behavior), knowledge level, touring habits, and historical interaction records.

[0094] The personalized recommendation engine 142 uses a hybrid recommendation algorithm that combines content-based recommendation (matching user portraits with attributes of entities in the knowledge graph), collaborative filtering recommendation (based on the behavior of similar users), and knowledge graph path-based recommendation (exploring deep connections between user interest points and potential recommended items in the knowledge graph) to generate personalized recommendations for attractions, exhibits, activities, dining, and touring routes for users.

[0095] The dynamic narrative generation unit 143 converts structured knowledge graph information and personalized recommendation results into narrative speeches that are fluent, vivid, and image-based in specific contexts, and can be flexibly adjusted in terms of narrative style, tone, and semantic detail.

[0096] The dynamic narrative generation unit 143 uses a narrative generation method based on template filling and rule selection or a generation method based on a neural network BART model, and integrates style transfer algorithms to achieve more creative, novel, and diverse high-quality narrative texts.

[0097] The immersive multi-modal interaction module 150 also includes the following functional modules:

[0098] The natural language generation unit 151 is configured to convert semantic representations (such as dialogue actions and knowledge fragments) within the system into natural language text, and synthesize it into speech output through a text-to-speech (TTS) engine that supports different voice tones, speech rates, and emotional speech synthesis.

[0099] The content presentation unit 152 includes information related to a scene obtained by the multi-modal real-time environment perception module (110) and the scene understanding result, which is fused and superimposed on a real scene and a virtual scene on a display interface of a user terminal device.

[0100] As shown in Figure 4 , the companion tour dialogue method based on real-time environment perception and knowledge graph enhancement provided by the present application comprises the following steps:

[0101] Step 1: Multi-modal real-time environment perception and context data generation

[0102] This step is performed by the multi-modal real-time environment perception module (EPM) 110, aiming to accurately capture the physical environment and user behavior of the user, and generate structured real-time context data S t .

[0103] 1.1, Data acquisition and preliminary processing: using the cameras, microphones, GPS, IMU, etc. sensors integrated in the user terminal (such as smart phones, AR glasses), to respectively collect visual information: image / video stream V raw , auditory information voice A raw , environmental sound E raw , location information L raw and attitude information P raw .

[0104] 1.2, Visual information processing: V raw applies a model based on Faster RCNN. The model contains a region proposal network (RPN) and a Fast RCNN detection network. Its joint loss function FasterRCNN can be summarized as the sum of RPN loss and Fast RCNN loss, and the complete FasterRCNN loss function is:

[0105] L FasterRCNN ({p i , t i , u i , v i ,}) = L RPN + L FastRCNN

[0106] Among them, the RPN loss is:

[0107]

[0108] Among them, pi: the probability that the predicted anchor i is foreground. The true label of anchor i (1 = foreground, 0 = background). t i : predicted bounding box offset. The actual bounding box offset. L cls Classification loss (usually binary cross-entropy). L reg : Regression loss (usually Smooth L1 loss). The normalization term for the classification loss (usually the size of a mini-batch). The normalized term of the regression loss (usually the number of anchors). λ RPN : Balance the weights of classification and regression losses (default 1).

[0109] The FastRCNN loss is:

[0110]

[0111] Among them, u j : Predicted category probability distribution (multi-class classification). Real category labels. j : The predicted bounding box correction value. The target for realistic bounding box correction. cls Classification loss (usually multi-class cross-entropy). L reg : Regression loss (usually Smooth L1 loss). The normalization term for classification loss. The normalized term of the regression loss. λ RPN : The weights that balance the classification and regression losses.

[0112] Scene text recognition (OCR): Employs a CRNN+CTC model. CRNN (Convolutional Recurrent Neural Network) first extracts image features using a CNN, then models the sequence features using an RNN (such as LSTM), and finally performs end-to-end training using the CTC (Connectionist Temporal Classification) loss function, eliminating the need for precise character-level alignment. The CTC loss function L... CTC Defined as the sum of the negative log-likelihoods of all paths π that can map to the true label sequence 1:

[0113]

[0114] Where, y: the model's output (e.g., the predicted probability distribution of an RNN or Transformer). l: the true label sequence (e.g., "cat"). ∏: a possible alignment path, i.e., an intermediate sequence of the model's output (e.g., "caat"). B: the Blank function (B function), used to remove duplicate characters and whitespace.-1(l) P(π|y): Probability of path π given model output y.

[0115] 1.3, Auditory information processing: Speech recognition (ASR) Conformer-based model

[0116] A raw A Conformer-based model is applied. Conformer model combines the ability of convolutional neural network to capture local features and the ability of self-attention mechanism of Transformer to capture global context. Its training usually adopts CTC loss, Transformer loss or hybrid CTC / Attention loss.

[0117] Model architecture, the encoder layer of Conformer consists of the following modules in series:

[0118] ConformerBlock = FFN → MultiHeadSelfAttention → Conv → FFN

[0119] FFN (Feed-Forward Network):

[0120] FNN(x) = GeLu(xW1 + b1)W2 + b2

[0121] Multi-Head Self-Attention (MHSA):

[0122]

[0123] Convolution module:

[0124] Conv(x) = DeptwiseConv1D(x) → Swish → PointwiseConv1D

[0125] Loss function (hybrid CTC / Attention):

[0126] L ASR = aL CTC (y, l) + (1 - a)L Attention (y, l)

[0127] CTC loss (alignment-independent):

[0128]

[0129] Attention loss (alignment-dependent):

[0130]

[0131] where a e [0, 1] is a weight hyper-parameter (usually set to 0.3).

[0132] Output text processing: recognition result U txt Post-processing (e.g., punctuation restoration, entity normalization) and injection S t .

[0133] 1.4, Environmental sound event detection: on E raw Perform environmental sound event detection to identify sounds E related to tourism scenarios such as "bell", "fountain", etc. cat .

[0134] Model architecture (CNN + Transformer hybrid)

[0135] Audio features = Log_Mel spectrum -> local features -> global context features -> E cat

[0136] Spectrum extraction:

[0137] X mel = 10log 10 (||STFT(E raw )|| 2 ·M mel )

[0138] where M mel is a Mel filter bank.

[0139] Classification loss:

[0140]

[0141] C is the number of event categories (e.g., "bell", "fountain", "crowd", "music")

[0142] Semantic category mapping, detected sound event E cat mapped to standardized labels:

[0143] E cat e {bell, fountain, crowd, music,...}

[0144] 1.5, Positioning and pose perception: Fusion of GPS, Wi-Fi fingerprinting, iBeacon, etc. data to obtain accurate outdoor / indoor geographic location L of the user fused .

[0145] Outdoor positioning (GPS dominant):

[0146] L GPS = (lat GPS , lon GPS , altGPS )

[0147] Indoor positioning (Wi-Fi / iBeacon fusion):

[0148]

[0149] Fusion strategy (least square method):

[0150]

[0151] Weights are dynamically adjusted according to signals.

[0152] 1.6, Pose estimation: Estimate the real-time pose of the user terminal by IMU data using an extended Kalman filter (EKF). The state vector x of the EKF k may include the pose quaternion and the gyroscope bias.

[0153] State vector x k :

[0154]

[0155] where the quaternion q represents the pose and β is the gyroscope bias.

[0156] Prediction step (IMU angular velocity driven):

[0157]

[0158] where represents the quaternion multiplication and f represents the discretized integration model.

[0159] Observation update (accelerometer + magnetometer):

[0160]

[0161] where R(q): quaternion to rotation matrix, g ref = [0, 0, 1], m ref represents the reference direction of the geomagnetic field.

[0162] Kalman gain and covariance update:

[0163]

[0164] where Jacobian matrix of the observation model h, R k : observation noise covariance.

[0165] State prediction:

[0166]

[0167] Where f is the state transition function, ω imu,k This is the gyroscope reading.

[0168] Error covariance prediction:

[0169]

[0170] Among them, F k Q is the Jacobian matrix of f. k Let be the process noise covariance.

[0171] Through the above steps, a precise orientation estimate of the device (such as AR glasses) relative to the geographic coordinate system is obtained. Used to determine the user's focal point.

[0172] 1.7 Contextual Data Fusion and Structuring: Contextual data fusion and structuring involves fusing the information obtained from the above processing into a structured real-time contextual data vector S. t :

[0173] ts∈R +

[0174] L fused =(lat,lon,alt,floor)∈R 4

[0175] O dev =(q w q x q y q z )∈S 3

[0176]

[0177] U text ∈T *

[0178] E cat ∈{e1,…,e N}

[0179] (Optional): A audio ∈R d

[0180] The S t As input to subsequent modules, it lays the foundation for understanding user intent and providing contextualized services.

[0181] Step 2: Dynamic association, querying, and augmented reasoning using the knowledge graph

[0182] like Figure 3As shown, this step is performed by a knowledge graph construction and enhancement module (KGM) 120, based on S t Interact with the tourism knowledge graph.

[0183] 2.1 Tourism knowledge graph construction: A large-scale tourism knowledge graph KG core = (E, R, F) is constructed in advance, where E is the entity set, R is the relation set, and F is the set of fact triples (h, r, t).

[0184]

[0185] Each entity is associated with an attribute vector:

[0186] φ(e i ) = [loc(e i ), type(e i ), time_peirod(e i ), popularity(e i ),...] ∈ R d

[0187] where loc(e i ) = (lat i , lon i ) represents the entity's geographic location.

[0188] 2.2 Context-driven subgraph activation: According to the geographic location L t and visual focus V fused in S focus , relevant entity nodes are located in KG core , and a dynamic subgraph KG active highly related to the current context is activated, with the activation degree or node relevance Rel(e i , S t ).

[0189] Dynamic subgraph generation:

[0190] KG active = {(h, r, t) ∈ KG core} | Rel(h, S t ) ≥ τVRel(t, S t ) ≥ τ

[0191] where τ is the relevance threshold (e.g., 0.7);

[0192] Entity relevance score:

[0193]

[0194] Key similarity function:

[0195] Spatial Similarity:

[0196]

[0197] Visual Focus Similarity:

[0198]

[0199] Semantic Similarity (based on Text Embedding):

[0200] sim sem (e i , S t ) = cos(BERT(desc(e i )), BERT(U text ))

[0201] Weight w k Learned Dynamically by Attention Mechanism:

[0202]

[0203] 2.3. Knowledge Query and Reasoning: The knowledge reasoning engine utilizes the TransE model based on graph embedding for complex query and relation reasoning. The TransE model embeds entities and relations into the same low-dimensional vector space, expecting that for a fact triple (h, r, t), the entity embedding: e h , e t ∈ R d , the relation embedding: r r ∈ R d , the scoring function and loss are:

[0204] f r (h, t) = -||e h +r r -e t ||L2 + b h +b t

[0205] where b h , b t represent entity bias terms.

[0206] Improved Loss Function (Adversarial Negative Sampling):

[0207]

[0208] Negative Sampling Probability:

[0209]

[0210] where α represents the sampling temperature parameter.

[0211] Multi-hop reasoning example: Query "Other works by Designer A":

[0212] PathScore = fr1(e A , e B ) + fr2(e B , e C )

[0213] Where r1 = design, r2 = similar works.

[0214] 2.4 Dynamic knowledge enhancement: Update information in KG or KG active , such as entity attributes, relationship weights or confidence, according to user interaction feedback and external real-time data sources (such as weather API, temporary exhibition information). core

[0215] Entity attribute update:

[0216] φ(e i )←βφ old (e i )+(1-β)φ new (API)

[0217] Where β is the forgetting factor (take 0.9).

[0218] Relationship confidence adjustment:

[0219] conf(h, r, t)←conf(h, r, t) + η·user feedback score

[0220] Where η is the learning rate;

[0221] Temporary event injection, add new dynamic triple:

[0222] F←F∪{(“A”, “B”, “C”)}

[0223] Step 3: Natural language understanding and dialogue management

[0224] This step is executed by the natural language understanding and dialogue management module (NLU&DM) 130, which processes user input and manages the dialogue process.

[0225] 3.1 Natural language understanding (NLU): Use fine-tuned pre-trained language model BERT. Encode user input U text and context S t , the [CLS] output vector of BERT is used for intent recognition, and the word-level output vector is used for slot filling.

[0226] BERT context encoding: ​

[0227]

[0228] where, [CLS] vector (global intent representation), Context representation of the jth token, Context(S t ): textualized context extracted from S (e.g., “current location: the Forbidden City; focus of sight: the Hall of Supreme Harmony”).

[0229] Intent recognition (classification task):

[0230]

[0231] where, K: number of intent categories (e.g., I ∈ {QUERY_FACT, RECOMMEND_ATTRACTION, …}), where QUERY_FACT (ask for facts), RECOMMEND_ATTRACTION (request to recommend attractions), and other intents.

[0232] Slot filling (sequence labeling):

[0233]

[0234] where, C: number of slot types (e.g., Slot k ∈ {ticket, opening time, …}).

[0235] Joint training loss:

[0236]

[0237] 3.2, Dialog state tracking (DST): maintain and update the state DS t of the current dialogue, including historical intents, confirmed slots, system historical actions, current environmental focus entities, and context information of KG_{active}.

[0238] State representation:

[0239]

[0240] State update function:

[0241] DS t+1 = Update(DS t , A sys , U text , S t )

[0242] Focus entity update rule:

[0243]

[0244] If a new entity is detected, otherwise inherit by relevance.

[0245] Slot filling:

[0246] Slots confirmed ←Slots confirmed ∪{(s,ExtractValue(s,U text ,KG active ))}

[0247] 3.3, Dialogue policy generation: DQN (Deep Q-Network) algorithm based on deep reinforcement learning is adopted to optimize the dialogue policy. DQN uses a neural network Q to approximate the state-action value function.

[0248] Q network architecture:

[0249]

[0250] State encoding:

[0251]

[0252] Action space:

[0253]

[0254] Bellman optimization objective:

[0255]

[0256] Reward function:

[0257]

[0258] Exploration policy:

[0259]

[0260] where ds is the current dialogue state, a is the action taken by the system (e.g. "provide attraction information", "ask user preference"), rew is the reward received (e.g. task completion degree, dialogue fluency), γ is the discount factor, ds' is the next state, θ i is the parameter of the current Q network, θ i ' is the parameter of the target Q network (copied periodically from θ i ), and U(D) represents sampling from the experience replay pool D. The system selects actions according to the ε-greedy policy: with probability ε, randomly select, and with probability 1-ε, select argmax a Q(ds, a'; θ i ).

[0261] Step 4: Personalized recommendation and narrative generation

[0262] This step is performed by the Personalized Recommendation and Narrative Generation Module (PNM) 140.

[0263] 4.1, User Profile Construction (UPM): Construct and maintain the user profile UP u , containing the user's long-term preferences (such as interest level in history, art, etc.), short-term intent (inferred from the current conversation), knowledge level , etc.

[0264]

[0265] Dynamic updating mechanism

[0266] Long-term preference update (based on implicit feedback):

[0267]

[0268] Short-term intent inference (based on current conversation):

[0269]

[0270] where C NLU : [CLS] vector of NLU module.

[0271] 4.2, Personalized Recommendation Engine: Combine UP u , S t and KG active , use hybrid recommendation algorithm to recommend scenic spots, exhibits, activities, etc. to users. Hybrid recommendation score Score rec (item j |UP u , S t ):

[0272]

[0273] The similarity function is divided into three types: content-based, knowledge graph path-based, and context matching-based.

[0274] A, content-based:

[0275]

[0276] B, Knowledge Graph Path: Use meta-path-based similarity calculation. Given user profile entity e u ∈KG and target scenic spot entity e∈KG, define the meta-path set P=(p1, p1, …p M) (e.g., p1: user likes architecture style belongs to the attraction).

[0277] Path similarity is:

[0278]

[0279] where II(·) is the path existence indicator function, w p is the path weight (learned from historical interaction data)

[0280] Dynamic path activation:

[0281] According to the context S t Dynamic filtering of relevant paths:

[0282]

[0283] where the relevance score is:

[0284] Rcl(p, S t ) = MLP([Emb(p); Emb(V focus ); Emb(E cat )])

[0285] C, Context matching: encode St into a context vector:

[0286]

[0287] Entity-context matching degree: for an entity e j , calculate its matching degree with the current context:

[0288]

[0289] where φ(e j ) is the entity feature vector (fusion of attributes, categories, etc.). W e , W s : learnable projection matrix.

[0290] Mixed recommendation score (complete formula):

[0291]

[0292] Weight dynamic calculation:

[0293]

[0294] where sim k represents the similarity / matching degree score of different recommendation strategies (based on content, collaborative filtering, based on knowledge graph path), w k is the weight of each strategy.

[0295] 4.3, Dynamic narrative generation: Fine-tuned BART model. BART is a Transformer-based encoder-decoder model, pre-training tasks include text denoising (e.g., sentence shuffling, token deletion / masking). For narrative generation, the input can be structured knowledge triples (from KGM), user profile features UP u , context S t , and style controls Style. The model learns to generate coherent, vivid, personalized narrative text N text . Its generation process can be represented as:

[0296] Input structured knowledge:

[0297]

[0298] where Desc(·): textual description of entity / relationship.

[0299] Conditional generation probability:

[0300]

[0301] Encoder input:

[0302]

[0303] where Style: style controls.

[0304] Loss function:

[0305]

[0306] where K input, is the information input to the BART encoder.

[0307] Step 5: Immersive multi-modal interaction presentation and feedback reception

[0308] This step is executed by the immersive multi-modal interaction module (IIM) 150.

[0309] 4.4, Natural Language Generation (NLG) and Text-to-Speech (TTS): Convert the system's internal semantic representation A sys , N text , into natural language text and synthesize it into speech output A out through a TTS engine (supports multiple voices, speech rates, emotions).

[0310] Semantic-to-text generation (based on BART):

[0311]

[0312] Style control:

[0313]

[0314] Speech synthesis (Neural TTS):

[0315] A out = WaveNet(Prosody(A text , EmoScore(U text ))

[0316] Prosody parameters:

[0317]

[0318] 4.5, Augmented Reality (AR) content presentation: The 6DoF pose of the user terminal (e.g. AR glasses) in the world coordinate system is provided by the EKF of the EPM module and the VSLAM (Visual Simultaneous Localization and Mapping) algorithm. VSLAM estimates camera motion and constructs a sparse / dense environment map by extracting and matching visual feature points ORB in consecutive video frames. VSLAM pose estimation:

[0319] Feature point matching (ORB descriptor):

[0320]

[0321] Pose optimization (PnP problem):

[0322]

[0323] where ∏: camera projection model, Map point world coordinates. For a specific object (e.g. "Forbidden City Hall of Supreme Harmony") recognized in the real scene, its anchor point in the world coordinate system is provided by the EPM.

[0324] 4.6, 3D model projection and rendering: If you want to superimpose the restored 3D model of the Hall of Supreme Harmony, it is defined by a series of vertices V model_local in its local coordinate system. First, transform the model to the world coordinate system: V model_world = T ma V model_local where T ma is the transformation of the model to the anchor point (may include scaling, rotation to match the real object).

[0325] Model transformation:

[0326]

[0327] where s: scaling factor, R align : rotation to align the real object.

[0328] Then, project the model vertices in the world coordinate system to the 2D display plane of the AR glasses:

[0329]

[0330] where K: camera intrinsic matrix, is the world-to-camera transformation.

[0331] Use the graphics rendering pipeline to render the 3D model with texture and fuse it with the real scene image captured by the camera in real time according to the projection point P img and depth information. The text label will be rendered as a 2D texture, or anchored to a specific part of the 3D model and always facing the user (billboarding).

[0332] Depth test and fusion:

[0333]

[0334] Text label rendering (billboarding):

[0335]

[0336] where δ: offset distance, n cam : camera orientation vector

[0337] 4.7, Multi-modal input processing: Receive and process various inputs of the user, such as voice commands (which have been processed by the EPM), touch screen gestures, air gestures, head pose, eye movement input, etc., and convert them into instructions or feedback that the system can understand, forming a closed-loop interaction.

[0338] Ray detection (gesture interaction):

[0339] Ray: r(t) = o cam + td cam , t ≥ 0

[0340] AABB intersection detection:

[0341] t hit = min{t | r(t) ∈ Box(V model_world )}

[0342] Eye movement input processing:

[0343]

[0344] Through the close cooperation of the above steps, the system can understand the user's situation and needs, intelligently apply the knowledge graph for question and answer explanation, realize personalized service based on user portrait, and provide immersive interaction using AR technology, thereby significantly improving the knowledge, interest and convenience of tourism. The system also contains a continuous learning and optimization mechanism to collect user feedback and continuously iterate and improve the performance of each module.

Claims

1. A companion dialogue system based on real-time environmental awareness and knowledge graph enhancement, characterized in that, The system includes a multimodal real-time environment perception module (110), a knowledge graph construction and enhancement module (120), a natural language understanding and dialogue management module (130), a personalized recommendation and narrative generation module (140), and an immersive multimodal interaction module (150). The multimodal real-time environment perception module (110) collects information about the physical environment of the user terminal device and the user's own physiological or behavioral information, and analyzes the information to generate structured real-time contextual data. The knowledge graph construction and enhancement module (120) stores and manages a tourism knowledge graph containing entities, concepts, relationships, and attributes in the tourism field, and enhances the knowledge graph based on the real-time contextual data. The system performs dynamic association, query, and reasoning; the Natural Language Understanding and Dialogue Management module (130) receives the user's natural language input, combines real-time contextual data and the query reasoning results of the knowledge graph, and performs intent recognition, semantic understanding, dialogue state tracking, and dialogue strategy generation; the Personalized Recommendation and Narrative Generation module (140) generates personalized tourism information recommendations and dynamic narrative content based on user profiles and real-time contextual data; the Immersive Multimodal Interaction module (150) displays the responses generated by the Natural Language Understanding and Dialogue Management module or the content generated by the Personalized Recommendation and Narrative Generation module to the user through one or more output modalities and accepts multimodal input from the user.

2. The companion dialogue system based on real-time environmental awareness and knowledge graph enhancement according to claim 1, characterized in that, The multimodal real-time environment perception module (110) includes a visual perception unit (111), an auditory perception unit (112), and a positioning and attitude perception unit (113), which acquires the geographical location information of the user terminal device and acquires the attitude information of the user terminal device, including orientation, pitch angle, and roll angle, through the inertial measurement unit.

3. The companion dialogue system based on real-time environmental awareness and knowledge graph enhancement according to claim 1, characterized in that, The knowledge graph construction and enhancement module (120) includes a core tourism knowledge graph (121), a dynamic knowledge enhancement unit (122), and a knowledge reasoning engine (123). The knowledge reasoning engine (123) is configured to perform rule-based reasoning and description logic-based reasoning in the core tourism knowledge graph (121) to respond to complex queries and generate deep knowledge associations.

4. The companion dialogue system based on real-time environmental awareness and knowledge graph enhancement according to claim 1, characterized in that, The natural language understanding and dialogue management module (130) includes a natural language understanding unit (131), a dialogue management unit (132), and a dialogue module (133). The dialogue management unit (132) uses a dialogue state tracker to track and update all information states of the current dialogue, and infers the next best action based on the model algorithm and outputs it to the dialogue module (133).

5. A companion dialogue method based on real-time environmental awareness and knowledge graph enhancement, characterized in that, Includes the following steps: (1) Collect information about the physical environment of the user terminal device and the user's own physiological or behavioral information through multimodal sensors, and analyze and generate structured real-time contextual data; (2) Receive natural language input from users or trigger proactive interaction based on real-time contextual data; The system performs natural language understanding on user input to identify user intent and key information; and combines the real-time contextual data and user intent to perform dynamic queries, associations, and inferences in a pre-constructed tourism knowledge graph to acquire relevant knowledge. (3) Based on the user's profile information, current context data, knowledge obtained from the knowledge graph, and the current dialogue status, generate information or scenario-based content that matches the interaction between the two parties in the current dialogue, as well as the dialogue decisions that the system will make. (4) An output modality that presents generated recommendations, narrative content or dialogue responses to users in the form of text, voice, augmented reality overlay, or virtual reality scene; (5) Immersive multimodal interactive presentation and feedback reception.

6. The companion dialogue method based on real-time environment awareness and knowledge graph enhancement according to claim 5, characterized in that, Step (1) is implemented through a multimodal real-time environment perception module, including the following steps: (1.1) Data Acquisition and Preliminary Processing: Visual information, including images or video streams, is acquired using sensors integrated into the user terminal. raw Auditory information, speech A raw Ambient sound E raw Location information L raw and attitude information P raw ; (1.2) Visual information processing, for V raw The application uses a Faster R-CNN-based model, which includes a Region Proposal Network (RPN) and a Fast R-CNN detection network. (1.3) Auditory information processing, speech recognition (ASR) is based on the Conformer model, and performs A... raw Applying Conformer-based models combines the ability of convolutional neural networks to capture local features with the ability of Transformer's self-attention mechanism to capture global context. Its training typically employs CTC loss, Transformer loss, or a hybrid CTC-Attention loss. (1.4) Environmental sound event detection, for E raw Perform environmental sound event detection; (1.5) Positioning and attitude perception: Integrating GPS, Wi-Fi fingerprint, and iBeacon data to obtain the user's precise outdoor or indoor geographic location L fused : The weights are dynamically adjusted based on the signals. (1.6) Attitude estimation: The real-time attitude of the user terminal is estimated using IMU data and an extended Kalman filter; the state vector x of the EKF. k This may include attitude quaternions and gyroscope bias; (1.7) Contextual data fusion and structuring: The information obtained from the above processing is fused into a structured real-time contextual data vector S. t : S t As input for subsequent modules.

7. The companion dialogue method based on real-time environment awareness and knowledge graph enhancement according to claim 5, characterized in that, Step (2) is implemented through the knowledge graph construction and enhancement module, including the following steps: (2.1) Knowledge Graph Construction: Construct a large-scale tourism knowledge graph KG core = (E, R, F), where E is the entity set, R is the relation set, and F is the set of fact triples (h, r, t). Each entity is associated with an attribute vector: φ(e i )=[loc(e i ),type(e i ),time_peirod(e i ),popularity(e i ),...]∈R d Where, loc(e) i )=(lat i lon i () represents the geographical location of an entity; (2.2) Context-driven subgraph activation, based on S t Geographical location L fused and visual focus V focus In KG core Locate relevant entity nodes in the graph and activate a dynamic subgraph KG that is highly relevant to the current context. active Activation level or node correlation Rel(e i S t ); (2.3) Knowledge Query and Reasoning: The knowledge reasoning engine utilizes the TransE model based on graph embedding for complex queries and relational reasoning. The TransE model embeds entities and relations into the same low-dimensional vector space, expecting that for a fact triple (h, r, t), the entity embedding is: e h e t ∈R d Relational embedding: r r ∈R d The scoring function and loss are: f r (h,t)=-||e h +r r -e t ||L2+b h +b t Among them, b h b t Indicates entity bias term; Improved loss function: Negative sampling probability: Where α represents the sampling temperature parameter; (2.4) Dynamic knowledge enhancement: dynamically update KG based on user interaction feedback and external real-time data sources. active or KG core The information in the middle.

8. The companion dialogue method based on real-time environment awareness and knowledge graph enhancement according to claim 5, characterized in that, Step (3) is executed through the Natural Language Understanding and Dialogue Management module, including the following steps: (3.1) Natural Language Understanding: Utilizing a finely tuned pre-trained language model BERT to process user input U text and Context S t Encoding is performed, and BERT's [CLS] output vector is used for intent recognition, while the word-level output vector is used for slot filling; (3.2) Dialogue state tracking, maintaining and updating the current dialogue state DS t This includes historical intent, confirmed slots, system historical actions, current environment focus entities, and context information for KG_{active}; (3.3) Dialogue strategy generation: The dialogue strategy is optimized by using the DQN algorithm based on deep reinforcement learning. DQN uses a neural network Q to approximate the state-action value function.

9. The companion dialogue method based on real-time environment awareness and knowledge graph enhancement according to claim 5, characterized in that, Step (4) is executed through the personalized recommendation and narrative generation module, including the following steps: (4.1) User profile construction, building and maintaining user profiles UP u Including users' long-term preferences Short-term intentions knowledge level (4.2) Personalized recommendation engine, combined with UP u S t and KG active It employs a hybrid recommendation algorithm to recommend attractions, exhibits, and activities to users. The hybrid recommendation score is... rec (item j |UP u S t ): Similarity functions are of three types: content-based, knowledge graph path-based, and context-based matching. (4.3) Dynamic narrative generation employs a fine-tuned BART model. BART is a Transformer-based encoder-decoder model. Pre-training tasks include text denoising. For narrative generation, the input consists of structured knowledge triples and user profile features. u Context S t And the style control character. The model learns to generate coherent, vivid, and personalized narrative text N. text : Input structured knowledge: Where Desc(·) represents a textual description of an entity or relation; Conditional generation probability: Encoder input: Style: Style control character; Loss function: Among them, K input , is the information input to the BART encoder.

10. The companion dialogue method based on real-time environment awareness and knowledge graph enhancement according to claim 5, characterized in that, Step (5) is executed through the immersive multimodal interaction module and includes the following steps: (5.1) Natural language generation and speech synthesis, which will transform the semantic representation of the system into A sys N text It is converted into natural language text and then synthesized into speech output A through a TTS engine. out ; (5.2) Augmented reality content presentation: The 6DoF pose of the user terminal in the world coordinate system is obtained. The camera to the world is provided by the EKF and VSLAM algorithms of the EPM module. VSLAM estimates the camera motion and constructs a sparse or dense environment map by extracting and matching visual feature points ORB in consecutive video frames. (5.3) Multimodal input processing: receiving and processing various user inputs, such as voice commands, touch screen gestures, air gestures, head postures, and eye tracking inputs, converting them into instructions or feedback that the system can understand, forming a closed-loop interaction.

Citation Information

Patent Citations

  • AR (Augmented Reality) guide system for enhancing stone forest tourism experience

    CN118314306A

  • Blogger live video AR glasses scenery labeling method and labeling system

    CN120281967A

Cited By

  • Accompanying robot control method and system

    CN121515199A

  • Multi-modal scene interaction control method and device based on pre-generated resource map

    CN122019714A