A scene intelligent perception intention recognition method and system for scenic area accompanying

CN122549446BActive Publication Date: 2026-09-11ZHEJIANG XIAOYOU DIGITAL INNOVATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611039151.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-11
Estimated Expiration
2046-07-14

AI Technical Summary

Technical Problem

[0004]本发明提供一种面向景区伴游的场景智能感知意图识别方法及系统,用以解决现有技术中存在的意图识别缺乏游览上下文、位置感知仅停留在坐标映射、待确认状态易被覆盖、服务鲁棒性不足等问题

Benefits of technology

[0015]This invention provides a scene-based intelligent perception and intent recognition method and system for scenic area escort services. It ensures the continuity of the interaction process through a priority handling mechanism for pending confirmation states, preventing user confirmation responses from being misidentified as new requests. Employing a three-level intent recognition strategy that combines the tour context, it improves the accuracy of semantic understanding while ensuring service continuity through a degradation processing mechanism. Through event-driven location awareness and a dual-channel complementary perception mechanism, it achieves semantic conversion and conflict resolution of location data. Simultaneously, through unified global state management and multimodal response output, it significantly improves the accuracy and response speed of intent recognition in scenic area escort scenarios, reduces the overhead of large language model calls, and significantly improves the intelligent escort experience for tourists.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549446B_ABST
    Figure CN122549446B_ABST
Patent Text Reader

Abstract

The application provides a scene intelligent perception intention recognition method and system for scenic area accompanying, and belongs to the technical field of scenic area accompanying. The method comprises the following steps: receiving and separating the text and position data of the multi-modal request of a user; preferentially checking three groups of to-be-confirmed state flags and adopting a differentiated confirmation strategy for processing; executing a three-level intention recognition strategy when there is no to-be-confirmed state; performing event-driven semantic conversion on the position data and adopting a double-channel complementary perception mechanism; and updating the state and generating a multi-modal response according to the processing result. The system comprises a request access module, a data separation module, a global state storage module, a core scheduling module, a data processing module and a response output module. The application greatly improves the accuracy and response speed of intention recognition in a scenic area, reduces the calling cost of a large language model, and guarantees the continuity and robustness of the service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of scenic area escort technology, specifically to a scene-based intelligent perception and intent recognition method and system for scenic area escorts. Background Technology

[0002] With the continuous advancement of smart tourism development and the popularization of intelligent interactive technologies, intelligent tour guide systems for scenic spots have become the core support for tourist services. These systems typically combine user text interaction, real-time location positioning, and basic intent recognition capabilities to provide tourists with routine services such as attraction explanations, route planning, facility inquiries, and navigation guidance. Some systems can also connect to large language models to complete natural language question answering, and are widely used to improve the efficiency of scenic spot services and the tourist experience.

[0003] Existing scenic area tour guide systems suffer from several technical deficiencies that make them difficult to adapt to various scenarios in practical applications: First, intent recognition relies solely on shallow matching based on user text, failing to incorporate contextual information such as tour stage, historical dialogue, and navigation status for disambiguation, which easily leads to semantic misunderstandings. Second, there is no priority processing mechanism for system-initiated requests for mid-trip assistance, facility navigation, and route adjustments awaiting confirmation; user responses are often treated as new requests and re-identified, resulting in the overwriting of pending confirmation statuses and disruption of the interaction flow. Third, location awareness only performs a simple mapping between coordinates and points of interest, and the dual-channel results of automatic location reporting and user text input lack conflict resolution mechanisms, resulting in insufficient perception accuracy. Fourth, intent recognition relies on a single model or rule, compromising system robustness and service continuity. Summary of the Invention

[0004] This invention provides a scene-based intelligent perception and intent recognition method and system for scenic area tour guides, which solves the problems existing in the prior art, such as lack of tour context in intent recognition, location perception only remaining at coordinate mapping, easy overwriting of pending confirmation status, and insufficient service robustness.

[0005] To achieve the above objectives, one embodiment of the present invention provides a scene-based intelligent perception and intent recognition method for tour guides in scenic areas, characterized by comprising the following steps: Step S1: Receive user multimodal requests and preprocess them to separate text message data and location coordinate data; Step S2: Obtain the three sets of pending confirmation status flags currently being maintained. The three sets of pending confirmation status flags include pending confirmation of midway requests, pending confirmation of facility navigation, and pending confirmation of route adjustments. If any pending confirmation status flag is not empty, then based on the text message data, use a differentiated confirmation strategy to determine the user's intention to confirm or reject the message, execute the corresponding confirmation or rejection operation, generate the pending confirmation status processing result, clear the processed pending confirmation status flags, and record the processing result. Step S3: If all pending confirmation status flags are empty, then perform a three-level intent recognition strategy on the text message data to generate text intent recognition results and key entities; Step S4: Calculate the geographical distance between the user and the target point of interest based on the location coordinate data, and trigger arrival events in stages according to a preset distance threshold; determine the user's location change trend and dwell time based on continuous location coordinate data, and determine departure events by combining the location change trend and dwell time, and generate location event recognition results; Step S5: Based on the results of the pending status processing, text intent recognition, and location event recognition, update the user's browsing status, dialogue memory, and navigation memory, and generate and push multimodal response content.

[0006] Furthermore, the differentiated confirmation strategy in step S2 is as follows: when the length of the user message is less than or equal to a preset length threshold, keyword matching is used to determine the confirmation or rejection intent; when the length of the user message is greater than the preset length threshold, explicit action phrase matching is used to determine the confirmation or rejection intent.

[0007] Furthermore, the three-level intent recognition strategy in step S3 includes: matching the corresponding preset rule set according to the current user's browsing state for recognition, and triggering the corresponding intent with preset trigger keywords; if the rule set is not matched, then calling the language model and injecting browsing context information to assist semantic understanding, wherein the browsing context information comes from the currently maintained user browsing state, dialogue memory and navigation memory; if the language model call is abnormal, then using the preset core rule set for recognition.

[0008] Furthermore, step S4 also includes: when the location coordinate data is detected as automatically reported data without text and the current user is in navigation mode, it is determined as a location update event and a location event recognition result is generated, without executing the three-level intent recognition strategy for text message data.

[0009] Furthermore, step S4 also includes: simultaneously using the automatic location reporting channel and the user-initiated text input channel for location perception; when the perception results of the two channels are inconsistent, the result of the user-initiated text input channel shall prevail.

[0010] On the other hand, a scene-based intelligent perception and intent recognition system for scenic area tour guides is also provided to implement the aforementioned scene-based intelligent perception and intent recognition method. The system is characterized by comprising: a request access module for receiving user multimodal requests; a data separation module for parsing the user multimodal requests and separating text message data and location coordinate data; a global state storage module for storing three sets of pending confirmation status flags, user tour status, dialogue memory, and navigation memory; and a core scheduling module for prioritizing the reading of pending confirmation status flags from the global state storage module. If any flag is not empty, the system is scheduled to proceed to the pending confirmation status processing flow; if all flags are empty, the system is then scheduled according to data type. The system is scheduled to either a regular intent recognition process or a location event processing process. A data processing module, including an intent processing unit and a location processing unit, executes the pending confirmation state processing process or the regular intent recognition process, outputs the pending confirmation state processing result or text intent recognition result and key entities, and writes the result back to the global state storage module to update the state. The location processing unit processes the location coordinate data, generates a location event recognition result, and writes the result back to the global state storage module to update the state. A response output module generates and pushes multimodal response content based on the updated state data, text intent recognition result, and location event recognition result.

[0011] Furthermore, the global state storage module includes a pending confirmation state storage area, a tour state storage area, a dialogue memory storage area, and a navigation memory storage area; wherein, the pending confirmation state storage area is used to store three independent pending confirmation state flags.

[0012] Furthermore, the intent processing unit includes a rule matching subunit, a language model invocation subunit, and a degradation processing subunit connected in sequence to execute the three-level intent recognition strategy in sequence.

[0013] Furthermore, the location processing unit includes an automatic reporting processing subunit, an active input processing subunit, and an event fusion subunit; the automatic reporting processing subunit is used to process automatically reported location coordinate data without text; the active input processing subunit is used to process text message data containing location semantics actively sent by the user; the event fusion subunit is used to merge the output results of the two processing subunits, and when the output results of the two processing subunits are inconsistent, the result of the active input processing subunit shall prevail.

[0014] Furthermore, the response output module includes a text generation subunit, a voice generation subunit, and a push control subunit; the push control subunit is used to schedule the text generation subunit or the voice generation subunit to generate response content according to the event type, and control the push timing.

[0015] This invention provides a scene-based intelligent perception and intent recognition method and system for scenic area escort services. It ensures the continuity of the interaction process through a priority handling mechanism for pending confirmation states, preventing user confirmation responses from being misidentified as new requests. Employing a three-level intent recognition strategy that combines the tour context, it improves the accuracy of semantic understanding while ensuring service continuity through a degradation processing mechanism. Through event-driven location awareness and a dual-channel complementary perception mechanism, it achieves semantic conversion and conflict resolution of location data. Simultaneously, through unified global state management and multimodal response output, it significantly improves the accuracy and response speed of intent recognition in scenic area escort scenarios, reduces the overhead of large language model calls, and significantly improves the intelligent escort experience for tourists. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 This is a flowchart of the scene intelligent perception and intent recognition method provided in the embodiments of the present invention; Figure 2 This is a flowchart of the pending confirmation status differentiation processing provided in an embodiment of the present invention; Figure 3 This is a flowchart of the intent recognition strategy provided in an embodiment of the present invention; Figure 4 This is a block diagram of the scene intelligent perception and intent recognition system architecture provided in the embodiments of the present invention. Detailed Implementation

[0017] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0018] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that such solutions have been or necessarily used.

[0019] With the development of smart tourism and the upgrading of tourists' demand for personalized escort services, existing intelligent escort systems in scenic spots suffer from problems such as lack of context disambiguation in intent recognition, easy overwriting of pending confirmation states, inability to drive location awareness through events, and lack of anomaly degradation mechanisms. Therefore, it is crucial to develop a scenario-based intelligent intent recognition solution that integrates multimodal perception and state-priority processing.

[0020] To address this issue, this invention proposes a scene-based intelligent perception and intent recognition method and system for scenic area tour guides. It ensures interactive continuity through three sets of independent pending confirmation status flags and differentiated confirmation strategies. A three-level intent recognition strategy combining the tour context achieves accurate semantic understanding and anomaly degradation. Event-driven semantic conversion is completed through a dual-channel complementary position perception mechanism. Multimodal responses are generated through unified global state management, forming a complete closed-loop processing flow. This significantly improves intent recognition accuracy, reduces the overhead of large language model calls, and enhances system robustness and service stability.

[0021] The following is combined with Figures 1-4 This invention is described in detail.

[0022] like Figure 1 As shown in the figure, this invention provides a scene-based intelligent perception and intent recognition method for scenic area guides, characterized by the following steps: Step S1: Receive user multimodal requests and preprocess them to separate text message data and location coordinate data; Step S2: Obtain the three sets of pending confirmation status flags currently being maintained. The three sets of pending confirmation status flags include pending confirmation of midway requests, pending confirmation of facility navigation, and pending confirmation of route adjustments. If any pending confirmation status flag is not empty, then based on the text message data, use a differentiated confirmation strategy to determine the user's intention to confirm or reject the message, execute the corresponding confirmation or rejection operation, generate the pending confirmation status processing result, clear the processed pending confirmation status flags, and record the processing result. Step S3: If all pending confirmation status flags are empty, then perform a three-level intent recognition strategy on the text message data to generate text intent recognition results and key entities; Step S4: Calculate the geographical distance between the user and the target point of interest based on the location coordinate data, and trigger arrival events in stages according to a preset distance threshold; determine the user's location change trend and dwell time based on continuous location coordinate data, and determine departure events by combining the location change trend and dwell time, and generate location event recognition results; Step S5: Based on the results of the pending status processing, text intent recognition, and location event recognition, update the user's browsing status, dialogue memory, and navigation memory, and generate and push multimodal response content.

[0023] Multimodal requests refer to interactive requests sent by users to the scenic area escort service via smart terminals, containing various data types and covering four core request forms: text input, voice input, automatic location reporting, and proactive location sharing. The pending confirmation status flag is a temporary status flag maintained by the service after initiating a confirmation request to the user and before the user replies. It records the type, content, and contextual information of the pending confirmation, ensuring that the user's reply is correctly routed to the corresponding processing flow. Specifically, it includes three categories: pending confirmation of mid-trip requests, pending confirmation of facility navigation, and pending confirmation of route adjustments. Location event recognition results refer to the conversion of raw location coordinate data into semantically meaningful event information, including three categories: arrival events, departure events, and location update events. User tour status refers to the stage the user is in during their tour of the scenic area; the service dynamically adjusts its intent recognition rules and response strategies based on this status. Dialogue memory refers to the service's stored records of recent interactions between the user and the service, providing contextual support for intent recognition. Navigation memory refers to all information stored by the service related to the user's navigation, including the current route, visited attractions, remaining attractions, and target points of interest. Multimodal response content refers to interactive content that a service outputs to a user, including text, voice, maps, and other forms, to meet the information display needs of different tour scenarios.

[0024] Specifically, in step S1, user multimodal requests are received and preprocessed to separate text message data and location coordinate data. Multimodal requests are uniformly encapsulated in JSON format, containing a unique request identifier in 32-bit UUID format, a unique user identifier, a millisecond-level request timestamp, a request type enumeration value, and corresponding data fields. Voice request preprocessing calls an open-source speech recognition model to convert to text, with a sampling rate set to 16 kHz mono, outputting UTF-8 encoded plain text. A prompt to repeat is returned when the text conversion confidence score is below 0.7. Location data preprocessing converts the data to the WGS84 coordinate system, filtering out anomalies with instantaneous speeds exceeding 120 km / h or exceeding the scenic area's geographical boundaries. The scenic area boundaries are pre-stored as polygonal regions in GeoJSON format. Text data preprocessing removes special characters, redundant spaces, and emoticons, and converts the data to Simplified Chinese. Data separation is performed based on the request type field. Text message data contains plain text content, while location coordinate data contains double-precision floating-point numbers of latitude and longitude with 6 decimal places and a positioning accuracy field in meters.

[0025] Specifically, in step S2, three sets of pending confirmation status flags are obtained. If any pending confirmation status flag is not empty, a differentiated confirmation strategy is used based on the text message data to determine the user's intention to confirm or reject the message, and the corresponding confirmation or rejection operation is executed. A pending confirmation status processing result is generated, the processed pending confirmation status flags are cleared, and the processing result is recorded. All three sets of pending confirmation status flags are stored in the user session memory in JSON format, with a lifespan of 300 seconds. Upon timeout, they are automatically cleared and recorded as unconfirmed by the user. The pending mid-way request flag records the request type, associated point of interest name, original request message, and creation time; the pending facility navigation flag records the unique identifier of the target point of interest, distance, and confirmation status; and the pending route adjustment flag records the adjustment mode, original route, and adjusted route. Each time a user request is received, the three sets of flags are read first. If any flag contains non-empty content, it is determined that a pending confirmation status exists.

[0026] Specifically, in step S3, if all pending confirmation status flags are empty, a three-level intent recognition strategy is executed on the text message data to generate text intent recognition results and key entities. The text intent recognition results are output in JSON format, containing a unique intent identifier, intent name, a confidence floating-point number between 0 and 1, and an array of key entities. Key entities include four fields: entity type, entity value, start position in the text, and end position. Entity types cover core semantic categories of scenic spots such as attractions, facilities, time, and routes.

[0027] Specifically, in step S4, the geographical distance between the user and the target point of interest is calculated based on location coordinate data, and arrival events are triggered in stages according to preset distance thresholds. The user's location change trend and dwell time are determined based on continuous location coordinate data, and departure events are determined by combining these factors, generating location event recognition results. The geographical distance between the user and the target point of interest is calculated using the Haversine spherical distance formula, with the Earth's average radius taken as 6,371,000 meters, and latitude and longitude uniformly converted to radians for calculation. Three distance thresholds are set for arrival events: a precise arrival event is triggered and navigation ends when the distance is less than or equal to 10 meters; an automatic voice explanation event is triggered when the distance is between 10 and 30 meters; and an impending arrival prompt event is triggered when the distance is between 30 and 100 meters. The location change trend is determined by calculating the average speed of three consecutive location points at 1-second intervals. An average speed below 0.5 meters per second indicates a stationary state, while an average speed above 0.5 meters per second indicates a moving state. Dwell time is counted from the start of the stationary state, and a cumulative duration exceeding 30 seconds is considered a dwell state. A departure event is triggered when the user is more than 100 meters from the target point of interest and three consecutive location points are in a moving state. The location event identification results include event type, associated point of interest identifier, event timestamp, and additional data fields.

[0028] Specifically, in step S5, based on the results of the pending confirmation status processing, text intent recognition, and location event recognition, the user's tour status, dialogue memory, and navigation memory are updated, and multimodal response content is generated and pushed. The user's tour status is defined in five ways: Idle, Planning, Navigating, Touring, and Completed. State transitions are triggered by events. The Idle state supports route planning intent, the Planning state supports route adjustment intent, the Navigation state supports navigation adjustment and facility query intent, the Touring state supports attraction explanations and surrounding facility query intent, and the Completed state supports return navigation and evaluation intent. The Dialogue Memory stores the most recent 20 rounds of interaction records, using a first-in, first-out (FIFO) strategy to discard old data. Each record includes roles, content, and timestamp fields. The Navigation Memory is automatically archived to the user's history after navigation ends. In the multimodal response, text responses are generated using templates, with a length controlled between 50 and 200 characters; the voice response speed is set to 150 words per minute, and the volume is 80% of the default value; the map response generates a 720×1280 pixel PNG card. The push priority decreases sequentially according to user request response, arrival event, departure event, and location update event.

[0029] The scene intelligent perception intent recognition method provided in this embodiment of the invention ensures the continuity of the interaction process through a pending confirmation state priority processing mechanism. Combined with multimodal data fusion and event-driven location perception, it realizes accurate recognition and rapid response of intent in scenic scene, effectively improving the intelligent tour companion experience for tourists.

[0030] like Figure 2 As shown, preferably, the differentiated confirmation strategy in step S2 is as follows: when the length of the user message is less than or equal to a preset length threshold, keyword matching is used to determine the confirmation or rejection intent; when the length of the user message is greater than the preset length threshold, explicit action phrase matching is used to determine the confirmation or rejection intent.

[0031] Among them, the differentiated confirmation strategy refers to a processing mechanism that uses different matching rules to determine the confirmation or rejection intent based on the length characteristics of the user's reply message, in order to improve the accuracy and response speed of intent recognition in the pending confirmation state. The keyword matching method refers to a rapid identification method that determines the user's intent by matching a preset set of short confirmation or rejection keywords, suitable for scenarios with short user replies. The explicit action phrase matching method refers to a semantic recognition method that determines the user's intent by matching phrases containing explicit action directives, suitable for scenarios with longer user replies.

[0032] Specifically, a preset length threshold is set to 5 characters. When the user message length is less than or equal to 5 characters, keyword matching is used. The preset confirmation keyword set includes words like "yes," "okay," "can," "confirm," "correct," and "okay." The preset rejection keyword set includes words like "no," "don't need," "don't want," "cancel," "never mind," and "can't." Punctuation and spaces are ignored during matching; the message is considered to contain any of these keywords. When the user message length is greater than 5 characters, explicit action phrase matching is used. Preset confirmation action phrases include "I confirm going," "I need to adjust the route," and "Navigate me there." Preset rejection action phrases include "I'm not going," "No need to adjust the route," and "I'll go myself." Fuzzy matching is used during matching, allowing other modifiers before and after the phrase. After matching, a corresponding confirmation or rejection result is generated for subsequent processing of pending confirmation statuses.

[0033] For example, when a user is prompted with a confirmation request, "Should I navigate you to the visitor center restroom?", if the user replies "OK" (a message length of 2 characters, less than or equal to a preset threshold of 5 characters), the method matches the confirmation keyword set to "OK", determines it as a confirmation intent, and immediately initiates navigation. If the user replies "I don't want to go now, let's go see attraction A first" (a message length of 13 characters, greater than the preset threshold), the method matches the rejection action phrase "I don't want to go" (an explicit action phrase), determines it as a rejection intent, cancels the confirmation request, and simultaneously identifies new attraction query requests for further processing.

[0034] In a preferred embodiment of the present invention, a differentiated confirmation strategy based on message length ensures both a fast response speed for short user replies and improves the accuracy of intent recognition for complex and long replies, effectively reducing semantic misjudgments in the pending confirmation state and significantly enhancing the smoothness of the interaction process.

[0035] like Figure 3 As shown, preferably, the three-level intent recognition strategy in step S3 includes: matching the corresponding preset rule set according to the current user's browsing state for recognition, and triggering the corresponding intent with preset trigger keywords; if the rule set is not matched, then calling the language model and injecting browsing context information to assist semantic understanding, wherein the browsing context information comes from the currently maintained user browsing state, dialogue memory and navigation memory; if the language model call is abnormal, then using the preset core rule set for recognition.

[0036] Rule-based matching recognition refers to a method of rapid intent judgment based on a pre-defined intent-keyword mapping rule set. It does not require calling a language model, offering fast response and definitive results. The language model refers to a natural language processing model pre-trained using a mixture of large-scale general text and scenic area-specific text. It can understand text semantics, recognize user intent, and resolve ambiguities, making it a core component for handling complex semantic expressions in this embodiment. Context-enhanced model recognition refers to a method of semantic understanding that inputs the user's current text and the tour context information into the language model, capable of handling complex semantics and ambiguous expressions. Core rule degradation recognition refers to a fallback recognition method activated when the language model service is unavailable. It uses only the most core, high-frequency intent rule set for judgment, ensuring basic service capabilities.

[0037] Specifically, this embodiment preferably uses a large language model with 7B parameters, compressed and deployed using NT4 quantization, which can meet the concurrent request requirements during peak tourist seasons. The model inference timeout is set to 3 seconds; exceeding this timeout is considered a model service anomaly, automatically triggering the third-level core rule degradation recognition process. The model input uses a structured prompt word format, explicitly limiting the output to only include the intent name and confidence score, prohibiting the generation of irrelevant content, thereby improving inference speed and result stability.

[0038] Specifically, the three-level intent recognition strategy strictly follows the order of first-level rule matching recognition, second-level context-enhanced model recognition, and third-level core rule degradation recognition. If the previous level recognition is successful and the confidence level reaches the preset threshold, the result is returned directly without proceeding to the next level. The first-level rule matching recognition uses a combination of exact matching and fuzzy matching. The preset rule set includes 20 high-frequency intents in scenic area escort scenarios. Each intent corresponds to 5-10 core keywords and 3-5 extended keywords. After a successful match, the confidence level is uniformly set to 0.95, and the preset threshold is 0.9.

[0039] Specifically, the second-level context-enhanced model is triggered when the first-level matching fails. The input includes the user's current text, the last 5 rounds of dialogue, the current tour status, the list of visited attractions, and the current target point of interest. A lightweight large language model is called for inference, the model temperature is set to 0.1, the maximum generated length is 64 characters, and the output includes the intent name and confidence score. A confidence score of 0.85 or higher is considered a successful recognition.

[0040] Specifically, the third-level core rule degradation recognition is triggered when the language model call times out or returns an error. The core rule set only includes four basic intents: attraction query, navigation request, facility query, and route planning. Each intent corresponds to three core keywords. The exact matching method is used, and the confidence level is uniformly set to 0.8 after a successful match to ensure that basic services are not interrupted.

[0041] In a preferred embodiment of the present invention, a three-tiered intent recognition architecture is adopted, which balances recognition speed, accuracy and service robustness. In normal scenarios, the frequency of language model calls is significantly reduced, while basic service capabilities can still be provided in abnormal scenarios, effectively improving the practicality and stability of the method.

[0042] Preferably, step S4 further includes: when the location coordinate data is detected as automatically reported data without text and the current user is in navigation mode, it is determined as a location update event and a location event recognition result is generated, without executing the three-level intent recognition strategy for text message data.

[0043] In a further preferred embodiment, step S4 also includes: simultaneously using the automatic location reporting channel and the user-initiated text input channel for location perception; when the perception results of the two channels are inconsistent, the result of the user-initiated text input channel shall prevail.

[0044] The automatic location reporting channel refers to the sensing channel through which the user terminal automatically sends location coordinate data at fixed time intervals, continuously tracking changes in the user's location and triggering background events. The active location input channel refers to the sensing channel through which the user actively inputs location-related information via text or voice, expressing the user's explicit location intent. The dual-channel complementary sensing mechanism refers to a mechanism that simultaneously utilizes both automatic location reporting and active location input channels for location event recognition, improving sensing accuracy and semantic richness through the complementary advantages of both. The conflict resolution strategy refers to the processing rules that determine which channel's result should be prioritized when location information from two channels contradicts each other, used to resolve recognition errors caused by inconsistencies in multi-source data.

[0045] Specifically, the automatic location reporting channel has a reporting interval of 1 second and is only enabled when the user is in navigation or browsing mode. It is automatically disabled in idle and planning modes to reduce terminal power consumption. The active location input channel remains enabled in all states, receiving text content containing location information input by the user, such as "I am at A," "I have arrived at the visitor center," or "I am currently at entrance B." The location data from the two channels are processed independently, generating corresponding location event recognition results for each channel before being fused.

[0046] Specifically, when the location information from two channels points to the same point of interest and the distance difference is less than 50 meters, the information is considered consistent and directly merged to generate the final location event recognition result. When the location information from two channels points to different points of interest or the distance difference is greater than or equal to 50 meters, a conflict resolution strategy is triggered, prioritizing the result from the active location input channel, as active input more accurately reflects the user's true location and intent. The validity period of the active location input result is 300 seconds; after this time, it automatically reverts to the result from the automatic location reporting channel. For example, if the automatic location reporting shows the user is 50 meters away from A, but the user actively inputs "I am at B," then the user's actively input location will be used, updating the user's current location and triggering the automatic voice narration event on Baidi.

[0047] In a preferred embodiment of the present invention, a dual-channel complementary sensing mechanism of automatic location reporting and active location input, combined with a conflict resolution strategy that prioritizes user active input, effectively solves the problems of insufficient accuracy and semantic loss in complex scenic environments caused by single location reporting. This significantly improves the accuracy of location event recognition and the matching degree of user intent, and avoids interference with the tour experience caused by erroneous events.

[0048] like Figure 4 As shown, this embodiment of the invention also provides a scene intelligent perception and intent recognition system for scenic area tour guides, used to implement the above-mentioned scene intelligent perception and intent recognition method. It is characterized by comprising: a request access module for receiving user multimodal requests; a data separation module for parsing the user multimodal requests and separating text message data and location coordinate data; a global state storage module for storing three sets of pending confirmation state flags, user tour status, dialogue memory, and navigation memory; and a core scheduling module for prioritizing the reading of pending confirmation state flags from the global state storage module. If any flag is not empty, the system is scheduled to proceed to the pending confirmation state processing flow; if all flags are empty, the system proceeds according to the data... The system is configured to schedule the data to either a regular intent recognition process or a location event processing process. The data processing module includes an intent processing unit and a location processing unit. The intent processing unit executes the pending state processing process or the regular intent recognition process, outputs the pending state processing result or text intent recognition result and key entities, and writes the result back to the global state storage module to update the state. The location processing unit processes the location coordinate data, generates a location event recognition result, and writes the result back to the global state storage module to update the state. The response output module generates and pushes multimodal response content based on the updated state data, text intent recognition result, and location event recognition result.

[0049] The global state storage module includes a pending state storage area, a tour state storage area, a dialogue memory storage area, and a navigation memory storage area; wherein, the pending state storage area is used to store three independent pending state flags.

[0050] The scene-intelligent perception and intent recognition system is deployed on the scenic area's cloud server, communicating with user smart terminals to process interaction requests and provide service responses. The request access module is the front-end module responsible for receiving and validating multimodal requests from user terminals; it is the sole entry point for system-user interaction. The data separation module is a preprocessing module responsible for breaking down multimodal requests into text message data and location coordinate data, providing a data foundation for subsequent parallel processing. The global state storage module is the core module responsible for persistently storing all state information of the user session, forming the basis for state-priority processing and context awareness. The core scheduling module is the central module that schedules different processing flows based on global state information, determining the processing path and priority of requests. The intent processing module is the functional module responsible for performing text intent recognition, converting text content into semantic intent. The location processing module is the functional module responsible for performing location event perception, converting coordinate data into semantic events. The response output module is the back-end module responsible for generating and pushing multimodal response content, completing the information interaction between the system and the user.

[0051] Specifically, the system's modules communicate via RESTful interfaces, employing a microservice architecture for horizontal scaling to handle concurrent requests during peak tourist seasons. The request access module receives user requests via HTTP / 2 persistent connections, performing signature and format verification. Invalid requests are returned with error codes, while valid requests are forwarded to the data separation module. The data separation module splits the data based on the request type field: text message data is forwarded to the intent processing module, and location coordinate data is forwarded to the location processing module. Both types of data are simultaneously sent to the core scheduling module.

[0052] Specifically, the global state storage module uses Redis as the storage medium. Each user has a unique session key, and the session timeout is 1800 seconds, after which all state information is automatically cleared. Internally, the global state storage module is divided into four independent areas: a pending state storage area, a browsing state storage area, a dialogue memory storage area, and a navigation memory storage area. These areas store corresponding types of state data, and while the data in each area is isolated, it can be accessed uniformly by the core scheduling module. The pending state storage area specifically stores three sets of pending state flags, and its data structure and lifecycle are consistent with those in the method embodiment.

[0053] Specifically, each time the core scheduling module receives a new request, it first reads the data in the pending status storage area of ​​the global state storage module. If a non-empty pending status exists, the request is directly scheduled to the pending status processing branch of the intent processing module; otherwise, the intent processing module and the location processing module are scheduled to execute in parallel. After collecting the processing results from each module, the core scheduling module updates the global state storage module and forwards the results to the response output module.

[0054] Specifically, the internal implementation logic of the intent processing module and the location processing module is consistent with the method embodiment, respectively executing a three-level intent recognition strategy and a dual-channel complementary perception mechanism. The response output module generates corresponding multimodal response content based on the processing result type, pushes it to the user terminal through the long connection established by the request access module, and records the push status to the dialogue memory storage area after the push is completed.

[0055] Preferably, the intent processing unit includes a rule matching subunit, a language model invocation subunit, and a degradation processing subunit connected in sequence to execute the three-level intent recognition strategy in sequence.

[0056] The rule matching subunit is responsible for performing the first-level rule matching and recognition, quickly determining user intent based on a pre-defined intent-keyword mapping rule set, characterized by fast response speed and high result certainty. The language model invocation subunit is responsible for performing the second-level context-enhanced model recognition, improving the understanding of complex semantics by injecting contextual information. The degradation processing subunit is responsible for performing the third-level core rule degradation recognition, providing backup recognition capabilities when the language model service fails, ensuring the continuity of basic services.

[0057] Specifically, the three sub-units are executed sequentially in the following order: rule matching sub-unit first, language model calling sub-unit second, and degradation processing sub-unit last. If the previous sub-unit successfully identifies an item and its confidence level reaches a preset threshold, the result is returned directly without calling subsequent sub-units. The rule matching sub-unit has a built-in rule set of 20 high-frequency intents in the scenic area escort scenario. Each intent contains 5 to 10 core keywords and 3 to 5 extended keywords. It uses a combination of exact matching and fuzzy matching. After a successful match, the confidence level is uniformly set to 0.95, and the preset success threshold is 0.9.

[0058] Specifically, the language model invocation subunit is triggered when the rule matching subunit fails to recognize the data. It retrieves the last five rounds of dialogue records, the current tour status, the list of visited attractions, and the current target point of interest from the global state storage module, combining these with the user's current text to form structured input for invoking the language model. The subunit has a built-in 3-second timeout control mechanism; if no result is received from the model within this time, it is considered a service error, automatically triggering a fallback subunit. The model output undergoes format validation, extracting only the intent name and confidence score; a confidence score of 0.85 or higher is considered a successful recognition.

[0059] Specifically, the downgrade processing subunit only contains the core rule set for four basic intents: attraction search, navigation request, facility search, and route planning. Each intent corresponds to three core keywords, and a strict, precise matching method is used. After a successful match, the confidence level is uniformly set to 0.8. The recognition result of the downgrade processing subunit will be marked as downgraded mode, and the current service status will be indicated to the user in the response content, guiding the user to use simpler commands for interaction.

[0060] In a preferred embodiment of the present invention, hierarchical intent recognition is achieved by setting three types of functional sub-units in a hierarchical manner. This can achieve fast and low-cost recognition by relying on rule matching, understand complex semantics by using a language model, and run as a fallback mechanism when the model is abnormal, thus taking into account recognition timeliness, semantic understanding ability and service operation stability.

[0061] Preferably, the location processing unit includes an automatic reporting processing subunit, an active input processing subunit, and an event fusion subunit; the automatic reporting processing subunit is used to process automatically reported location coordinate data without text; the active input processing subunit is used to process text message data containing location semantics actively sent by the user; the event fusion subunit is used to merge the output results of the two processing subunits, and when the output results of the two processing subunits are inconsistent, the result of the active input processing subunit shall prevail.

[0062] The automatic reporting processing subunit specifically handles a set of location events automatically generated from location coordinate data, determining arrival, departure, and location update events based on location change trends and dwell time. The active input processing subunit specifically handles a set of semantic events based on user-inputted location information, parsing text and associated points of interest to generate events with definite location semantics. The event fusion subunit combines automatically reported location events and actively input parsed events, merging and solidifying the event data according to preset conflict resolution rules and a 50-meter distance threshold for dual-channel data.

[0063] Specifically, the automatic reporting processing subunit continuously acquires coordinate sequences at fixed time intervals, compares the distance between the current location and the target point of interest in real time using spherical distance calculation, and determines event types such as "imminent arrival" and "precise arrival" according to tiered distance thresholds. Simultaneously, based on the movement speed and dwell time of consecutive points, it distinguishes between moving and stationary states, thereby identifying the user's behavior of leaving visited locations. The active input processing subunit performs entity parsing on place names, attraction names, and facility names contained in the user's text, matches them against the scenic area's preset point of interest database, maps the corresponding location coordinates and semantic information, and directly generates high-priority location events. The event fusion subunit simultaneously receives two event results. When the spatial deviation between the two location information is less than a set threshold, it directly merges and outputs them uniformly; when the spatial deviation exceeds the threshold or points to different points of interest, it follows the rule of higher priority for active input to complete conflict resolution, ultimately outputting a uniquely determined location event identification result for subsequent synchronous updates of the tour status and memory data.

[0064] In a preferred embodiment of the present invention, location data is processed collaboratively by three sub-units: automatic reporting, active input, and event fusion. This enables standardized parsing and intelligent fusion of dual-channel location events, effectively resolving conflicts between dual-channel location data, reducing the probability of misjudgment in complex scenic environments, and further ensuring the uniqueness, accuracy, and logic of location event identification.

[0065] Preferably, the response output module includes a text generation subunit, a voice generation subunit, and a push control subunit; the push control subunit is used to schedule the text generation subunit or the voice generation subunit to generate response content according to the event type, and control the push timing.

[0066] The tour status storage area is a data storage area used to record and maintain the user's current tour progress status. It is used to intuitively represent the user's tour stage and provide state constraints for intent recognition and location event determination. The dialogue memory storage area is a storage area used to temporarily retain the content of the user's dialogue during the interaction process. It is used to supplement contextual information, assist in understanding ambiguous statements, and associate historical interaction intents. The navigation memory storage area is a data storage area specifically for storing the tour links related to the user's navigation process. It is used to record the travel trajectory, the order of attractions visited, and navigation task information. The pending confirmation status storage area is a dedicated storage area for temporarily caching three types of pending confirmation items. It is used to retain interaction tasks that have not been confirmed, ensuring that the user's second response accurately corresponds to the previous confirmation request.

[0067] Specifically, the four storage areas are independent and isolated from each other, and are updated in a unified time sequence. All use memory caching for temporary storage, taking effect synchronously with the user session lifecycle. The retention period for a single user session is uniformly set to 1800 seconds; if no new interaction occurs within this timeout, all cached data is automatically cleared. The tour status storage area records the user's current tour status in real time, retaining only one valid status. Statuses switch unidirectionally based on location events and user intent, and cannot be changed across states. The dialogue memory storage area uses a first-in, first-out (FIFO) discard rule, retaining a maximum of the most recent 20 rounds of interaction text. It only stores plain text interaction content, excluding redundant data such as location and images, reducing cache usage. The navigation memory storage area records visited attractions, unvisited attractions, the current navigation point, and route offset information in an orderly manner. It is automatically archived after each navigation task and no longer participates in real-time logical operations. The pending confirmation status storage area caches a maximum of three pending confirmation items simultaneously. Similar pending confirmation items cannot be stacked. The maximum retention period for a single pending confirmation status is 300 seconds; it automatically expires and is cleared after this timeout.

[0068] In a preferred embodiment of the present invention, different types of tour data are classified and cached by dividing the data into four independent storage areas. This enables refined management of user status, interaction records, navigation information, and pending confirmation items, avoids interference between different types of data, ensures simple and efficient data retrieval, and further improves the consistency of interaction and the accuracy of logic.

[0069] In summary, the present invention provides a scene-based intelligent perception and intent recognition method and system for scenic area tour companions. It achieves standardized parsing of tour data through a multimodal data preprocessing process, avoids interaction logic breaks by prioritizing pending confirmation states, and employs a three-level hierarchical intent recognition architecture that balances recognition speed, accuracy, and service fault tolerance. It also optimizes the accuracy of location event determination in complex scenic environment environments by combining dual-channel location awareness and conflict resolution strategies. Simultaneously, it provides refined management of user tour status, dialogue memory, navigation memory, and temporary pending confirmation information by dividing the storage area into independent storage regions. Intent recognition, location awareness, status updates, and multimodal response generation are completed through the collaborative linkage of various functional units. This invention effectively solves the technical defects of traditional scenic area intelligent tour companion interactions, such as insufficient context utilization, easy overwriting of confirmation messages, low semantic recognition accuracy, limited location awareness, and weak adaptability to abnormal scenarios. It reduces the frequency of large language model calls and computational consumption, improves the continuity of human-computer interaction, the accuracy of intent recognition, and the reliability of location event determination, making the scenic area intelligent tour companion service more aligned with tourists' tour habits and enhancing the stability, intelligence, and environmental adaptability of the tour companion service.

[0070] The above description is merely a preferred embodiment of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A scene-based intelligent perception and intent recognition method for tour guides in scenic areas, characterized in that, Includes the following steps: Step S1: Receive user multimodal requests and preprocess them to separate text message data and location coordinate data; Step S2: Obtain the three sets of pending confirmation status flags currently being maintained. The three sets of pending confirmation status flags include pending confirmation of midway requests, pending confirmation of facility navigation, and pending confirmation of route adjustments. If any pending confirmation status flag is not empty, then based on the text message data, use a differentiated confirmation strategy to determine the user's intention to confirm or reject the message, execute the corresponding confirmation or rejection operation, generate the pending confirmation status processing result, clear the processed pending confirmation status flags, and record the processing result. Step S3: If all pending confirmation status flags are empty, then perform a three-level intent recognition strategy on the text message data to generate text intent recognition results and key entities; Step S4: Calculate the geographical distance between the user and the target point of interest based on the location coordinate data, and trigger arrival events in stages according to preset distance thresholds; Based on continuous location coordinate data, the user's location change trend and dwell time are determined, and the departure event is determined by combining the location change trend and dwell time to generate location event recognition results; Step S5: Based on the results of the pending status processing, text intent recognition, and location event recognition, update the user's browsing status, dialogue memory, and navigation memory, and generate and push multimodal response content.

2. The scene intelligent perception and intent recognition method according to claim 1, characterized in that, The differentiated confirmation strategy described in step S2 is as follows: when the length of the user message is less than or equal to a preset length threshold, keyword matching is used to determine the confirmation or rejection intent; when the length of the user message is greater than the preset length threshold, explicit action phrase matching is used to determine the confirmation or rejection intent.

3. The scene intelligent perception and intent recognition method according to claim 1, characterized in that, The three-level intent recognition strategy in step S3 includes: matching the corresponding preset rule set according to the current user's browsing state for recognition, and triggering the corresponding intent with preset trigger keywords; if the rule set is not matched, calling the language model and injecting browsing context information to assist semantic understanding, wherein the browsing context information comes from the currently maintained user browsing state, dialogue memory and navigation memory; if the language model call is abnormal, using the preset core rule set for recognition.

4. The scene intelligent perception and intent recognition method according to claim 1, characterized in that, Step S4 further includes: when the location coordinate data is detected as automatically reported data without text and the current user is in navigation mode, it is determined as a location update event and a location event recognition result is generated, and the three-level intent recognition strategy for text message data is not executed.

5. The scene intelligent perception and intent recognition method according to claim 1, characterized in that, Step S4 also includes: simultaneously using the automatic location reporting channel and the user-initiated text input channel for location perception; when the perception results of the two channels are inconsistent, the result of the user-initiated text input channel shall prevail.

6. A scene intelligent perception and intent recognition system for scenic area tour guides, used to implement the scene intelligent perception and intent recognition method according to any one of claims 1-5, characterized in that, include: The access request module is used to receive multimodal user requests; The data separation module is used to parse the user's multimodal request and separate the text message data and location coordinate data. The global state storage module is used to store three sets of pending confirmation status flags, user browsing status, dialogue memory, and navigation memory; The core scheduling module is used to prioritize reading the pending status flags in the global state storage module. If any flag is not empty, it will schedule to the pending status processing flow. If all flags are empty, then the process is directed to the regular intent recognition flow or the location event handling flow, depending on the data type. The data processing module includes an intent processing unit and a location processing unit. The intent processing unit is used to execute the pending confirmation state processing flow or the regular intent recognition flow, output the pending confirmation state processing result or the text intent recognition result and key entities, and write the result back to the global state storage module to update the state. The location processing unit is used to process the location coordinate data, generate location event recognition results, and write the results back to the global state storage module to update the state. The response output module is used to generate and push multimodal response content based on the updated status data, text intent recognition results, and location event recognition results.

7. The scene intelligent perception and intent recognition system according to claim 6, characterized in that, The global state storage module includes a pending state storage area, a tour state storage area, a dialogue memory storage area, and a navigation memory storage area; wherein, the pending state storage area is used to store three independent pending state flags.

8. The scene intelligent perception and intent recognition system according to claim 6, characterized in that, The intent processing unit includes a rule matching subunit, a language model invocation subunit, and a degradation processing subunit connected in sequence, to execute the three-level intent recognition strategy in sequence.

9. The scene intelligent perception and intent recognition system according to claim 6, characterized in that, The location processing unit includes an automatic reporting processing subunit, an active input processing subunit, and an event fusion subunit. The automatic reporting processing subunit is used to process automatically reported location coordinate data without text. The active input processing subunit is used to process text message data containing location semantics sent by the user. The event fusion subunit is used to merge the output results of the two processing subunits. When the output results of the two processing subunits are inconsistent, the result of the active input processing subunit shall prevail.

10. The scene intelligent perception and intent recognition system according to claim 6, characterized in that, The response output module includes a text generation subunit, a voice generation subunit, and a push control subunit; the push control subunit is used to schedule the text generation subunit or the voice generation subunit to generate response content according to the event type, and control the push timing.

Citation Information

Patent Citations

  • Multi-modal man-machine interaction system and method based on life support

    CN117591636A

  • Multi-mode AI intelligent agent system for personalized service of tourist attraction

    CN120910344A