Self-service voice guidance system and method

By monitoring the location and behavior of visitors in real time through the self-service voice guide system, generating prompts, and solving the problem of unbalanced visitor flow in the exhibition hall, the system achieves dynamic balance and efficient voice guide service within the exhibition hall, thereby enhancing the visitor experience.

CN122093456BActive Publication Date: 2026-07-21SOYO TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-24
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

There is an imbalance in visitor flow distribution in the exhibition hall. Popular exhibit areas are overcrowded while less popular exhibit areas are deserted. The existing exhibition hall lacks the ability to proactively sense and intelligently control crowds, making it difficult to accurately guide and manage dense crowds, which affects the smoothness of the visit and the satisfaction of the experience.

Method used

The system employs a self-service voice guide system, which uses voice guide terminals, positioning base stations, and edge servers to monitor the location and behavior of visitors in real time, generate prompts, guide visitors to avoid peak hours, optimize crowd distribution, and achieve a dynamic balance of visitors within the exhibition hall.

Benefits of technology

It effectively alleviates congestion at exhibitions, improves visitor flow and experience satisfaction, enhances the exhibition hall's carrying capacity and space utilization, and ensures the timeliness and efficiency of the audio guide service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093456B_ABST
    Figure CN122093456B_ABST
Patent Text Reader

Abstract

The application discloses a self-service voice guide system and method, the system comprises a voice guide terminal, a positioning base station and an edge server; the positioning base station is used for detecting the event that the voice guide terminal enters the electronic fence area of the first exhibit; the edge server is used for responding to the first introduction request of the first voice guide terminal in the first quantity, outputting the first reply according to the first introduction request; controlling the first voice guide terminal in the second quantity to output the first reply and / or the preset introduction information of the first exhibit; and the edge server is also used for determining the first exhibition state of the visiting user according to the position information and the entering time; generating prompt information according to the first exhibition state, and controlling at least one second voice guide terminal to output the prompt information. The application can realize the accurate guidance of the exhibition hall passenger flow.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent guidance technology for exhibition halls, and in particular to a self-service voice guidance system and method. Background Technology

[0002] With the continuous upgrading of cultural consumption demands, more and more people are visiting various exhibition halls. Exhibition halls need to rely on high-quality exhibition content and efficient service experiences to continuously attract visitors. However, there is a significant imbalance in visitor distribution in current exhibition hall operations. Popular exhibit areas are often overcrowded with excessive crowds and poor mobility, while less popular exhibit areas are deserted and rarely visited. Existing exhibition halls only provide passive guided tours, lacking proactive perception and intelligent control capabilities, making it difficult to accurately guide and manage dense crowds. Summary of the Invention

[0003] This application provides a self-service voice tour guide system and method to improve the user's tour experience and satisfaction.

[0004] In a first aspect, embodiments of this application provide a self-service voice navigation system, including a voice navigation terminal, a positioning base station, and an edge server;

[0005] The positioning base station is used to detect events where the voice guide terminal enters the electronic fence area of ​​the first exhibit; when the event is detected, the corresponding event information is transmitted to the edge server, and the event information includes the location information and entry time of the voice guide terminal;

[0006] The voice guide terminal is used to receive a first introduction request from a visitor regarding the first exhibit based on voiceprint recognition, and to transmit the first introduction request to the edge server.

[0007] The edge server is configured to respond to a first introduction request from a first number of first voice-guided terminals, output a first response based on the first introduction request, wherein the first voice-guided terminals are those whose entry time falls within a preset time window; control a second number of first voice-guided terminals to output the first response and / or preset introduction information for the first exhibit, wherein the second number is greater than or equal to the first number; and...

[0008] The edge server is also used to determine the first exhibition status of the visitor based on the location information and the entry time. The first exhibition status is used to characterize the real-time behavior of the visitor during the exhibition. Based on the first exhibition status, the server generates prompt information and controls at least one second voice guide terminal to output the prompt information. The prompt information is used to guide the visitor to visit during off-peak hours. The second voice guide terminal is a voice guide terminal associated with the off-peak visit guidance target within the electronic fence area.

[0009] Wherein, the first exhibition status includes distribution information and dwell time. After determining the first exhibition status of the visitor, the edge server is used to generate prompt information based on the first exhibition status and control at least one second voice guide terminal to output the prompt information, specifically for:

[0010] If the distribution information is found to not meet the density distribution constraints, the spatial pattern of the population distribution is determined based on the location information.

[0011] Based on the spatial form, multiple target users are identified, and the target users are visitors who do not have visual obstruction between them and the first exhibit.

[0012] Determine the viewing time of the multiple target users;

[0013] Based on the exhibition viewing time and / or the dwell time, at least one second voice guide terminal is determined;

[0014] Determine the degree of association between at least one second exhibit and the first exhibit, wherein the second exhibit and the first exhibit are spatially related;

[0015] Determine the second participation status of the visiting users within the electronic fence area of ​​each second exhibit;

[0016] Based on the degree of association, the second exhibition status, and the distribution information, a first prompt is generated, and the at least one second voice guide terminal is controlled to output the first prompt.

[0017] The first exhibition status further includes first flow information. When the edge server determines at least one second voice guide terminal based on the dwell time, it is specifically used for:

[0018] If the dwell time is detected to be greater than a first preset time or less than a second preset time, the corresponding visitor's voice guide terminal is identified as the second voice guide terminal.

[0019] The presence of queuing people is determined based on the dwell time, the spatial pattern, and the first flow information.

[0020] Upon detecting the presence of the queuing crowd, the queuing time is determined based on the exhibition viewing time, the third number of target users, and the location information.

[0021] If the queuing time is detected to be greater than the third preset time, the corresponding visitor's voice guide terminal is identified as the second voice guide terminal.

[0022] Wherein, after determining the existence of the queuing crowd, the edge server, when determining the queuing time based on the viewing time, the third number of target users, and the location information, is specifically used for:

[0023] The replacement time for the target user is determined based on the exhibition viewing time and the third quantity.

[0024] The fourth number of visitors in the queue is determined based on the location information.

[0025] Determine the second flow information of the queue;

[0026] A reference queuing time is determined based on the replacement time and the second quantity;

[0027] The travel time of the visiting user is determined based on the second flow information;

[0028] The queuing time is determined based on the reference queuing time and the movement time.

[0029] Specifically, when the edge server determines whether a queuing crowd exists based on the dwell time, the spatial pattern, and the first flow information, it is used for:

[0030] Based on the spatial shape and the distance between the visitor and the first exhibit, the electronic fence area is divided into multiple sub-areas;

[0031] Based on the initial flow information and dwell time of the visiting users described in each sub-region, determine the movement characteristics;

[0032] The presence of a queuing crowd is determined based on the movement characteristics.

[0033] Specifically, when the edge server determines at least one second audio guide terminal based on the exhibition viewing time, it is used for:

[0034] The dwell time of multiple third-party voice guide terminals after completing the guided tour related to the first exhibit is obtained, wherein the third-party voice guide terminal is the voice guide terminal of the target user;

[0035] If the dwell time is greater than the fourth preset time or the viewing time is greater than the fifth preset time, then the voice guide terminal corresponding to the target user will be determined as the second voice guide terminal.

[0036] The edge server is further configured to respond to a second introduction request from multiple voice guide terminals corresponding to the target users, output a second response based on the second introduction request, and control the multiple voice guide terminals corresponding to the target users to output the second response.

[0037] The system further includes a central server, and the edge servers establish a communication connection with the central server. When the edge servers respond to first introduction requests from multiple first voice guide terminals and output a first response based on the first introduction request, they are specifically used for:

[0038] Determine the request type of the first introductory request;

[0039] Match the corresponding target device according to the request type and trigger the execution of the corresponding target task. The target device includes the edge server and the central server. The target task is used to determine the first response that matches the first introduction request.

[0040] Wherein, after determining the request type of the first introduction request, the edge server, when matching the corresponding target device according to the request type and triggering the execution of the corresponding target task, is specifically used for:

[0041] If the request type is the first type, then the interaction information between the visiting user and the first voice guide terminal is obtained, as well as the identity feature information of the visiting user is obtained. The first type is used to indicate a deep information introduction request with multiple nested intentions. The interaction information, the identity feature information and the first introduction request are transmitted to the central server. After the central server determines the first response, the first response transmitted by the central server is received.

[0042] If the request type is the second type, then the first introduction intent of the first introduction request is determined, and the second type is used to indicate the basic attribute information introduction request of a single intent; the first response is determined according to the first introduction intent.

[0043] The central server is used to receive the interaction information, the identity feature information, and the first introduction request transmitted by the edge server.

[0044] Based on the interaction information, the identity feature information, and the first introduction request, an intent is identified to obtain a second introduction intent;

[0045] Generate a reference response based on the second description of the intended meaning;

[0046] Based on the detection of sensitive information, the reference response is corrected to obtain the first response;

[0047] The first response is transmitted to the edge server.

[0048] Secondly, embodiments of this application provide a self-service voice guidance method, applied to an edge server of a self-service voice guidance system, wherein the self-service voice guidance system further includes a voice guidance terminal and a positioning base station, and the method includes:

[0049] Obtain event information detected by the positioning base station, the event information including the entry time of the electronic fence area of ​​the first exhibit and the location information of the voice guide terminal;

[0050] The system acquires a first number of first introduction requests for the first exhibit collected by the first voice guide terminal, wherein the first voice guide terminal is a voice guide terminal whose entry time is within a preset time window.

[0051] The first response is requested based on the first description.

[0052] The second number of the first voice guide terminals outputs the first response and / or the preset introduction information of the first exhibit, wherein the second number is greater than or equal to the first number;

[0053] Based on the location information and the entry time, the first exhibition status of the visiting user is determined, and the first exhibition status is used to characterize the real-time behavior of the visiting user during the exhibition process.

[0054] Based on the first exhibition status, a prompt message is generated, and at least one second voice guide terminal is controlled to output the prompt message. The prompt message is used to guide the visitors to visit during off-peak hours. The second voice guide terminal is a voice guide terminal associated with the off-peak visit guidance target within the electronic fence area.

[0055] Thirdly, embodiments of this application provide a self-service voice navigation device, including:

[0056] The first acquisition unit is used to acquire event information detected by the positioning base station, the event information including the entry time of the electronic fence area of ​​the first exhibit and the location information of the voice guide terminal;

[0057] The second acquisition unit is used to acquire a first number of first introduction requests for the first exhibit collected by the first voice guide terminal, wherein the first voice guide terminal is a voice guide terminal whose entry time is within a preset time window.

[0058] The output unit is used to output a first response based on the first introduction request;

[0059] A first control unit is configured to control a second number of the first voice guide terminals to output the first response and / or preset introduction information of the first exhibit, wherein the second number is greater than or equal to the first number;

[0060] The determining unit is used to determine the first exhibition status of the visiting user based on the location information and the entry time. The first exhibition status is used to characterize the real-time behavior of the visiting user during the exhibition process.

[0061] The second control unit is used to generate prompt information based on the first exhibition status and control at least one second voice guide terminal to output the prompt information. The prompt information is used to guide the visitors to visit during off-peak hours. The second voice guide terminal is a voice guide terminal associated with the off-peak visit guidance target within the electronic fence area.

[0062] Fourthly, embodiments of this application provide an electronic device, including a memory, a processor, and executable program code stored in the memory and executable on the processor, wherein the processor executes the executable program code and performs the steps described in the first aspect.

[0063] Fifthly, embodiments of this application provide a computer-readable storage medium storing executable program code, the executable program code including execution instructions for performing the steps described in the first aspect.

[0064] Sixthly, embodiments of this application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of embodiments of this application. The computer program product may be a software installation package.

[0065] As can be seen, in this embodiment, the parallel response of the voice guide terminal through time-sliding windows can effectively avoid problems such as response timeouts and service interruptions caused by business congestion, reduce peak business pressure, and improve the response time and operational efficiency of the exhibition hall voice guide service; and based on the user's real-time exhibition status, targeted output of prompt information assists the exhibition hall in achieving refined management of the exhibition order, effectively alleviating exhibition congestion, achieving dynamic balance of visitor distribution within the exhibition hall, improving the user's visit smoothness and experience satisfaction, and at the same time improving the exhibition hall's carrying capacity and space utilization. Attached Figure Description

[0066] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0067] Figure 1 This is a system architecture diagram of a self-service voice guide system provided in an embodiment of this application;

[0068] Figure 2 This is a system architecture diagram of another self-service voice guide system provided in the embodiments of this application;

[0069] Figure 3 This is a schematic diagram illustrating the division of an electronic fence area for multiple exhibits in an exhibition hall, provided in an embodiment of this application.

[0070] Figure 4 This is a schematic diagram of the distribution of people in an exhibition hall provided in an embodiment of this application;

[0071] Figure 5 This is a flowchart illustrating a self-service voice navigation method provided in an embodiment of this application;

[0072] Figure 6 This is a block diagram of the functional units of a self-service voice guide device provided in an embodiment of this application;

[0073] Figure 7 This is a block diagram of the functional units of another self-service voice guide device provided in the embodiments of this application;

[0074] Figure 8 This is a schematic diagram of the structure of an electronic device proposed in an embodiment of this application. Detailed Implementation

[0075] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0076] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0077] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0078] With the continuous upgrading of cultural consumption demands, more and more people are visiting various exhibition halls. Exhibition halls need to rely on high-quality exhibition content and efficient service experiences to continuously attract visitors. However, there is a significant imbalance in visitor distribution in current exhibition hall operations. Popular exhibit areas are often overcrowded, with excessive crowds and poor mobility, while less popular exhibit areas are deserted. Exhibition halls struggle to accurately manage and guide dense crowds, and in crowded environments, visitors often find it difficult to immerse themselves in appreciating the details of exhibits and understanding their cultural significance and historical value, severely impacting their viewing experience.

[0079] To address the aforementioned issues, this application provides a self-service voice navigation system and method. The embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0080] Please see Figure 1 , Figure 1 This is a system architecture diagram of a self-service voice guide system provided in an embodiment of this application. Figure 1 As shown, the self-service voice guide system 100 includes a voice guide terminal 101, a positioning base station 102, and an edge server 103. A communication connection is established between the voice guide terminal 101 and the positioning base station 102, and between the positioning base station 102 and the edge server 103, as well as between the voice guide terminal 101 and the edge server 103.

[0081] The voice-guided terminal 101 first completes the device initialization configuration, pre-installs the display driver, audio playback and microphone acquisition modules, and completes the identity binding with the edge server 103. The voice-guided terminal is referred to as the terminal.

[0082] The voice-guided terminal 101 includes a display screen, headphones, and a microphone. The display screen visually displays exhibit-related guide information, the terminal's operating interface, the path guidance interface, fence area prompts, and equipment operating status, providing users with an intuitive means of operation and information acquisition. The headphones are the output component for the voice-guided content, used to play audio explanations of exhibits, while also effectively isolating environmental noise and enhancing the user's auditory experience. The microphone has the function of picking up user voice information, collecting user voice commands, voice inquiries, and other voice content for voice interaction, voice control, and other functions.

[0083] During the system initialization phase, positioning base stations 102 are deployed according to the exhibition hall zones to ensure signal coverage of the electronic fence area of ​​each exhibit. Communication tests are also conducted with the voice guide terminal 101, with a communication delay of less than or equal to 100ms, to ensure the real-time performance and stability of data transmission between the positioning base station 102 and the voice guide terminal 101.

[0084] The exhibition hall can be divided into zones based on the type of exhibits, such as bronze ware exhibition area and calligraphy and painting exhibition area; it can also be divided based on spatial structure, such as large exhibition hall, corridor, corner area, etc. Different spaces have different wall obstruction and signal reflection characteristics. After zoning, the base station deployment plan can be adjusted accordingly. In corner areas with many wall obstructions and complex signal reflection, the base station deployment can be increased. In open and unobstructed large exhibition halls and corridors, a distributed base station layout can be adopted to ensure uniform coverage of positioning signals in each spatial area.

[0085] The edge server 103 is an edge server cluster. Its service area is divided according to the principle that a single edge server cluster covers multiple exhibition areas; for example, a single edge server cluster covers 3 to 5 exhibition areas. A bidirectional communication link is established between the edge server 103 and the positioning base station 102, and a load balancing strategy is configured to ensure that the maximum number of voice-guided terminals a single cluster can support during peak hours is greater than or equal to 500, thereby effectively handling the business processing pressure under peak conditions.

[0086] Specifically, the voice guide terminal 101 is a handheld interactive device used by users to interact with the self-service voice guide system, and it is also the location object detected by the positioning base station 102. After the positioning base station 102 is deployed, it collects the location data of the voice guide terminal 101 in real time according to a preset fixed period, for example, once every 200ms. Then, the positioning base station 102 compares the real-time location coordinates of the voice guide terminal 101 with the preset coordinate range of the electronic fence area of ​​each exhibit one by one. The error of the coordinate comparison is controlled within a preset threshold, for example, the preset threshold is 1m. If the comparison determines that the voice guide terminal 101 has entered the electronic fence area of ​​the first exhibit, the positioning base station 102 uploads the location-triggered event information to the edge server 103 of the exhibit area to which the fence belongs, completing the real-time reporting of the fence entry information. The event information includes the terminal identifier, the actual entry time into the fence, and the corresponding exhibit fence number.

[0087] The voice guide terminal 101 uses voiceprint recognition technology to identify the first introduction request made by a visitor for the current first exhibit; and synchronously transmits the identified first introduction request to the edge server 103, waiting for the server's response. Specifically, during voiceprint recognition, invalid voices are filtered out to confirm that the introduction request is directed to the current first exhibit, avoiding erroneous requests for different exhibits.

[0088] When the edge server 103 receives the terminal entry event reported by the positioning base station 102, it starts the time sliding window mechanism, presets a fixed sliding window duration, aggregates the terminal entry events of the first exhibit in time, and summarizes all terminals that enter the exhibit fence within the sliding window period into parallel trigger events. The terminals that arrive at the electronic fence area of ​​the first exhibit in the same time period are named the first voice guide terminal. The edge server responds to the first introduction requests of multiple first voice guide terminals at the same time.

[0089] The edge server collects all first introduction requests uploaded by the first voice guide terminals and outputs the corresponding first response for all introduction requests during this time period.

[0090] As can be seen, in this embodiment, the cost of repetitive interaction and processing under multi-terminal triggering is greatly reduced, the overall response latency of the system is effectively reduced, the voice guide experience is guaranteed when multiple users visit the same exhibit at the same time, and the processing efficiency of edge servers for concurrent events triggered by multiple terminals is improved.

[0091] Please refer to Figure 2 , Figure 2 This is a system architecture diagram of another self-service voice guide system provided in an embodiment of this application. For example... Figure 2As shown, the self-service voice guide system also includes a central server 104. The central server 104 establishes a two-way communication link with the edge server 103, and the unique identifier and device parameters of the terminal are uploaded to the central server 104.

[0092] In one possible embodiment, an electronic fence area for each exhibit is defined using a positioning base station. Specifically, first, the display format of each exhibit within the exhibition hall area is obtained, such as display case type, stand type, wall-mounted type, etc.; then, the shape of the corresponding electronic fence for the exhibit is determined based on the display format, such as a rectangle for a display case type and a circle for a stand type. For an example, please refer to [link to example]. Figure 3 , Figure 3 This is a schematic diagram illustrating the division of an electronic fence area for multiple exhibits in an exhibition hall, as provided in an embodiment of this application. Figure 3 As shown, for the first, second, and third exhibits in the diagram, their display formats are first determined, and the corresponding electronic fence shapes are then identified. The first exhibit is fitted with a circular electronic fence, while the second and third exhibits are fitted with rectangular electronic fences. The initial dimensions of the electronic fences are determined based on the exhibition hall space constraints, including the width of the exhibition hall aisles, the safety distance between exhibits, and the layout of surrounding facilities. The initial dimensions are adjusted according to the exhibits' popularity to obtain target dimensions. If the popularity is high, the fence area is appropriately expanded to ensure viewing space; if the popularity is low, it is reduced as needed to make reasonable use of the exhibition hall space. For example, the third exhibit's popularity is lower than that of the first and second exhibits, so its corresponding electronic fence is smaller than those of the first and second exhibits. Combining the determined electronic fence shapes with the adjusted target dimensions, electronic fences for each exhibit within the exhibition hall are generated.

[0093] In one possible embodiment, when the edge server responds to a first introduction request from multiple first voice navigation terminals and outputs a first response based on the first introduction request, it is specifically used to: determine the request type of the first introduction request; match the corresponding target device according to the request type and trigger the execution of a corresponding target task, wherein the target device includes the edge server and the central server, and the target task is used to determine the first response that matches the first introduction request.

[0094] Among them, the edge server 103 extracts the request in each first introduction request, determines the characteristics such as the problem direction, and filters out invalid and redundant information, such as users' colloquial expressions and irrelevant interjections. For example, if the first introduction request is "Can you tell me what this exhibit is for?", the feature extraction outputs the introduction request as "Get the basic functional information of the exhibit".

[0095] Based on the extracted features, the complexity of processing the request is determined, and the request type is output. The request type includes a first type for indicating a deep information description request with multiple nested intents, and a second type for indicating a basic attribute information description request with a single intent.

[0096] For example, if the first introduction request is "recommend relevant exhibits based on my profession", this request is nested with multiple interrelated intents, such as extracting user profession characteristics, matching the related attributes of exhibits, filtering suitable exhibits and completing targeted recommendation introductions. The processing not only requires decomposing the multi-layered nested intents, but also integrating user identity feature information, multi-dimensional related information of exhibits, etc. The overall processing complexity is high, which meets the characteristics of deep information introduction requirements with multiple nested intents, and is the first type.

[0097] For example, if the first introduction request is "to introduce the background of this exhibit", the request only points to the single intent of the exhibit's basic attributes. There is no need to break down the intent during processing. The response can be completed by simply retrieving the corresponding data. The overall processing complexity is low, which meets the characteristics of the basic attribute information introduction request with a single intent. Therefore, it is the second type.

[0098] Specifically, according to the preset matching rules between the request type and the target device, a corresponding processing device is assigned to the current introduction request, and a target task is sent to the processing device. The target device executes the task and finally generates a first response that highly matches the first introduction request.

[0099] In one possible embodiment, after determining the request type of the first introduction request, the edge server, when matching the corresponding target device according to the request type and triggering the execution of the corresponding target task, specifically performs the following steps: if the request type is a first type, it acquires the interaction information between the visiting user and the first voice guide terminal, and acquires the identity feature information of the visiting user, wherein the first type is used to indicate a deep information introduction request with multiple nested intents; it transmits the interaction information, the identity feature information, and the first introduction request to the central server; after the central server determines the first response, it receives the first response transmitted by the central server; if the request type is a second type, it determines the first introduction intent of the first introduction request, wherein the second type is used to indicate a basic attribute information introduction request with a single intent; and it determines the first response based on the first introduction intent.

[0100] Specifically, differentiated model deployments are implemented to meet the intent analysis needs at different levels. A lightweight intent extraction model is deployed on the edge server side. This model is pre-trained and built with a basic question-and-answer database for exhibits. The database includes multiple lightweight tasks such as exhibit origin, cultural history, and geographical background, which are adapted to the processing needs of the edge. On the central server side, a large intent classification model with a large language model architecture is deployed. This model integrates user profile association algorithms and contextual semantic analysis modules to perform in-depth user intent classification.

[0101] Specifically, a dual-model architecture is adopted in the user's interactive voice information processing mechanism to reduce peak business processing pressure.

[0102] The lightweight intent extraction model performs intent recognition and lightweight task category matching based on the user's current voice information, and executes the corresponding lightweight task based on the matching results, such as answering questions about the source of exhibits, answering questions about the cultural and historical background of exhibits, and answering questions about the natural geographical background information of exhibits.

[0103] The interactive information consists of contextually continuous interactive data, including historical continuous voice dialogue records between the user and the terminal, and records of the user's terminal operation behavior. The identity feature information includes the identity information registered and entered by the user in the self-guided tour system, historical visitor profiles such as areas of interest in exhibits, frequently asked question types, and tagged characteristics such as professional visitor, general tourist, and student.

[0104] In one possible embodiment, the central server is configured to receive the interaction information, the identity feature information, and the first introduction request transmitted by the edge server; perform intent recognition based on the interaction information, the identity feature information, and the first introduction request to obtain a second introduction intent; generate a reference response based on the second introduction intent; modify the reference response based on sensitive information detection to obtain the first response; and transmit the first response to the edge server.

[0105] The central server 104 receives data transmitted from the edge server 103 and then calls the intent classification model to perform intent recognition processing. Based on interaction information, identity feature information, and the first introduction request, it classifies the user's dialogue intent, such as "personalized exhibit recommendation" and "in-depth historical interpretation". Then, based on the classification results, it creates original response content, which includes user profile-related information, such as "As a history major student, you can focus on a certain detail of this exhibit". It also performs content sensitivity review on the original response content. If no sensitive information is detected, the response is converted into voice and pushed to the edge server, which then pushes it to the terminal. If sensitive information is detected, the response is regenerated and reviewed a second time, and a response without sensitive information is output and transmitted to the edge server. The edge server then controls the corresponding terminal to output the first response.

[0106] After receiving the first response, the edge server calls its own audio playback module to push playback instructions to multiple first voice guide terminals in parallel, ensuring that the playback latency of all terminals is less than a preset threshold, thus achieving synchronous playback across multiple terminals.

[0107] Specifically, for terminals that receive a user's introduction request, the control terminal plays the corresponding first response; for terminals that do not receive a user's introduction request, the control terminal plays the preset introduction information of the first exhibit, that is, the fixed standardized introduction content of the exhibit, such as pre-recorded audio of the exhibit explanation.

[0108] When the terminal outputs the first response or preset introductory information, it can adopt a multimodal fusion output method to simultaneously realize the output forms of real-time voice broadcast, text interface display and presentation of relevant visual content of exhibits.

[0109] As can be seen, in this embodiment of the application, by assigning tasks for different request types, the computing power of edge servers and central servers can be efficiently utilized, resource waste can be avoided, system response latency can be reduced, the real-time and accuracy of responses can be ensured, the efficiency of system voice interaction processing and response quality can be guaranteed, and the voice guide interaction experience of visiting users can be optimized.

[0110] In one possible embodiment, upon detecting that a user has entered the electronic fence area corresponding to an exhibit, the system can determine whether there is visual obstruction between the exhibit and the user based on the user's real-time location information, combined with the exhibit's display layout and the spatial structure of the exhibition hall. This identifies all users within the electronic fence area who do not visually obstruct the exhibit, and then pushes the pre-set introductory information corresponding to the exhibit directly to these users. This achieves precise matching and targeted delivery of exhibit information to the user's actual viewing area, avoiding invalid information pushes and ensuring synchronization between information push and the user's actual viewing progress. It satisfies the user's need to receive corresponding explanations while viewing the exhibit, while also avoiding information interference caused by invalid pushes, thus improving the fit between the user's viewing and guided tour experience.

[0111] In one possible embodiment, the self-service voice guide system can monitor and optimize the flow and order of visitors to popular exhibits, thus solving the problem of overcrowding in waiting areas due to excessive crowds and slow movement in the vicinity of exhibits.

[0112] Specifically, the edge server can determine the real-time exhibition status of visitors every second. According to the system's preset congestion control rules, it can issue instructions to at least one second voice guide terminal at the entrance of popular exhibit areas, surrounding passageways, and adjacent exhibit areas, controlling them to output corresponding prompts. The prompts may include real-time congestion conditions in popular exhibit areas, optimal detour routes, and suggestions for off-peak visits. This guides visitors to adjust their visit routes, alleviates congestion, and improves flow efficiency, thereby reducing congestion in popular exhibit areas.

[0113] In one possible embodiment, the edge server determines the visitor's first exhibition participation status based on the location information and entry time of the voice guide terminal. This first participation status characterizes the visitor's real-time behavioral state during the exhibition. Specifically, the first participation status includes distribution information, dwell time, initial flow information, and the number of visitors.

[0114] Specifically, based on the location information of the voice guide terminals, the number of viewers of the voice guide terminals within the electronic fence area is determined; based on the area of ​​the electronic fence area and the number of viewers of the voice guide terminals, the crowd distribution density of the first exhibit is determined, and distribution information is obtained; based on the entry time and location information, the dwell time of the voice guide terminals is determined; based on the location information, the movement trajectory of the voice guide terminals within the electronic fence area is generated; based on the dwell time and movement trajectory, the movement speed of the visitors is determined, and first flow information is obtained.

[0115] The crowd density is calculated by dividing the number of voice guide terminals within the electronic fence area by the fence area. For example, if the fence area is 10 square meters and there are 20 terminals within the area, the crowd density is 2 people / square meter. Movement speed is calculated as the ratio of the displacement difference to the time difference between two consecutive location data points of the voice guide terminal. For example, if a terminal moves 1 meter within 5 seconds, the movement speed is 0.2 m / s. Movement speed refers to the average movement speed of visitors. Dwell time is directly recorded as the duration the voice guide terminal remains within the electronic fence area. For example, terminals with a dwell time exceeding 10 minutes can be marked as "excessive dwell time".

[0116] The edge server analyzes the initial exhibition status in real time and, based on the need for off-peak visit control, generates prompts to guide visitors to visit during off-peak hours. These prompts can be generated by considering the actual exhibition status within the electronic fence area, such as suggested times for visiting the area, directions to nearby exhibits to prioritize, and suggested movement directions within the area. The edge server then selects voice guide terminals within the electronic fence area that are associated with the preset off-peak visit guidance goals and designates them as secondary voice guide terminals. The association can be determined based on the terminal's location within the electronic fence area, the visitor's viewing trajectory, and movement direction. For example, voice guide terminals at the entrance of the electronic fence area, those with long dwell times, and slow movement speeds are included in the scope associated with the off-peak visit guidance goals. The edge server then sends instructions to at least one of the selected secondary voice guide terminals to broadcast the generated prompts, allowing visitors to understand the current exhibition status of the electronic fence area and the off-peak visit suggestions. This enables them to adjust their viewing plans, routes, and stop arrangements independently, thus effectively guiding visitors to visit during off-peak hours.

[0117] As can be seen, in this embodiment, the parallel response of the voice guide terminal through time-sliding windows can effectively avoid problems such as response timeouts and service interruptions caused by business congestion, reduce peak business pressure, and improve the response time and operational efficiency of the exhibition hall's voice guide service. Furthermore, based on the user's real-time exhibition status, dynamic monitoring of crowd flow and circulation can be performed, and targeted prompts can be output to assist the exhibition hall in achieving refined management of the exhibition order, effectively alleviating exhibition congestion, achieving dynamic balance in the distribution of visitors within the exhibition hall, improving the user's visit smoothness and experience satisfaction, and simultaneously increasing the exhibition hall's carrying capacity and space utilization.

[0118] In one possible embodiment, after determining the first exhibition status of the visiting user, the edge server is used to generate prompt information based on the first exhibition status and control at least one second voice guide terminal to output the prompt information. Specifically, this involves: detecting that the distribution information does not meet the density distribution constraint; determining the spatial pattern of the crowd distribution based on the location information; determining multiple target users based on the spatial pattern, wherein the target users are visiting users who do not have visual obstruction from the first exhibit; determining the viewing time of the multiple target users; determining at least one second voice guide terminal based on the viewing time and / or the dwell time; determining the degree of association between at least one second exhibit and the first exhibit, wherein the second exhibit and the first exhibit have a spatial location association; determining the second exhibition status of the visiting user within the electronic fence area of ​​each second exhibit; generating a first prompt based on the degree of association, the second exhibition status, and the distribution information, and controlling the at least one second voice guide terminal to output the first prompt.

[0119] The distribution information not meeting the density distribution constraint refers to a crowd density exceeding a preset threshold, indicating that the current electronic fence area is experiencing excessive crowd gathering, exceeding the exhibition hall's preset reasonable visitor capacity, and potentially causing problems such as area congestion, obstructed views, and reduced visitor efficiency. For an example, please refer to... Figure 4 , Figure 4 This is a schematic diagram of the distribution of people in an exhibition hall provided in an embodiment of this application, such as... Figure 4 As shown, the black circles represent visitors. The crowd density distribution of the second and third exhibits meets the density distribution constraint, while the crowd density distribution of the first exhibit does not meet the density distribution constraint and requires crowd management.

[0120] The edge server uses multiple discrete terminal location data, combined with the spatial boundary of the electronic fence area of ​​the first exhibit and the location coordinates of the exhibit, to visualize the location distribution of people in the area through spatial coordinate mapping and crowd clustering analysis, and to determine the real-time distribution pattern of people in the physical space.

[0121] The edge server establishes a visible spatial area based on the spatial layout information, viewing angle, and field of view of the first exhibit. It then performs visual occlusion checks on user positions in different areas of the space, filtering out all visitors who are within the exhibit's visible range, have no physical obstructions, and can directly view the first exhibit. These users are then identified as target users. For example, such as... Figure 4 As shown, in the electronic fence area of ​​the first exhibit, the red circle represents the target user.

[0122] The viewing time refers to the time when target users within the electronic fence area of ​​the first exhibit, where there are no visual obstructions, actually begin to view the exhibit. It is a time indicator used to distinguish whether users have truly entered the effective viewing stage.

[0123] The dwell time refers to the cumulative time that visitors spend in the electronic fence area of ​​the first exhibit. By analyzing the viewing time of target users and the dwell time of ordinary visitors, the actual viewing status of users is analyzed, and users who have completed their viewing progress, stayed for too long, and are suitable for receiving off-peak guidance are selected. The voice guide terminal they carry is the second voice guide terminal, which is the targeted target for off-peak guidance prompts.

[0124] Specifically, if the stay time is greater than the first preset time, or the stay time is less than the second preset time, or the exhibition viewing time is greater than the fourth preset time, then the corresponding user's voice guide terminal is determined to be the second voice guide terminal.

[0125] The second exhibit is a peripheral exhibit that has a spatial locational connection with the first exhibit, such as exhibits in the same exhibition area or adjacent exhibition areas. The degree of connection between the first and second exhibits can be determined based on spatial distance and exhibition theme. This connection determines the recommendation priority of peripheral second exhibits; if they share the same theme, are in the same exhibition area, or are located in adjacent areas, the higher the degree of connection, the higher they will be included in the recommendation list, ensuring that the guidance aligns with the user's current viewing interests.

[0126] The second exhibition status represents the real-time behavior of people in the second exhibition area, such as crowd distribution density, user movement status, and mobility. It is used to determine whether each second exhibition area is in a state of relaxed visitor flow and can accommodate the crowds dispersed by the first exhibition. For example, second exhibitions with low visitor density, good mobility, and no congestion are recommended as effective items to avoid guiding users to new gathering areas and causing secondary congestion.

[0127] The distribution information, namely the real-time crowd density distribution in the first exhibit area, is used to determine the way the prompts are presented. If the first exhibit has a high degree of crowding and obvious congestion, the prompts will focus more on guiding the flow of people; if the crowding is mild, the prompts will focus on gentle guidance.

[0128] For example, the first suggestion could be that the current first exhibit area has a high visitor density, and to improve your viewing experience, we recommend that you visit the second exhibit area which is adjacent to this exhibit area. This exhibit has the same theme as this one, and the area is currently less crowded with visitors and has a good view. You can reach it by walking straight for 50 meters from your current location to the north side of the exhibit area.

[0129] In this process, the edge server, based on the output prompts, sends playback commands to the second voice guide terminals carried by users within the designated first exhibit area who are suitable for receiving off-peak guidance. The commands control the terminals to output the first prompts in a multimodal manner, while simultaneously displaying location guidance icons for the second exhibit on the terminal screen. This ensures precise delivery of guidance information, guarantees the actual effectiveness of off-peak guidance, avoids irrelevant information from interfering with other visitors, and improves the efficiency of information reception for users by using targeted output, thus promoting the orderly dispersal of crowds gathered around the first exhibit.

[0130] In one possible embodiment, when the edge server determines at least one second voice guide terminal based on the dwell time, it specifically performs the following steps: if the dwell time is greater than a first preset time or less than a second preset time, the corresponding visitor's voice guide terminal is identified as the second voice guide terminal; if the dwell time, spatial morphology, and first flow information are used to determine whether a queuing crowd exists; if the queuing crowd is detected, the queuing time is determined based on the exhibition time, the third number of target users, and the location information; if the queuing time is detected to be greater than a third preset time, the corresponding visitor's voice guide terminal is identified as the second voice guide terminal.

[0131] The edge server can push voice prompts to "terminals with excessively long dwell times" or "newly added terminals within the fence." Specifically, if a user's dwell time exceeds a first preset time or is less than a second preset time, a prompt message will be pushed to them.

[0132] When the crowd density in the first exhibit area exceeds the threshold, queuing is a typical orderly form of crowd gathering. Identifying the queuing pattern can distinguish it from irregular crowding, providing a basis for subsequent precise staggered peak guidance. At the same time, it can effectively reduce the waiting cost of queuing people, ensure the exhibition experience, and achieve efficient crowd dispersal.

[0133] If the crowd exhibits an orderly spatial distribution along the visible direction of the exhibits, the overall dwell time continues to increase, and the flow information shows an extremely low movement rate or no obvious displacement, it can be determined that there are queuing people in the area.

[0134] In one possible embodiment, when the edge server determines whether there is a queuing crowd based on the dwell time, the spatial pattern, and the first flow information, it specifically performs the following steps: dividing the electronic fence area into multiple sub-areas based on the spatial pattern and the distance between the visitor and the first exhibit; determining movement characteristics based on the first flow information and dwell time of the visitor in each sub-area; and determining whether there is a queuing crowd based on the movement characteristics.

[0135] Among them, taking the spatial pattern of crowd distribution as a reference, such as the overall direction of the gathering and the distribution outline, and taking the actual physical distance between the visitor and the first exhibit as the dividing dimension, the electronic fence area is divided into multiple clearly defined sub-areas. Each sub-area can be divided into "near exhibit area, medium distance area, and far exhibit area", resulting in multiple sub-areas.

[0136] Specifically, the first flow information and dwell time of users in each sub-region are extracted. The first flow information is used to determine the user's movement speed, movement direction, and whether the movement path is fixed, while the dwell time is used to determine the duration distribution characteristics. Through the first flow information and dwell time, the movement characteristics of the population in each sub-region are determined, such as the synchronicity of movement speed, the consistency of movement direction, and the differences in the distribution of dwell time.

[0137] Among them, the unidirectional orderly slow movement along a fixed route is the spatial movement characteristic of queuing, and the obvious gradient of the dwell time along the route is the temporal dwelling characteristic of queuing. By distinguishing between the orderly aggregation of queuing and the irregular aggregation of crowds through spatial movement characteristics and temporal dwelling characteristics, the queuing crowd can be accurately identified.

[0138] In one possible embodiment, after determining the existence of the queuing crowd, the edge server, when determining the queuing time based on the exhibition viewing time, the third number of target users, and the location information, specifically performs the following: determining the replacement time of the target users based on the exhibition viewing time and the third number; determining the fourth number of previously visiting users in the queue based on the location information; determining the second flow information of the queue; determining a reference queuing time based on the replacement time and the fourth number; determining the movement time of the visiting users based on the second flow information; and determining the queuing time based on the reference queuing time and the movement time.

[0139] Among them, the viewing time is the actual viewing time of the target user who is currently viewing the exhibition without visual obstruction, and the third quantity is the total number of target users who are currently viewing the exhibition. By summing the viewing times of all target users and dividing by the third quantity, the average viewing time per user, i.e. the replacement time, is calculated. It is used to characterize the baseline time for each user to complete the viewing and be replaced by the next user.

[0140] Among them, the spatial length of the queue and the specific position of the current user in the queue are determined based on the location information, and then the specific number of other visitors waiting in line ahead of the current user is determined, which is the fourth quantity.

[0141] The process involves calculating the product of the replacement time and the fourth quantity to determine the total waiting time for the current user to wait for all the users in the queue ahead of them to complete their exhibition viewing, which is the reference queuing time.

[0142] The second flow information of the queue consists of dynamic traffic indicators, including the average movement speed, distance traveled per unit time, and traffic flow efficiency. Based on the spatial orientation of the queue's fixed movement path and the straight-line distance between the current user's real-time location and the viewing area, the actual travel distance of the user along the fixed movement path to the viewing area is determined. The average movement speed in the second flow information is corrected using the traffic flow efficiency to obtain the queue's actual effective movement speed, which is a correction coefficient for actual movement in the queuing scenario. The basic movement time is determined using the actual travel distance and the actual effective movement speed.

[0143] The movement time is verified by measuring the distance traveled per unit time in the second flow information. Specifically, the verification movement time is determined by measuring the actual travel distance and the distance traveled per unit time. The edge server performs a consistency check between the basic movement time and the verification movement time. If the difference between the two is within a preset error threshold, the average of the two is taken as the final movement time. If the difference exceeds the error threshold, the actual observed verification movement time is determined as the final movement time.

[0144] The estimated queuing time for each user in the queue is determined by adding the queuing time to the movement time.

[0145] If a user's queuing time exceeds the third estimated time, it indicates that their waiting cost is too high. In this case, the user's voice guide terminal is designated as the second voice guide terminal and is used as the target of the first push notification for off-peak guidance, thus achieving accurate screening of users with high waiting costs.

[0146] As can be seen, in this embodiment, the system accurately identifies queuing crowds and irregularly clustered crowds. Based on the determined queuing time, the system identifies the terminals of users with high waiting costs as the second voice guide terminals, allowing the guidance information to reach the users who truly need it. This avoids indiscriminate guidance interfering with the information of normal visitors and those queuing for short periods, improves the efficiency and conversion effect of guidance information, and ensures the user's exhibition experience.

[0147] At the same time, it not only effectively disperses the long queues in popular exhibit areas, but also directs people to the surrounding secondary exhibit areas that are closely related to the primary exhibits and have more space, thus redistributing visitor flow within the exhibition hall, avoiding uneven distribution of visitors, and improving the exhibition hall's capacity and space utilization.

[0148] In one possible embodiment, when the edge server determines at least one second voice guide terminal based on the exhibition viewing time, it is specifically configured to: obtain the dwell time of multiple third voice guide terminals after completing the guided tour related to the first exhibit, wherein the third voice guide terminal is the voice guide terminal of the target user; if the dwell time is greater than a fourth preset time or the exhibition viewing time is greater than a fifth preset time, then the voice guide terminal corresponding to the target user is determined as the second voice guide terminal.

[0149] Dwell time refers to the time a target user spends in an area after completing the guided tour of the first exhibit, such as the time spent after the terminal plays the target response or preset introductory information.

[0150] The system filters users within the first exhibit area based on their viewing time and dwell time after the guided tour. Specifically, viewing time exceeding the fifth preset time indicates that users have fully grasped the exhibit information, and continuing to stay would occupy area space; dwell time exceeding the fourth preset time after the guided tour indicates that users are engaging in meaningless and ineffective dwelling. The edge server then identifies the voice guide terminals of users meeting the above filtering criteria as the second voice guide terminals, achieving accurate identification of crowds in popular exhibit areas.

[0151] As can be seen, in this embodiment, the target is unobstructed users. The system identifies the crowds in the area based on the viewing progress and dwell time, which complements the screening of queuing crowds. This achieves full-dimensional coverage of all crowds in popular exhibit areas, further improving the accuracy of crowd management and improving crowd flow efficiency.

[0152] In one possible embodiment, the edge server is further configured to respond to a second introduction request from the voice guide terminals corresponding to the multiple target users, output a second response based on the second introduction request, and control the voice guide terminals corresponding to the multiple target users to output the second response.

[0153] The second introduction request is a personalized, supplementary guide consultation request initiated by the target user through their own voice guide terminal, such as detailed explanations of exhibits or consultation on extended knowledge. If multiple target users in the same exhibit area initiate the second introduction request within the same time window, the edge server responds synchronously. Based on the specific consultation content in the request, the request type is determined. If the request type is a single intent, it is parsed and matched against the pre-stored first exhibit information database to generate a second response highly matching the user's request. If the request type is a nested multi-intent request, the relevant information is transmitted to the central server, which generates the response content and performs sensitive information detection on the response content. If no sensitive information is detected, the response is converted to voice and pushed to the edge server, which controls the corresponding terminal to output the response information. If sensitive information is detected, the response is regenerated and reviewed a second time, finally outputting a second response without sensitive information, which is then converted to voice and pushed to the edge server, which controls the corresponding terminal to output the response information.

[0154] As can be seen, in this embodiment of the application, the edge server supports concurrent responses from multiple terminals of the target user, ensuring the synchronization between the response output and the user's actual viewing progress. This meets the user's guide needs to view exhibits and receive corresponding explanations during the viewing process, while avoiding information interference caused by invalid push notifications and improving the fit between the user's viewing and guide experience.

[0155] In one possible embodiment, if the visitor's movement speed does not meet the movement speed constraint, then the duration for which the visitor's movement speed is less than a preset speed is determined; the changing trend of visitors in the electronic fence area is predicted, and it is determined whether the number of visitors will increase or decrease in the next time period; a second prompt is generated based on the changing trend, the second exhibition status, and the movement speed. The second prompt includes traffic efficiency guidance information and exhibition route optimization information, and at least one second voice guide terminal is controlled to output the second prompt. The at least one second voice guide terminal may include a voice guide terminal with a duration less than a sixth preset time, and / or a voice guide terminal with a dwell time greater than a first preset time, and / or a voice guide terminal with a dwell time less than a second preset time.

[0156] In one possible embodiment, the total number of visitors within the electronic fence area is the first number of visitors. A second number of visitors is determined based on the number of voice guide terminals whose dwell time exceeds a first preset time. If both the first and second numbers exceed the first preset number, crowd control measures are required. Specifically, the difference between the first and second numbers is determined. If the difference exceeds the first preset number, at least one third voice guide terminal is identified, located near the boundary of the electronic fence area. A third prompt is generated, including information on dispersing and pausing, and information on optimizing the visitor route. The at least one second and at least one third voice guide terminal are controlled to output the third prompt. The at least one second voice guide terminal includes voice guide terminals whose dwell time exceeds the first preset time, and / or voice guide terminals whose dwell time is less than the second preset time.

[0157] In one possible embodiment, if the distribution information satisfies the density distribution constraint, the user's stopping and viewing behavior can be determined by the user's movement speed, and then value-added information and services beyond basic navigation can be pushed, avoiding the inefficiency of information reception caused by pushing complex content during the user's movement.

[0158] In one possible embodiment, after the push notification message is sent, the edge server continuously monitors the location changes of the target terminal. If the terminal leaves the current electronic fence area within 1 minute, it is marked as "intervention successful"; if the terminal does not leave, the notification will be pushed again after 1.5 minutes, and the notification will be pushed a maximum of 2 times.

[0159] In one possible embodiment, the self-service voice guidance system optimizes system parameters and model performance based on operational data. Specifically, the central server first collects operational data in real time at each stage, including event summary data such as sliding window trigger success rate and parallel response latency, as well as voice processing data such as lightweight model recognition accuracy, intent classification large model review pass rate, and user request response time, and crowd optimization data such as intervention success rate, average dwell time in popular exhibit areas, and peak crowd density. Then, the operational data is analyzed and optimized. Specifically, data review and targeted optimization are initiated during off-peak hours each day, which can be 22:00-6:00 the next day; if the sliding window parallel response latency is greater than 500ms, the sliding window duration is adjusted or edge server nodes are added; if the lightweight model recognition accuracy is less than 90%, training data is supplemented, such as adding exhibit question-and-answer samples; if the crowd intervention success rate is less than 70%, the prompts or dwell time thresholds are adjusted, such as shortening the dwell time threshold from 10 minutes to 8 minutes. Finally, based on the data analysis results, the system undergoes a comprehensive iterative update every month, optimizing the parameters of the two models. For example, the contextual association weights in the intent classification model are adjusted, the preset audio library is updated (e.g., adding seasonally limited introductions to exhibits), and the electronic fence range is adjusted according to actual exhibit changes (e.g., updating coordinates based on exhibit movement). The specific data mentioned above are example data and can be adjusted according to actual circumstances; they are not limited here.

[0160] Please refer to Figure 5 , Figure 5 This is a flowchart illustrating a self-service voice navigation method provided in an embodiment of this application, such as... Figure 5 As shown, this method is applied to the edge server of a self-service voice guide system. The self-service voice guide system also includes voice guide terminals and positioning base stations. Communication connections are established between the voice guide terminals and the positioning base stations, as well as between the voice guide terminals and the edge server. The method includes the following steps:

[0161] S510, Obtain event information detected by the positioning base station.

[0162] The event information includes the entry time of the electronic fence area of ​​the first exhibit and the location information of the voice guide terminal.

[0163] The positioning base station collects the location data of the voice guide terminal in real time and compares the real-time location coordinates of the voice guide terminal with the preset coordinate range of the electronic fence area of ​​each exhibit. If the comparison determines that the voice guide terminal has entered the electronic fence area of ​​an exhibit, the location-triggered event information is uploaded to the edge server of the exhibit area to which the fence belongs, thus completing the real-time reporting of the fence entry information.

[0164] S520: Obtain a first number of first introduction requests for the first exhibit collected by the first voice guide terminal.

[0165] The first audio guide terminal is the audio guide terminal whose entry time is within a preset time window.

[0166] Among them, the voice guide terminal uses voiceprint recognition technology to identify the first introduction request made by the visitor for the current first exhibit; and synchronously transmits the identified first introduction request to the edge server.

[0167] After receiving the corresponding first introduction request uploaded by the voice guide terminal, the edge server summarizes multiple first introduction requests within the same preset time window and processes all introduction requests within that time period in a unified manner based on the summary results.

[0168] S530, output the first response according to the first description request.

[0169] The edge server processes all introduction requests collected within the same time window and outputs the first response corresponding to the introduction request, thus achieving centralized response and information feedback for multiple terminal requests within the same time period.

[0170] S540, control a second number of the first voice guide terminals to output the first response and / or the preset introduction information of the first exhibit.

[0171] Wherein, the second quantity is greater than or equal to the first quantity.

[0172] In this process, after generating the corresponding first response, the edge server sends it to the first voice guide terminal, controlling the corresponding terminal to play the first response synchronously, or to play the first response and preset introductory information synchronously.

[0173] If the first voice guide terminal does not detect an introduction request within the same time window, it will directly control it to play preset introduction information.

[0174] S550, based on the location information and the entry time, determine the first exhibition status of the visiting user.

[0175] The first exhibition status is used to characterize the real-time behavior of the visiting user during the exhibition process.

[0176] The first stage of participation includes crowd density, movement speed, and dwell time.

[0177] Specifically, the system determines the number of visitors to the audio guide terminals within the electronic fence area based on their location information; it determines the crowd density based on the area of ​​the electronic fence area and the number of visitors to the audio guide terminals; it determines the dwell time of the audio guide terminals based on their entry time and location information; it generates the movement trajectory of the audio guide terminals within the electronic fence area based on their location information; and it determines the movement speed of the visitors based on their dwell time and movement trajectory.

[0178] S560, generate prompt information based on the first exhibition status, and control at least one second voice guide terminal to output the prompt information.

[0179] The prompt information is used to guide visitors to visit during off-peak hours, and the at least one second voice guide terminal is a voice guide terminal associated with the off-peak visit guidance target within the electronic fence area.

[0180] The edge server compares the real-time exhibition data with preset thresholds. If any of the following conditions are met, a congestion intervention mechanism is immediately triggered: the crowd density exceeds a preset range, the average movement speed is lower than a preset range, or there are at least a preset number of terminals whose dwell time exceeds a preset time. For example, specific conditions could be a crowd density exceeding 1.5 people / m², an average movement speed below 0.3 m / s, or at least five terminals whose dwell time exceeds 10 minutes.

[0181] During the intervention, the edge server will push customized voice prompts to terminals that stay for extended periods within the electronic fence area or to newly added terminals within the fence. The prompts may include messages such as "There are currently many visitors in the first exhibit area. We suggest you go to the second exhibit area, where there is supplementary information about the same series of exhibits." At the same time, the location guide icon for the second exhibit will be displayed on the corresponding terminal screen.

[0182] As can be seen, in this embodiment of the application, crowd monitoring optimizes flow efficiency, effectively improves the effectiveness of crowd management in the system, and significantly improves user experience and service quality during peak periods.

[0183] For examples consistent with the above embodiments, please refer to... Figure 6 , Figure 6 This is a functional unit block diagram of a self-service voice guide device provided in an embodiment of this application, such as... Figure 6As shown, the self-service voice guide device 60 includes: a first acquisition unit 61, used to acquire event information detected by the positioning base station, the event information including the entry time of the electronic fence area of ​​the first exhibit and the location information of the voice guide terminal; a second acquisition unit 62, used to acquire a first introduction request for the first exhibit collected by a first number of first voice guide terminals, the first voice guide terminals being voice guide terminals whose entry time is within a preset time window; an output unit 63, used to output a first response according to the first introduction request; and a first control unit 64, used to control a second number of the first voice guide terminals to output the first response. The second quantity is greater than or equal to the first quantity, and the second quantity is combined with / or the preset introduction information of the first exhibit. The determining unit 65 is used to determine the first exhibition status of the visitor based on the location information and the entry time. The first exhibition status is used to characterize the real-time behavior of the visitor during the exhibition. The second control unit 66 is used to generate prompt information based on the first exhibition status and control at least one second voice guide terminal to output the prompt information. The prompt information is used to guide the visitor to visit during off-peak hours. The second voice guide terminal is a voice guide terminal associated with the off-peak visit guidance target within the electronic fence area.

[0184] It is understood that since the method embodiments and the device embodiments are different presentations of the same technical concept, the content of the method embodiment section in this application should be adapted to the device embodiment section in a synchronous manner, and will not be repeated here.

[0185] In the case of using integrated units, please refer to Figure 7 , Figure 7 This is a functional unit block diagram of another self-service voice guide device provided in the embodiments of this application, such as... Figure 7 As shown, the self-service audio guide device 60 includes a processing module 602 and a communication module 601. The processing module 602 controls and manages the actions of the self-service audio guide device 60, for example, executing the steps of the first acquisition unit 61, the second acquisition unit 62, the output unit 63, the first control unit 64, the determination unit 65, and the second control unit 66, and / or performing other processes of the technology described herein. The communication module 601 is used for interaction between the self-service audio guide device 60 and other devices.

[0186] Among them, such as Figure 7 As shown, the self-service audio guide device 60 may also include a storage module 603, which is used to store the program code and data of the self-service audio guide device 60.

[0187] The processing module 602 can be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication module 601 can be a transceiver, RF circuitry, or a communication interface, etc. The storage module 603 can be a memory.

[0188] All relevant content for each scenario involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here. The above-mentioned self-service voice guidance device 60 can perform the above-mentioned... Figure 2 The self-service voice guide system shown illustrates the self-service voice guide method implemented.

[0189] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device proposed in an embodiment of this application, as shown below. Figure 8 As shown, the electronic device 800 includes a processor 810, a memory 820, a communication interface 830, and one or more programs 821. The one or more programs 821 are stored in the memory and configured to be executed by the processor. When the program is executed, it includes some or all of the steps of any of the methods described in the above method embodiments. The processor, memory, and communication interface are interconnected and complete communication between them.

[0190] The memory can be volatile memory such as Dynamic Random Access Memory (DRAM) or non-volatile memory such as a hard disk drive. The memory stores a set of executable program code, and the processor calls the executable program code stored in the memory to execute some or all of the steps of the self-service voice guidance method implemented by any self-service voice guidance system described in any of the above embodiments.

[0191] As can be seen, the electronic device 800 described in this application embodiment first acquires event information detected by the positioning base station, the event information including the entry time of the electronic fence area of ​​the first exhibit and the location information of the voice guide terminal; then acquires a first introduction request for the first exhibit collected by a first number of first voice guide terminals, the first voice guide terminals being voice guide terminals whose entry time is within a preset time window; then outputs a first response according to the first introduction request; next, controls a second number of first voice guide terminals to output the first response and / or preset introduction information of the first exhibit, the second number being greater than or equal to the first number; then, determines the first exhibition status of the visitor based on the location information and the entry time, the first exhibition status being used to characterize the real-time behavior of the visitor during the exhibition; finally, generates prompt information based on the first exhibition status and controls at least one second voice guide terminal to output the prompt information, the prompt information being used to guide the visitor to visit during off-peak hours, the second voice guide terminal being a voice guide terminal associated with the off-peak visit guidance target within the electronic fence area.

[0192] This application achieves parallel response of the voice guide terminal through time-sliding windows, which can effectively avoid problems such as response timeouts and service interruptions caused by business congestion, reduce peak business pressure, and improve the response time and operational efficiency of the exhibition hall's voice guide service; and outputs targeted prompts based on the user's exhibition status, assisting the exhibition hall in achieving refined management of the exhibition order, effectively alleviating exhibition congestion, achieving dynamic balance in the distribution of visitors within the exhibition hall, improving the smoothness of the visit and the satisfaction of the user experience, while also improving the exhibition hall's carrying capacity and space utilization.

[0193] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes an electronic device.

[0194] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include an electronic device.

[0195] It should be noted that, for the sake of simplicity, the aforementioned methods are described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are optional, and the actions and modules involved are not necessarily essential to this application.

[0196] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0197] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0198] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0199] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software program module.

[0200] If the integrated unit is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0201] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc.

[0202] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A self-service voice navigation system, characterized in that, This includes voice navigation terminals, positioning base stations, and edge servers; The positioning base station is used to detect events where the voice guide terminal enters the electronic fence area of ​​the first exhibit; When the event is detected, the corresponding event information is transmitted to the edge server. The event information includes the location information and entry time of the voice guide terminal. The voice guide terminal is used to receive a first introduction request from a visitor regarding the first exhibit based on voiceprint recognition, and to transmit the first introduction request to the edge server. The edge server is configured to respond to a first introduction request from a first number of first voice-guided terminals, output a first response based on the first introduction request, wherein the first voice-guided terminals are those whose entry time falls within a preset time window; control a second number of first voice-guided terminals to output the first response and / or preset introduction information for the first exhibit, wherein the second number is greater than or equal to the first number; and... The edge server is also used to determine the first exhibition status of the visiting user based on the location information and the entry time. The first exhibition status is used to characterize the real-time behavior of the visiting user during the exhibition process. Based on the first exhibition status, a prompt message is generated, and at least one second voice guide terminal is controlled to output the prompt message. The prompt message is used to guide the visitors to visit during off-peak hours. The second voice guide terminal is a voice guide terminal associated with the off-peak visit guidance target within the electronic fence area. The system further includes a central server, and the edge servers establish a communication connection with the central server. When the edge servers respond to first introduction requests from multiple first voice guide terminals and output a first response based on the first introduction request, they are specifically used for: Determine the request type of the first introductory request; Match the corresponding target device according to the request type and trigger the execution of the corresponding target task. The target device includes the edge server and the central server. The target task is used to determine the first response that matches the first introduction request. Wherein, after determining the request type of the first introduction request, the edge server, when matching the corresponding target device according to the request type and triggering the execution of the corresponding target task, is specifically used for: If the request type is the first type, then the interaction information between the visiting user and the first voice guide terminal is obtained, as well as the identity feature information of the visiting user is obtained. The first type is used to indicate a deep information introduction request with multiple nested intentions. The interaction information, the identity feature information and the first introduction request are transmitted to the central server. After the central server determines the first response, the first response transmitted by the central server is received. If the request type is the second type, then the first introduction intent of the first introduction request is determined, and the second type is used to indicate the basic attribute information introduction request of a single intent; the first response is determined according to the first introduction intent.

2. The system according to claim 1, characterized in that, The first exhibition participation status includes distribution information and dwell time. After determining the first exhibition participation status of the visitor, the edge server is used to generate prompt information based on the first exhibition participation status and control at least one second voice guide terminal to output the prompt information. Specifically, it is used to: If the distribution information is found to not meet the density distribution constraints, the spatial pattern of the population distribution is determined based on the location information. Based on the spatial form, multiple target users are identified, and the target users are visitors who do not have visual obstruction between them and the first exhibit. Determine the viewing time of the multiple target users; Based on the exhibition viewing time and / or the dwell time, at least one second voice guide terminal is determined; Determine the degree of association between at least one second exhibit and the first exhibit, wherein the second exhibit and the first exhibit are spatially related; Determine the second participation status of the visiting users within the electronic fence area of ​​each second exhibit; Based on the degree of association, the second exhibition status, and the distribution information, a first prompt is generated, and the at least one second voice guide terminal is controlled to output the first prompt.

3. The system according to claim 2, characterized in that, The first exhibition status also includes first flow information. The edge server, based on the dwell time, determines when at least one second voice guide terminal is available, specifically for: If the dwell time is detected to be greater than a first preset time or less than a second preset time, the corresponding visitor's voice guide terminal is identified as the second voice guide terminal. The presence of queuing people is determined based on the dwell time, the spatial pattern, and the first flow information. Upon detecting the presence of the queuing crowd, the queuing time is determined based on the exhibition viewing time, the third number of target users, and the location information. If the queuing time is detected to be greater than the third preset time, the corresponding visitor's voice guide terminal is identified as the second voice guide terminal.

4. The system according to claim 3, characterized in that, After determining the existence of the queuing crowd, the edge server, when determining the queuing time based on the viewing time, the third number of target users, and the location information, specifically uses the following methods: The replacement time for the target user is determined based on the exhibition viewing time and the third quantity. The fourth number of visitors in the queue is determined based on the location information. Determine the second flow information of the queue; A reference queuing time is determined based on the replacement time and the fourth quantity; The travel time of the visiting user is determined based on the second flow information; The queuing time is determined based on the reference queuing time and the movement time.

5. The system according to claim 3, characterized in that, When the edge server determines whether a queuing crowd exists based on the dwell time, the spatial pattern, and the first flow information, it is specifically used for: Based on the spatial shape and the distance between the visitor and the first exhibit, the electronic fence area is divided into multiple sub-areas; Based on the initial flow information and dwell time of the visiting users described in each sub-region, determine the movement characteristics; The presence of a queuing crowd is determined based on the movement characteristics.

6. The system according to claim 2, characterized in that, When the edge server determines at least one second audio guide terminal based on the exhibition viewing time, it is specifically used for: The dwell time of multiple third-party voice guide terminals after completing the guided tour related to the first exhibit is obtained, wherein the third-party voice guide terminal is the voice guide terminal of the target user; If the dwell time is greater than the fourth preset time or the viewing time is greater than the fifth preset time, then the voice guide terminal corresponding to the target user will be determined as the second voice guide terminal.

7. The system according to claim 2, characterized in that, The edge server is also configured to respond to a second introduction request from the voice guide terminals corresponding to the multiple target users, output a second response based on the second introduction request, and control the voice guide terminals corresponding to the multiple target users to output the second response.

8. A self-service voice navigation method, characterized in that, An edge server for a self-service voice guide system, the self-service voice guide system further including a voice guide terminal and a positioning base station, the method comprising: Obtain event information detected by the positioning base station, the event information including the entry time of the electronic fence area of ​​the first exhibit and the location information of the voice guide terminal; The system acquires a first number of first introduction requests for the first exhibit collected by the first voice guide terminal, wherein the first voice guide terminal is a voice guide terminal whose entry time is within a preset time window. The first response is requested based on the first description. The second number of the first voice guide terminals outputs the first response and / or the preset introduction information of the first exhibit, wherein the second number is greater than or equal to the first number; Based on the location information and the entry time, the first exhibition status of the visiting user is determined, and the first exhibition status is used to characterize the real-time behavior of the visiting user during the exhibition process. Based on the first exhibition status, a prompt message is generated, and at least one second voice guide terminal is controlled to output the prompt message. The prompt message is used to guide the visitors to visit during off-peak hours. The second voice guide terminal is a voice guide terminal associated with the off-peak visit guidance target within the electronic fence area. The system further includes a central server, and the edge servers establish a communication connection with the central server. The step of outputting a first response based on the first description request includes: Determine the request type of the first introductory request; Match the corresponding target device according to the request type and trigger the execution of the corresponding target task. The target device includes the edge server and the central server. The target task is used to determine the first response that matches the first introduction request. The method further includes: If the request type is the first type, then the interaction information between the visiting user and the first voice guide terminal is obtained, as well as the identity feature information of the visiting user is obtained. The first type is used to indicate a deep information introduction request with multiple nested intentions. The interaction information, the identity feature information and the first introduction request are transmitted to the central server. After the central server determines the first response, the first response transmitted by the central server is received. If the request type is the second type, then the first introduction intent of the first introduction request is determined, and the second type is used to indicate the basic attribute information introduction request of a single intent; the first response is determined according to the first introduction intent.

Citation Information

Patent Citations

  • Accurate navigation system used for guiding indoor mall shopping, exhibition and sightseeing

    CN101694524A

  • AR-based tourist attraction intelligent guide method and related equipment

    CN120894186A