A method for automatically pushing scenic spot interpretation content by fusing multi-source perception

By identifying changes in tourists' emotions and abrupt changes in the sound field through multi-source sensing data, selecting samples with consistent environments, and calculating the delay in pushing explanation content, the dynamic adaptive optimization problem of the scenic area explanation system under acoustic interference is solved, thereby improving tourists' content reception and comprehension.

CN121434397BActive Publication Date: 2026-04-21WUXI LANNA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUXI LANNA TECHNOLOGY CO LTD
Filing Date
2025-11-07
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

The existing scenic area interpretation system cannot dynamically perceive changes in tourists' emotions and sudden changes in environmental acoustics, resulting in poor reception and understanding of the interpretation content, especially when attention is interrupted by sudden changes in the sound field energy of adjacent exhibition areas.

Method used

Significant emotional changes are identified by multi-source sensory data. By combining the sound field energy rise rate of abrupt sound field events in adjacent exhibition areas, sample data with consistent environmental backgrounds are selected, a comprehensive effectiveness score is calculated, and the push delay parameter of the explanation content is determined to achieve adaptive optimization.

Benefits of technology

It improves the responsiveness of the explanation content to interference in multi-exhibition-area scenarios and the accuracy of emotional response, thereby enhancing the content reception and immersive experience for visitors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121434397B_ABST
    Figure CN121434397B_ABST
Patent Text Reader

Abstract

This invention relates to the field of guided tour content delivery technology, providing a method for automatically delivering guided tour content to tourist attractions by integrating multi-source perception. The method includes: selecting the sample with the highest comprehensive effectiveness score as a reference sample; obtaining the corresponding guided tour content delivery delay parameter; and using this delay parameter as a standard for automatically delivering guided tour content to subsequent visitors in the target exhibition area. This invention has advantages such as strong interference adaptability, high accuracy in capturing emotional responses, and dynamically transferable delivery strategies, significantly improving the content reception and immersive experience for visitors in multi-exhibition area scenarios, and possessing broad application prospects in cultural exhibitions and smart tour guiding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio guide content delivery technology, and in particular to a method for automatically delivering audio guide content for tourist attractions that integrates multi-source perception. Background Technology

[0002] In existing scenic area tour guide systems, the mainstream approach to delivering guide content typically relies on location-triggered or time-triggered mechanisms. This means that when a visitor enters a pre-defined area or reaches a set time, the system automatically plays the corresponding guide content. While this method achieves basic automation, its triggering logic is simplistic and fails to consider changes in the visitor's psychological state or external environment. When there are significant acoustic differences between exhibit areas or sudden acoustic interference (such as a sudden increase in audio playback from adjacent areas, stage sound effects, or crowd noise), visitors' attention is easily interrupted, making it difficult for them to concentrate in the initial stages of entering a new exhibit area, thus affecting their reception and comprehension of the guide content. Current technology lacks the dynamic perception and response capabilities to this chain of "sudden environmental acoustic changes—emotional changes—content reception," and cannot adaptively optimize the delivery of guides to address sudden interference.

[0003] Some studies have attempted to identify visitor emotions using multi-source sensory data (such as facial expression recognition, voice emotion analysis, and body posture detection) to improve the relevance of explanation content. However, existing solutions often focus only on classifying emotional states and personalizing content recommendations, without incorporating the causes of emotional changes into the analysis, especially neglecting the immediate psychological reactions caused by sudden changes in sound field energy between adjacent exhibition areas. In fact, a sudden increase in sound pressure can trigger short-term emotional fluctuations and distractions in visitors, directly affecting their understanding and memory of previous explanations in subsequent exhibition areas. However, current technology has not yet established a quantitative correlation between changes in environmental acoustics and visitor cognitive performance, making it difficult to optimize the timing of explanation recommendations in response to acoustic disturbances.

[0004] This invention proposes an improvement solution to address the aforementioned technical problems. The inventors discovered that an increase in sound field energy in adjacent exhibition areas can trigger significant emotional changes in visitor groups, leading to a decrease in their comprehension of the preceding explanations in the next exhibition area. Summary of the Invention

[0005] The purpose of this invention is to provide a method for automatically pushing scenic spot explanation content that integrates multi-source perception, in order to solve the problems mentioned in the background art.

[0006] This invention is implemented as follows: a method for automatically pushing scenic spot explanation content that integrates multi-source perception, the method comprising:

[0007] When a significant emotional change is detected in the visitor group in the target exhibition area, it is determined whether the significant emotional change is caused by a sudden change in the sound field of an adjacent exhibition area. If it is determined to be so, the rise rate of the sound field energy of the sudden change in the sound field is obtained.

[0008] Several samples were retrieved from the database that had a consistent environmental background, exhibited significant emotional changes caused by abrupt changes in the sound field of adjacent exhibition areas, and had a consistent rate of increase in sound field energy.

[0009] Each sample was analyzed sequentially to determine the corresponding visitor feedback data, and the visitor group's understanding of the pre-exhibition explanation content was calculated based on the feedback data.

[0010] Further, the intensity index of emotional change and its corresponding intensity value were extracted for each sample, and the comprehensive effectiveness score of each sample was calculated based on the correspondence between comprehension and intensity of emotional change.

[0011] The sample with the highest overall effectiveness score is selected as the reference sample. The push delay parameter of the explanation content corresponding to the reference sample is obtained, and this push delay parameter is used as the standard for automatic push of explanation content to subsequent visitors in the target exhibition area.

[0012] As a further limitation of the technical solution of the present invention, the significant emotional change refers to the sudden change in the emotional state of the tourist group identified based on multi-source perception data; the multi-source perception data includes at least one or more of the tourist's facial expression changes, voice tone changes, and body movement characteristics changes, and when the magnitude of the change in the emotional state exceeds a preset threshold, it is determined to be a significant emotional change.

[0013] As a further limitation of the technical solution of the present invention, the sound field mutation event of the adjacent exhibition area refers to the situation in which the sound field energy level changes significantly compared with the previous exhibition area after the visitor enters the next exhibition area; wherein, when the average sound field energy detected within a preset time period after the visitor enters the next exhibition area increases by more than a preset amount compared with the average sound field energy detected within a preset time period before the visitor leaves the previous exhibition area, it is determined to be a sound field mutation event.

[0014] As a further limitation of the technical solution of this invention embodiment, the steps of sequentially analyzing each sample, determining the corresponding visitor feedback data for each sample, and calculating the visitor group's understanding of the pre-exhibition explanation content based on the feedback data include:

[0015] Analyze each sample to obtain feedback data from a preset number of visitors in each sample during a preset pre-entry time period in the next exhibition area;

[0016] Based on the feedback data, determine the extent to which visitors have grasped the content of the previous explanations for the next exhibition area, and quantify it as a comprehension index;

[0017] The feedback data specifically includes any one or more of the following: voice feedback information from visitors after entering the next exhibition area, operation interaction records, and stay duration data.

[0018] As a further limitation of the technical solution of this invention, the steps of further extracting the intensity index of emotional change of each sample and its corresponding intensity value, and calculating the comprehensive effectiveness score of each sample based on the correspondence between comprehension and intensity of emotional change include:

[0019] Each sample is analyzed to determine one or more of the following: changes in facial expression, tone of voice, and body movement characteristics of a preset number of tourists in the front and back exhibition areas. The changes are then quantified to obtain the corresponding emotional change intensity index and its intensity value.

[0020] The intensity values ​​of various emotional changes are normalized, and weight values ​​are assigned according to the relative magnitude of each emotional change intensity value in the overall intensity distribution. The higher the intensity value of the emotional change, the greater the corresponding weight value.

[0021] The weight values ​​of each emotion change intensity index are combined with the sample's comprehension level to calculate the overall validity score of each sample.

[0022] As a further limitation of the technical solution of this invention, when calculating the comprehensive validity score of the sample, the weight value corresponding to the intensity index of emotional change of the sample is multiplied by the degree of comprehension.

[0023] As a further limitation of the technical solution of the present invention, the consistent environmental background means that the exhibition areas corresponding to each sample are consistent in key environmental elements such as acoustic environment, lighting conditions, spatial layout, and the type and specifications of exhibits, and the variation range of each environmental element is within the preset allowable deviation range.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] This invention proposes an intelligent push notification method based on a linkage mechanism of "significant emotional change - abrupt sound field events - delay in delivery of explanation content" by integrating multi-source sensory data, achieving dynamic adaptive optimization of the scenic area's explanation system. Unlike existing push notification methods that rely solely on location or time triggers, this invention introduces "the rate of increase in sound field energy in adjacent exhibition areas" as a trigger feature. Combined with the facial expressions, voice tones, and movement characteristics of the visitor group, it identifies the phenomenon of group emotional resonance caused by abrupt changes in environmental acoustics. Furthermore, by establishing a correlation model of "emotional change intensity index - comprehension index," it comprehensively scores the effectiveness of the samples. Based on this, the system selects the optimal reference sample to determine the delay in delivery of explanation content, achieving precise delivery of explanation content during the attention recovery period.

[0026] This invention has the advantages of strong interference adaptability, high accuracy in capturing emotional responses, and dynamic transferability of push strategies, which significantly improves the content reception effect and immersive experience of tourists in multi-exhibition area scenarios, and has broad application prospects in cultural exhibitions and smart tours. Attached Figure Description

[0027] Figure 1 A flowchart of the method provided in the embodiments of the present invention;

[0028] Figure 2 This is a flowchart illustrating the tourist comprehension calculation process in the method provided in this embodiment of the invention;

[0029] Figure 3 This is a flowchart illustrating the calculation of the comprehensive effectiveness score in the method provided in this embodiment of the invention. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0031] Figure 1 A flowchart of the method provided by an embodiment of the present invention is shown.

[0032] Specifically, a method for automatically pushing scenic spot explanation content that integrates multi-source perception includes the following steps:

[0033] Step S100: When a significant emotional change is detected in the visitor group in the target exhibition area, determine whether the significant emotional change is caused by a sound field mutation event in an adjacent exhibition area. If it is determined to be so, obtain the sound field energy rise rate of the sound field mutation event.

[0034] The significant emotional change refers to a sudden change in the emotional state of the tourist group identified based on multi-source sensory data. The multi-source sensory data includes at least one or more of the following: changes in facial expressions, changes in voice tone, and changes in body movement characteristics. When the magnitude of the change in the emotional state exceeds a preset threshold, it is determined to be a significant emotional change.

[0035] The sonic field mutation event in adjacent exhibition areas refers to a situation where the sonic field energy level changes significantly after a visitor enters a subsequent exhibition area compared to the previous exhibition area. Specifically, a sonic field mutation event is defined as an event where the average sonic field energy detected within a preset time period after a visitor enters a subsequent exhibition area increases by more than a preset amount compared to the average sonic field energy detected within a preset time period before the visitor leaves the previous exhibition area.

[0036] In this embodiment of the invention, the target exhibition area can typically be an exhibition hall, experience cabin, or themed room with an independent physical space and enclosed acoustic boundaries, such as the "Ancient Civilization Exhibition Hall" in a history museum, the "Energy Interaction Zone" in a science and technology museum, or the "Immersive Image Hall" in an art museum. These spaces are relatively enclosed, with clear sound wave reflection paths and easily detectable sound pressure changes, providing a reliable physical basis for sound field energy monitoring. To ensure the statistical stability of emotion perception data, the size of the visitor group should reach a preset minimum threshold, such as no less than 5 people or the minimum number of people the system sets for recognition, to avoid misjudgments caused by individual emotional biases.

[0037] Changes in sound field energy can be caused by a variety of factors, such as the activation of high-power sound effects devices in adjacent exhibition areas, the playback of high sound pressure videos, the operation of large interactive mechanical equipment, or the amplification of reflected sound due to crowd gathering. When these events occur, the sound field energy level will rise significantly in a short period of time, forming a perceptible abrupt change in the sound field.

[0038] Significant emotional changes are often the result of multiple factors, including not only changes in the sound field but also factors such as lighting, spatial density, dynamic visuals of exhibits, and changes in ambient temperature and humidity. This invention focuses on studying changes in sound field energy because acoustic energy parameters are more objectively measurable, and their changes have the most direct and short-term significant impact on visitors' physiological and emotional responses (such as surprise, tension, and attention shift). This allows for the effective elimination of interference from other non-acoustic factors in subsequent algorithm modeling, thereby improving the stability of trigger judgments.

[0039] The identification of significant emotional changes can be achieved through multi-source perception technology. The system integrates data from facial expression recognition, speech intonation analysis, and body movement monitoring, utilizing emotion recognition algorithms such as Convolutional Neural Networks (CNN) or Long Short-Term Memory Networks (LSTM) to assess the emotional state of tourist groups in real time. Facial images can be acquired via cameras and processed by existing emotion recognition models; speech intonation features can be extracted using Mel-frequency cepstral coefficients and combined with an emotion classification model; and body movement features can be detected using inertial sensors or posture recognition algorithms. These technologies all fall within the scope of existing mature emotion recognition technologies.

[0040] Sound field energy can be acquired using sound pressure level sensors or array microphones placed within the exhibition area. The system collects sound pressure data at different time periods and obtains the average sound field energy value through weighted averaging. To determine whether a sound field abrupt change event has occurred, a preset amplitude threshold needs to be set as a judgment benchmark. When the average sound field energy value of the subsequent exhibition area increases by more than the average value of the previous exhibition area compared to the previous exhibition area, it indicates that a sudden change has occurred in the environmental sound field. The threshold value can be determined based on the acoustic structure and background noise statistics of the exhibition hall, ensuring that the judgment result reflects the actual sound field changes while avoiding false triggers caused by slight fluctuations, thereby ensuring the reliability of the system response.

[0041] Furthermore, the method for automatically pushing scenic spot explanation content that integrates multi-source perception also includes the following steps:

[0042] Step S200: Retrieve several samples from the database that have a consistent environmental background, exhibit significant emotional changes caused by abrupt changes in the sound field of adjacent exhibition areas, and have a consistent rate of increase in sound field energy.

[0043] The consistency of the environmental background means that the exhibition areas corresponding to each sample are consistent in key environmental elements such as acoustic environment, lighting conditions, spatial layout, and the type and specifications of exhibits, and the variation of each environmental element is within the preset allowable deviation range.

[0044] In this embodiment of the invention, the database can originate from a multi-source data acquisition system. This system can consist of data collected and accumulated over a long period by the target exhibition hall's own sensing devices, or it can be expanded by interconnecting and sharing data with other exhibition halls that have similar exhibition themes or spatial layouts. By aggregating sample data across exhibition halls, a larger-scale statistical foundation can be formed, providing rich data support for subsequent model calculations and feature matching, thereby improving the accuracy of the correlation between emotional changes and sound field characteristics. The data types in the database should include multi-dimensional perceptual information and visitor behavior data, specifically including: acoustic environment parameters of the exhibition area (such as sound pressure level, reverberation time, sound field energy change curve), lighting condition parameters (illuminance, color temperature, change frequency), spatial layout characteristics (area, exhibit arrangement, viewing path length), exhibit type and specification information (size, material, theme attributes), and multi-source perceptual data of visitors within the corresponding time period (facial expression changes, voice tone characteristics, movement trajectory, dwell time and interaction frequency, etc.); and visitor feedback data.

[0045] Setting a "consistent environmental background" constraint during sample selection ensures the comparability of selected data samples and prevents interference from external factors in the analysis of the relationship between sound field changes and emotional changes. Since absolute consistency in real-world environments is difficult to achieve, this invention employs a matching method with a controllable deviation range. This allows for fluctuations in various environmental parameters within a preset allowable deviation range; as long as the overall environmental conditions are similar, the environmental background is considered consistent. This ensures both the scientific rigor of data matching and a sufficient sample size, avoiding overly stringent selection criteria that could lead to sample scarcity.

[0046] The significance of sample selection lies in establishing a unified comparative benchmark. By selecting samples with similar environmental backgrounds, sound field variation characteristics, and emotional response characteristics, biases caused by differences in exhibition area themes, spatial variations, or environmental noise can be eliminated. This makes the subsequently extracted emotional change intensity and comprehension indicators more representative and reliable, resulting in more stable and repeatable analytical results when calculating the comprehensive effectiveness score and determining the push delay parameters. This selection process essentially standardizes the data samples, enabling the system to adaptively push information based on a unified logic even when facing different exhibition areas and visitor groups.

[0047] Furthermore, the method for automatically pushing scenic spot explanation content that integrates multi-source perception also includes the following steps:

[0048] Step S300: Analyze each sample in sequence, determine the tourist feedback data corresponding to each sample, and calculate the tourist group's understanding of the pre-exhibition explanation content based on the feedback data.

[0049] Specifically, Figure 2 A flowchart for calculating tourist comprehension is shown.

[0050] The process of analyzing each sample sequentially to determine the corresponding visitor feedback data and calculating the visitor group's understanding of the pre-exhibition explanations based on the feedback data includes the following steps:

[0051] Step S301: Analyze each sample and obtain feedback data of a preset number of tourists in each sample during a preset pre-entry time period after entering the next exhibition area.

[0052] Step S302: Determine the extent to which tourists have grasped the previous explanation content of the next exhibition area based on the feedback data, and quantify it as a comprehension index;

[0053] Step S303, the feedback data specifically includes any one or more of the following: voice feedback information of tourists after entering the next exhibition area, operation interaction records, and stay duration data.

[0054] In this embodiment of the invention, when quantifying the tourist's understanding of the content explained in the previous section of the next exhibition area in step S302, a comprehension index calculation model can be established based on multi-dimensional tourist feedback data. Specifically, within a preset pre-exhibition period after the tourist enters the next exhibition area, the system obtains the tourist's real-time feedback information through voice analysis, behavioral data statistics, and interaction record detection, and calculates the tourist's understanding of the content explained in the previous exhibition area by combining it with a historical sample model.

[0055] Voice feedback information can be extracted using voice recognition technology to measure the semantic matching rate and response latency when visitors interact with the explanations or answer questions. When visitors can accurately answer questions related to the previous exhibit within a short time, or when many highly relevant keywords appear in discussions, it indicates a high level of understanding. Interaction records can include the frequency of clicks, replay behavior, or self-search behavior on multimedia terminals. Higher interaction frequency and more concentrated operation time related to the previous exhibit indicate stronger memory retention. Dwell time data reflects the degree of immersion of visitors in the initial areas; if visitors spend a longer time in the initial areas and their observation behavior is stable, the corresponding comprehension index value is higher.

[0056] The system normalizes the aforementioned multidimensional features and then calculates a weighted average based on the relevance of each feature to the level of understanding in the sample data, thus obtaining the overall comprehension index for the tourist group. For example, it can be calculated as follows:

[0057] The comprehension index = a × semantic matching rate + b × interaction frequency + c × standardized dwell time, where a, b, and c are feature weights automatically trained by the system based on historical samples.

[0058] This level of mastery reflects the level of knowledge retention and cognitive continuity of visitors regarding the content explained in the previous exhibition area after being affected by a sudden change in the sound field of an adjacent exhibition area. Since sudden changes in the sound field can cause short-term distraction or psychological disturbance to visitors, thus affecting their information processing efficiency and memory retention, identifying the level of mastery is significant in objectively assessing the impact of this disturbance on learning and comprehension. By calculating the comprehension index, the system can determine whether visitors experience cognitive biases due to changes in the external sound field, thereby providing an adaptive adjustment basis for subsequent explanations. This allows the explanation strategy to be automatically optimized based on the visitor's current comprehension level, achieving personalized and dynamically adjusted explanation content delivery.

[0059] Furthermore, the method for automatically pushing scenic spot explanation content that integrates multi-source perception also includes the following steps:

[0060] Step S400: Further extract the intensity index of emotional change and its corresponding intensity value for each sample, and calculate the comprehensive effectiveness score for each sample based on the correspondence between comprehension and intensity of emotional change.

[0061] Step S500: Select the sample with the highest comprehensive effectiveness score as the reference sample, obtain the push delay parameter of the explanation content corresponding to the reference sample, and use the push delay parameter as the standard for automatic push of explanation content to subsequent visitors in the target exhibition area.

[0062] Specifically, Figure 3 A flowchart for calculating the comprehensive effectiveness score is shown.

[0063] The process of further extracting the intensity index of emotional change and its corresponding intensity value for each sample, and calculating the comprehensive effectiveness score for each sample based on the correlation between comprehension and intensity of emotional change, specifically includes the following steps:

[0064] Step S401: Analyze each sample, determine any one or more of the changes in facial expressions, voice tone, and body movement characteristics of a preset number of tourists in the front and back exhibition areas, and quantify the changes to obtain the corresponding emotional change intensity index and its intensity value.

[0065] Step S402: Normalize the intensity values ​​of various emotional changes and assign weight values ​​according to the relative magnitude of each emotional change intensity value in the overall intensity distribution. The higher the intensity value of the emotional change, the greater the corresponding weight value.

[0066] Step S403: Combine the weight values ​​of each emotion change intensity index with the comprehension of the sample to calculate the comprehensive validity score of each sample.

[0067] When calculating the overall validity score of the sample, the weight value corresponding to the intensity of the emotional change index of the sample is multiplied by the comprehension level.

[0068] In this embodiment of the invention, in step S300, directly selecting the sample with the highest comprehension index as the reference sample is feasible in engineering, because this sample has the best memory retention and cognitive continuity of the previous explanation content of the exhibition area during the collection period, and can provide a usable time delay reference for the explanation content push. However, the above method is not scientific enough. The inventors have found that even under the constraint of consistent environmental background, there may still be subtle differences between different exhibition areas. The abrupt change events of the sound field in adjacent exhibition areas are not exactly the same in terms of amplitude, spectrum composition and duration, resulting in inconsistent emotional changes among different visitor groups. If the intensity value of the emotional change intensity index is large and the comprehension index is still high, it indicates that the sample still shows good cognitive stability under strong interference, and is more representative of a robust push strategy; conversely, if the intensity value of the emotional change intensity index is small and the comprehension index is high, the sample is not representative of interference and is less persuasive as a reference sample. Therefore, in step S400, the present invention introduces a comprehensive validity score to consider the intensity values ​​of both the comprehension index and the emotional change intensity index, thereby improving the rationality and robustness of the reference sample selection.

[0069] Step S401 is used to extract quantifiable emotional change intensity indicators and their intensity values ​​for each sample. The system analyzes each sample, determines one or more of the following: facial expression changes, voice tone changes, and body movement characteristics changes of a preset number of tourists in the front and back exhibition areas. These characteristics are then quantified to obtain emotional change intensity indicators and their intensity values ​​that characterize the amplitude of group emotional fluctuations. This step ensures that the emotional level measurements of samples from different sources are comparable and weighted.

[0070] Step S402 maps the intensity values ​​of different emotion channels to weight values ​​under a unified dimension. The system normalizes the intensity values ​​of various emotion changes and assigns weight values ​​according to the relative magnitude of each emotion change intensity value in the overall intensity distribution. The higher the intensity value of the emotion change, the larger the corresponding weight value. This design allows samples with stronger interference and greater fluctuations to gain higher discourse power in the comprehensive score under the same consistent environmental background, thus more closely aligning with the goal of "maintaining comprehensibility even in highly interfering situations."

[0071] Step S403 is used to generate a comprehensive effectiveness score. The system combines the weight values ​​corresponding to each emotion change intensity index with the sample's comprehension index to calculate the comprehensive effectiveness score for each sample. When calculating the comprehensive effectiveness score of a sample, the weight value corresponding to the sample's emotion change intensity index is multiplied by the comprehension index. Through the sequential processing of steps 401-403 above, this invention achieves a coupled measurement of "emotional interference intensity" and "cognitive retention level." Compared to a method that only ranks samples based on the comprehension index, it can prioritize the selection of samples that still obtain higher comprehension indices under stronger acoustic interference, exhibiting better representativeness, transferability, and anti-interference capabilities. This comprehensive selection mechanism reflects the synergistic utilization of emotion change intensity and comprehension indices, demonstrating novelty and practicality.

[0072] Step S500 is used for the transfer and application of the reference sample. The system selects the sample with the highest comprehensive effectiveness score as the reference sample, obtains the corresponding explanation content push delay parameter, and uses this push delay parameter as a standard for the automatic push of explanation content to subsequent visitors in the target exhibition area. The reference sample selected based on the comprehensive effectiveness score can provide more robust control over the timing of explanation content push in actual scenarios with the same sound field energy rise rate and environmental background. This ensures that the explanation content reaches visitors at a reasonable delay point after a sudden sound field event, reducing push loss when information is obscured or attention has not returned, and increasing the probability of the explanation content being perceived and understood.

[0073] The beneficial effects of this invention are reflected in the following aspects. First, in trigger recognition, by linking significant emotional changes with abrupt changes in the sound field of adjacent exhibition areas, and combining this with the sound field energy rise rate, quantifiable and reproducible trigger conditions are achieved, reducing false triggers caused by non-acoustic factors. Second, in sample selection, by constraining consistent environmental backgrounds and consistent sound field energy rise rates, the comparability and statistical power of cross-venue and cross-exhibition area data comparisons are improved. Third, in reference sample selection, by coupling the comprehension index with the intensity value of the emotional change intensity index through a comprehensive effectiveness score, samples with high comprehension indices even under strong interference are prioritized, significantly enhancing the robustness of the explanation content delivery strategy. Fourth, in execution control, by standardizing the application of explanation content delivery delay parameters, the system can adaptively deliver information with the optimal delay after abrupt changes in the sound field, improving the audibility and comprehensibility of the explanation information for visitors. This method is applicable to multi-hall exhibition halls and immersive experience spaces with adjacent exhibition areas. It can be promoted and applied in museums, science and technology museums, art galleries and large-scale themed exhibitions, and has good engineering feasibility and expansion prospects.

[0074] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0075] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0076] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0077] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

[0078] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for automatically pushing scenic spot explanation content integrating multi-source perception, characterized in that, The method includes: When a significant emotional change is detected in the visitor group in the target exhibition area, it is determined whether the significant emotional change is caused by a sudden change in the sound field of an adjacent exhibition area. If it is determined to be so, the rise rate of the sound field energy of the sudden change in the sound field is obtained. The sonic field mutation event in the adjacent exhibition area refers to the situation where the sonic field energy level changes significantly after a visitor enters the next exhibition area compared to the previous exhibition area; in particular, when the average sonic field energy detected within a preset time period after a visitor enters the next exhibition area increases by more than a preset amount compared to the average sonic field energy detected within a preset time period before the visitor leaves the previous exhibition area, it is determined to be a sonic field mutation event. Several samples were retrieved from the database that had a consistent environmental background, exhibited significant emotional changes caused by abrupt changes in the sound field of adjacent exhibition areas, and had a consistent rate of increase in sound field energy. Each sample was analyzed sequentially to determine the corresponding visitor feedback data, and the visitor group's understanding of the pre-exhibition explanation content was calculated based on the feedback data. Further, the intensity index of emotional change and its corresponding intensity value were extracted for each sample, and the comprehensive effectiveness score of each sample was calculated based on the correspondence between comprehension and intensity of emotional change. The sample with the highest overall effectiveness score is selected as the reference sample. The push delay parameter of the explanation content corresponding to the reference sample is obtained, and this push delay parameter is used as the standard for automatic push of explanation content to subsequent visitors in the target exhibition area.

2. The method for automatically pushing scenic spot explanation content by integrating multi-source perception as described in claim 1, characterized in that, The significant emotional change refers to a sudden change in the emotional state of the tourist group identified based on multi-source sensory data. The multi-source sensory data includes at least one or more of the following: changes in facial expressions, changes in voice tone, and changes in body movement characteristics. When the magnitude of the change in the emotional state exceeds a preset threshold, it is determined to be a significant emotional change.

3. The method for automatically pushing scenic spot explanation content by integrating multi-source perception as described in claim 1, characterized in that, The steps involved in analyzing each sample sequentially, determining the corresponding visitor feedback data for each sample, and calculating the visitor group's understanding of the pre-exhibition explanations based on the feedback data include: Analyze each sample to obtain feedback data from a preset number of visitors in each sample during a preset pre-entry time period in the next exhibition area; Based on the feedback data, determine the extent to which visitors have grasped the content of the previous explanations for the next exhibition area, and quantify it as a comprehension index; The feedback data specifically includes any one or more of the following: voice feedback information from visitors after entering the next exhibition area, operation interaction records, and stay duration data.

4. The method for automatically pushing scenic spot explanation content by integrating multi-source perception as described in claim 2, characterized in that, The steps to further extract the intensity index of emotional change and its corresponding intensity value for each sample, and to calculate the comprehensive validity score for each sample based on the correlation between comprehension and intensity of emotional change, include: Each sample is analyzed to determine one or more of the following: changes in facial expression, tone of voice, and body movement characteristics of a preset number of tourists in the front and back exhibition areas. The changes are then quantified to obtain the corresponding emotional change intensity index and its intensity value. The intensity values ​​of various emotional changes are normalized, and weight values ​​are assigned according to the relative magnitude of each emotional change intensity value in the overall intensity distribution. The higher the intensity value of the emotional change, the greater the corresponding weight value. The weight values ​​of each emotion change intensity index are combined with the sample's comprehension level to calculate the overall validity score of each sample.

5. The method for automatically pushing scenic spot explanation content by integrating multi-source perception as described in claim 4, characterized in that, When calculating the overall validity score of the sample, the weight value corresponding to the intensity of the emotional change index of the sample is multiplied by the comprehension level.

6. The method for automatically pushing scenic spot explanation content by integrating multi-source perception as described in claim 1, characterized in that, The consistency of the environmental background means that the exhibition areas corresponding to each sample are consistent in key environmental elements such as acoustic environment, lighting conditions, spatial layout, and the type and specifications of exhibits, and the variation of each environmental element is within the preset allowable deviation range.

Citation Information

Patent Citations

  • Tourism mobile terminal information pushing method based on medial multi-dimensional content expression

    CN102984219A

  • Travel itinerary navigation method, apparatus and device, and storage medium

    CN120509993A