Scenic spot automatic voice explanation method and system based on position awareness

By analyzing users' camera behavior data, the content of the scenic area's automatic voice guide system is dynamically adjusted, solving the problem of mismatch between the guide content and the user's focus of observation, and achieving a personalized and intelligent guide experience.

CN121350299AActive Publication Date: 2026-01-16BEIJING GRAVITY YUHUA FILM & TELEVISION CULTURE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511515444.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-16
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing automated audio guide systems in scenic areas cannot dynamically adjust the semantic level of the explanation based on the user's self-filming behavior, resulting in explanations that are either too superficial or too in-depth, failing to match the user's focus of observation.

Method used

By acquiring users' camera behavior data, analyzing and combining camera engagement and shooting alignment deviation, calculating correction factors, and dynamically adjusting the semantic level of the narration content to match the user's cognitive level.

Benefits of technology

It achieves adaptive matching between the content of the explanation and the user's observation behavior, enhances the personalization and intelligence of the visiting experience, and optimizes the visual capture and shooting composition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121350299A_ABST
    Figure CN121350299A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of scenic spot intelligent explanation, and provides a scenic spot automatic voice explanation method and system based on position awareness, and the method comprises the steps: carrying out the statistical analysis based on the photographing alignment deviation degree corresponding to each exhibition object, and when the photographing alignment deviation degree is in a preset deviation threshold value interval and exceeds a preset specific value, carrying out the voice explanation of the exhibition object. And comprehensively evaluating the alignment deviation degree of each time of shooting so as to calculate a correction factor related to semantic hierarchy correction. According to the invention, by introducing a user shooting behavior analysis mechanism based on location awareness, the conversion of a voice explanation system from static content presentation to dynamic semantic self-adaptive adjustment is realized. After the system detects the shooting behavior of a user on an exhibition object, the shooting alignment deviation degree is calculated, a correction factor is generated according to the shooting alignment deviation degree, and a semantic hierarchy adaptability numerical value of explanation data content is dynamically corrected, so that subsequent explanation content is adaptively matched with a user cognition hierarchy in semantic depth and visual guidance intensity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent tour guide technology for scenic spots, and in particular to a method and system for automatic voice tour guides for scenic spots based on location awareness. Background Technology

[0002] With the development of mobile terminals and voice recognition technology, location-aware automated audio guide systems are becoming increasingly common in scenic spots, museums, and exhibition halls. These systems typically use Bluetooth beacons, GPS, or indoor positioning modules to identify the visitor's location and automatically play audio guide content corresponding to that location, replacing human guides. While existing audio guide methods can improve visitor efficiency and information coverage to some extent, they still remain at the level of static content matching. That is, regardless of the differences in users' interests, observation methods, or interactive behaviors, the system outputs consistent audio guide content, lacking the ability to dynamically respond to individualized user experiences.

[0003] In recent years, some improvement solutions have attempted to incorporate user behavior data, such as dwell time, movement trajectory, or eye tracking, to help optimize the order in which explanatory content is presented. However, these solutions often only reflect the user's area of ​​focus at a macro level and cannot reflect the user's active observation and photography behaviors during the visit. For example, when a user takes multiple photos of exhibits, adjusts the framing, or changes the composition, existing systems cannot identify the underlying visual attention patterns behind these behaviors, thus failing to create a semantic fit between the explanatory content and the user's depth of observation.

[0004] This demonstrates a technical flaw in existing automated audio guide technologies for scenic spots: a disconnect between the content and the user's actual perception. The system cannot dynamically adjust the semantic level of the explanation based on the user's autonomous photography behavior of exhibits, resulting in explanations that are either too superficial to arouse interest or too in-depth to capture the user's current focus. Summary of the Invention

[0005] The purpose of this invention is to provide a location-aware automatic voice explanation method and system for scenic spots, aiming to solve the problems mentioned in the background art.

[0006] This invention is implemented as follows: a location-aware automatic voice narration method for scenic spots, the method comprising: Acquire camera behavior data from users who enter the scenic area's exhibition hall and continuously traverse multiple exhibition areas, and conduct a comprehensive analysis of users' camera engagement. When the overall camera engagement exceeds a preset threshold, the camera behavior data is analyzed to identify the degree of alignment deviation of the user's shooting of the key display area of ​​the exhibit in the previous display areas of each exhibit. Statistical analysis is performed on the shooting alignment deviation of each exhibit. When the proportion of shooting alignment deviation within the preset deviation threshold range exceeds the preset ratio, a comprehensive evaluation is performed on the shooting alignment deviation of each time to calculate the correction factor related to semantic level correction. After determining that a user has entered a new exhibit display area based on location awareness, the semantic level adaptability value of the preset explanatory data content corresponding to the exhibit is obtained; The semantic level adaptability value is dynamically corrected based on the correction factor to obtain a new semantic level adaptability value. Based on the new semantic level adaptability value, the explanation data content that is adapted to the user's cognitive level is matched from the voice explanation database.

[0007] As a further limitation of the technical solution of the present invention, the comprehensive camera engagement refers to a comprehensive characteristic value determined based on the professionalism of the camera equipment, the duration of composition adjustment, the number of shots, and the frequency of changes in the imaging parameters of the equipment when the user performs camera operations in the display area of ​​the exhibit. It is used to characterize the user's attention and operational engagement in the camera behavior.

[0008] As a further limitation of the technical solution of this invention, when the overall camera engagement exceeds a preset threshold, the step of analyzing camera behavior data to identify the user's alignment deviation in shooting the key display area of ​​the exhibit within the previous exhibition areas includes: Based on the camera behavior data, the user's camera engagement level in each exhibit display area is calculated, and the camera engagement level is statistically averaged to obtain the comprehensive camera engagement level; If the overall camera engagement is greater than a preset threshold, the camera behavior data is analyzed to determine the user's framing posture information when shooting the exhibit. The framing posture information is then compared with a preset standard shooting reference model to calculate the user's shooting alignment deviation within the corresponding exhibit display area. The shooting alignment deviation is used to characterize the accuracy with which the user captures the key content of the exhibit during the shooting process.

[0009] As a further limitation of the technical solution of the embodiment of the present invention, the framing posture information includes the spatial posture parameters of the shooting device, the framing direction angle, the shooting distance and the field of view coverage. The preset standard shooting reference model is a standard perspective model established based on the spatial three-dimensional structural data of the exhibits in the scenic exhibition hall and the characteristics of the key display areas. It is used to represent the standard framing posture information of the exhibits under ideal shooting conditions.

[0010] As a further limitation of the technical solution of this invention embodiment, the calculation of the shooting alignment deviation specifically includes: The spatial parameters of the user's framing posture information are compared with those of the standard framing posture information in the standard shooting reference model, and the differences in framing direction angle, shooting distance, and field of view coverage are calculated. The shooting alignment deviation is determined based on the weighted result of the differences of each parameter, which is used to characterize the accuracy of the user's alignment of the key display area of ​​the exhibit during the shooting process.

[0011] As a further limitation of the technical solution of this invention, the step of statistically analyzing the shooting alignment deviation of each exhibit, and comprehensively evaluating the shooting alignment deviation of each time when the proportion of shooting alignment deviation within a preset deviation threshold range exceeds a preset ratio, to calculate the correction factor related to semantic level correction, includes: Obtain the alignment deviation of each photographed exhibit by the user, and determine whether each alignment deviation falls within the preset deviation threshold range; When the proportion of shooting alignment deviations falling within the preset deviation threshold range exceeds a preset ratio, all shooting alignment deviations are statistically averaged, and combined with a preset correction intensity coefficient, a correction factor related to semantic level correction is calculated.

[0012] As a further limitation of the technical solution of the present invention, the preset deviation threshold range refers to the deviation range that can be accepted for alignment error in the key display area of ​​the exhibit. It represents the shooting deviation range in which the user can capture the approximate range of the important display area of ​​the exhibit but fails to accurately align with its key content.

[0013] As a further limitation of the technical solution of this invention, the steps of dynamically correcting the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and matching and generating explanation data content adapted to the user's cognitive level from the voice explanation database according to the new semantic level adaptability value include: A correction factor is used to dynamically adjust the semantic level adaptability value of the preset explanatory data content to obtain a new semantic level adaptability value. Access the audio guide database of the scenic area's exhibition hall and extract the set of audio guide data content corresponding to the current exhibit display area; Based on the new semantic level adaptability value, the explanatory data content with matching semantic level is selected from the explanatory data content set, and the explanatory data content is output as voice so that the explanatory content is adapted to the user's cognitive level.

[0014] A location-aware automated audio guide system for scenic spots, the system comprising: The data acquisition module is used to acquire the camera behavior data of users who enter the scenic area exhibition hall and continuously pass through multiple exhibition areas, and to conduct a comprehensive analysis of users' camera engagement. The deviation recognition module is used to analyze camera behavior data when the overall camera engagement exceeds a preset threshold in order to identify the user's alignment deviation when shooting the key display area of ​​the exhibit in the previous exhibition areas. The correction factor calculation module is used to perform statistical analysis based on the shooting alignment deviation of each exhibit. When the proportion of shooting alignment deviation within the preset deviation threshold range exceeds the preset ratio, the shooting alignment deviation of each shooting is comprehensively evaluated to calculate the correction factor related to semantic level correction. The semantic level acquisition module is used to acquire the semantic level adaptability value of the preset explanatory data content corresponding to the exhibit after determining that the user has entered a new exhibit display area based on location awareness. The content matching module is used to dynamically correct the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and then match and generate narration data content that is adapted to the user's cognitive level from the voice narration database based on the new semantic level adaptability value.

[0015] As a further limitation of the technical solution of the present invention, the comprehensive camera engagement refers to a comprehensive characteristic value determined based on the professionalism of the camera equipment, the duration of composition adjustment, the number of shots, and the frequency of changes in the imaging parameters of the equipment when the user performs camera operations in the display area of ​​the exhibit. It is used to characterize the user's attention and operational engagement in the camera behavior.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention transforms the voice narration system from static content presentation to dynamic semantic adaptive adjustment by introducing a location-aware user shooting behavior analysis mechanism. After detecting the user's camera behavior on exhibits, the system calculates the shooting alignment deviation and generates a correction factor accordingly. This dynamically adjusts the semantic level adaptability of the narration data, ensuring that subsequent narration content adaptively matches the user's cognitive level in terms of semantic depth and visual guidance intensity. This mechanism forms a closed-loop process of "behavioral perception—semantic adjustment—cognitive feedback," which not only improves the consistency between the narration content and the user's observation behavior but also subtly optimizes visual capture and shooting composition without interfering with user operations.

[0017] This invention effectively solves the problem of the disconnect between the narration content and user perception (autonomous camera aspect) in traditional scenic area automatic guide systems, and realizes a dynamic narration experience that emphasizes both personalization and intelligence, which has high practical value and promotion prospects. Attached Figure Description

[0018] Figure 1 A flowchart of the method provided in the embodiments of the present invention; Figure 2 This is a flowchart illustrating the method for identifying shooting alignment deviation based on comprehensive camera engagement in the embodiments of the present invention. Figure 3 This is a flowchart illustrating the calculation of semantic hierarchy correction factors based on shooting alignment deviation in the method provided in this embodiment of the invention; Figure 4 This is a flowchart illustrating the method for dynamically adjusting the semantic hierarchy adaptation of the explanation content based on a correction factor in the embodiments of the present invention. Figure 5 The application architecture diagram of the system provided in the embodiments of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0020] Figure 1 A flowchart of the method provided by an embodiment of the present invention is shown.

[0021] Specifically, a location-aware automatic voice narration method for scenic spots includes the following steps: Step S100: Obtain the camera behavior data of users who enter the scenic area exhibition hall and continuously pass through multiple exhibition areas, and conduct a comprehensive camera engagement analysis of the users.

[0022] The comprehensive camera engagement refers to a comprehensive characteristic value determined based on the professionalism of the camera equipment, the duration of composition adjustment, the number of shots, and the frequency of changes in the equipment's imaging parameters when a user performs camera operations within the exhibition area. It is used to characterize the user's level of attention and operational engagement in the camera behavior.

[0023] In this embodiment of the invention, the scenic area exhibition hall can be a cultural exhibition venue with fixed or temporary exhibition space, including museums, science and technology museums, art galleries, memorial halls, special exhibition halls, heritage site exhibition halls, intangible cultural heritage experience halls, corporate exhibition halls, themed exhibition halls in theme parks, display points in historical districts, and immersive performance spaces, etc. The aforementioned spaces can be enclosed exhibition areas within a single building, semi-open exhibition areas such as connecting corridors, lobbies, mezzanines, and sunken plazas, or independent outdoor exhibition tents or enclosed exhibition areas within a park. As long as they provide a relatively clear visitor flow and resting area for viewing and explaining the exhibits, they all fall under the category of scenic area exhibition halls in this embodiment of the invention.

[0024] The scenic area's exhibition hall is equipped with an audio guide database to store multiple sets of audio guide content corresponding to different exhibit display areas. Each exhibit display area can have multiple types of audio guides, with differences in language structure, information depth, and visual description precision. For example, the audio guide content for the same exhibit can include basic guided tours, explanations of key areas, and detailed composition guidance, to meet the diverse needs of different users in their viewing and photography behaviors. The system assigns a semantic level adaptability value to each type of audio guide content in the audio guide database, representing the comprehensive level of the audio guide content in terms of semantic complexity, visual guidance intensity, and information refinement.

[0025] Understandably, the semantic level fit value not only characterizes the semantic depth and information complexity of the explanatory content, but also determines the degree to which the explanatory content guides visual perception in information expression. When the semantic level fit value is low, the explanatory content tends to be more holistic and general, emphasizing the thematic background and general knowledge of the exhibits, enabling users to form a perceptual impression within a broader scope. When the semantic level fit value is high, the description of the explanatory content is more structured and detailed in semantic organization, and its language implicitly guides the focus on key display areas, compositional levels, or spatial features of the exhibits, thereby influencing the distribution of users' visual attention at the cognitive level.

[0026] The delineation of exhibition areas can be determined based on a combination of location awareness and exhibition elements. Specifically, a spatial enclosure matching the physical occupancy of each exhibit is established: for exhibits inside display cases, a two-dimensional polygonal enclosure parallel to the front edge of the display case can be used, with several equidistant buffer zones in front; for three-dimensional sculptures or large objects, a three-dimensional envelope can be established around the outer edge of the exhibit, with fan-shaped viewing zones and minimum safety distance zones set in the main viewing directions; for wall-mounted paintings or calligraphy, a rectangular viewing corridor can be set along the normal direction of the work, with lateral buffer zones superimposed. The above-mentioned enclosures are calibrated using location-aware base stations such as Bluetooth beacons, ultra-wideband positioning, Wi-Fi RTT, infrared door magnets, or visual anchors. The system uses entry / stay / leave events as criteria to determine whether a user is within a specific exhibit's display area.

[0027] The video recording behavior data is primarily acquired by smart cameras installed within the exhibition hall, with user terminals serving only as a supplementary information source. Several smart cameras are deployed in each exhibit display area to detect the shooting behavior of users entering that area. These smart cameras are equipped with personnel detection and posture recognition functions, used to identify user actions such as framing, raising the camera, and shutter triggering without capturing the content being photographed, and to record corresponding time and spatial posture information. By detecting the user's dwell time within the exhibit display area, the duration of framing, the number of shooting actions, and changes in shooting angle, corresponding video recording behavior data can be generated.

[0028] User terminals (such as mobile phones or tablets) can provide auxiliary data, including timestamps of shutter trigger events, number of camera parameter changes, and device model and type information, which are used to correlate and verify with the behavioral data collected at the exhibition hall. The smart cameras at the exhibition hall and user terminals can be linked and matched via Bluetooth or a wireless network to confirm the same user's shooting behavior events.

[0029] Camera behavior data includes objectively detectable basic data fields, such as: framing start and end times, shutter trigger time, dwell time within the exhibit's display area, range of framing posture angle changes, number of shooting actions, number of camera parameter changes, and device type identification. All of the above data can be directly obtained from the smart camera's posture detection module and the user terminal's system event recording module; it is objectively existing data and does not involve image content analysis.

[0030] The calculation process for the overall camera engagement score is as follows: After detecting that a user is taking photos within a certain exhibition area, the system counts the number of times the user takes photos within that area, the duration of the shot, the number of times the device parameters are adjusted, and the device type rating. These data are then normalized and weighted according to preset weights to obtain the corresponding overall camera engagement score. This overall camera engagement score characterizes the user's level of attention and operational intensity during the photo-taking process, providing a basis for identifying subsequent alignment deviations.

[0031] Furthermore, the location-aware automatic voice narration method for scenic spots also includes the following steps: Step S200: When the overall camera engagement exceeds a preset threshold, analyze the camera behavior data to identify the user's alignment deviation when shooting the key display areas of the exhibits in the previous exhibition areas.

[0032] Specifically, Figure 2 A flowchart is shown to identify the shooting alignment deviation based on the comprehensive camera engagement.

[0033] Specifically, when the overall camera engagement exceeds a preset threshold, the camera behavior data is analyzed to identify the user's alignment deviation when shooting towards key display areas of the exhibits within the previous exhibition areas. This includes the following steps: Step S201: Calculate the user's camera engagement in each exhibit display area based on the camera behavior data, and statistically average the camera engagement to obtain a comprehensive camera engagement score. Step S202: Determine whether the overall camera engagement is greater than a preset threshold. If so, parse the camera behavior data, determine the user's framing posture information when shooting the exhibit, and compare the framing posture information with a preset standard shooting reference model to calculate the user's shooting alignment deviation within the corresponding exhibit display area. The shooting alignment deviation is used to characterize the accuracy with which the user captures the key content of the exhibit during the shooting process.

[0034] The framing posture information includes the spatial pose parameters of the shooting device, the framing direction angle, the shooting distance, and the field of view coverage; The preset standard shooting reference model is a standard perspective model established based on the spatial three-dimensional structural data of the exhibits in the scenic exhibition hall and the characteristics of the key display areas. It is used to represent the standard framing posture information of the exhibits under ideal shooting conditions.

[0035] The calculation of the alignment deviation during shooting specifically includes: The spatial parameters of the user's framing posture information are compared with those of the standard framing posture information in the standard shooting reference model, and the differences in framing direction angle, shooting distance, and field of view coverage are calculated. The shooting alignment deviation is determined based on the weighted result of the differences of each parameter, which is used to characterize the accuracy of the user's alignment of the key display area of ​​the exhibit during the shooting process.

[0036] In this embodiment of the invention, in step S201, the overall camera engagement score is obtained by statistically averaging the camera engagement scores of each exhibit's display area. This averaging method effectively reduces the impact of occasional fluctuations in individual shooting behavior, making the overall result more stable and reliable, while ensuring comparability between different exhibit display areas. The overall camera engagement score obtained in this way can truly reflect the overall shooting intensity of users throughout the visit, avoiding abnormal deviations caused by excessively short stay times or unexpected interruptions in individual exhibit display areas, thereby improving the accuracy and robustness of subsequent behavior analysis.

[0037] In step S202, when the overall camera engagement exceeds a preset threshold, it indicates that the user's shooting behavior during the visit is relatively proactive and continuous, demonstrating high attention and willingness to take photos. At this point, their shooting behavior characteristics are meaningful for analysis, so the system only activates the shooting alignment deviation identification module under this condition to ensure the stability and representativeness of the analysis sample. If the user's overall camera engagement is below the preset threshold, it usually means that the user only took random photos or temporarily posed for shots; such samples are not valuable for reference, and the system does not perform deviation analysis. The preset threshold can be determined based on the statistical results of historical sample data from the exhibition hall, for example, by taking the 70%–80% quantile of the overall sample distribution as the initial judgment line; it can also be adaptively calibrated based on the exhibition hall type, exhibit density, and visitor flow characteristics to balance trigger rate and analysis accuracy.

[0038] After the overall camera engagement meets the judgment criteria, the system analyzes the camera behavior data and extracts the user's framing posture information. This framing posture information includes the spatial pose parameters of the shooting device, the framing direction angle, the shooting distance, and the field of view coverage. These parameters can be directly obtained through the collaborative perception of the exhibition hall's intelligent camera and the user terminal's sensing module, possessing high observability and physical accuracy. In addition to the above parameters, in some implementations, further extended parameters such as framing stability, camera posture angular velocity, and movement path curvature can be selected to enhance the multi-dimensional expressive capability of the framing posture information.

[0039] The establishment of a standard shooting reference model is based on the spatial three-dimensional structural data of the exhibit and the characteristics of key display areas. For exhibits with three-dimensional model data, the geometric outline, main viewing direction, and center point of key display areas of the model can be directly extracted, and ideal shooting pose parameters can be generated through spatial calibration. For exhibits without a three-dimensional model, a standard shooting reference model can be established based on the display case layout parameters, the planar coordinates of the exhibits, and the viewing path, using regularized spatial calculations or manual calibration. In some embodiments, the system can use a lightweight algorithm based on sample learning to fit the standard shooting samples calibrated by experts to obtain standard framing posture information that more closely matches the actual viewing experience.

[0040] The calculation of shooting alignment deviation involves two steps: parameter comparison and weighted summation. First, the system compares the user's framing posture information with the standard framing posture information in the standard shooting reference model, calculating the framing direction angle difference (reflecting the consistency between the user's shooting direction and the standard direction), shooting distance difference (characterizing the offset of the shooting point position), and field of view coverage difference (reflecting the difference in the coverage ratio of the key display area in the composition). Then, the system assigns preset weights based on the importance of each parameter and sums the differences to obtain the final shooting alignment deviation result. The smaller the shooting alignment deviation value, the higher the alignment between the user's shooting angle and distance and the standard model.

[0041] In an extended implementation, to further improve evaluation accuracy, complex factors such as illumination difference, occlusion rate difference, or visibility coverage difference can be introduced. Illumination difference reflects the deviation in brightness conditions between the user's framing position and the standard shooting position; occlusion rate difference determines the difference between the proportion of exhibits obscured in the captured image and the standard baseline; and visibility coverage difference measures the difference in the visible range of key display areas from the user's perspective. By incorporating these additional factors into the comprehensive calculation with lower weights, the robustness and accuracy of the system in complex environments can be improved while maintaining the simplicity of the main judgment dimensions.

[0042] Furthermore, the location-aware automatic voice narration method for scenic spots also includes the following steps: Step S300: Statistical analysis is performed based on the shooting alignment deviation of each exhibit. When the proportion of shooting alignment deviation within the preset deviation threshold range exceeds the preset ratio, a comprehensive evaluation is performed on the shooting alignment deviation of each time to calculate the correction factor related to semantic level correction.

[0043] Specifically, Figure 3 A flowchart is shown to calculate the semantic hierarchy correction factor based on the shooting alignment deviation.

[0044] The process involves statistical analysis of the alignment deviation for each exhibit. When the proportion of alignment deviations within a preset threshold exceeds a preset ratio, a comprehensive evaluation of each alignment deviation is performed to calculate the correction factor related to semantic level correction. This includes the following steps: Step S301: Obtain the alignment deviation of each photographed exhibit by the user, and determine whether each alignment deviation falls within the preset deviation threshold range. Step S302: When the proportion of shooting alignment deviations falling within the preset deviation threshold range exceeds a preset ratio, the average of all shooting alignment deviations is calculated, and a correction factor related to semantic level correction is obtained by combining the preset correction intensity coefficient.

[0045] The preset deviation threshold range refers to the deviation range that can be accepted for alignment error in the key display area of ​​the exhibit. It represents the shooting deviation range in which the user can capture the approximate range of the important display area of ​​the exhibit but fails to accurately align it with its key content.

[0046] In this embodiment of the invention, in step S300, the system performs statistical analysis based on the shooting alignment deviation of each exhibit to evaluate the user's overall framing stability and shooting accuracy. When the proportion of shooting alignment deviation within a preset deviation threshold exceeds a preset ratio, it indicates that although the user's shooting behavior in multiple exhibit display areas is not completely precise, their shooting orientation is basically stable, and they can capture the approximate composition range of the key display areas of the exhibits, reflecting that the user's shooting ability is in a state of "correctable but not ideal". At this time, if the system still outputs the explanation content according to the original semantic level, it may lead to the explanation information level being too high, and the expression of key content not matching the user's visual attention area, thereby affecting the explanation experience. Therefore, this step establishes a correction factor related to semantic level adaptability correction by statistically analyzing the overall performance of the user's shooting alignment deviation, so that the subsequent explanation content can be adaptively adjusted to achieve dynamic optimization of personalized explanation accuracy.

[0047] The preset deviation threshold range is used to characterize the acceptable alignment error range within the key display area of ​​the exhibit. Its lower limit corresponds to the deviation boundary when the user is completely out of alignment with the key area, and its upper limit corresponds to the upper limit range where the user can capture the main outline of the key display area of ​​the exhibit but fails to accurately align with its core content.

[0048] In step S302, when the proportion of shooting alignment deviations within a preset deviation threshold range exceeds a preset ratio, the system determines that the user's overall shooting deviation behavior is consistent and correctable, and therefore calculates the average of all shooting alignment deviations. The reason for using the average value as a basic component of the correction factor is that the average value can smooth out individual differences in shooting accuracy among different exhibits, extract the overall deviation trend, and avoid over-correction caused by individual excessive deviation samples. At the same time, the average value reflects the average deviation magnitude under stable user conditions, and more objectively represents the user's framing ability.

[0049] To achieve flexible adjustment, this invention incorporates a preset correction intensity coefficient when calculating the correction factor. This correction intensity coefficient controls the magnitude of semantic level adjustment to avoid over-correction that could lead to an imbalance in the explanation hierarchy. The setting of the correction intensity coefficient can be comprehensively determined based on factors such as the complexity of the exhibition content, the density of exhibit information, and the granularity of the explanation hierarchy. For example, in art exhibitions with a deep hierarchy of exhibit information, a higher correction intensity coefficient can be set to amplify the impact of deviations; while in science and technology exhibitions with simple structures and intuitive content, a lower coefficient can be used to maintain the stability of the explanation hierarchy. In practice, the correction intensity coefficient can be obtained through system pre-calibration or automatically updated based on historical explanation feedback data through weighted learning.

[0050] Furthermore, the location-aware automatic voice narration method for scenic spots also includes the following steps: Step S400: After determining that the user has entered a new exhibition area based on location awareness, obtain the semantic level adaptability value of the preset explanation data content corresponding to the exhibition.

[0051] Step S500: Dynamically correct the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and match and generate explanation data content that is adapted to the user's cognitive level from the voice explanation database according to the new semantic level adaptability value.

[0052] Specifically, Figure 4 A flowchart is shown to dynamically adjust the semantic hierarchy of the explanation content based on the correction factor.

[0053] The process of dynamically correcting the semantic level fit value based on a correction factor to obtain a new semantic level fit value, and then matching and generating explanation data content that matches the user's cognitive level from the voice explanation database based on the new semantic level fit value, specifically includes the following steps: Step S401: The semantic level adaptability value of the preset explanatory data content is dynamically adjusted upward by using a correction factor to obtain a new semantic level adaptability value. Step S402: Retrieve the audio guide database of the scenic area exhibition hall and extract the set of audio guide data content corresponding to the current exhibition area. Step S403: Based on the new semantic level adaptability value, select the explanatory data content with matching semantic level from the explanatory data content set, and output the explanatory data content by voice so that the explanatory content is adapted to the user's cognitive level.

[0054] In this embodiment of the invention, the semantic level adaptability value is dynamically corrected based on a correction factor because the aforementioned correction factor is quantified from the user's shooting behavior data and the statistical characteristics of shooting alignment deviation. This factor can directly reflect the individual differences of users in terms of visual capture accuracy, composition stability, and focus performance in key areas. By combining this behavioral feedback quantification parameter with the semantic level adaptability value of the explanation content, the system can achieve adaptive adjustment of the user's perceptual behavior at the semantic level, ensuring that the expression depth, information density, and visual guidance intensity of the explanation content remain dynamically consistent with the user's current observation characteristics. Compared to static explanation matching methods, the correction mechanism adopted in this invention not only considers the user's historical shooting characteristics but also achieves gradual optimization of the semantic level during continuous visits, thereby forming a self-closed-loop structure of "behavioral feedback—semantic adjustment—cognitive matching".

[0055] The advantage of this correction method lies in its ability to implicitly optimize the user's visual attention pattern through dynamic adaptation of the audio narration content without requiring explicit intervention or additional operations. By adjusting the semantic level adaptation value, the system can adjust the distribution of key points in the narration at the semantic level, allowing users to spontaneously adjust their visual attention direction during auditory cognition, thereby improving the accuracy of subsequent shooting composition and observation efficiency. Furthermore, this correction process is scalable and can be flexibly configured according to different types of exhibition halls, exhibit characteristics, or audience group characteristics, achieving a balance between personalized narration and shooting guidance.

[0056] In addition to directly adjusting the semantic level fit value using correction factors, other implementations of the system can also combine time-series weighting, user group behavior clustering analysis, or semantic fuzzy matching to correct the explanation level. For example, an individual shooting feature model can be established based on the user's behavior records from multiple visits, and the semantic level fit value can be adjusted by the dynamic coefficients output by the model; or a group correction template can be generated based on the behavior patterns of similar user groups to perform initial calibration of the semantic level for new users.

[0057] In summary, this invention transforms the voice narration system from passive content matching to proactive behavior optimization by introducing a location-aware user behavior analysis mechanism. This method can adaptively adjust the narration content without requiring additional user intervention, dynamically coordinating the narration with the user's visual perception features, cognitive level, and observation methods, thereby improving the visitor experience and information delivery efficiency in scenic areas and exhibition halls.

[0058] The overall beneficial effects of this invention are as follows: by perceiving the shooting behavior, analyzing the deviation, and dynamically correcting the semantic level, it achieves bidirectional adaptation between voice narration and user behavior characteristics, solving the problem of the disconnect between narration content and user observation behavior in traditional automatic narration systems for scenic spots. This allows the narration content to proactively guide the user's visual attention. This solution can be widely applied in various scenarios such as museums, art galleries, science and technology museums, and outdoor scenic spots, and has good scalability and practical application prospects.

[0059] Furthermore, Figure 5 An application architecture diagram of the system provided in an embodiment of the present invention is shown.

[0060] In another preferred embodiment of the present invention, a location-aware automatic voice guide system for scenic spots includes: The data acquisition module 100 is used to acquire the camera behavior data of users who enter the scenic area exhibition hall and continuously pass through multiple exhibition areas, and to conduct a comprehensive analysis of the user's camera engagement.

[0061] The comprehensive camera engagement refers to a comprehensive characteristic value determined based on the professionalism of the camera equipment, the duration of composition adjustment, the number of shots, and the frequency of changes in the equipment's imaging parameters when a user performs camera operations within the exhibition area. It is used to characterize the user's level of attention and operational engagement in the camera behavior.

[0062] Furthermore, the location-aware automatic voice guide system for scenic spots also includes: The deviation recognition module 200 is used to analyze camera behavior data when the overall camera engagement exceeds a preset threshold in order to identify the user's alignment deviation when shooting the key display area of ​​the exhibit in the previous exhibition areas.

[0063] Furthermore, the location-aware automatic voice guide system for scenic spots also includes: The correction factor calculation module 300 is used to perform statistical analysis based on the shooting alignment deviation of each exhibit. When the proportion of shooting alignment deviation within the preset deviation threshold range exceeds the preset ratio, the alignment deviation of each shooting is comprehensively evaluated to calculate the correction factor related to semantic level correction.

[0064] Furthermore, the location-aware automatic voice guide system for scenic spots also includes: The semantic level acquisition module 400 is used to acquire the semantic level adaptability value of the preset explanatory data content corresponding to the exhibit after determining that the user has entered a new exhibit display area based on location awareness.

[0065] Furthermore, the location-aware automatic voice guide system for scenic spots also includes: The content matching module 500 is used to dynamically correct the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and to match and generate narration data content that is adapted to the user's cognitive level from the voice narration database based on the new semantic level adaptability value.

[0066] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0067] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0068] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0069] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

[0070] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A location-aware-based automatic voice guide method for a scenic spot, characterized in that, The method comprises: acquiring camera shooting behavior data of a user entering an exhibition hall of a scenic spot and continuously passing through a plurality of exhibition display areas, and performing comprehensive camera shooting input degree analysis on the user; when the comprehensive camera shooting input degree is greater than a preset threshold, analyzing the camera shooting behavior data to identify a shooting alignment deviation of the user in a previous exhibition display area with respect to a key display area of an exhibition; based on a shooting alignment deviation corresponding to each exhibition, when a proportion of the shooting alignment deviation within a preset deviation threshold interval exceeds a preset ratio, performing comprehensive evaluation on each shooting alignment deviation to calculate a correction factor related to semantic level correction; after determining that the user enters a new exhibition display area based on location awareness, acquiring a semantic level adaptability value of preset explanation data content corresponding to the exhibition; based on the correction factor, dynamically correcting the semantic level adaptability value to obtain a new semantic level adaptability value, and matching and generating explanation data content adapted to a cognitive level of the user from a voice explanation database according to the new semantic level adaptability value.

2. The method for automatic voice interpretation of scenic spots based on location awareness according to claim 1, characterized in that, The comprehensive camera shooting input degree refers to a comprehensive characteristic value determined based on a professional degree of a camera shooting device, a composition adjustment duration, a shooting frequency, and a frequency of changes in imaging parameters of the device when the user performs camera shooting in the exhibition display area, and is used to represent a degree of attention and an operation input intensity of the user in the camera shooting behavior. 3.The method according to claim 2, wherein, When the comprehensive camera shooting input degree is greater than a preset threshold, the step of analyzing the camera shooting behavior data to identify a shooting alignment deviation of the user in a previous exhibition display area with respect to a key display area of an exhibition comprises: calculating a camera shooting input degree of the user in each exhibition display area based on the camera shooting behavior data, and statistically averaging the camera shooting input degree to obtain a comprehensive camera shooting input degree; determining whether the comprehensive camera shooting input degree is greater than a preset threshold, and if so, analyzing the camera shooting behavior data, determining a framing posture information of the user when shooting the exhibition, and comparing the framing posture information with a preset standard shooting reference model to calculate a shooting alignment deviation of the user in the corresponding exhibition display area; The shooting alignment deviation is used to represent a capture accuracy of the user for key display content of the exhibition in the camera shooting process.

4. The method according to claim 3, wherein, The framing posture information comprises a spatial pose parameter, a framing direction angle, a shooting distance, and a field of view coverage range of the shooting device; The preset standard shooting reference model is a standard view angle model established based on exhibition spatial three-dimensional structure data and key display area features of the exhibition hall of the scenic spot, and is used to represent standard framing posture information of the exhibition under ideal shooting conditions.

5. The method according to claim 3, wherein, The calculation of the shooting alignment deviation specifically comprises: comparing the framing posture information of the user with the standard framing posture information in the standard shooting reference model to calculate a framing direction angle difference value, a shooting distance difference value, and a field of view coverage difference value; determining the shooting alignment deviation based on a weighted result of each parameter difference value, which is used to represent an alignment accuracy of the user for the key display area of the exhibition in the shooting process.

6. The method according to claim 3, wherein, The step of performing statistical analysis on the photograph alignment deviation degrees corresponding to each exhibition object and performing comprehensive evaluation on the photograph alignment deviation degrees when the proportion of the photograph alignment deviation degrees within the preset deviation threshold interval exceeds the preset proportion, to calculate the correction factor related to the semantic level correction, comprises: Obtaining photograph alignment deviation degrees of each photographed exhibition object by the user, and determining whether each photograph alignment deviation degree falls within a preset deviation threshold interval; When the proportion of the photograph alignment deviation degrees within the preset deviation threshold interval exceeds the preset proportion, performing statistical averaging on all photograph alignment deviation degrees, and combining a preset correction intensity coefficient to calculate a correction factor related to the semantic level correction.

7. The method according to claim 6, wherein, The preset deviation threshold interval refers to a deviation interval for representing the exhibition object focus display area that can accept alignment errors, which represents the approximate range of the important display area of the exhibition object that the user can capture, but the photograph deviation range of the important content of the exhibition object is not accurate. 8.The method of claim 6, wherein, The step of dynamically correcting the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and matching and generating explanation data content adapted to the user's cognitive level from the voice explanation database according to the new semantic level adaptability value, comprises: Dynamically up-regulating the semantic level adaptability value of the preset explanation data content by using the correction factor to obtain a new semantic level adaptability value; Accessing the voice explanation database of the scenic exhibition hall, and extracting a set of explanation data content corresponding to the current exhibition object display area from the database; Based on the new semantic level adaptability value, filtering the explanation data content with a matching semantic level from the set of explanation data content, and outputting the explanation data content in voice, so that the explanation content is adapted to the user's cognitive level.

9. A location-aware based automatic voice guide system for scenic spots, characterized in that, The system comprises: A data acquisition module for obtaining camera behavior data of a user entering a scenic exhibition hall and continuously passing through a plurality of exhibition object display areas, and performing comprehensive camera input degree analysis on the user; A deviation identification module for analyzing the camera behavior data to identify the photograph alignment deviation degrees of the user in the previous exhibition object display areas when the comprehensive camera input degree is greater than a preset threshold; A correction factor calculation module for performing statistical analysis on the photograph alignment deviation degrees corresponding to each exhibition object, and performing comprehensive evaluation on the photograph alignment deviation degrees when the proportion of the photograph alignment deviation degrees within the preset deviation threshold interval exceeds the preset proportion, to calculate the correction factor related to the semantic level correction; A semantic level acquisition module for obtaining the semantic level adaptability value of the preset explanation data content corresponding to the exhibition object after determining that the user enters a new exhibition object display area based on location sensing; A content matching module for dynamically correcting the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and matching and generating explanation data content adapted to the user's cognitive level from the voice explanation database according to the new semantic level adaptability value.

10. The location-aware based automatic voice guide system for scenic spots according to claim 9, characterized in that, The comprehensive camera shooting input degree refers to a comprehensive characteristic value determined based on the professionalism of the camera shooting equipment, the composition adjustment duration, the shooting times, and the frequency of changes of the imaging parameters of the equipment, and is used to represent the attention degree and operation input intensity of the user in the camera shooting behavior.

Citation Information

Patent Citations

  • Intelligent shooting method for scene understanding and script analysis driven by large science and technology movie and television model

    CN120786172A

  • Photographing interaction method and apparatus, storage medium and terminal device

    WO2019218879A1