A scenic spot automatic voice interpretation method and system based on location awareness

By analyzing user behavior data captured in the exhibition hall, the content of the automatic voice guide system in the scenic area is dynamically adjusted, solving the problem of the disconnect between the guide content and user observation behavior, and realizing a personalized and intelligent guide experience.

CN121350299BActive Publication Date: 2026-03-27BEIJING GRAVITY YUHUA FILM & TELEVISION CULTURE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing automated audio guide systems in scenic areas cannot dynamically adjust the semantic level of the explanation based on the user's self-filming behavior, resulting in a disconnect between the explanation content and the user's observation behavior, and failing to meet personalized needs.

Method used

By acquiring user video recording behavior data within the exhibition hall, analyzing video recording engagement and shooting alignment deviation, calculating correction factors, and dynamically adjusting the semantic level of the explanatory content to match the user's cognitive level.

Benefits of technology

It achieves adaptive matching between the content of the explanation and the user's observation behavior, enhances the personalization and intelligence of the visiting experience, and optimizes the visual capture and shooting composition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121350299B_ABST
    Figure CN121350299B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of intelligent explanation of scenic spots, and provides a scenic spot automatic voice explanation method and system based on position sensing, which comprises the following steps: based on the shooting alignment deviation corresponding to each exhibition object, when the proportion of the shooting alignment deviation within the preset deviation threshold interval exceeds the preset ratio, the shooting alignment deviation of each time is comprehensively evaluated to calculate the correction factor related to semantic level correction. The application realizes the transformation of the voice explanation system from static content presentation to dynamic semantic self-adaptive adjustment by introducing a user shooting behavior analysis mechanism based on position sensing. After detecting the user's camera behavior on the exhibition object, the system calculates the shooting alignment deviation and generates a correction factor accordingly, dynamically corrects the semantic level adaptability value of the explanation data content, and makes the subsequent explanation content adaptively match the user's cognitive level in terms of semantic depth and visual guidance intensity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent explanation of scenic spots, in particular to a scenic spot automatic voice explanation method and system based on location perception. BACKGROUND

[0002] With the development of mobile terminals and speech recognition technology, automatic voice explanation systems based on location perception are gradually and commonly used in scenic spots, museums and exhibition halls. Such systems usually identify the location of tourists through Bluetooth beacons, GPS or indoor positioning modules, and automatically play the explanation content corresponding to the current location to replace manual explanation. The existing voice explanation method can improve the visiting efficiency and information coverage to a certain extent, but the whole still stays at the level of matching static content, that is, no matter how different users' interests, observation methods or interactive behaviors are, the explanation content output by the system remains the same, lacking the dynamic response ability to individualized experience of users.

[0003] In recent years, some improvement schemes have tried to introduce user behavior data such as stay time, movement trajectory or gaze tracking to assist in optimizing the push order of explanation content. However, these schemes often only reflect the user's attention area at a macro level and cannot reflect the user's active observation and shooting behavior characteristics during the visit. For example, when the user takes multiple shots of the exhibition, adjusts the shooting direction or composition, the existing system cannot identify the visual attention pattern implied behind these behaviors, so it still cannot form semantic adaptation on the explanation content that matches the user's observation depth.

[0004] As can be seen, the existing automatic explanation technology of scenic spots has the technical defect that the explanation content is out of touch with the user's real perception behavior. The system cannot dynamically adjust the explanation semantic level according to the user's autonomous shooting behavior characteristics of the exhibition, resulting in explanation content that is either too superficial to stimulate interest or too in-depth to deviate from the user's current observation focus. SUMMARY

[0005] The purpose of the present application is to provide a scenic spot automatic voice explanation method and system based on location perception, which aims to solve the problems raised in the background art.

[0006] The present application is implemented as follows: a scenic spot automatic voice explanation method based on location perception, the method comprising:

[0007] Obtaining the camera shooting behavior data of a user who enters an exhibition hall of a scenic spot and successively experiences a plurality of exhibition display areas, and performing comprehensive camera shooting input degree analysis on the user;

[0008] When the comprehensive camera shooting input degree is greater than a preset threshold, analyzing the camera shooting behavior data to identify the shooting alignment deviation of the user in the previous exhibition display areas with respect to the key display area of the exhibition;

[0009] statistical analysis is performed on the shooting alignment deviation corresponding to each exhibition object, and when the proportion of the shooting alignment deviation within the preset deviation threshold interval exceeds a preset ratio, the shooting alignment deviation of each time is comprehensively evaluated to calculate a correction factor related to semantic level correction;

[0010] After determining that the user enters a new exhibition object display area based on location awareness, a semantic level adaptability value of preset explanation data content corresponding to the exhibition object is obtained;

[0011] The semantic level adaptability value is dynamically corrected based on the correction factor to obtain a new semantic level adaptability value, and the explanation data content adapted to the user's cognitive level is matched and generated from the voice explanation database according to the new semantic level adaptability value.

[0012] As a further limitation of the technical scheme of the embodiment of the application, the comprehensive shooting input degree refers to a comprehensive characteristic value determined based on the professionalism of the shooting device, the composition adjustment duration, the shooting times and the frequency of change of the device imaging parameters when the user performs the shooting operation in the exhibition object display area, and is used to represent the attention degree and operation input intensity of the user in the shooting behavior.

[0013] As a further limitation of the technical scheme of the embodiment of the application, when the comprehensive shooting input degree is greater than a preset threshold, the step of analyzing the shooting behavior data to identify the shooting alignment deviation of the user in the previous exhibition object display area for the key exhibition object display area includes:

[0014] The shooting input degree of the user in each exhibition object display area is calculated based on the shooting behavior data, and the shooting input degree is statistically averaged to obtain a comprehensive shooting input degree;

[0015] It is judged whether the comprehensive shooting input degree is greater than a preset threshold, if yes, the shooting behavior data is analyzed, the framing posture information of the user when shooting the exhibition object is determined, and the framing posture information is compared with a preset standard shooting reference model to calculate the shooting alignment deviation of the user in the corresponding exhibition object display area;

[0016] The shooting alignment deviation is used to represent the capture accuracy of the user for the key exhibition object display content in the shooting process.

[0017] As a further limitation of the technical scheme of the embodiment of the application, the framing posture information includes the spatial pose parameters, the framing direction angle, the shooting distance and the field of view coverage range of the shooting device.

[0018] The preset standard shooting reference model is a standard view angle model established based on the three-dimensional structure data of the exhibition space of the scenic spot exhibition hall and the feature of the key display area, and is used to represent the standard framing posture information of the exhibition under ideal shooting conditions.

[0019] As a further limitation of the technical scheme of the embodiment of the present application, the calculation of the shooting alignment deviation specifically includes:

[0020] The framing posture information of the user is compared with the standard framing posture information in the standard shooting reference model in terms of spatial parameters, and the framing direction angle difference, the shooting distance difference and the field of view coverage difference are calculated.

[0021] The shooting alignment deviation is determined based on the weighted results of the parameter differences, and is used to represent the alignment accuracy of the user to the key display area of the exhibition in the shooting process.

[0022] As a further limitation of the technical scheme of the embodiment of the present application, based on the statistical analysis of the shooting alignment deviations corresponding to each exhibition, when the proportion of the shooting alignment deviations within the preset deviation threshold interval exceeds the preset ratio, the steps of comprehensively evaluating each shooting alignment deviation to calculate the correction factor related to the semantic level correction include:

[0023] The shooting alignment deviations of the user for each photographed exhibition are obtained, and it is determined whether each shooting alignment deviation falls within the preset deviation threshold interval.

[0024] When the proportion of the number of shooting alignment deviations within the preset deviation threshold interval exceeds the preset ratio, the statistical average of all shooting alignment deviations is calculated, and the correction factor related to the semantic level correction is calculated in combination with the preset correction strength coefficient.

[0025] As a further limitation of the technical scheme of the embodiment of the present application, the preset deviation threshold interval refers to a deviation interval that can accept alignment errors in the key display area of the exhibition, which represents the shooting deviation range in which the user can capture the approximate range of the important display area of the exhibition but cannot accurately align the key content.

[0026] As a further limitation of the technical scheme of the embodiment of the present application, the steps of dynamically correcting the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and matching and generating the explanation data content adapted to the user's cognitive level from the voice explanation database according to the new semantic level adaptability value include:

[0027] The semantic level adaptability value of the preset explanation data content is dynamically adjusted by using the correction factor to obtain a new semantic level adaptability value.

[0028] Call the voice guide database of the exhibition hall in the scenic area, and extract the guide data content set corresponding to the current exhibition display area from the database;

[0029] Based on the new semantic level adaptation value, the guide data content with the matching semantic level is filtered from the guide data content set, and the guide data content is output by voice, so that the guide content is adapted to the cognitive level of the user.

[0030] A location-aware-based automatic voice guide system for scenic areas, the system comprising:

[0031] A data acquisition module for acquiring camera behavior data of a user entering an exhibition hall in a scenic area and continuously passing through multiple exhibition display areas, and performing comprehensive camera input analysis on the user;

[0032] A deviation identification module for analyzing camera behavior data when the comprehensive camera input degree is greater than a preset threshold to identify the shooting alignment deviation of the user in the previous exhibition display area for the key exhibition display area;

[0033] A correction factor calculation module for statistical analysis based on the shooting alignment deviation of each exhibition, and when the proportion of the shooting alignment deviation within the preset deviation threshold interval exceeds the preset ratio, the shooting alignment deviation is comprehensively evaluated to calculate the correction factor related to the semantic level correction;

[0034] A semantic level acquisition module for obtaining the semantic level adaptation value of the preset guide data content corresponding to the exhibition after determining that the user enters a new exhibition display area based on location awareness;

[0035] A content matching module for dynamically correcting the semantic level adaptation value based on the correction factor to obtain a new semantic level adaptation value, and matching and generating guide data content adapted to the cognitive level of the user from the voice guide database according to the new semantic level adaptation value.

[0036] As a further limitation of the technical scheme of the embodiment of the application, the comprehensive camera input degree refers to the comprehensive characteristic value determined based on the professionalism of the camera equipment, the composition adjustment duration, the shooting frequency and the imaging parameter change frequency of the equipment when the user performs camera operation in the exhibition display area, which is used to represent the attention degree and operation input intensity of the user in the camera behavior.

[0037] Compared with the prior art, the present application has the following beneficial effects:

[0038] The application realizes the transformation of the voice guide system from static content presentation to dynamic semantic self-adaptive adjustment by introducing a user shooting behavior analysis mechanism based on position awareness. After detecting the user's camera behavior for the exhibition, the system calculates the shooting alignment deviation and generates a correction factor accordingly, dynamically corrects the semantic level adaptation value of the guide data content, and makes the subsequent guide content adaptively match the user's cognitive level in terms of semantic depth and visual guidance intensity. This mechanism forms a closed-loop process of "behavior awareness-semantic adjustment-cognitive feedback", which not only improves the consistency of the guide content and the user's observation behavior, but also optimizes the user's visual capture and shooting composition performance imperceptibly without interfering with the user's operation.

[0039] The application effectively solves the problem of disconnection between guide content and user perception behavior (autonomous camera shooting) in traditional scenic spot automatic guide systems, realizes a dynamic guide experience with both personalization and intelligence, and has high practical value and promotion prospects. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 The flowchart of the method provided for the embodiment of the application is shown.

[0041] Figure 2 The flowchart of identifying the shooting alignment deviation based on the comprehensive camera input degree in the method provided for the embodiment of the application is shown.

[0042] Figure 3 The flowchart of calculating the semantic level correction factor based on the shooting alignment deviation in the method provided for the embodiment of the application is shown.

[0043] Figure 4 The flowchart of dynamically adjusting the semantic level adaptation of the guide content based on the correction factor in the method provided for the embodiment of the application is shown.

[0044] Figure 5 The application architecture diagram of the system provided for the embodiment of the application is shown. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical scheme and advantages of the application clearer and more apparent, the application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and do not limit the application.

[0046] Figure 1 The flowchart of the method provided for the embodiment of the application is shown.

[0047] Specifically, a scenic spot automatic voice guide method based on position awareness, the method specifically comprises the following steps:

[0048] In step S100, the camera shooting behavior data of a user entering an exhibition hall of a scenic spot and continuously passing through a plurality of exhibition display areas is acquired, and comprehensive camera shooting input degree analysis is performed on the user.

[0049] The comprehensive camera shooting input degree refers to a comprehensive characteristic value determined based on the professionalism of a camera shooting device, the composition adjustment duration, the number of times of shooting, and the frequency of changes of imaging parameters of the device when the user performs camera shooting operation in the exhibition display area, and is used to represent the degree of attention and the intensity of operation input of the user in the camera shooting behavior.

[0050] In the embodiment of the present application, the exhibition hall of the scenic spot can be a cultural display place with fixed or temporary exhibition space, including a museum, a science and technology museum, an art museum, an art gallery, a memorial hall, a special exhibition hall, a site display hall, a non-heritage experience hall, an enterprise exhibition hall, a theme exhibition hall in a theme park, a display point in a historical block, and an immersive exhibition space, etc. The above-mentioned space can be a closed exhibition area in a single building, a semi-open exhibition area such as a corridor, a lobby, a mezzanine, a sunken square, or an outdoor independent exhibition booth or enclosed exhibition area in a park, as long as it can provide a relatively clear visiting route and a stay area for the appreciation and explanation of the exhibition, which belongs to the exhibition hall of the scenic spot in the embodiment of the present application.

[0051] The exhibition hall of the scenic spot is configured with a voice explanation database for storing a plurality of sets of explanation data contents corresponding to different exhibition display areas. Each exhibition display area can have a plurality of explanation types, and different types differ in language structure, information depth and visual description accuracy. For example, the explanation contents of the same exhibition can include basic guide explanation, key area pointing explanation and detail composition guidance explanation, etc., to meet the different needs of different users in the appreciation behavior and the shooting behavior. The system sets a semantic level adaptability value for each type of explanation data content in the voice explanation database, which is used to represent the comprehensive level of the explanation content in terms of semantic complexity, visual orientation intensity and information accuracy.

[0052] It can be understood that the semantic level adaptability value is not only used to represent the semantic depth and information complexity of the explanation content, but also determines the degree of guidance of visual perception in the information expression of the explanation content. When the semantic level adaptability value is low, the expression of the explanation content tends to be holistic and general, emphasizing the theme background and general knowledge of the exhibition, so that the user can form a perception impression in a wider range; when the semantic level adaptability value is high, the description of the explanation content is more structured and detailed in semantic organization, and the language expression implies the focus guidance of the key display area, the composition level or the spatial features of the exhibition, thereby affecting the visual attention distribution of the user at the cognitive level.

[0053] The exhibition object display area can be determined based on the location awareness and the exhibition elements. Specifically, a spatial fence matching the physical position of each exhibition object is established: for an in-cabinet exhibit, a two-dimensional polygon fence parallel to the front edge of the cabinet is used, and several equidistant buffer zones are set in front; for a three-dimensional sculpture or large object, a three-dimensional envelope is established around the exhibit, and a fan-shaped viewing zone and a minimum safety distance zone are set in the main viewing direction; for a wall-mounted painting or calligraphy, a rectangular viewing corridor is set in the normal direction of the work, and a lateral buffer zone is superimposed. The above fence is calibrated by location-aware base stations such as Bluetooth beacons, ultra-wideband positioning, Wi-Fi RTT, infrared door magnets, or visual anchor points. The system uses entry / stay / exit events as criteria to determine whether a user is in a certain exhibition object display area.

[0054] The camera behavior data is mainly obtained by intelligent cameras installed in the exhibition hall, and the user terminal is only used as an auxiliary information source. Several intelligent cameras are installed in the exhibition hall to sense the shooting behavior of users entering the area. The intelligent cameras can have personnel detection and posture recognition functions to identify the user's framing, camera holding, shutter triggering, and other actions without capturing the shooting content, and record the corresponding time information and spatial posture information. By detecting the user's stay time in the exhibition object display area, framing duration, number of shooting actions, and shooting angle changes, corresponding camera behavior data can be generated.

[0055] The user terminal (such as a mobile phone or tablet) can provide auxiliary data, including the timestamp of the shutter trigger event, the number of camera parameter changes, device model and type information, etc., which can be used to correspond and verify the behavior data collected by the exhibition hall side. The intelligent cameras and user terminals in the exhibition hall can be associated and matched through Bluetooth or wireless networks to confirm the shooting behavior events of the same user.

[0056] The camera behavior data includes objective and detectable basic data fields, such as: framing start and end time, shutter trigger time, stay time in the exhibition object display area, framing posture angle change range, number of shooting actions, number of camera parameter changes, and device type identification. The above data can be directly obtained by the posture detection module of the intelligent camera and the system event recording module of the user terminal, and belongs to objectively existing data without image content analysis.

[0057] The calculation process of the comprehensive camera shooting input degree is as follows: after the system detects that the user has a shooting behavior in a certain exhibition object display area, the shooting times, the shooting duration, the device parameter adjustment times and the device type grading score of the user in the area are counted, and after the above data is normalized and weighted according to the preset weight, the corresponding comprehensive camera shooting input degree is obtained. The comprehensive camera shooting input degree is used to represent the attention degree and operation input intensity of the user in the shooting behavior, and provides a basis for subsequent identification of the shooting alignment deviation.

[0058] Further, the automatic voice guide method based on location awareness in the scenic spot further includes the following steps:

[0059] Step S200, when the comprehensive camera shooting input degree is greater than a preset threshold, the camera shooting behavior data is analyzed to identify the shooting alignment deviation of the user in the previous exhibition object display area for the exhibition object key display area.

[0060] Specifically, Figure 2 A flowchart for identifying the shooting alignment deviation based on the comprehensive camera shooting input degree is shown.

[0061] When the comprehensive camera shooting input degree is greater than a preset threshold, the camera shooting behavior data is analyzed to identify the shooting alignment deviation of the user in the previous exhibition object display area for the exhibition object key display area, which specifically includes the following steps:

[0062] Step S201, based on the camera shooting behavior data, the camera shooting input degree of the user in each exhibition object display area is calculated, and the camera shooting input degree is statistically averaged to obtain a comprehensive camera shooting input degree;

[0063] Step S202, it is judged whether the comprehensive camera shooting input degree is greater than a preset threshold, if yes, the camera shooting behavior data is analyzed, the shooting posture information of the user when shooting the exhibition object is determined, and the shooting posture information is compared with a preset standard shooting reference model to calculate the shooting alignment deviation of the user in the corresponding exhibition object display area;

[0064] The shooting alignment deviation is used to represent the capture accuracy of the user for the key display content of the exhibition object in the camera shooting process.

[0065] The shooting posture information includes the spatial pose parameters of the shooting device, the shooting direction angle, the shooting distance and the field of view coverage range.

[0066] The preset standard shooting reference model is a standard view angle model established based on the exhibition object space three-dimensional structure data and the key display area features of the exhibition hall in the scenic spot, and is used to represent the standard shooting posture information of the exhibition object under ideal shooting conditions.

[0067] The calculation of the shooting alignment deviation specifically includes:

[0068] The user's view posture information is compared with the standard view posture information in the standard shooting reference model in terms of spatial parameters, and a view direction angle difference value, a shooting distance difference value and a field of view coverage difference value are calculated;

[0069] A shooting alignment deviation degree is determined based on the weighted results of the parameter difference values, and is used to represent the alignment accuracy of the user to the key display area of the exhibition in the shooting process.

[0070] In the embodiment of the present application, in step S201, the comprehensive camera shooting input degree is obtained by statistically averaging the camera shooting input degrees of the display areas of the exhibitions. The average calculation method can effectively reduce the influence of occasional fluctuations of single shooting behavior, make the comprehensive result more stable and reliable, and ensure the comparability between different display areas of the exhibitions. The comprehensive camera shooting input degree obtained in this way can truly reflect the overall camera shooting input intensity of the user in the whole visiting process, avoid abnormal deviation caused by too short or accidental interruption of the stay time of individual display areas of the exhibitions, and thus improve the accuracy and robustness of subsequent behavior analysis.

[0071] In step S202, when the comprehensive camera shooting input degree is greater than a preset threshold, it indicates that the camera shooting behavior of the user in the visiting process is relatively active and continuous, and has high attention and shooting willingness. At this time, the camera shooting behavior characteristics have analysis significance, and therefore the system starts the shooting alignment deviation degree recognition module under this condition to ensure the stability and representativeness of the analysis sample. If the comprehensive camera shooting input degree of the user is lower than the preset threshold, it usually represents that the user only performs casual shooting or temporary view, and such sample does not have reference value, and the system does not perform deviation analysis. The preset threshold can be determined according to the statistical results of the historical sample data of the exhibition hall, for example, taking 70% to 80% quantile in the overall sample distribution as the initial determination line; or can be adaptively calibrated according to the exhibition hall type, the exhibition density and the guest flow characteristics, so as to balance the triggering rate and the analysis accuracy.

[0072] After the comprehensive camera shooting input degree meets the determination condition, the system analyzes the camera shooting behavior data and extracts the view posture information of the user. The view posture information includes the spatial pose parameters of the shooting device, the view direction angle, the shooting distance and the field of view coverage range, which can be directly obtained through the collaborative perception of the intelligent camera of the exhibition hall and the sensing module of the user terminal, and has high observability and physical accuracy. In addition to the above-mentioned parameters, in some embodiments, extended parameters such as view stability, camera posture angular velocity, moving path curvature can be further selected to enhance the multi-dimensional expression ability of the view posture information.

[0073] The establishment of the standard shooting reference model is based on the spatial three-dimensional structure data of the exhibits and the features of the key display areas. For exhibits with three-dimensional model data, the geometric outline, main view direction and key display area center point of the model can be directly extracted, and the ideal shooting pose parameters can be generated through spatial calibration. For exhibits without three-dimensional model data, the standard shooting reference model can be established by regularized space calculation or manual calibration according to the display cabinet arrangement parameters, exhibit plane coordinates and viewing lines. In some embodiments, the system can use a lightweight algorithm based on sample learning to fit the standard shooting samples calibrated by experts to obtain standard framing posture information that is more consistent with the actual viewing experience.

[0074] The calculation of the shooting alignment deviation includes two steps of parameter comparison and weighted synthesis. The system first compares the user framing posture information with the standard framing posture information in the standard shooting reference model in terms of spatial parameters, calculates the framing direction angle difference (used to reflect the consistency of the user shooting direction and the standard direction), the shooting distance difference (used to represent the position offset of the shooting point), and the field of view coverage difference (used to reflect the coverage proportion difference of the key display area in the composition). Then, the system sets preset weights according to the importance of each parameter, and performs weighted aggregation on each difference value to obtain the final shooting alignment deviation result. The smaller the shooting alignment deviation value, the higher the alignment degree of the user shooting angle, distance and the standard model.

[0075] In an extended embodiment, to further improve the evaluation accuracy, complex factors such as illumination difference, occlusion rate difference or visibility coverage difference can also be introduced. The illumination difference is used to reflect the deviation of the user framing position and the standard shooting position in terms of brightness conditions, the occlusion rate difference is used to determine the difference between the proportion of exhibits being occluded in the shooting picture and the standard benchmark, and the visibility coverage difference is used to measure the difference in the visible range of the key display area from the user's perspective. By including the above additional factors in the comprehensive calculation with a lower weight, the robustness and judgment accuracy of the system in complex environments can be improved while keeping the main judgment dimensions simple.

[0076] Further, the method for automatically providing voice guidance in a scenic area based on location awareness further comprises the following steps:

[0077] Step S300, based on the shooting alignment deviation of each exhibit, when the proportion of the shooting alignment deviation within the preset deviation threshold interval exceeds the preset ratio, the shooting alignment deviation of each time is comprehensively evaluated to calculate the correction factor related to the semantic level correction.

[0078] Specifically, Figure 3 A flowchart for calculating the semantic level correction factor based on the shooting alignment deviation is shown.

[0079] The statistical analysis of the shooting alignment deviation corresponding to each exhibition object is performed, and when the proportion of the shooting alignment deviation within the preset deviation threshold interval exceeds the preset ratio, the shooting alignment deviations are comprehensively evaluated to calculate a correction factor related to semantic level correction, which specifically includes the following steps:

[0080] In step S301, the shooting alignment deviation of each photographed exhibition object is obtained, and it is determined whether each shooting alignment deviation falls within a preset deviation threshold interval.

[0081] In step S302, when the number of shooting alignment deviations within the preset deviation threshold interval exceeds the preset ratio, the shooting alignment deviations are statistically averaged, and a correction factor related to semantic level correction is calculated in combination with a preset correction intensity coefficient.

[0082] The preset deviation threshold interval refers to a deviation interval that can accept alignment errors in the key display area of the exhibition object, which represents the approximate range of the important display area of the exhibition object that the user can capture but cannot accurately align the key content.

[0083] In the embodiment of the present application, in step S300, the system performs statistical analysis based on the shooting alignment deviation corresponding to each exhibition object to evaluate the overall shooting stability and shooting accuracy of the user. When the proportion of the shooting alignment deviation within the preset deviation threshold interval exceeds the preset ratio, it indicates that the user's shooting behavior in the exhibition area is not completely accurate, but the shooting orientation is basically stable, and the approximate composition range of the key display area of the exhibition object can be captured, reflecting that the user's shooting ability is in the state of "correctable but not ideal". At this time, if the system still outputs the explanation content according to the original semantic level, it may lead to high level of explanation information, mismatch between key content expression and user visual attention area, and affect the explanation experience. Therefore, this step establishes a correction factor related to semantic level adaptation correction by statistically analyzing the overall performance of the user's shooting alignment deviation, so that the subsequent explanation content can be adaptively adjusted to realize dynamic optimization of personalized explanation accuracy.

[0084] The preset deviation threshold interval is used to represent the acceptable alignment error range in the key display area of the exhibition object. The lower limit corresponds to the deviation boundary when the user does not align the key area at all, and the upper limit corresponds to the upper limit interval when the user can capture the main outline of the key display area of the exhibition object but cannot accurately align the core content.

[0085] In step S302, when the proportion of the photograph alignment deviation degree within the preset deviation threshold interval exceeds the preset ratio, the system determines that the user's overall photograph deviation behavior has consistency and correctability, and thus the overall photograph alignment deviation degree is statistically averaged. The reason for using the average value as a basic component of the correction factor is that the average value can smooth the individual differences in the photographing accuracy of the user among different exhibits, extract the overall deviation trend, and avoid excessive correction caused by individual excessive deviation samples. At the same time, the average value reflects the average deviation amplitude of the user at a stable level, and can more objectively represent the user's framing ability.

[0086] To achieve flexible adjustment, the present application combines a preset correction strength coefficient when calculating the correction factor. The correction strength coefficient is used to control the amplitude of semantic level adjustment to avoid over-correction causing imbalance of the explanation level. The correction strength coefficient can be determined comprehensively according to factors such as the complexity of the exhibition hall content, the information density of the exhibits, and the granularity of the explanation level. For example, in an artistic exhibition hall with a deep information level of exhibits, a higher correction strength coefficient can be set to amplify the influence of deviation differences; while in a simple structure and intuitive content science and technology exhibition hall, a lower coefficient can be used to maintain the stability of the explanation level. In specific implementation, the correction strength coefficient can be obtained by system pre-calibration, or automatically updated by weighted learning according to historical explanation feedback data.

[0087] Further, the automatic voice explanation method based on location awareness further includes the following steps:

[0088] Step S400, after determining that the user enters a new exhibit display area based on location awareness, the semantic level adaptability value of the preset explanation data content corresponding to the exhibit is obtained.

[0089] Step S500, dynamically correcting the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and matching and generating explanation data content adapted to the user's cognitive level from the voice explanation database according to the new semantic level adaptability value.

[0090] Specifically, Figure 4 A flow chart of dynamically adjusting the semantic level adaptability of explanation content based on the correction factor is shown.

[0091] The dynamic correction of the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and the matching and generation of explanation data content adapted to the user's cognitive level from the voice explanation database according to the new semantic level adaptability value specifically includes the following steps:

[0092] Step S401, dynamically up-regulating the semantic level adaptability value of the preset explanation data content using the correction factor to obtain a new semantic level adaptability value;

[0093] Step S402, call the voice guide database of the scenic exhibition hall, and extract the guide data content set corresponding to the current exhibition area from the guide data content set;

[0094] Step S403, based on the new semantic level adaptability value, filter out the guide data content with matched semantic level from the guide data content set, and output the guide data content in voice, so that the guide content is adapted to the cognitive level of the user.

[0095] In the embodiment of the application, the semantic level adaptability value is dynamically corrected based on the correction factor. The correction factor is quantified from the user's shooting behavior data and the shooting alignment deviation statistical characteristics. The factor can directly reflect the individual differences of the user in visual capture accuracy, composition stability and key area focus performance. By combining the behavior feedback quantitative parameter with the semantic level adaptability value of the guide content, the system can adaptively adjust the user's perception behavior at the semantic level, so that the expression depth, information density and visual guidance intensity of the guide content remain dynamically consistent with the user's current observation characteristics. Compared with the static guide matching method, the correction mechanism adopted by the application not only considers the user's historical shooting characteristics, but also realizes the gradual optimization of the semantic level in the continuous visit process, thereby forming a self-closing loop structure of "behavior feedback-semantic adjustment-cognitive matching".

[0096] The advantage of this correction method is that it can implicitly optimize the user's visual attention mode through dynamic adaptation of voice guide content without explicit intervention or additional operation of the user. By adjusting the semantic level adaptability value, the system can adjust the focus distribution of the guide content at the semantic level, so that the user spontaneously adjusts the visual attention direction in the auditory cognitive process, thereby improving the composition accuracy and observation efficiency of subsequent shooting. In addition, the correction process is scalable and can be flexibly configured according to different types of exhibition halls, exhibition features or audience group characteristics, to achieve a balance between personalized guidance and shooting guidance.

[0097] In addition to directly adjusting the semantic level adaptability value by using the correction factor, in other embodiments, the system can also correct the guide level by combining time series weighting, user group behavior clustering analysis or semantic fuzzy matching methods. For example, an individual shooting feature model can be established according to the behavior records of multiple visits of the user, and the semantic level adaptability value is adjusted by the model output dynamic coefficient; or a group correction template is generated based on the behavior patterns of similar user groups, which is used to initially calibrate the semantic level of new users.

[0098] Overall, the present application realizes the transformation of the voice guide system from passive content matching to active behavior optimization by introducing a location-aware user behavior analysis mechanism. This method can adaptively adjust the guide content without additional user operations, making the guide dynamically coordinated with the user's visual capture features, cognitive level, and observation methods, thereby improving the visit experience and information transmission efficiency of the scenic exhibition hall.

[0099] The overall benefit of the present application is that through the perception of the shooting behavior, the analysis of the deviation degree and the dynamic correction of the semantic level, the bidirectional adaptation of the voice guide and the user behavior characteristics is realized, the problem of disconnection between the guide content and the user observation behavior in the traditional scenic automatic guide system is solved, and the guide content can actively guide the user's visual attention. The scheme can be widely applied to museums, art galleries, science and technology museums and outdoor scenic spots, and has good popularization and practical application prospect.

[0100] Further, Figure 5 The application architecture diagram of the system provided by the embodiment of the present application is shown.

[0101] Among them, in another preferred embodiment provided by the present application, a location-aware scenic automatic voice guide system comprises:

[0102] The data acquisition module 100 is configured to acquire the camera shooting behavior data of the user who enters the scenic exhibition hall and continuously experiences multiple exhibition object display areas, and to analyze the comprehensive camera shooting input degree of the user.

[0103] The comprehensive camera shooting input degree refers to a comprehensive feature value determined based on the professionalism of the camera shooting device, the composition adjustment duration, the shooting frequency and the imaging parameter change frequency of the device when the user performs the camera shooting operation in the exhibition object display area, and is used to represent the attention degree and operation input intensity of the user in the camera shooting behavior.

[0104] Further, the location-aware scenic automatic voice guide system further comprises:

[0105] The deviation identification module 200 is configured to analyze the camera shooting behavior data to identify the shooting alignment deviation degree of the user in the previous exhibition object display area with respect to the key exhibition object display area when the comprehensive camera shooting input degree is greater than the preset threshold.

[0106] Further, the location-aware scenic automatic voice guide system further comprises:

[0107] The correction factor calculation module 300 is configured to statistically analyze the photograph alignment deviation corresponding to each exhibition object, and when the proportion of the photograph alignment deviation within the preset deviation threshold interval exceeds a preset ratio, the photograph alignment deviation of each time is comprehensively evaluated to calculate a correction factor related to the semantic level correction.

[0108] Further, the location-aware scenic area automatic voice guide system further comprises:

[0109] The semantic level acquisition module 400 is configured to acquire the semantic level adaptability value of the preset guide data content corresponding to the exhibition object after determining that the user enters a new exhibition object display area based on location awareness.

[0110] Further, the location-aware scenic area automatic voice guide system further comprises:

[0111] The content matching module 500 is configured to dynamically correct the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and match and generate guide data content adapted to the user's cognitive level from the voice guide database according to the new semantic level adaptability value.

[0112] It should be understood that although each step in the flowchart of each embodiment of the present application is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in each embodiment can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.

[0113] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0114] Any combination of the technical features of the above-mentioned embodiments can be combined. In order to make the description simple, all possible combinations of the technical features in the above-mentioned embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0115] The above-mentioned embodiments only express several embodiments of the present application, and the description is more specific and detailed, but it should not be understood as limiting the scope of the present application. It should be noted that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

[0116] The above-mentioned is only the preferred embodiment of the present application, and does not limit the present application. Any modification, equivalent replacement, and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A method for automatic voice narration in scenic areas based on location awareness, characterized in that, The method includes: Acquire camera behavior data from users who enter the scenic area's exhibition hall and continuously traverse multiple exhibition areas, and conduct a comprehensive analysis of users' camera engagement. When the overall camera engagement exceeds a preset threshold, the camera behavior data is analyzed to identify the degree of alignment deviation of the user's shooting of the key display area of ​​the exhibit in the previous display areas of each exhibit. Statistical analysis is performed on the shooting alignment deviation of each exhibit. When the proportion of shooting alignment deviation within the preset deviation threshold range exceeds the preset ratio, a comprehensive evaluation is performed on the shooting alignment deviation of each time to calculate the correction factor related to semantic level correction. The steps include: obtaining the alignment deviation of the user's shots for each exhibit, and determining whether each alignment deviation falls within a preset deviation threshold range; When the proportion of shooting alignment deviations falling within the preset deviation threshold range exceeds a preset ratio, all shooting alignment deviations are statistically averaged and combined with a preset correction intensity coefficient to calculate a correction factor related to semantic level correction. The preset deviation threshold range refers to the deviation range that can be used to characterize the acceptable alignment error in the key display area of ​​the exhibit. It represents the shooting deviation range in which the user can capture the approximate range of the important display area of ​​the exhibit but fails to accurately align it with its key content. After determining that a user has entered a new exhibit display area based on location awareness, the semantic level adaptability value of the preset explanatory data content corresponding to the exhibit is obtained; The semantic level adaptability value is dynamically corrected based on the correction factor to obtain a new semantic level adaptability value. Based on the new semantic level adaptability value, the explanation data content that is adapted to the user's cognitive level is matched from the voice explanation database. The steps include: dynamically adjusting the semantic level adaptability value of the preset explanatory data content using a correction factor to obtain a new semantic level adaptability value. Access the audio guide database of the scenic area's exhibition hall and extract the set of audio guide data content corresponding to the current exhibit display area; Based on the new semantic level adaptability value, the explanatory data content with matching semantic level is selected from the explanatory data content set, and the explanatory data content is output as voice so that the explanatory content is adapted to the user's cognitive level.

2. The location-aware automatic voice narration method for scenic spots according to claim 1, characterized in that, The comprehensive camera engagement refers to a comprehensive characteristic value determined based on the professionalism of the camera equipment, the duration of composition adjustment, the number of shots, and the frequency of changes in the equipment's imaging parameters when a user performs camera operations within the exhibition area. It is used to characterize the user's level of attention and operational engagement in the camera behavior.

3. The location-aware automatic voice narration method for scenic spots according to claim 2, characterized in that, When the overall camera engagement exceeds a preset threshold, the steps for analyzing camera behavior data to identify the user's alignment deviation from the key display areas of the exhibits in the previous exhibition areas include: Based on the camera behavior data, the user's camera engagement level in each exhibit display area is calculated, and the camera engagement level is statistically averaged to obtain the comprehensive camera engagement level; If the overall camera engagement is greater than a preset threshold, the camera behavior data is analyzed to determine the user's framing posture information when shooting the exhibit. The framing posture information is then compared with a preset standard shooting reference model to calculate the user's shooting alignment deviation within the corresponding exhibit display area. The shooting alignment deviation is used to characterize the accuracy with which the user captures the key content of the exhibit during the shooting process.

4. The location-aware automatic voice narration method for scenic spots according to claim 3, characterized in that, The framing posture information includes the spatial pose parameters of the shooting device, the framing direction angle, the shooting distance, and the field of view coverage; The preset standard shooting reference model is a standard perspective model established based on the spatial three-dimensional structural data of the exhibits in the scenic exhibition hall and the characteristics of the key display areas. It is used to represent the standard framing posture information of the exhibits under ideal shooting conditions.

5. The location-aware automatic voice narration method for scenic spots according to claim 3, characterized in that, The calculation of the alignment deviation during shooting specifically includes: The spatial parameters of the user's framing posture information are compared with those of the standard framing posture information in the standard shooting reference model, and the differences in framing direction angle, shooting distance, and field of view coverage are calculated. The shooting alignment deviation is determined based on the weighted result of the differences of each parameter, which is used to characterize the accuracy of the user's alignment of the key display area of ​​the exhibit during the shooting process.

6. A location-aware automatic voice guide system for scenic spots, characterized in that, The system is used to implement the method of any one of claims 1-5, the system comprising: The data acquisition module is used to acquire the camera behavior data of users who enter the scenic area exhibition hall and continuously pass through multiple exhibition areas, and to conduct a comprehensive analysis of users' camera engagement. The deviation recognition module is used to analyze camera behavior data when the overall camera engagement exceeds a preset threshold in order to identify the user's alignment deviation when shooting the key display area of ​​the exhibit in the previous exhibition areas. The correction factor calculation module is used to perform statistical analysis based on the shooting alignment deviation of each exhibit. When the proportion of shooting alignment deviation within the preset deviation threshold range exceeds the preset ratio, the shooting alignment deviation of each shooting is comprehensively evaluated to calculate the correction factor related to semantic level correction. The semantic level acquisition module is used to acquire the semantic level adaptability value of the preset explanatory data content corresponding to the exhibit after determining that the user has entered a new exhibit display area based on location awareness. The content matching module is used to dynamically correct the semantic level adaptability value based on the correction factor to obtain a new semantic level adaptability value, and then match and generate narration data content that is adapted to the user's cognitive level from the voice narration database based on the new semantic level adaptability value.

7. The location-aware automatic voice guide system for scenic spots according to claim 6, characterized in that, The comprehensive camera engagement refers to a comprehensive characteristic value determined based on the professionalism of the camera equipment, the duration of composition adjustment, the number of shots, and the frequency of changes in the equipment's imaging parameters when a user performs camera operations within the exhibition area. It is used to characterize the user's level of attention and operational engagement in the camera behavior.

Citation Information

Patent Citations

  • Intelligent shooting method for scene understanding and script analysis driven by large science and technology movie and television model

    CN120786172A

  • Photographing interaction method and apparatus, storage medium and terminal device

    WO2019218879A1