Intelligent voice explanation method and system based on positioning trigger

By combining spatial location with content segmentation mapping and visitor interaction information, a location-triggered intelligent voice guide system achieves personalized and flexible content matching, solving the problems of content mismatch and accidental triggering in existing systems, and improving visitor experience and system efficiency.

CN120910305BActive Publication Date: 2025-12-12CHENGDU ZHIYUAN YIYI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511438273.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-12-12
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing location-based intelligent voice guide systems cannot match personalized content according to the interests and needs of different tourists, and are prone to accidental triggering in crowded scenarios, resulting in guide content that does not meet the needs of tourists and insufficient interactivity.

Method used

By mapping spatial location to content segments, exhibits are deconstructed into multiple content items. By combining real-time visitor location and interactive information, the marking range and broadcast content are dynamically adjusted. High-precision scanning and AI segmentation algorithms are used to identify structural features, and infrared and camera modules are used to analyze visitor body language and eye direction, so as to achieve personalized and flexible matching of explanation content.

Benefits of technology

It improves content matching efficiency, reduces false trigger rate, enhances the interactivity between the system and tourists, enables personalized explanations based on the interests of different tourists, and optimizes the utilization of hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910305B_ABST
    Figure CN120910305B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent voice explanation method and system based on positioning trigger, it is related to intelligent voice explanation technical field. Including: the mark area of target explanation point is obtained, target mark range item is obtained, the specific explanation content of target explanation point is obtained, contrast content item is obtained, contrast content item is used to indicate the voice explanation information matched with target explanation point;Segmentation splitting based on contrast content item, obtain at least one content disassembly item, content disassembly item is used to indicate the classification explanation information of target explanation point.The application is coupled to the mapping of space position and content segment, and the exhibits are deconstructed into multiple content disassembly items, the corresponding depth explanation is triggered by the real-time position of tourist and the stay duration threshold of block, the content matching efficiency is improved, different tourists are adapted, the target mark range is automatically adjusted with tourist density by real-time quantity grade division and the intelligent matching of shrinkage scale value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent voice interpretation, in particular to an intelligent voice interpretation method and system based on positioning triggering. BACKGROUND

[0002] Intelligent voice interpretation is a system or service that provides information, commentary, education or guided tours through voice interaction with users under the support of voice technology and artificial intelligence. It can not only convert text into natural speech, but also understand user questions and context and provide relevant interpretation. Intelligent voice interpretation based on positioning triggering is a technology that uses modern positioning technologies such as GPS, Bluetooth beacons and Wi-Fi positioning to provide automated, personalized and scenario-based tour services for tourists, commonly used in scenic spots and museums.

[0003] The patent publication No. CN112967641A discloses an automatic identification and interpretation and augmented display method for scenic spots based on AR technology. It maintains scenic spot information for identification through internet public data, including acquisition of sample pictures for identification, maintenance of title pictures, brief introduction, source, location, audio and video resources, etc. of scenic spots. It realizes identification of landscapes in scenic spots through AR technology, realizes data and image augmentation display of identified objects, greatly improves the viewing experience of tourists and the dissemination effect of historical and cultural information.

[0004] The above and similar technical solutions provide location-based content interpretation based on scenic spots and museums. The system automatically identifies the location and triggers the corresponding voice interpretation content when the tourist enters the preset electronic fence area by acquiring the location relationship between the tourist and the interpretation point. For example, when approaching a certain showcase, the system will play its forging process and legendary story of excavation. However, different tourists are interested in different content of the same interpretation point, so it is impossible to provide classified interpretation of different content for different tourists. In addition, the interaction between the broadcast system and the tourist is insufficient. When the tourist is not interested in the interpretation content of the current interpretation point or is interested in other interpretation points in the same exhibition hall, the current interpretation point cannot provide adaptive interpretation content, resulting in that the broadcast content cannot meet different real-time situations. SUMMARY

[0005] The present application aims to provide an intelligent voice interpretation method and system based on positioning triggering to solve the problems raised in the background.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solution: an intelligent voice interpretation method and system based on positioning triggering, comprising:

[0007] The marker area of the target explanation point is acquired, a target marker range item is obtained, the specific explanation content of the target explanation point is acquired, and a comparison content item is obtained, which is used to represent the voice explanation information matched with the target explanation point;

[0008] The comparison content item is segmented and split based on the comparison content item, and at least one content split item is obtained, which is used to represent the classified explanation information of the target explanation point;

[0009] The real-time position item is obtained by acquiring the actual position information of the tourists in the target marker range item, and the matched content split item is selected as the broadcast content item based on the mapping information between the real-time position item and the content split item;

[0010] The broadcast component is set, the broadcast component is a voice output component, the output of the broadcast content item is performed based on the broadcast component, the voice explanation is performed, the explanation information item is obtained, and thus the split of the current explanation point explanation content and the classified matching type tracking explanation according to the interesting content of different tourists are realized.

[0011] The body interaction information and the language interaction information of the tourists are acquired by the interaction information acquisition module, and the interaction information item is obtained;

[0012] The information positioning is performed based on the interaction information item, the landing position of the interaction information item is acquired, the explanation point corresponding to the landing position is taken as the expanded explanation, the comparison content item of the expanded explanation is acquired in a data transmission and data sharing manner, and thus the expanded explanation of the explanation points other than the current explanation point according to the interaction information of the tourists is realized.

[0013] Further, the target marker range item acquisition method comprises:

[0014] The position information of the target explanation point is acquired, a target position item is obtained, the target position item is taken as a center marker point, an initial marker range is set, and the initial marker range item is obtained based on the combination of the initial marker range and the center marker point.

[0015] The search range is set, the search range is a fixed range value, the number of tourists is searched based on the search range, and a real-time number item is obtained.

[0016] The number of tourists is classified based on the number of tourists, a number grade item is obtained, a contraction ratio value is set based on the number grade item, a target contraction value is obtained based on the comparison result of the real-time number item and the number grade item, and the target marker range is obtained based on the combination result of the target contraction value and the initial marker range item.

[0017] Further, the content split item acquisition method comprises:

[0018] Acquire high-precision scanning data of the target explanation point, obtain scanning information items, and deconstruct the scanning information items into digital units based on the scanning information items for calculation to obtain an explanation information set;

[0019] Based on the explanation information set, identify structural features through an AI segmentation algorithm to obtain a structural division set, which at least includes one category of structural features;

[0020] Based on the mapping information of the structural division set and the explanation information set, segment and split the comparison content items to obtain content disassembly items.

[0021] Further, the mapping information acquisition method of the real-time position item and the content disassembly item includes:

[0022] Based on the content disassembly item, segment and divide the target marking range item to obtain at least one range marking item corresponding to the content disassembly item;

[0023] Based on the comparison data of the real-time position item and the range marking item, including position comparison and time comparison, set a retrieval threshold, which is a fixed time value. When the comparison data reaches the retrieval threshold, determine the corresponding marking range item as a mapping item, and further obtain the mapping information of the real-time position item and the content disassembly item.

[0024] Further, the acquisition method of the explanation information item includes:

[0025] The broadcast component is composed of at least one directional transmission module. Based on the real-time position item, head data of the visitor is acquired, including head orientation, to obtain a target position item. Based on the target position item, the orientation of the broadcast component is adjusted;

[0026] Acquire the number of directional transmission modules corresponding to the broadcast content item, and set a limit number value. When the number of corresponding directional transmission modules exceeds the limit number value, respectively reduce and prolong the voice broadcast speed of the directional transmission modules through speed adjustment until the broadcast content items of the directional transmission modules output the same content to obtain a matching node;

[0027] Based on the matching node, the directional transmission modules are handed over until the number of directional transmission modules is lower than the limit number value.

[0028] Further, the interaction information includes language keywords, and the acquisition method of the interaction information item includes: based on the voice acquisition module, language data of the visitor is acquired to obtain language input items, and keyword extraction and semantic analysis are performed based on the language input items to obtain language interaction items, and further obtain the interaction information item.

[0029] Further, the landing position acquisition method of the interaction information item includes:

[0030] Based on the camera acquisition module, the scene feature information is acquired, and the scene feature information includes the appearance feature of the explanation point. Based on the comparison result of the language interaction item and the scene feature information, the adaptive feature in the scene feature information is acquired, the position information of the adaptive feature is acquired, and then the landing point position of the interaction information item is obtained.

[0031] Further, the interaction information includes the limb direction and the line of sight direction, and the acquisition method of the interaction information item includes:

[0032] The limb information and the line of sight information of the tourist are acquired by the camera acquisition module, and the limb information item and the line of sight information item are obtained.

[0033] Based on the limb information item, the direction of the fingertips of the tourist is acquired by the infrared component, and the limb direction is obtained. Based on the line of sight information item, the direction of the line of sight of the tourist is acquired by the camera component, and the line of sight direction is obtained. The limb direction and the line of sight direction are combined to obtain the interaction information item.

[0034] Further, the landing point position acquisition method of the interaction information item includes:

[0035] Based on the camera acquisition module, the scene environment information is acquired, and the scene information item is obtained. Based on the combination result of the line of sight direction in the interaction information item and the scene information item, the landing point range of the line of sight of the tourist is acquired, and the landing point range item is obtained.

[0036] Based on the limb direction in the interaction information item, the extension line of the limb direction is acquired, and the pointing mark line is obtained. The landing point position of the pointing mark line in the landing point range item is acquired, and then the landing point position of the interaction information item is obtained.

[0037] An intelligent voice explanation system based on positioning triggering uses the above-mentioned intelligent voice explanation method based on positioning triggering, which includes:

[0038] The division module acquires the marking area of the target explanation point to obtain the target marking range item, acquires the specific explanation content of the target explanation point to obtain the comparison content item, and the comparison content item is used to represent the voice explanation information matched with the target explanation point.

[0039] The analysis module performs segmented splitting based on the comparison content item to obtain at least one content disassembly item, acquires the actual position information of the tourist in the target marking range item to obtain the real-time position item, and selects the matched content disassembly item as the broadcast content item based on the mapping information of the real-time position item and the content disassembly item.

[0040] The broadcast module sets a broadcast component, the broadcast component is a voice output component, performs the output of the broadcast content item based on the broadcast component, and performs voice explanation to obtain the explanation information item, so that the splitting of the explanation content of the current explanation point and the classified matching type tracking explanation according to the interested content of different tourists are realized.

[0041] The expansion module: the body interaction information and the language interaction information of the tourist are acquired through the interaction information acquisition module, the interaction information items are obtained, the feature extraction is carried out based on the interaction information items, the comprehensive feature items are obtained, the feature positioning is carried out based on the comprehensive feature items, the landing position of the comprehensive feature item is acquired, and the explanation point corresponding to the landing position is taken as the expanded explanation.

[0042] Compared with the prior art, the beneficial effects of the present application are:

[0043] The intelligent voice explanation method and system based on positioning triggering, through the strong coupling mapping of space position and content segmentation, the exhibits are deconstructed into a plurality of content disassembled items, the corresponding depth explanation is triggered through the real-time position of the tourist and the stay time threshold of the block, compared with the traditional "whole area unified broadcast" mode, the content matching efficiency is improved, and different tourists are adapted, and through the intelligent matching of the real-time quantity level division and the contraction ratio value, the target marking range is automatically adjusted with the density of the tourist, the design solves the mis-triggering pain point of the traditional fixed triggering area in the crowded scene, and the mis-triggering rate is reduced.

[0044] Meanwhile, through the analysis of the fingertip orientation and the line of sight focus by the infrared and camera modules, the cross-exhibit reference positioning is realized in combination with the voice keywords, the ray projection space positioning method is adopted, the body pointing extension line intersects with the line of sight landing point range, the landing position is positioned, the expanded explanation is automatically associated, and through the directional sound source speed coordination algorithm, when a plurality of tourists trigger the same block, the sound wave is synchronized through the speed of the adjusting module, finally, the single broadcast source is merged, and the hardware resource occupation is reduced. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 It is the overall flowchart of the present application;

[0046] Figure 2 It is the target marking range item acquisition flowchart of the present application;

[0047] Figure 3 It is the initial marking range item and search range diagram of the present application;

[0048] Figure 4 It is the target marking range item diagram of the present application;

[0049] Figure 5 It is the mapping information diagram of the real-time position item and the content disassembled item of the present application. DETAILED DESCRIPTION

[0050] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0051] In scenic spots and museums, location-based content explanation systems are increasingly popular, such systems obtain the location relationship between the visitor and the explanation point, and after the visitor enters the preset electronic fence area, the location is automatically identified and the corresponding voice explanation content is triggered, for example, when the visitor approaches a certain showcase, the system will play its forging process, legendary story, etc., however, first of all, different visitors may have different interests in the content of the same explanation point, and the existing system cannot classify and explain different content for different visitors, the system usually presets a set of standard explanation content, which cannot be adjusted according to the age, knowledge background, interest and other factors of the visitor, for some visitors who have a deep understanding of the historical background, the basic explanation provided by the system may be too simple, and for some visitors who are interested in specific technical details, the standard explanation may not meet their desire for knowledge, this "one-size-fits-all" approach results in a low matching degree between the explanation content and the visitor's needs, affecting the visitor's visit experience, secondly, the existing broadcast system lacks interaction with the visitor, which makes it unable to flexibly respond to various real-time situations, and the intelligent voice explanation method based on positioning triggering provided by the present application maps the exhibits into multiple content disassembled items through strong coupling of spatial position and content segmentation, triggers corresponding in-depth explanation through the real-time position of the visitor and the stay time threshold of the block, compared with the traditional "whole area unified broadcast" mode, the content matching efficiency is improved, and it is adapted to different visitors, at the same time, through real-time quantity level division and intelligent matching of shrinkage ratio value, the target marking range is automatically adjusted with the visitor density, this design solves the mis-triggering pain point of traditional fixed triggering area in crowded scenes, and the mis-triggering rate is reduced, as shown in Figure 1 The steps S100-S700 are included.

[0052] Step S100: Obtain the marking area of the target explanation point to obtain the target marking range item.

[0053] It should be noted that, as Figure 2As shown, the target marking range item acquisition method includes: acquiring the position information of the target explanation point to obtain a target position item, taking the target position item as a center marking point, setting an initial marking range, the initial marking range is 3m, based on the combination of the initial marking range and the center marking point, obtaining an initial marking range item; set the search range, the search range is a fixed range value, 5m, based on the search range, the number of tourists is searched to obtain a real-time number item; based on the number of tourists, the level is divided to obtain a number level item, which is divided into three levels, namely rare, normal and crowded, and the corresponding number of tourists is <3, 3-6 and >6 respectively, based on the number level item, the shrinkage ratio value is set, the shrinkage ratio value is 10%, 20% and 30% respectively, based on the comparison result of the real-time number item and the number level item, the target shrinkage value is obtained, based on the combination result of the target shrinkage value and the initial marking range item, the target marking range is obtained.

[0054] In the specific implementation process, as shown in Figures 3-4 As shown, there are A, B, C three exhibition points in a certain exhibition room, which are respectively taken as target explanation points, at this time, the position information of the three exhibition points is acquired to obtain target position items, taking the target position items as center marking points, according to the set initial marking range, three initial marking range items are obtained, at this time, according to the set search range, the number of people in the search range of A is 5, the number of people in the search range of B is 2, and the number of people in the search range of C is 8, then the number level items of the tourists of A, B and C three exhibition points are normal, rare and crowded respectively, according to the set shrinkage ratio value, the target marking range items of A, B and C three exhibition points are 2.4m, 2.7m and 2.1m respectively, and then the effect of automatically adjusting the marking area size according to the number of tourists is realized, which prevents the marking area from being accidentally touched due to too many tourists.

[0055] Step S200: acquiring the specific explanation content of the target explanation point to obtain a comparison content item.

[0056] It should be noted that the comparison content item is used to represent the voice explanation information matched with the target explanation point, and the voice explanation information corresponding to different scenic spots is different, for example, in a bronze ware target explanation point, there are not only forging process, historical information, but also historical event information, therefore, the specific explanation content of the target explanation point is acquired by big data acquisition.

[0057] Step S300: based on the comparison content item, the segmented splitting is performed to obtain at least one content splitting item.

[0058] It should be noted that the content disassembly item is used to represent the classified explanation information of the target explanation point, and the method for obtaining the content disassembly item includes: obtaining high-precision scanning data of the target explanation point to obtain scanning information items, decomposing the scanning information items into digital units based on the scanning information items for calculation to obtain an explanation information set; based on the explanation information set, structural features are identified through an AI segmentation algorithm to obtain a structural division set, and the structural division set includes at least one category structural feature; based on the mapping information of the structural division set and the explanation information set, the contrast content item is segmented and split to obtain the content disassembly item.

[0059] Specifically, first, the scanning data of the target explanation point is obtained through laser radar point cloud scanning, including a phase laser radar, a structured light scanner and a terahertz imager. For example, it takes 48 hours to scan the kneeling terracotta of the Qin Dynasty, generating 1.2 billion point cloud data. Then, through AI-enabled semantic segmentation, the algorithm architecture is:

[0060] import torch

[0061] from mmdet3d.apis import init_model

[0062] # Load pre-trained cultural relic segmentation model (Qin culture special weight)

[0063] model = init_model(

[0064] config='configs / voxelnet / voxelnet_terracotta.py',

[0065] checkpoint='terracotta_voxelnet_2025.pth' )

[0067] # Voxel processing (0.5mm³ voxel size)

[0068] voxel_size = [0.005, 0.005, 0.005] # Unit: meter

[0069] voxel_features = voxelize(pointcloud, voxel_size)

[0070] preds = model(voxel_features) # Output semantic label map

[0071] The structure division set includes a structure feature area, a material variation area, a symbolic inscription area, and a micro-damage area, wherein the structure feature area is a physical structure such as a joint of the piece or a chignon line, the material variation area includes a bronze corrosion or a pigment peeling area, the symbolic inscription area includes an inscription carving or a painting title and postscript position, and the micro-damage area includes an area with a crack depth greater than 0.1 mm. Based on the mapping information of the structure division set and the explanation information set, the comparison content item is segmented and split to obtain the content split item.

[0072] Step S400: Obtain a real-time position item by acquiring actual position information of the tourist in the target marker range item. Based on the mapping information of the real-time position item and the content split item, a matched content split item is selected as a broadcast content item.

[0073] It should be noted that the mapping information of the real-time position item and the content split item includes: based on the content split item, the target marker range item is segmented and divided to obtain at least one range marker item corresponding to the content split item; based on the comparison data of the real-time position item and the range marker item, including position comparison and time comparison, a retrieval threshold is set, the retrieval threshold is a fixed time value, which is 5s, when the comparison data reaches the retrieval threshold, the corresponding marker range item is determined as a mapping item, and then the mapping information of the real-time position item and the content split item is obtained.

[0074] In the specific implementation process, as shown in Figure 5 , an existing painting craft exhibit is split into three content split items from left to right, and the presented contents are painting origin, painting material, and painting history, respectively, which account for one third of the painting craft exhibit. An existing tourist a is in the middle position of the painting craft exhibit and stays for 6s. At this time, according to the set retrieval threshold, the marker range item corresponding to the tourist a is determined as a mapping item, that is, the painting material, and then the mapping information of the real-time position item and the content split item is obtained.

[0075] Step S500: Set a broadcast component, output the broadcast content item based on the broadcast component, and obtain an explanation information item.

[0076] It should be noted that the broadcast component is a voice output component, the output of the broadcast content item based on the broadcast component is voice explanation, and the explanation information item is obtained, so that the current explanation point explanation content is split and classified matching tracking explanation is realized according to the interested content of different tourists.

[0077] It should be noted that the method for obtaining the explanation information item includes: the broadcasting component is composed of at least one directional transmission module, based on the real-time position item, the head data of the visitor is obtained, including the head direction, the target position item is obtained, and the orientation of the broadcasting component is adjusted based on the target position item; the number of directional transmission modules corresponding to the broadcasting content item is obtained, the limit number value is set, the limit number value is 1, when the number of corresponding directional transmission modules exceeds the limit number value, the voice broadcasting speed of the directional transmission module is respectively reduced and prolonged by adjusting the speech speed, until the output content of the broadcasting content item of the directional transmission module is the same, and a matching node is obtained; based on the matching node, the directional transmission module is handed over until the number of directional transmission modules is lower than the limit number value.

[0078] In the specific implementation process, the existing certain painting craft exhibit is split into two content disassembly items from left to right, and the presented contents are the origin of the painting and the history of the painting, respectively accounting for one half of the painting craft exhibit, and the existing visitor b is located at the left position of the painting craft exhibit and stays for 6s, at this time, according to the set search threshold, the marking range item corresponding to the visitor b is determined as the mapping item, that is, the origin of the painting, and then the mapping information of the real-time position item and the content disassembly item is obtained, according to the head position and head direction of the visitor b, one directional transmission module bb in the broadcasting component is used for broadcasting, at this time, another visitor c also comes to the left position and stays for 5s, at this time, according to the set search threshold, the marking range item corresponding to the visitor c is determined as the mapping item, that is, the origin of the painting, and then the mapping information of the real-time position item and the content disassembly item is obtained, according to the head position and head direction of the visitor c, one directional transmission module cc in the broadcasting component is used for broadcasting, since the number of corresponding directional transmission modules at this time exceeds the limit number value, at this time, the voice broadcasting speed of bb is reduced, and the voice broadcasting speed of cc is simultaneously increased, until the output content of the broadcasting content item of bb and cc is the same, a matching node is obtained, at this time, the directional transmission modules are handed over, and b and c are simultaneously broadcasted by bb or cc, so that the working mode optimization effect of the directional transmission module is realized.

[0079] Step S600: obtaining the body interaction information and the language interaction information of the visitor through the interaction information acquisition module to obtain the interaction information item.

[0080] It should be noted that the interaction information includes language keywords, and the method for obtaining the interaction information item includes: obtaining the language data of the visitor based on the voice acquisition module to obtain the language input item, performing keyword extraction and semantic analysis based on the language input item to obtain the language interaction item, and then obtaining the interaction information item.

[0081] Specifically, by acquiring language data of the visitor such as "look at this", "look at that bronze sword", and the like, keyword extraction and semantic analysis are performed to obtain the interaction information item.

[0082] It should be noted that the interaction information includes body pointing and line-of-sight direction, and the method for obtaining the interaction information item includes: acquiring body information and line-of-sight information of the visitor by the camera acquisition module to obtain the body information item and the line-of-sight information item; based on the body information item, acquiring the direction of the visitor's fingertips by the infrared component to obtain the body pointing, and based on the line-of-sight information item, acquiring the direction of the visitor's line of sight by the camera component to obtain the line-of-sight pointing, and the body pointing and the line-of-sight pointing are combined to obtain the interaction information item.

[0083] Specifically, the body information and the line-of-sight information of the visitor are acquired by the camera acquisition module, the direction of the visitor's fingertips is acquired by the infrared component, and the direction of the visitor's line of sight is acquired by the camera component to obtain the interaction information item.

[0084] Step S700: Based on the interaction information item, the landing position of the interaction information item is acquired, and the explanation point corresponding to the landing position is taken as the expanded explanation.

[0085] It should be noted that after the landing position is acquired, the contrast content item of the expanded explanation is acquired in a data transmission and data sharing manner, and the expanded explanation of the explanation point other than the current explanation point is realized according to the interaction information of the visitor.

[0086] It should be noted that the method for obtaining the landing position of the interaction information item includes: based on the camera acquisition module, acquiring scene feature information, the scene feature information includes appearance features of the explanation points, based on the contrast result of the language interaction item and the scene feature information, acquiring the adaptive feature in the scene feature information, acquiring the position information of the adaptive feature, and then obtaining the landing position of the interaction information item.

[0087] Specifically, the feature information in the scene is acquired by the camera acquisition module, including appearance features, color features, and the like of the explanation points, such as one red exhibit, one yellow exhibit, one white exhibit, one disc-shaped exhibit, and the like, then the most adaptive feature is acquired according to the contrast result of the language interaction item and the scene feature information, and the landing position is determined according to the position information of the feature.

[0088] It should be noted that the drop point position acquisition method of the interaction information item includes: based on the camera acquisition module, acquiring the scene environment information to obtain the scene information item, based on the combination result of the line of sight pointing in the interaction information item and the scene information item, acquiring the line of sight drop point range of the tourist to obtain the drop point range item; based on the body pointing in the interaction information item, acquiring the body pointing extension line to obtain the pointing mark line, acquiring the drop point position of the pointing mark line in the drop point range item, and then obtaining the drop point position of the interaction information item.

[0089] Specifically, after the line of sight pointing and the body pointing of the tourist are acquired, the line of sight drop point range is acquired according to the line of sight pointing through the camera acquisition module, that is, the camera, and the extension line position of the body pointing is acquired according to the infrared component, the drop point position of the extension line position in the drop point range item is acquired, and then the drop point position of the interaction information item is obtained.

[0090] An intelligent voice explanation system based on positioning triggering uses the above-mentioned intelligent voice explanation method based on positioning triggering, which includes: a division module: acquiring the marking area of the target explanation point to obtain the target marking range item, acquiring the specific explanation content of the target explanation point to obtain the contrast content item, and the contrast content item is used to represent the voice explanation information matched with the target explanation point; an analysis module: based on the contrast content item, segmenting and splitting to obtain at least one content splitting item, acquiring the actual position information of the tourist in the target marking range item to obtain the real-time position item, based on the mapping information of the real-time position item and the content splitting item, selecting the matched content splitting item as the broadcast content item; a broadcast module: setting a broadcast component, the broadcast component is a voice output component, based on the broadcast component, outputting the broadcast content item to perform voice explanation to obtain the explanation information item, thereby realizing the splitting of the current explanation point explanation content and the classified matching tracking explanation according to the interesting content of different tourists; an expansion module: acquiring the body interaction information and the language interaction information of the tourist through the interaction information acquisition module to obtain the interaction information item, based on the interaction information item, extracting the features to obtain the comprehensive feature item, based on the comprehensive feature item, positioning the features to acquire the drop point position of the comprehensive feature item, and taking the explanation point corresponding to the drop point position as the expanded explanation, in the mode of data transmission and data sharing, acquiring the broadcast of the contrast content item of the expanded explanation.

[0091] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended embodiments and their equivalents.

Claims

1. A positioning trigger-based intelligent voice guide method, comprising: obtaining a target guide point marking area to obtain a target marking range item, obtaining specific guide content of the target guide point to obtain a contrast content item, and the contrast content item being used to represent voice guide information matched with the target guide point; characterized in that: segmenting and splitting the contrast content item based on the contrast content item to obtain at least one content disassembly item, and the content disassembly item being used to represent classified guide information of the target guide point; obtaining real-time position information of a tourist in the target marking range item to obtain a real-time position item, selecting a matched content disassembly item as a broadcast content item based on mapping information of the real-time position item and the content disassembly item; setting a broadcast component, the broadcast component being a voice output component, outputting the broadcast content item based on the broadcast component, performing voice guide to obtain guide information items, thereby realizing disassembly of guide content of a current guide point and classified matching type tracking guide according to interesting content of different tourists; obtaining body interaction information and language interaction information of the tourist through an interaction information acquisition module to obtain an interaction information item; positioning information based on the interaction information item, obtaining a landing point position of the interaction information item, taking a guide point corresponding to the landing point position as an expanded guide, and obtaining a contrast content item of the expanded guide in a data transmission and data sharing manner, thereby realizing expanded guide of guide points other than the current guide point according to the interaction information of the tourist. 2.The method of claim 1, wherein the method further comprises: The target marking range item acquisition method comprises: obtaining position information of the target guide point to obtain a target position item, taking the target position item as a central marking point, setting an initial marking range, and obtaining an initial marking range item based on the combination of the initial marking range and the central marking point; setting a search range, the search range being a fixed range value, searching the number of tourists based on the search range to obtain a real-time number item; grading based on the number of tourists to obtain a number grade item, setting a contraction ratio value based on the number grade item, obtaining a target contraction value based on the contrast result of the real-time number item and the number grade item, and obtaining a target marking range based on the combination result of the target contraction value and the initial marking range item. 3.The method of claim 1, wherein the method further comprises: The content disassembly item acquisition method comprises: obtaining high-precision scanning data of the target guide point to obtain a scanning information item, decomposing the scanning information item into digital units based on the scanning information item for calculation to obtain a guide information set; identifying structural features based on the guide information set through an AI segmentation algorithm to obtain a structure division set, the structure division set including at least one category structural feature; segmenting and splitting the contrast content item based on the mapping information of the structure division set and the guide information set to obtain the content disassembly item. 4.The method of claim 1, wherein the method further comprises: The mapping information acquisition method of the real-time position item and the content disassembly item comprises: segmenting and dividing the target marking range item based on the content disassembly item to obtain at least one range marking item corresponding to the content disassembly item; Based on the comparison data of real-time location items and range marker items, including location comparison and time comparison, a retrieval threshold is set, the retrieval threshold is a fixed time value, when the comparison data reaches the retrieval threshold, the corresponding marker range item is determined as a mapping item, and then the mapping information of the real-time location item and the content disassembly item is obtained.

5. The method of claim 1, wherein the method further comprises: The obtaining method of the explanation information item includes: The broadcast component is composed of at least one directional transmission module, based on the real-time location item, the head data of the tourist is obtained, including the head direction, the target location item is obtained, and the direction of the broadcast component is adjusted based on the target location item; The number of directional transmission modules corresponding to the broadcast content item is obtained, and a limit number value is set, when the number of corresponding directional transmission modules exceeds the limit number value, the voice broadcast speed of the directional transmission module is respectively reduced and prolonged by adjusting the speech speed, until the output content of the broadcast content item of the directional transmission module is the same, and a matching node is obtained; Based on the matching node, the directional transmission module is handed over until the number of directional transmission modules is lower than the limit number value. 6.The method of claim 1, wherein the method further comprises: The interactive information includes language keywords, and the obtaining method of the interactive information item includes: based on the voice acquisition module, the language data of the tourist is obtained to obtain a language input item, keyword extraction and semantic analysis are performed based on the language input item to obtain a language interaction item, and then the interactive information item is obtained.

7. The method of claim 6, wherein the method further comprises: The drop point position obtaining method of the interactive information item includes: Based on the camera acquisition module, scene feature information is obtained, the scene feature information includes the appearance characteristics of the explanation point, based on the comparison result of the language interaction item and the scene feature information, the adaptive feature in the scene feature information is obtained, the position information of the adaptive feature is obtained, and then the drop point position of the interactive information item is obtained. 8.The method of claim 1, wherein the method further comprises: determining whether the user is in the predetermined area; and if the user is in the predetermined area, providing the user with a voice guide. The interactive information includes limb pointing and line of sight direction, and the obtaining method of the interactive information item includes: The limb information item and the line of sight information item are obtained by the camera acquisition module to obtain the limb information and the line of sight information of the tourist; Based on the limb information item, the fingertip direction of the tourist is obtained by the infrared component to obtain the limb pointing, based on the line of sight information item, the line of sight direction of the tourist is obtained by the camera component to obtain the line of sight pointing, and the limb pointing and the line of sight pointing are combined to obtain the interactive information item. 9.The method of claim 8, wherein the method further comprises: The drop point position obtaining method of the interactive information item includes: Based on the camera acquisition module, scene environment information is obtained to obtain a scene information item, based on the combination result of the line of sight pointing in the interactive information item and the scene information item, the line of sight drop point range of the tourist is obtained to obtain a drop point range item; Based on the limb pointing in the interactive information item, the limb pointing extension line is obtained to obtain a pointing mark line, the drop point position of the pointing mark line in the drop point range item is obtained, and then the drop point position of the interactive information item is obtained.

10. A positioning trigger-based intelligent voice guide system, characterized in that: A kind of intelligent voice explanation method based on positioning trigger is used, including: The division module: the marker area of target explanation point is obtained to obtain a target marker range item, the specific explanation content of target explanation point is obtained to obtain a comparison content item, and the comparison content item is used to represent the voice explanation information matched with the target explanation point; The analysis module: based on the control content item is segmented split, get at least one content analysis item, by getting the actual position information of the tourist in the target marker range item, get real-time location item, based on real-time location item and content analysis item mapping information, select the matching content analysis item as the broadcast content item; The broadcast module: set the broadcast component, the broadcast component is a voice output component, based on the broadcast component to output the broadcast content item, voice explanation, get the explanation information item, so as to realize the current explanation point explanation content split and according to different tourist's interest content classification matching type tracking explanation; The expansion module: through the interactive information acquisition module to obtain the body interaction information and language interaction information of the tourist, get the interactive information item, based on the interactive information item to extract features, get the comprehensive feature item, based on the comprehensive feature item to locate the feature, get the landing position of the comprehensive feature item, take the explanation point corresponding to the landing position as the expansion explanation, in the form of data transmission and data sharing, get the broadcast of the control content item of the expansion explanation.

Citation Information

Patent Citations

  • Automatic identification, explanation and enhanced display method for scenic spot based on AR technology

    CN112967641A

  • Display device for cultural and artistic communication planning

    CN114259155A

  • Method and apparatus for associating audio objects with content and geo-location

    US20140079225A1