Exhibit guidance methods and related devices, mobile terminals and storage media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]然而,无论是扫码的方式,还是语音讲解器的方式,在导览过程中均鲜有交互,互动体验较差
[0026]上述方案,响应于在移动终端的拍摄画面中识别到目标展品,对拍摄画面进行检测,得到拍摄画面中目标展品的关键部位的图像位置,且目标展品为预先设有讲解数据的展品,讲解数据包括目标展品各个关键部位的讲解信息,基于此在图像位置处显示AR指示标识,并响应于用户触发AR指示标识,输出关键部位的讲解信息,故在展品导览过程中,用户能够通过移动终端拍摄展品进行识别,并在识别为目标展品的情况下,通过目标展品的关键部位在拍摄画面中图像位置上显示AR指示标识来与用户进行交互互动,从而能够通过互动触发讲解的交互形式实现展品导览,并提升讲解效率,进而能够提升导览体验。
Smart Images

Figure CN115345927B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to an exhibit navigation method and related devices, mobile terminals and storage media. Background Technology
[0002] Currently, guided tours are typically conducted by human guides, which is costly. With the development of electronic information technology, methods such as QR code scanning and audio guides are becoming increasingly popular for providing audio explanations.
[0003] However, neither the QR code scanning method nor the audio guide method offers much interaction during the guided tour, resulting in a poor user experience. Therefore, improving the guided tour experience has become an urgent issue to be addressed. Summary of the Invention
[0004] This application provides an exhibit guidance method and related devices, mobile terminals, and storage media.
[0005] The first aspect of this application provides an exhibit guidance method, comprising: in response to identifying a target exhibit in a shooting screen of a mobile terminal, detecting the shooting screen to obtain the image position of key parts of the target exhibit in the shooting screen; wherein the target exhibit is an exhibit with pre-set explanatory data, and the explanatory data includes explanatory information of each key part of the target exhibit; displaying an AR indicator at the image position; and in response to the user triggering the AR indicator, outputting explanatory information of the key parts.
[0006] Therefore, in response to the identification of a target exhibit in the mobile terminal's camera view, the camera view is detected to obtain the image position of the key parts of the target exhibit in the camera view. The target exhibit is an exhibit with pre-set explanatory data, which includes explanatory information for each key part of the target exhibit. Based on this, an AR indicator is displayed at the image position. In response to the user triggering the AR indicator, the explanatory information for the key parts is output. Thus, during the exhibit tour, the user can identify the exhibit by taking a picture of it with the mobile terminal. When the exhibit is identified as the target exhibit, the AR indicator is displayed at the image position of the key parts of the target exhibit in the camera view to interact with the user. This interactive form of triggering explanations can realize exhibit tours, improve explanation efficiency, and enhance the tour experience.
[0007] Before detecting the target exhibit in the captured image of the mobile terminal and obtaining the image position of the key part of the target exhibit in the captured image, the method further includes: matching the captured image with each exhibit model to obtain matching results; wherein the matching results include: the matching degree of the current exhibit with each exhibit model; and analyzing the matching results to obtain the recognition result of the captured image; wherein the recognition result includes: whether the current exhibit is the target exhibit.
[0008] Therefore, the captured images are matched with each exhibit model to obtain matching results, which include the matching degree between the current exhibit and each exhibit model. Based on the matching results, the recognition results of the captured images are obtained, which include whether the current exhibit is the target exhibit. Thus, the presence of the target exhibit in the captured images can be identified through model matching, which can help improve the accuracy of exhibit recognition.
[0009] The process involves analyzing the matching results to obtain the recognition results of the captured images, including: in response to the maximum matching degree being higher than a preset threshold, taking the exhibit model corresponding to the maximum matching degree as the target exhibit, and determining that the recognition results include the current exhibit as the target exhibit.
[0010] Therefore, when the maximum matching degree is greater than the preset threshold, the exhibit to which the exhibit model corresponding to the maximum matching degree belongs is taken as the target exhibit, and the recognition result is determined to include the current exhibit as the target exhibit. Thus, in the exhibit recognition process, the exhibit model with the maximum matching degree with the current exhibit and the matching degree is greater than the preset threshold can be selected, and the target exhibit and recognition result can be determined based on this, which is conducive to further improving the accuracy of identifying the target exhibit.
[0011] The exhibit model contains several key parts, and the exhibit model has explanatory information for the key parts marked on the model position corresponding to the key parts.
[0012] Therefore, the exhibit model contains several key parts, and the exhibit model marks the explanatory information of the key parts at the corresponding model positions. This allows the key parts in the captured image to be detected directly based on the model positions marked with explanatory information, and the explanatory information of the key parts can be obtained directly from the exhibit model. In other words, both the detection of key parts and the acquisition of explanatory information can be achieved based on the exhibit model, which helps to improve the efficiency of exhibit tours.
[0013] The target exhibit has multiple key parts, and the method further includes: in response to the existence of undetected key parts, outputting a first prompt; wherein the first prompt is used to prompt the mobile terminal to adjust the shooting posture in order to capture the undetected key parts.
[0014] Therefore, the target exhibit has multiple key parts, and in response to the existence of undetected key parts, a first prompt is output. The first prompt is used to prompt the user to adjust the shooting position of the mobile terminal to capture the undetected key parts. Thus, even when the target photo has multiple key parts and there are still undetected key parts, the first prompt can prompt the user to understand all the key parts of the target exhibit during the visit, which is conducive to improving the interactive experience of exhibit tour.
[0015] Before outputting the first prompt, the method also includes: obtaining the first position of the undetected key part on the target exhibit and obtaining the current pose of the mobile terminal; and analyzing the first position and the current pose to obtain the shooting pose that the mobile terminal needs to adjust.
[0016] Therefore, before outputting the first prompt, the first position of the undetected key part on the target exhibit is obtained, and the current pose of the mobile terminal is obtained. Based on the first position and the current pose, the shooting pose that the mobile terminal needs to adjust is analyzed. Thus, during the exhibit tour, even if there is an undetected key part, the shooting pose that the mobile terminal needs to adjust can be determined by combining the first position of the undetected key part on the target exhibit and the current pose of the mobile terminal, thereby improving the accuracy of the shooting pose.
[0017] The method further includes: in response to the user adjusting the mobile terminal to a new shooting position and recognizing the target exhibit in the new shooting screen of the mobile terminal, displaying an AR indicator in the new shooting position.
[0018] Therefore, in response to the user adjusting the mobile terminal to a new shooting position and recognizing the target exhibit in the new shooting screen of the mobile terminal, an AR indicator is displayed in the new shooting position. This enables the AR indicator to follow the user's shooting position after adjustment, thereby improving the interactive experience of exhibit guidance.
[0019] The method further includes: in response to the existence of related exhibits in the exhibition hall that are related to the target exhibit, outputting a second prompt; wherein the second prompt is used to prompt the user to visit the related exhibits.
[0020] Therefore, in response to the existence of related exhibits in the exhibition hall that are related to the target exhibit, a second prompt is output, which is used to prompt the user to visit the related exhibits. In other words, when there are related exhibits in the exhibition hall that are related to the target exhibit that the user is currently visiting, the second prompt is output to prompt the user to visit the related exhibits, thereby satisfying the user's tour needs for exhibits of interest as much as possible and improving the interactive experience of exhibit tour.
[0021] The method further includes, after outputting the second prompt, the following steps: in response to receiving a confirmation instruction from the user regarding visiting the related exhibits, obtaining the second location of the related exhibits in the exhibition hall and obtaining the current pose of the mobile terminal; and displaying an AR navigation icon on the current screen of the mobile terminal based on the second location and the current pose.
[0022] Therefore, after outputting the second prompt, in response to receiving the user's confirmation instruction to visit the related exhibits, the system obtains the second location of the related exhibits in the exhibition hall and the current pose of the mobile terminal. Based on the second location and the current pose, an AR navigation icon is displayed on the current screen of the mobile terminal. Thus, after the user confirms that they are visiting the related exhibits, the system can combine the second location of the related exhibits in the exhibition hall with the current pose of the mobile terminal to guide the user by displaying an AR navigation icon on the current screen of the mobile terminal, which helps to improve the interactive experience of exhibit tours.
[0023] A second aspect of this application provides an exhibit guidance device, comprising: a detection module, an identification module, and an interaction module. The detection module is used to detect the captured image in response to the identification of a target exhibit in the captured image of a mobile terminal, and to obtain the image position of key parts of the target exhibit in the captured image; wherein the target exhibit is an exhibit with pre-set explanatory data, and the explanatory data includes explanatory information for each key part of the target exhibit; the identification module is used to display AR indicator identification at the image position; and the interaction module is used to output explanatory information for the key parts in response to the user triggering the AR indicator identification.
[0024] A third aspect of this application provides a mobile terminal, including a camera, a display screen, a memory, and a processor. The camera, display screen, and memory are respectively coupled to the processor, which executes program instructions stored in the memory to implement the exhibit guidance method in the first aspect.
[0025] The fourth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the exhibit guidance method of the first aspect described above.
[0026] The above solution, in response to identifying a target exhibit in the mobile terminal's camera view, detects the image of the captured image to obtain the image location of key parts of the target exhibit. The target exhibit is pre-loaded with explanatory data, including information on each key part of the target exhibit. Based on this, an AR indicator is displayed at the image location. In response to the user triggering the AR indicator, the explanatory information for the key parts is output. Therefore, during exhibit tours, users can identify exhibits by taking pictures with their mobile terminals. If the exhibit is identified as a target exhibit, AR indicators are displayed at the image location of the key parts of the target exhibit in the captured image to interact with the user. This interactive approach to triggering explanations improves exhibit tour efficiency and enhances the overall tour experience. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating one embodiment of the exhibit guidance method of this application;
[0028] Figure 2a This is a schematic diagram illustrating the effect of one embodiment of the exhibit guidance method of this application;
[0029] Figure 2b This is a schematic diagram illustrating the effect of another embodiment of the exhibit guidance method of this application;
[0030] Figure 2c This is a schematic diagram illustrating the effect of another embodiment of the exhibit guidance method of this application;
[0031] Figure 2d This is a schematic diagram illustrating the effect of another embodiment of the exhibit guidance method of this application;
[0032] Figure 3a This is a schematic diagram illustrating the effect of another embodiment of the exhibit guidance method of this application;
[0033] Figure 3b This is a schematic diagram illustrating the effect of another embodiment of the exhibit guidance method of this application;
[0034] Figure 4 This is a schematic diagram of the framework of an embodiment of the exhibit guidance device of this application;
[0035] Figure 5 This is a schematic diagram of the framework of an embodiment of the mobile terminal of this application;
[0036] Figure 6 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0037] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0038] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.
[0039] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper means two or more.
[0040] Please see Figure 1 , Figure 1 This is a flowchart illustrating one embodiment of the exhibit guidance method of this application.
[0041] Specifically, this may include the following steps:
[0042] Step S11: In response to the detection of the target exhibit in the shooting screen of the mobile terminal, the shooting screen is detected to obtain the image position of the key part of the target exhibit in the shooting screen.
[0043] In this embodiment of the disclosure, the target exhibit is an exhibit with pre-set explanatory data, and the explanatory data includes explanatory information for each key part of the target exhibit.
[0044] In one implementation scenario, all exhibits in the exhibition hall can have pre-set explanatory data, meaning that all exhibits in the exhibition hall can be target exhibits; or, only some exhibits in the exhibition hall can have pre-set explanatory data, meaning that only some exhibits in the exhibition hall can be target exhibits. For example, exhibits that are relatively rare, unique, or distinctive in the exhibition hall can be selected as target exhibits, and explanatory data can be pre-set for these target exhibits.
[0045] In one implementation scenario, the target exhibit can have at least one key part; that is, the target exhibit can have one key part, two key parts, or three or more key parts, without limitation. For example, taking a porcelain vase as an example, its key parts can include, but are not limited to: the mouth, neck, foot ring, and dragon-shaped handles, etc. Other cases can be deduced similarly, and will not be listed in detail here.
[0046] In one implementation scenario, the explanatory information for key parts can include, but is not limited to, text information, image information, audio information, and video information. Of course, it can also be a combination of these types of information. For example, videos explaining the key parts of the target exhibit by experts, scholars, or researchers in the field can be pre-recorded to serve as explanatory information for each key part. Alternatively, the explanatory videos can be condensed into text information, or only the audio information from the explanatory videos can be extracted as explanatory information; there are no limitations on this.
[0047] In one implementation scenario, to facilitate the identification of the presence of a target exhibit from the mobile terminal's captured image, multiple images of the target exhibit can be pre-captured from various angles. These images are then combined to form an image set of the target exhibit. For example, the target exhibit can be captured from multiple perspectives, such as frontal view, side view (e.g., left view, right view), and top view; this is not a limitation. Furthermore, to minimize subsequent recognition interference, other objects can be excluded from the frame during capture, or objects other than the target exhibit can be removed from the image after capture. Based on this, when recognizing the captured image, it is only necessary to compare the captured image with each target exhibit's image set to obtain a comparison score for each target exhibit. If the maximum comparison score is greater than a preset threshold, it can be determined that a target exhibit exists in the captured image, and the target exhibit in the captured image can be identified as the target exhibit corresponding to the maximum comparison score. Specifically, the image region of the current exhibit in the captured image can be detected first. Based on this, for each target exhibit's image set, the similarity between each image in the image set and the aforementioned image region can be obtained. The similarity between each image in the image set can be statistically analyzed (e.g., summed, averaged, weighted, etc.) to obtain the comparison score of the target exhibit.
[0048] In a specific implementation scenario, to improve the accuracy and efficiency of detecting the image region of the current exhibit in the captured image, an object detection network can be pre-trained to detect the image region of the current exhibit. Based on this, the object detection network can be used to detect the image region of the target exhibit in the captured image. For example, the object detection network can include, but is not limited to, convolutional neural networks, etc. During the training process of the object detection network, several sample images of exhibits can be pre-collected, and the sample regions of the exhibits in the sample images can be labeled. Based on this, the object detection network can be used to detect the sample images to obtain the predicted region of the exhibit in the sample images, and the network parameters of the object detection network can be adjusted based on the difference between the sample images and the predicted regions. It should be noted that the specific method for measuring the difference can refer to loss functions such as cross-entropy, and the specific method for adjusting the parameters can refer to optimization methods such as gradient descent, which will not be elaborated here.
[0049] In a specific implementation scenario, to improve the accuracy and efficiency of similarity measurement, a feature extraction network can be pre-trained to extract image features. Based on this, the feature extraction network can extract the first feature of the aforementioned image region, and then extract the second feature of each image within the image set. The similarity between each image and the aforementioned image region is then obtained based on the similarity between the first and second features (e.g., cosine similarity). For example, the feature extraction network can include, but is not limited to, convolutional layers; the network structure is not limited here. During the training of the feature extraction network, sample images of several exhibits can be pre-collected, and the feature extraction network can extract sample features from each sample image. For each sample image, images belonging to the same exhibit can be used as positive examples, and images belonging to different exhibits can be used as negative examples. The loss value of the feature extraction network can then be obtained based on the differences between the sample features of a sample image and the sample features of its positive and negative examples. Based on this, the network parameters of the feature extraction network can be adjusted. It should be noted that the specific methods for measuring the difference can be found in loss functions such as triplet loss, and the specific methods for adjusting the parameters can be found in optimization methods such as gradient descent, which will not be elaborated here.
[0050] In a specific implementation scenario, the image set of the target exhibit can include not only images taken from different angles, but also images of various key parts of the target exhibit. When the target exhibit is identified in the captured image, feature point matching can be performed between the image region of the current exhibit in the captured image and the images of each key part of the target exhibit, obtaining the matching degree of the key parts and their image positions in the captured image. Based on this, key parts with matching degrees greater than a preset threshold can be selected, confirming the presence of these selected key parts in the captured image. It should be noted that the aforementioned feature point matching can include, but is not limited to, ORB (Oriented Fast and Rotated BRIEF) and SIFT (Scale Invariant Feature Transform), etc., and is not limited here.
[0051] In one implementation scenario, to improve the accuracy of identifying target exhibits, instead of pre-collecting image sets of each target exhibit as mentioned above, exhibit models for each target exhibit can be pre-constructed. Based on this, the captured image can be matched against the exhibit models of each target exhibit to obtain matching results. These matching results include the degree of matching between the current exhibit in the captured image and each exhibit model. Analysis can then be performed based on the matching results to obtain the identification result of the captured image, including whether the current exhibit is a target exhibit. This method can identify the presence of target exhibits in the captured image through model matching, thereby improving the accuracy of exhibit identification.
[0052] In a specific implementation scenario, the target exhibit can be photographed from different angles to obtain a series of two-dimensional images containing visual motion information. These two-dimensional images are then reconstructed into a three-dimensional model based on methods such as SFM (Structure From Motion) to obtain the exhibit model. The specific process of three-dimensional reconstruction can be found in the technical details of three-dimensional reconstruction methods such as SFM, and will not be elaborated here. Furthermore, in this embodiment, the exhibit model can undergo model rendering and other operations to give it the texture, color, and other features of the target exhibit. For example, still taking the target exhibit "porcelain vase" as an example, the final generated exhibit model not only has the shape features of the target exhibit "porcelain vase," but also its texture features, such as lines of varying thickness on the surface of the "porcelain vase," and its color features, such as various glazes on the surface of the "porcelain vase," etc., which are not limited here.
[0053] In a specific implementation scenario, as mentioned earlier, the image region of the current exhibit in the captured image can be detected first, and then matched with the exhibit models of each target exhibit based on this image region. The specific detection method for the image region can be found in the aforementioned descriptions and will not be repeated here. Based on this, for each target exhibit's exhibit model, it can be matched with the aforementioned image region to obtain the matching degree between the target exhibit's exhibit model and the current exhibit. Specifically, based on the two-dimensional coordinates and depth values of pixels in the image region, the camera pose of the mobile terminal, and camera intrinsic parameters, the pixels can be projected into three-dimensional space to obtain projection points, and point cloud data composed of these projection points can be acquired. This allows for comparison of the matching degree between the point cloud data and the exhibit models of each target exhibit. It should be noted that the aforementioned camera pose can be determined based on visual positioning methods such as SLAM (Simultaneous Localization and Mapping). For details on the technical details of visual positioning methods such as SLAM, please refer to the technical details of SLAM and other visual positioning methods; these will not be repeated here.
[0054] In a specific implementation scenario, unlike the previous approach of first detecting the image region of the current exhibit and then matching it with the exhibit model, we can also first extract feature points from the captured image and then match each exhibit model based on the extracted feature points. It should be noted that feature points can include, but are not limited to, structural feature points and texture feature points, etc., and are not limited here. For the specific process of feature point extraction, please refer to the technical details of feature point detection methods such as ORB (Oriented Fast and Rotated BRIEF) and SURF (Speeded-Up Robust Features), which will not be elaborated upon here.
[0055] In a specific implementation scenario, after obtaining the matching degree between the exhibit models of each target exhibit and the current exhibit, the highest matching degree can be obtained. It is then checked whether the highest matching degree is greater than a preset threshold. If so, the exhibit to which the exhibit model corresponding to the highest matching degree belongs can be taken as the target exhibit, and the recognition result includes the current exhibit as a target exhibit. In other words, if the highest matching degree is greater than the preset threshold, the exhibit to which the exhibit model corresponding to the highest matching degree belongs can be directly identified as the current exhibit appearing in the captured image. Conversely, if the highest matching degree is less than the preset threshold, it can be considered that there is no target exhibit in the captured image. It should be noted that the preset threshold can be set according to the actual application requirements. For example, if the accuracy requirement for target exhibit recognition is high, the preset threshold can be set slightly larger, while if the accuracy requirement for target exhibit recognition is relatively relaxed, the preset threshold can be set appropriately smaller. No specific limitation is made here. This method can filter out exhibit models with the highest matching degree to the current exhibit and a matching degree greater than the preset threshold during the exhibit recognition process, and determine the target exhibit and recognition result based on this, thus helping to further improve the accuracy of target exhibit recognition.
[0056] In a specific implementation scenario, the exhibit model contains several key parts, and explanatory information for these key parts can be marked on the corresponding model positions. Taking a "porcelain vase" as an example, key parts may include, but are not limited to: the mouth, neck, foot, and dragon-shaped handles. Based on this, explanatory information for the "mouth," neck, foot, and dragon-shaped handles can be marked on the exhibit model corresponding to the key parts "mouth," "neck," "foot," and "dragon-shaped handles." Other cases can be deduced similarly, and will not be listed here.
[0057] In a specific implementation scenario, when a target exhibit is identified in the captured image, images of key parts can be extracted from the exhibit's model. Feature point matching can be performed between the image region of the current exhibit in the captured image and the images of each key part of the target exhibit to obtain the matching degree of the key parts and their image positions in the captured image. Based on this, key parts with matching degrees greater than a preset threshold can be selected, confirming the presence of these selected key parts in the captured image.
[0058] Step S12: Display the AR indicator at the image location.
[0059] Specifically, after identifying the target exhibit in the captured image and detecting the image location of key parts of the target exhibit, AR indicator icons can be displayed at various image locations in the captured image. It should be noted that AR indicator icons can be represented by icons with the following shapes, such as magnifying glasses, horns, etc., without limitation.
[0060] Step S13: In response to the user triggering the AR indicator, output the explanation information of the key parts.
[0061] In an implementation scenario, the triggering methods for AR indicators may include, but are not limited to, single click, long press, and multiple consecutive clicks, etc., without any restrictions.
[0062] In one implementation scenario, please refer to the following: Figure 2a and Figure 2b , Figure 2a This is a schematic diagram illustrating the effect of one embodiment of the exhibit guidance method of this application. Figure 2b This is a schematic diagram illustrating the effect of another embodiment of the exhibit guidance method of this application. For example... Figure 2a and Figure 2b As shown, the captured image contains two key parts of the two target exhibits: one is the "bottle handle," and the other is the "pattern." These two key parts are marked with AR indicators (e.g., ...) at their respective locations in the image. Figure 2a , Figure 2b (As shown by the magnifying glass in the image), the user can trigger any AR indicator. For example... Figure 2b As shown, the explanatory information can be text. Therefore, after triggering the AR indicator corresponding to the "pattern," a text box can be overlaid on the captured image, and the explanatory information for the "pattern" (e.g., Figure 2b The text box will display the explanation information as follows: "The pattern is..., and its meaning is...". Alternatively, the explanation information can be audio, in which case the explanation information can be played after the AR indicator corresponding to the pattern is triggered. Or, the explanation information can be video, in which case the explanation information can be played in a floating window after the AR indicator corresponding to the pattern is triggered.
[0063] In one implementation scenario, the target exhibit may have multiple key parts. Taking the target exhibit "porcelain vase" as an example, such as... Figure 2a or Figure 2b As shown, the target exhibit, the "porcelain vase," has two key parts from the current perspective (namely, the "vase handles" and the "patterns"). Please refer to the relevant documentation for further information. Figure 2c , Figure 2cThis is a schematic diagram illustrating another embodiment of the exhibit guidance method of this application. The target exhibit, a "porcelain vase," has one more key part on its back (i.e., a notched "foot"). Other cases can be deduced similarly, and will not be listed here. In this case, in response to the existence of an undetected key part, a first prompt can be output, and the first prompt is used to suggest adjusting the shooting position of the mobile terminal to capture the undetected key part. The above method can prompt the user by outputting a first prompt when the target photo has multiple key parts and there are still undetected key parts, so as to enable the user to understand all the key parts of the target exhibit during the visit, which is beneficial to improving the interactive experience of exhibit guidance.
[0064] In a specific implementation scenario, as mentioned earlier, images of key parts of the target exhibit can be pre-captured. After matching the captured images with these images and detecting the key parts in the captured images, any unmatched key parts of the target exhibit can be considered as undetected key parts, and a first prompt can be output. Alternatively, as mentioned earlier, an exhibit model of the target exhibit can be pre-built, and explanatory information for each key part can be marked on the model corresponding to that key part. After matching the captured images with images extracted from the key parts of the exhibit model and detecting the key parts in the captured images, any unmatched key parts of the target exhibit can be considered as undetected key parts, and a first prompt can be output.
[0065] In a specific implementation scenario, before outputting the first prompt, the initial position of the undetected key parts on the target exhibit can be obtained, along with the current pose of the mobile terminal. Based on these initial positions and the current pose, analysis can be performed to determine the required adjustment of the mobile terminal's shooting pose. Then, the first prompt can be generated and output based on this shooting pose. It should be noted that the initial position of the undetected key parts on the target exhibit can be determined using a pre-built exhibit model, or by pre-capturing images of various key parts of the target exhibit and marking their positions. The initial position of the undetected key parts can then be determined using these marked positions. Furthermore, the current pose of the mobile terminal can be obtained through visual positioning methods such as SLAM, which are not limited here. After obtaining the initial position and current pose, the required adjustment of the shooting pose can be estimated, such as a 20-degree right rotation or a 30-degree left rotation, which are not limited here. Please refer to the relevant documentation. Figure 2d , Figure 2d This is a schematic diagram illustrating the effect of another embodiment of the exhibit guidance method of this application. Figure 2dAs shown, the calculations above determine that the required shooting posture adjustment is a 180-degree right rotation. This generates the first prompt: "Please rotate 180 degrees to your right to photograph the key parts on the back." This method, during exhibit tours, can determine the necessary shooting posture adjustment for the mobile terminal by combining the initial position of the undetected key parts on the target exhibit with the current posture of the mobile terminal, even when undetected key parts exist. Therefore, it improves the accuracy of the shooting posture adjustment.
[0066] In a specific implementation scenario, in response to the user adjusting the mobile terminal to a new shooting position and recognizing the target exhibit in the new shooting frame, an AR indicator can be displayed in the new shooting position. For example, the image positions of key parts of the target exhibit in the new shooting frame can be detected, and AR indicators can be displayed at each detected image position in the new shooting frame. Alternatively, for example, the image positions of untriggered parts in the new shooting frame can be detected, and AR indicators can be displayed at the image positions of untriggered parts, where the untriggered parts are key parts other than those already triggered, and the triggered parts are the key parts corresponding to the AR indicator already triggered by the user. Still using... Figures 2a to 2d Taking the situation shown as an example, after the user rotates 180 degrees to the right as instructed in the first step, the image can be captured as shown. Figure 2c The above detection process can be repeated in the image shown, and the key parts "bottle ear" and "foot ring" will be detected in the captured image. Since the user has already triggered the AR indicator corresponding to the key part "bottle ear", it can be determined that the only untriggered part is the "foot ring". Therefore, an AR indicator can be displayed at the image location of the key part "foot ring" in the captured image (e.g., Figure 2c (As shown in the medium magnifying glass). Other cases can be deduced similarly, and will not be listed here. The above method, after the user adjusts their shooting posture, avoids repeatedly detecting the triggered parts in the new shooting frame, thereby reducing the interference of the triggered parts on user interaction after posture adjustment, and thus improving the interactive experience of exhibit guidance.
[0067] In one implementation scenario, an exhibition hall may contain various exhibits. Some exhibits are related to the user's current target exhibit. For example, some exhibits may share the same or similar key features as the target exhibit, or some exhibits may have historical continuity with the target exhibit in key features. These exhibits can then be considered related exhibits of the target exhibit. In this case, in response to the existence of related exhibits in the exhibition hall, a second prompt can be output, which is used to encourage the user to visit the related exhibits. Please refer to [reference needed]. Figure 3a , Figure 3aThis is a schematic diagram illustrating the effect of another embodiment of the exhibit guidance method of this application. Figure 3a As shown, in the user and undetected key areas (such as...) Figure 2c After interacting with the exhibit (circled foot), if it detects that there are related exhibits in area C1 of the exhibition hall that are related to the currently viewed exhibit, a second prompt can be generated: "There are related exhibits in area C1 of the exhibition hall that are related to the currently viewed exhibit. Do you want to visit them?" This second prompt can be displayed as a pop-up window (e.g., a top pop-up). Of course, other forms (e.g., voice prompts) can also be used to display the second prompt, which is not limited here. In this way, when there are related exhibits in the exhibition hall that are related to the currently viewed exhibit, the second prompt is displayed to encourage the user to visit the related exhibits. This can satisfy the user's need for guided tours of exhibits of interest as much as possible, thereby improving the interactive experience of exhibit tours.
[0068] In a specific implementation scenario, if there are multiple related exhibits in the exhibition hall that are related to the target exhibit, the second prompt may also include options corresponding to each of the multiple related exhibits, so that the user can choose when they need to visit the related exhibits.
[0069] In a specific implementation scenario, in response to receiving a user's confirmation instruction to visit related exhibits, the system can obtain the second location of the related exhibits within the exhibition hall and the current pose of the mobile terminal. Based on this, AR navigation markers can be displayed on the current screen of the mobile terminal, using the second location and the current pose. For example, a 3D map of the exhibition hall can be pre-constructed, allowing the second location of the related exhibits within the exhibition hall to be obtained based on the 3D map. Furthermore, the current pose of the mobile terminal can be obtained using visual positioning methods such as SLAM. Based on this, path planning can be performed using the second location and the current pose, and the AR navigation markers can be displayed using the planned navigation path and the current pose obtained through real-time positioning until the user navigates to the related exhibit within the exhibition hall. Please refer to [reference needed]. Figure 3a and Figure 3b , Figure 3b This is a schematic diagram illustrating the effect of another embodiment of the exhibit guidance method of this application. Figure 3a As shown, when the user selects "Yes," it can be considered that the user's confirmation instruction to visit related exhibits has been received. At this time, through the above process, AR navigation icons can be displayed on the current screen (such as...). Figure 3b (Using a right-turn arrow) to navigate users to related exhibits. This method, after a user decides to visit a related exhibit, combines the exhibit's secondary location within the exhibition hall with the mobile device's current position, displaying AR navigation icons on the mobile device's current screen to guide the user, thus enhancing the interactive experience of exhibit navigation.
[0070] The above solution, in response to identifying a target exhibit in the mobile terminal's camera view, detects the image of the captured image to obtain the image location of key parts of the target exhibit. The target exhibit is pre-loaded with explanatory data, including information on each key part of the target exhibit. Based on this, an AR indicator is displayed at the image location. In response to the user triggering the AR indicator, the explanatory information for the key parts is output. Therefore, during exhibit tours, users can identify exhibits by taking pictures with their mobile terminals. If the exhibit is identified as a target exhibit, AR indicators are displayed at the image location of the key parts of the target exhibit in the captured image to interact with the user. This interactive approach to triggering explanations improves exhibit tour efficiency and enhances the overall tour experience.
[0071] Please see Figure 4 , Figure 4 This is a schematic diagram of the framework of an embodiment of the exhibit guidance device 40 of this application. The exhibit guidance device 40 includes: a detection module 41, an identification module 42, and an interaction module 43. The detection module 41 is used to detect the captured image in response to the recognition of a target exhibit in the captured image of a mobile terminal, and obtain the image position of the key parts of the target exhibit in the captured image; wherein, the target exhibit is an exhibit with pre-set explanatory data, and the explanatory data includes explanatory information of each key part of the target exhibit; the identification module 42 is used to display an AR indicator at the image position; the interaction module 43 is used to output the explanatory information of the key parts in response to the user triggering the AR indicator.
[0072] The above solution, in response to identifying a target exhibit in the mobile terminal's camera view, detects the image of the captured image to obtain the image location of key parts of the target exhibit. The target exhibit is pre-loaded with explanatory data, including information on each key part of the target exhibit. Based on this, an AR indicator is displayed at the image location. In response to the user triggering the AR indicator, the explanatory information for the key parts is output. Therefore, during exhibit tours, users can identify exhibits by taking pictures with their mobile terminals. If the exhibit is identified as a target exhibit, AR indicators are displayed at the image location of the key parts of the target exhibit in the captured image to interact with the user. This interactive approach to triggering explanations improves exhibit tour efficiency and enhances the overall tour experience.
[0073] In some disclosed embodiments, the exhibit guidance device 40 further includes a matching module for matching the captured image with each exhibit model to obtain a matching result; wherein the matching result includes the degree of matching between the current exhibit in the captured image and each exhibit model; the exhibit guidance device 40 further includes a result analysis module for analyzing the matching result to obtain a recognition result of the captured image; wherein the recognition result includes whether the current exhibit is the target exhibit.
[0074] Therefore, the captured images are matched with each exhibit model to obtain matching results. The matching results include the matching degree between the current exhibit in the captured image and each exhibit model. Based on the matching results, the recognition results of the captured images are obtained. The recognition results include whether the current exhibit is the target exhibit. Thus, the presence of the target exhibit in the captured images can be identified through model matching, which can help improve the accuracy of exhibit recognition.
[0075] In some disclosed embodiments, the analysis module includes a selection submodule, used to select the exhibit to which the exhibit model corresponding to the maximum matching degree belongs as the target exhibit in response to the maximum matching degree being higher than a preset threshold. The analysis module also includes a determination submodule, used to determine that the identification result includes the current exhibit as the target exhibit.
[0076] Therefore, when the maximum matching degree is greater than the preset threshold, the exhibit to which the exhibit model corresponding to the maximum matching degree belongs is taken as the target exhibit, and the recognition result is determined to include the current exhibit as the target exhibit. Thus, in the exhibit recognition process, the exhibit model with the maximum matching degree with the current exhibit and the matching degree is greater than the preset threshold can be selected, and the target exhibit and recognition result can be determined based on this, which is conducive to further improving the accuracy of identifying the target exhibit.
[0077] In some publicly disclosed embodiments, the exhibit model to which the exhibit belongs includes several key parts, and the exhibit model marks the explanatory information of the key parts at the model positions corresponding to the key parts.
[0078] Therefore, the exhibit model contains several key parts, and the exhibit model marks the explanatory information of the key parts at the corresponding model positions. This allows the key parts in the captured image to be detected directly based on the model positions marked with explanatory information, and the explanatory information of the key parts can be obtained directly from the exhibit model. In other words, both the detection of key parts and the acquisition of explanatory information can be achieved based on the exhibit model, which helps to improve the efficiency of exhibit tours.
[0079] In some disclosed embodiments, the target exhibit has multiple key parts, and the exhibit guidance device 40 further includes a first prompt module, which is used to output a first prompt in response to the existence of undetected key parts; wherein, the first prompt is used to prompt the mobile terminal to adjust the shooting posture in order to capture the undetected key parts.
[0080] Therefore, the target exhibit has multiple key parts, and in response to the existence of undetected key parts, a first prompt is output. The first prompt is used to prompt the user to adjust the shooting position of the mobile terminal to capture the undetected key parts. Thus, even when the target photo has multiple key parts and there are still undetected key parts, the first prompt can prompt the user to understand all the key parts of the target exhibit during the visit, which is conducive to improving the interactive experience of exhibit tour.
[0081] In some disclosed embodiments, the exhibit guidance device 40 further includes a first acquisition module, used to acquire the first position of the undetected key part on the target exhibit and acquire the current pose of the mobile terminal; the exhibit guidance device 40 also includes a pose analysis module, used to analyze based on the first position and the current pose to obtain the shooting pose that the mobile terminal needs to adjust.
[0082] Therefore, before outputting the first prompt, the first position of the undetected key part on the target exhibit is obtained, and the current pose of the mobile terminal is obtained. Based on the first position and the current pose, the shooting pose that the mobile terminal needs to adjust is analyzed. Thus, during the exhibit tour, even if there is an undetected key part, the shooting pose that the mobile terminal needs to adjust can be determined by combining the first position of the undetected key part on the target exhibit and the current pose of the mobile terminal, thereby improving the accuracy of the shooting pose.
[0083] In some disclosed embodiments, the identification module 42 is also used to display an AR indicator in the new shooting pose in response to the user adjusting the mobile terminal to a new shooting position and recognizing the target exhibit in the new shooting screen of the mobile terminal.
[0084] Therefore, in response to the user adjusting the mobile terminal to a new shooting position and recognizing the target exhibit in the new shooting screen of the mobile terminal, an AR indicator is displayed in the new shooting position. This enables the AR indicator to follow the user's shooting position after adjustment, thereby improving the interactive experience of exhibit guidance.
[0085] In some publicly disclosed embodiments, the exhibit guidance device 40 further includes a second prompt module, which outputs a second prompt in response to the existence of related exhibits in the exhibition hall that are related to the target exhibit; wherein the second prompt is used to prompt the user to visit the related exhibits.
[0086] Therefore, in response to the existence of related exhibits in the exhibition hall that are related to the target exhibit, a second prompt is output, which is used to prompt the user to visit the related exhibits. In other words, when there are related exhibits in the exhibition hall that are related to the target exhibit that the user is currently visiting, the second prompt is output to prompt the user to visit the related exhibits, thereby satisfying the user's tour needs for exhibits of interest as much as possible and improving the interactive experience of exhibit tour.
[0087] In some disclosed embodiments, the exhibit guidance device 40 further includes a second acquisition module, which is used to acquire the second location of the related exhibit in the exhibition hall and acquire the current pose of the mobile terminal in response to receiving a confirmation instruction from the user regarding visiting the related exhibit; the exhibit guidance device 40 also includes a navigation module, which is used to display AR navigation icons on the current screen of the mobile terminal based on the second location and the current pose.
[0088] Therefore, after outputting the second prompt, in response to receiving the user's confirmation instruction to visit the related exhibits, the system obtains the second location of the related exhibits in the exhibition hall and the current pose of the mobile terminal. Based on the second location and the current pose, an AR navigation icon is displayed on the current screen of the mobile terminal. Thus, after the user confirms that they are visiting the related exhibits, the system can combine the second location of the related exhibits in the exhibition hall with the current pose of the mobile terminal to guide the user by displaying an AR navigation icon on the current screen of the mobile terminal, which helps to improve the interactive experience of exhibit tours.
[0089] Please see Figure 5 , Figure 5 This is a schematic diagram of a framework of an embodiment of the mobile terminal 50 of this application. The mobile terminal 50 includes a camera 51, a display screen 52, a memory 53, and a processor 54. The camera 51, display screen 52, and memory 53 are respectively coupled to the processor 54. The processor 54 is used to execute program instructions stored in the memory 53 to implement the steps in any of the above-described exhibit guidance method embodiments. Specifically, the mobile terminal 50 may include, but is not limited to, mobile phones, tablet computers, smart glasses, etc., and is not limited thereto.
[0090] Specifically, processor 54 controls itself, as well as camera 51, display screen 52, and memory 53, to implement the steps of any of the above-described exhibit guidance method embodiments. Processor 54 can also be referred to as a CPU (Central Processing Unit). Processor 54 may be an integrated circuit chip with signal processing capabilities. Processor 54 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 54 can be implemented using integrated circuit chips.
[0091] The above solution allows users to take photos of exhibits using their mobile devices for identification during the guided tour. When an exhibit is identified as a target, AR indicators are displayed on key parts of the target exhibit in the image, enabling interaction between the user and the exhibit. This interactive approach allows for guided tours with interactive explanations, improving efficiency and enhancing the overall tour experience.
[0092] Please see Figure 6 , Figure 6 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium 60 of this application. The computer-readable storage medium 60 stores program instructions 601 that can be executed by a processor. The program instructions 601 are used to implement the steps of any of the above-described exhibit guidance method embodiments.
[0093] The above solution allows users to take photos of exhibits using their mobile devices for identification during the guided tour. When an exhibit is identified as a target, AR indicators are displayed on key parts of the target exhibit in the image, enabling interaction between the user and the exhibit. This interactive approach allows for guided tours with interactive explanations, improving efficiency and enhancing the overall tour experience.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0095] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0096] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0097] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0098] This disclosure relates to the field of augmented reality (AR). It involves acquiring image information of target objects in a real-world environment and then using various visual algorithms to detect or identify the relevant features, states, and attributes of these objects, thereby achieving an AR effect that combines virtual and real elements to suit specific applications. For example, target objects may include human features such as faces, limbs, gestures, and movements; objects such as signs and markers; or venues such as sand tables, display areas, or displayed items. Visual algorithms may include visual localization, SLAM, 3D reconstruction, image registration, background segmentation, keypoint extraction and tracking of objects, and pose or depth detection. Specific applications can include interactive scenarios related to real-world scenes or objects, such as guided tours, navigation, explanations, reconstruction, and virtual effect overlay displays, as well as human-related special effects processing, such as makeup enhancement, body enhancement, special effects displays, and virtual model displays.
[0099] Convolutional neural networks (CNNs) can be used to detect or identify the relevant features, states, and attributes of target objects. The aforementioned CNNs are network models obtained through training using deep learning frameworks.
Claims
1. A method for guiding visitors through exhibits, characterized in that, include: In response to the identification of a target exhibit in the shooting screen of a mobile terminal, the shooting screen is detected to obtain the image position of the key parts of the target exhibit in the shooting screen; wherein, the target exhibit is an exhibit with pre-set explanatory data, and the explanatory data includes explanatory information of each of the key parts of the target exhibit; An AR indicator is displayed at the location of the image in the captured image; In response to the user triggering the AR indicator, explanatory information about the key parts is output; wherein, the exhibit to which the exhibit model belongs includes several key parts, and the exhibit model marks the explanatory information of the key parts at the model positions corresponding to the key parts; the step of detecting the captured image to obtain the image position of the key parts of the target exhibit in the captured image includes: Extract images of the key parts from the exhibit model of the target exhibit; The image area of the current exhibit in the captured image is matched with the image of each of the key parts of the target exhibit by feature point matching, so as to obtain the image position and matching degree of each of the key parts in the captured image; The key parts with a matching degree greater than a preset threshold are selected, and it is determined that the selected key parts exist in the captured image.
2. The method according to claim 1, characterized in that, Before detecting the target exhibit in the captured image of the mobile terminal and obtaining the image position of the key part of the target exhibit in the captured image, the method further includes: The captured images are matched with each exhibit model to obtain matching results; wherein, the matching results include: the matching degree between the current exhibit in the captured images and each of the exhibit models; Based on the matching results, an analysis is performed to obtain the recognition result of the captured image; wherein, the recognition result includes: whether the current exhibit is the target exhibit.
3. The method according to claim 2, characterized in that, The analysis based on the matching results to obtain the recognition result of the captured image includes: In response to the maximum matching degree being higher than a preset threshold, the exhibit to which the exhibit model corresponding to the maximum matching degree belongs is taken as the target exhibit, and the identification result is determined to include the current exhibit as the target exhibit.
4. The method according to any one of claims 1 to 3, characterized in that, The target exhibit has multiple key parts, and the method further includes: In response to the presence of undetected critical parts, a first prompt is output; wherein, the first prompt is used to prompt the mobile terminal to adjust its shooting posture in order to capture the undetected critical parts.
5. The method according to claim 4, characterized in that, Before outputting the first prompt, the method further includes: The first position of the undetected key part on the target exhibit is obtained, and the current pose of the mobile terminal is obtained; Based on the analysis of the first position and the current pose, the shooting pose that the mobile terminal needs to adjust is obtained.
6. The method according to claim 4, characterized in that, The method further includes: In response to the user adjusting the mobile terminal to a new shooting position and recognizing the target exhibit in the new shooting screen of the mobile terminal, an AR indicator is displayed in the new shooting position.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: In response to the existence of related exhibits in the exhibition hall that are related to the target exhibit, a second prompt is output; wherein, the second prompt is used to prompt the user to visit the related exhibits.
8. The method according to claim 7, characterized in that, After outputting the second prompt, the method further includes: In response to receiving a confirmation instruction from a user regarding visiting the associated exhibits, the system obtains the second location of the associated exhibits within the exhibition hall and the current pose of the mobile terminal. Based on the second location and the current pose, an AR navigation icon is displayed on the current screen of the mobile terminal.
9. An exhibit guidance device, characterized in that, include: The detection module is used to detect the target exhibit in the shooting screen of the mobile terminal in response to the detection of the target exhibit in the shooting screen and obtain the image position of the key parts of the target exhibit in the shooting screen; wherein the target exhibit is an exhibit with pre-set explanatory data, and the explanatory data includes explanatory information of each of the key parts of the target exhibit; The identification module is used to display an AR indicator at the image location in the captured image; The interaction module is used to respond to the user triggering the AR indicator and output explanatory information about the key parts; wherein, the exhibit model to which the exhibit belongs includes several key parts, and the exhibit model marks the explanatory information of the key parts at the model positions corresponding to the key parts; the step of detecting the captured image to obtain the image position of the key parts of the target exhibit in the captured image includes: Extract images of the key parts from the exhibit model of the target exhibit; The image area of the current exhibit in the captured image is matched with the image of each of the key parts of the target exhibit by feature point matching, so as to obtain the image position and matching degree of each of the key parts in the captured image; The key parts with a matching degree greater than a preset threshold are selected, and it is determined that the selected key parts exist in the captured image.
10. A mobile terminal, characterized in that, The device includes a camera, a display screen, a memory, and a processor. The camera, the display screen, and the memory are respectively coupled to the processor. The processor is used to execute program instructions stored in the memory to implement the exhibit guidance method according to any one of claims 1 to 8.
11. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the exhibit guidance method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Product feature acquisition method, terminal and storage medium
CN110021062A
Method and computer system for displaying identification result
CN112270297A
AR-based scene guide method, AR glasses, electronic device and storage medium
CN113409470A
Exhibit display method and device, guide method and device, electronic equipment and storage medium
CN114511671A
System and method for a virtual showroom
US20210304500A1