Active interactive guide system and active interactive guide method

By using translucent display devices and image acquisition devices in the interactive tour guide system, combined with facial feature and gaze recognition technology of the processing device, the problem of virtual information display in long-distance and multi-user situations is solved, achieving highly accurate virtual information display and a comfortable interactive experience.

CN116402990BActive Publication Date: 2026-07-24IND TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IND TECH RES INST
Filing Date
2023-01-05
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing interactive navigation systems cannot accurately determine the user's line of sight when the user is far from the display device, resulting in the inability to display virtual information correctly. Furthermore, in multi-user scenarios, it is impossible to determine which user's virtual information to display, causing reading difficulties.

Method used

By employing a light-transmitting display device, a target image acquisition device, and a user image acquisition device, combined with a processing unit, the system identifies users by recognizing their facial features and gaze, segments user images to identify users at a distance, calculates the virtual information display position, and optimizes the virtual-real fusion display position correction algorithm to improve recognition accuracy.

Benefits of technology

It achieves real-time tracking of the user's gaze and highly accurate display of virtual information, providing a comfortable contactless interactive experience, solving the problems of long-distance user identification and information display in multi-user situations, and optimizing the virtual-real fusion display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116402990B_ABST
    Figure CN116402990B_ABST
Patent Text Reader

Abstract

The application provides an active interactive guiding system, which comprises a display device, a target image acquisition device, a user image acquisition device and a processing device. The target image acquisition device acquires a dynamic target image. The user image acquisition device acquires a user image. The processing device identifies and selects a served object from the user image, and acquires facial features of the served object. If the facial features match facial feature points, the processing device detects a line of sight of the served object, identifies a target object gazed at by the served object, generates a three-dimensional coordinate corresponding to a face position of the served object, a three-dimensional coordinate corresponding to a position of the target object and depth width information, calculates a position of an intersection point of the line of sight passing through the display device, and displays virtual information of the target object at the intersection point of the display device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an interactive tour guide technology, and more particularly to an active interactive tour guide system and an active interactive tour guide method. Background Technology

[0002] With the development of image processing and spatial positioning technologies, the application of transparent displays has gradually gained attention. This technology allows display devices to be paired with dynamic objects, supplemented by virtual related information, and to generate interactive experiences according to user needs, presenting information in a more intuitive way. Furthermore, virtual information associated with dynamic objects can be displayed at specific locations on the transparent display device, allowing users to simultaneously view the dynamic objects and the virtual information superimposed on them through the transparent display device.

[0003] However, when the user is far from the display device, the device that captures the user's image may not be able to determine the user's gaze. As a result, the system will not be able to determine what the user is looking at, and therefore will not be able to display the correct virtual information on the display device, or even overlay the virtual information corresponding to the dynamic object being looked at onto the dynamic object.

[0004] Furthermore, when the system detects that multiple users are viewing a dynamic object at the same time, each user's gaze direction may be different. The system will then be unable to determine which dynamic object's related virtual information to display. As a result, the interactive guide system will be unable to present the virtual information corresponding to the dynamic object that the user is viewing, leading to difficulty and discomfort for the viewer in reading the virtual information. Summary of the Invention

[0005] This invention provides an active interactive tour guide system, including a light-transmitting display device, a target object image acquisition device, a user image acquisition device, and a processing unit. The light-transmitting display device is positioned between at least one user and multiple dynamic objects. The target object image acquisition device is coupled to the display device to acquire images of the dynamic objects. The user image acquisition device is coupled to the display device to acquire user images. The processing unit is coupled to the display device. The processing device is used to identify and track dynamic objects in dynamic object images. It is also used to identify at least one user in a user image and select a target object, acquire facial features of the target object, and determine whether the facial features match multiple facial feature points. If the facial features match these facial feature points, the processing device detects the target object's gaze, where the gaze passes through the display device to focus on the target object of the dynamic object. If the facial features do not match facial feature points, the processing device performs image segmentation to divide the user image into multiple images to be identified. The user image acquisition device performs user identification on each of the images to be identified. Furthermore, the processing device is used to identify the target object being gazed at by the target object based on the gaze, generate three-dimensional coordinates corresponding to the face position of the target object, three-dimensional coordinates corresponding to the position of the target object, and depth and width information of the target object, calculate the intersection position of the gaze passing through the display device, and display virtual information corresponding to the target object at the intersection position on the display device.

[0006] This invention provides an active interactive tour guide method applicable to an active interactive tour guide system having a light-transmitting display device, a target object image acquisition device, a user image acquisition device, and a processing device, wherein the display device is positioned between at least one user and multiple dynamic objects, and the processing device is used to execute the active interactive tour guide method. The proactive interactive tour guide method includes: acquiring dynamic object images using a target object image acquisition device, identifying dynamic objects in the dynamic object images, and tracking dynamic objects; acquiring user images using a user image acquisition device, identifying at least one user in the user images and selecting the target object, acquiring the facial features of the target object and determining whether the facial features match multiple facial feature points; if the facial features match facial feature points, detecting the target object's gaze, wherein the gaze passes through the display device to look at the target object of the dynamic object; if the facial features do not match facial feature points, performing image segmentation to segment the user image into multiple images to be identified, and performing user identification on each of the images to be identified separately; based on the gaze identification of the target object being looked at by the target object, generating three-dimensional coordinates corresponding to the face position of the target object and three-dimensional coordinates corresponding to the position of the target object, as well as the depth and width information of the target object, calculating the intersection position of the gaze passing through the display device, and displaying the virtual information corresponding to the target object at the intersection position of the display device.

[0007] This invention provides an active interactive tour guide system, including a light-transmitting display device, a target object image acquisition device, a user image acquisition device, and a processing device. The light-transmitting display device is positioned between at least one user and multiple dynamic objects. The target object image acquisition device is coupled to the display device to acquire images of the dynamic objects. The user image acquisition device is coupled to the display device to acquire user images. The processing device is coupled to the display device. The processing device is used to identify and track dynamic objects in the dynamic object images. Furthermore, the processing device is used to identify at least one user in the user images and select a target object based on a service area, detecting the gaze of the target object. The service area has an initial size, and the gaze passes through the display device to focus on the target object of the dynamic object. The processing device is further used to identify the target object being gazed at by the target object based on the gaze, generate three-dimensional coordinates corresponding to the face position of the target object and three-dimensional coordinates corresponding to the position of the target object, as well as depth and width information of the target object, calculate the intersection point of the gaze passing through the display device, and display virtual information corresponding to the target object at the intersection point on the display device.

[0008] Based on the above, the active interactive tour guide system and method described in this invention can track the viewing user's gaze direction in real time, stably track moving target objects, and actively display virtual information corresponding to the target object, providing highly accurate augmented reality information and a comfortable non-contact interactive experience. This invention also integrates internal and external perception recognition with virtual-real fusion and system virtual-real fusion pairing calculation cores, actively using internal perception to determine the viewing angle of the visitor and then matching it with external perception AI target object recognition to realize augmented reality applications. Furthermore, this invention optimizes the virtual-real fusion display position correction algorithm to perform offset correction, improves facial recognition for distant users, and prioritizes the selection of service recipients, greatly solving the problem of insufficient manpower and creating an interactive experience where knowledge and information are transmitted at zero distance.

[0009] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the present invention. Attached Figure Description

[0010] Figure 1 This is a block diagram illustrating an active interactive tour guide system according to an embodiment of the present invention;

[0011] Figure 2 This is a schematic diagram of an active interactive tour guide system according to an embodiment of the present invention;

[0012] Figure 3A This is a schematic diagram illustrating the performance of image segmentation to identify distant users according to an embodiment of the present invention;

[0013] Figure 3BThis is a schematic diagram illustrating the performance of image segmentation to identify distant users according to an embodiment of the present invention;

[0014] Figure 3C This is a schematic diagram illustrating the performance of image segmentation to identify distant users according to an embodiment of the present invention;

[0015] Figure 3D This is a schematic diagram illustrating the performance of image segmentation to identify distant users according to an embodiment of the present invention;

[0016] Figure 3E This is a schematic diagram illustrating the performance of image segmentation to identify distant users according to an embodiment of the present invention;

[0017] Figure 4 This is a schematic diagram illustrating the selection of a service recipient by an active interactive tour guide system according to an embodiment of the present invention;

[0018] Figure 5 This is a schematic diagram illustrating the adjustment of the service area range according to an embodiment of the present invention;

[0019] Figure 6 This is a flowchart illustrating an active interactive tour guide method according to an embodiment of the present invention;

[0020] Figure 7 This is a flowchart illustrating an active interactive tour guide method according to an embodiment of the present invention.

[0021] Explanation of icon numbers

[0022] 1: Active interactive navigation system; 110: Display device; 120: Target image acquisition device; 130: User image acquisition device; 140: Processing device; 150: Database; A1~A20: Temporary image blocks; Area1: Object field; Area2: Implementation field; Area3: Service field; CP: Intersection location; cut1~cut8: Dividing line; FarUser: Remote user; FR, FR': Recognition result; Img, Img': User image; Img1: Central image to be recognized; Img2~Img9: Surrounding images to be recognized; Obj: Dynamic object; P1: Focal point; SerUser: Subjected user Service objects; S1, S2, S3: line of sight; S610, S620, S630, S640, S650, S660, S670, S680, S711, S712, S713, S714, S715, S721, S722, S723, S724, S725, S726a, S726b, S727, S728, S740, S750: steps; Ser_Range: service area range; Ser_Range_L: left range; Ser_Range_R: right range; TarObj: target object; User: user; Vinfo: virtual information; Vf: display object box. Detailed Implementation

[0023] The structural and working principles of the present invention will be described in detail below with reference to the accompanying drawings:

[0024] The following description will detail some exemplary embodiments of the present invention with reference to the accompanying drawings. Component symbols used in the following description, when appearing in different drawings, are considered to be the same or similar components. These exemplary embodiments are only a part of the present invention and do not disclose all possible implementations of the invention. More precisely, these exemplary embodiments are merely examples of the methods, apparatus, and systems within the scope of the present invention's patent applications.

[0025] Figure 1 This is a block diagram illustrating an active interactive tour guide system 1 according to an embodiment of the present invention. Firstly, through... Figure 1 This section introduces the various components and their configuration relationships in the proactive interactive tour guide system 1. Detailed functions will be explained in conjunction with the flowcharts of subsequent embodiments. Figure 1 And invented.

[0026] Please refer to Figure 1The active interactive tour guide system 1 of the present invention includes a light-transmitting display device 110, a target object image acquisition device 120, a user image acquisition device 130, a processing device 140, and a database 150. The processing device 140 can be wirelessly, wiredly, or electrically connected to the display device 110, the target object image acquisition device 120, the user image acquisition device 130, and the database 150.

[0027] The display device 110 is disposed between at least one user and multiple dynamic objects. In practice, the display device 110 may be, for example, a transmissive light-transmitting display such as a liquid crystal display (LCD), a field sequential color liquid crystal display, a light emitting diode (LED) display, or an electrohumidification display, or a projection-type light-transmitting display.

[0028] The target object image acquisition device 120 and the user image acquisition device 130 can be coupled to and disposed on the display device 110, or they can be coupled to the display device 110 but disposed near the display device 110. The image acquisition directions of the target object image acquisition device 120 and the user image acquisition device 130 are respectively oriented towards different directions of the display device 110. That is, the image acquisition direction of the target object image acquisition device 120 is oriented towards the direction with multiple dynamic objects, while the image acquisition direction of the user image acquisition device 130 is oriented towards the direction of at least one user in the implementation field. The target object image acquisition device 120 is used to acquire dynamic object images of multiple dynamic objects, while the user image acquisition device 130 is used to acquire user images of at least one user in the implementation field.

[0029] In practice, the target object image acquisition device 120 includes an RGB image sensing module, a depth sensing module, an inertial sensing module, and a GPS positioning sensing module. The target object image acquisition device 120 can use the RGB image sensing module, or a combination of the RGB image sensing module, depth sensing module, inertial sensing module, or GPS positioning sensing module, to perform image recognition and positioning of multiple dynamic objects. The RGB image sensing module can include a visible light sensor or a non-visible light sensor such as an infrared sensor. Furthermore, the target object image acquisition device 120 can also use, for example, an optical locator to perform optical spatial positioning of the dynamic object. Any device or combination thereof capable of locating the position information of a dynamic object falls within the scope of the target object image acquisition device 120.

[0030] The user image acquisition device 130 includes an RGB image sensing module, a depth sensing module, an inertial sensing module, and a GPS positioning sensing module. The user image acquisition device 130 can perform image recognition and positioning of at least one user using the RGB image sensing module, or a combination of the RGB image sensing module, the depth sensing module, the inertial sensing module, or the GPS positioning sensing module. The RGB image sensing module may include a visible light sensor or a non-visible light sensor such as an infrared sensor. Any device or combination thereof capable of locating the position information of at least one user falls within the scope of the user image acquisition device 130.

[0031] In this embodiment of the invention, the image acquisition device described above can be used to acquire images and includes a camera lens with a lens and a photosensitive component. The depth sensor described above can be used to detect depth information, which can be implemented using active depth sensing technology and passive depth sensing technology. Active depth sensing technology can calculate depth information by actively emitting light sources, infrared rays, ultrasonic waves, lasers, etc., as signals, combined with time-of-flight ranging technology. Passive depth sensing technology can use two image acquisition devices to acquire two images in front of them from different perspectives, and use the parallax of the two images to calculate depth information.

[0032] Processing device 140 is used to control the operation of the active interactive tour guide system 1, and may include memory and a processor. Figure 1 (Not shown). Memory can be, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk or other similar device, integrated circuit, or combination thereof. Processor can be, for example, a central processing unit (CPU), application processor (AP), or other programmable general-purpose or special-purpose microprocessor, digital signal processor (DSP), image signal processor (ISP), graphics processing unit (GPU) or other similar device, integrated circuit, or combination thereof.

[0033] Database 150 is coupled to processing device 140 for storing data provided by processing device 140 for feature comparison. Database 150 can provide storage media for storing data or programs of any type, such as any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk or other similar devices, integrated circuits and combinations thereof.

[0034] In this embodiment, the processing device 140 may be a calculator device built into or connected to the display device 110. The target image acquisition device 120 and the user image acquisition device 130 may be respectively located on opposite sides of the active interactive tour guide system 1 relative to the display device 110, for locating users and dynamic objects, and transmitting information to the processing device 140 via their respective communication interfaces in a wired or wireless manner. In some embodiments, the target image acquisition device 120 and the user image acquisition device 130 may each have a processor and memory, and have the computing capability to perform object recognition and object tracking based on image data.

[0035] Figure 2 This is a schematic diagram illustrating an active interactive tour guide system 1 according to an embodiment of the present invention. Please refer to... Figure 2 One side of the display device 110 faces the object field Area 1, while the other side faces the implementation field Area 2. The target object image acquisition device 120 and the user image acquisition device 130 are both coupled to the display device 110. The image acquisition direction of the target object image acquisition device 120 faces the object field Area 1, while the image acquisition direction of the user image acquisition device 130 faces the implementation field Area 2. The implementation field Area 2 includes a service field Area 3. Users who wish to view the virtual information corresponding to the dynamic object Obj through the display device 110 can stand in the service field Area 3.

[0036] The dynamic object Obj resides in the object field Area1. Figure 2 The dynamic object Obj shown is for illustrative purposes only; there may be only one or more dynamic objects Obj. The user viewing the dynamic object Obj is located in either the implementation area Area2 or the service area Area3. Figure 2 The User shown is for illustrative purposes only; there may be only one User or multiple Users.

[0037] User can view the dynamic object Obj located in object area Area 1 via display device 110 in service area Area 3. In some embodiments, target object image acquisition device 120 is used to acquire dynamic object image of dynamic object Obj, processing device 140 identifies the spatial location information of dynamic object Obj in dynamic object image, and tracks dynamic object Obj. User image acquisition device 130 is used to acquire user image of user, processing device 140 identifies the spatial location information of user in user image, and selects the serviced object SerUser.

[0038] When User is standing in Service Area 3, User occupies a suitable proportion in the user image acquired by User Image Acquisition Device 130. Processing Device 140 can identify User and select SerUser as the service recipient using general face recognition methods. However, if User is not standing in Service Area 3 but in Implementation Area 2, then User is referred to as FarUser (far distance user). User Image Acquisition Device 130 can also capture FarUser's image. However, because FarUser occupies too small a proportion in the user image, Processing Device 140 may not be able to identify FarUser using general face recognition methods and select SerUser as the service recipient from among FarUsers.

[0039] In one embodiment, database 150 stores multiple facial feature points. When processing device 140 identifies user (User) in a user image and selects service recipient (SerUser), processing device 140 collects facial features of service recipient (SerUser) and determines whether the facial features match multiple facial feature points. These facial features include features on the face such as eyes, nose, mouth, eyebrows, and face shape. Generally, there are 468 facial feature points. Once the collected facial features match default facial feature points, user identification can be effectively performed.

[0040] If the processing device 140 determines that multiple facial feature points match, it means that the user (User) occupies a suitable proportion in the user image acquired by the user image acquisition device 130. The processing device 140 can then identify the user (User) and select the service recipient (SerUser) using a general face recognition method. At this time, the processing device 140 uses facial feature points to calculate the facial position of the service recipient (SerUser) to detect the gaze direction of the service recipient (SerUser)'s gaze S1, and generates a number (ID) corresponding to the service recipient (SerUser) and the three-dimensional coordinates (x, y, y) of the facial position. u , y u , z u ).

[0041] Wherein, gaze S1 means that when the gaze of the service object SerUser passes through the display device 110 and looks at a target object TarObj among multiple dynamic objects Obj, the eyes focus on a part of the target object TarObj. Figure 2 The line of sight S2 or line of sight S3 shown indicates that when the line of sight of the service object SerUser passes through the display device 110 and looks at a target object TarObj among multiple dynamic objects Obj, the eyes focus on other parts of the target object TarObj.

[0042] If the processing device 140 determines that the facial features do not match multiple facial feature points, it may be because no user is standing in the implementation area Area 2 and the service area Area 3, or because a distant user, FarUser, is standing in the implementation area Area 2. It could also be that the user image acquisition device 130 needs to implement a supplementary lighting mechanism to improve the clarity of the user image. When the processing device 140 detects a distant user, FarUser, in the implementation area Area 2, it first performs image segmentation to divide the user image into multiple images to be identified. At least one of these images will include the distant user, FarUser. This increases the proportion of the distant user in that image, which is beneficial for the processing device 140 to identify the distant user and recognize their spatial location information among the multiple images. The processing device 140 performs user identification on each of the multiple images to be identified. In the image to be identified that has the distant user FarUser, it collects the facial features of the distant user FarUser and uses the facial feature points to calculate the facial position of the service object SerUser in the distant user FarUser and the gaze direction of the gaze S1.

[0043] However, most common image segmentation techniques divide an image into multiple smaller images directly using multiple segmentation lines. If a common image segmentation technique is used to segment the user image described in this invention, the segmentation lines are very likely to fall exactly on the face of the distant user FarUser in the user image. As a result, the processing device 140 will be unable to effectively identify the distant user FarUser.

[0044] Therefore, in performing image segmentation, the processing apparatus 140 of one embodiment of the present invention temporarily divides the user image into multiple temporary image blocks using temporary segmentation lines, and then segments the user image into multiple images to be identified based on the temporary image blocks. Furthermore, one of the multiple images to be identified has an overlapping area with its adjacent counterpart; here, "adjacent" can mean vertically adjacent, horizontally adjacent, or diagonally adjacent. The overlapping area is to ensure that the face of the distant user FarUser in the user image can be completely preserved in the image to be identified. The following will describe in detail how the processing apparatus 140 of the present invention performs image segmentation to identify the distant user FarUser.

[0045] Figures 3A-3E This is a schematic diagram illustrating the performance of image segmentation to identify distant users according to an embodiment of the present invention. Please refer to [the original text]. Figure 3A , 3B First, the processing device 140 temporarily divides the user image Img into multiple temporary image blocks A1 to A20 using temporary segmentation lines cut1 to cut8. Then, the processing device 140 further segments the user image Img into multiple images to be recognized based on the temporary image blocks A1 to A20. Each of these multiple images to be recognized includes a central image to be recognized and multiple peripheral images to be recognized.

[0046] For example, such as Figure 3B , 3C As shown, the processing device 140 segments the central image to be recognized, Img1, based on temporary image blocks A7, A8, A9, A12, A13, A14, A17, A18, and A19; the processing device 140 segments the surrounding images to be recognized, Img2, based on temporary image blocks A4, A5, A9, and A10; the processing device 140 segments the surrounding images to be recognized, Img3, based on temporary image blocks A9, A10, A14, A15, A19, and A20; the processing device 140 segments the surrounding images to be recognized, Img4, based on temporary image blocks A19, A20, A24, and A25; and the processing device 140 segments the surrounding images to be recognized, Img4, based on temporary image blocks A1, A1, A20, A24, and A25. The processing device 140 segments the surrounding image to be identified Img5 based on temporary image blocks A6, A7, A11, A12, A16 and A17; the processing device 140 segments the surrounding image to be identified Img7 based on temporary image blocks A16, A17, A21 and A22; the processing device 140 segments the surrounding image to be identified Img8 based on temporary image blocks A2, A3, A4, A7, A8 and A9; and the processing device 140 segments the surrounding image to be identified Img9 based on temporary image blocks A17, A18, A19, A22, A23 and A24.

[0047] Taking the central image to be identified, Img1, as an example, the images to be identified that are vertically adjacent to Img1 are the surrounding images to be identified, Img8 and Img9. There is an overlapping area between the central image to be identified, Img1, and the surrounding images to be identified, Img8, including temporary image blocks A7, A8, and A9. There is also an overlapping area between the central image to be identified, Img1, and the surrounding images to be identified, Img9, including temporary image blocks A17, A18, and A19.

[0048] The images to be identified that are adjacent to the central image to be identified (Img1) on the left and right are the peripheral images to be identified (Img3 and Img6). There is an overlapping area between the central image to be identified (Img1) and the peripheral images to be identified that are adjacent to it on the left and right, including temporary image blocks A9, A14, and A19. There is also an overlapping area between the central image to be identified (Img1) and the peripheral images to be identified that are adjacent to it on the left and right, including temporary image blocks A7, A12, and A17.

[0049] The images diagonally adjacent to the central image to be identified, Img1, are the surrounding images to be identified, Img2, Img4, Img5, and Img7. There is an overlapping area between the central image to be identified, Img1, and the diagonally adjacent surrounding images to be identified, Img2, including temporary image block A9.

[0050] Furthermore, for example, the surrounding images to be identified, Img5 and Img6, are vertically adjacent to each other and also have overlapping areas between them, including temporary image blocks A6 and A7. Similarly, the surrounding images to be identified, Img5 and Img8, are horizontally adjacent to each other and also have overlapping areas between them, including temporary image blocks A2 and A7.

[0051] After the processing device 140 segments the user image Img into a central image to be recognized Img1 and surrounding images to be recognized Img2 to Img9, the user image acquisition device 130 performs face recognition on each of the central image to be recognized Img1 and the surrounding images to be recognized Img2 to Img9. Figure 3D As shown, the processing device 140 recognizes the user's face in the central image to be recognized, Img1, and generates a recognition result FR. After the processing device 140 performs face recognition on each image to be recognized and obtains the recognition result corresponding to each image, as... Figure 3EAs shown, the processing device 140 fuses the central image to be identified Img1 and the surrounding images to be identified Img2 to Img9 into a recognized user image Img', and identifies the spatial location information of the distant user FarUser based on the recognition result FR'.

[0052] In one embodiment, the database 150 stores multiple object feature points corresponding to each dynamic object Obj. When the processing device 140 identifies the target object TarObj being viewed by the serviced object SerUser based on the gaze S1 of the serviced object SerUser, the processing device 140 collects the pixel features of the target object TarObj and compares the pixel features with the object feature points; if the pixel features match the object feature points, the processing device 140 generates a number corresponding to the target object TarObj and three-dimensional coordinates (x, y, y) corresponding to the position of the target object TarObj. o , y o , z o ) and the depth and width information of the target object TarObj (w o , h o ).

[0053] The processing device 140 can determine the display position of the virtual information Vinfo on the display device 110 based on the spatial location information of the served object SerUser and the spatial location information of the target object TarObj. Specifically, the processing device 140 determines the display position of the virtual information Vinfo on the display device 110 based on the three-dimensional coordinates (x, y, x) of the face of the served object SerUser. u , y u , z u ) and the three-dimensional coordinates (x, y) of the target object TarObj. o , y o , z o ), depth and width information (h o , w o The system calculates the intersection point CP of the line of sight S1 of the served object SerUser across the display device 110, and displays the virtual information Vinfo corresponding to the target object TarObj at the intersection point CP of the display device 110. Figure 2 In this context, virtual information Vinfo can be displayed in a display object box Vf, with the center point of the display object box Vf being the intersection point CP.

[0054] Specifically, the display position of the virtual information Vinfo can be considered as the point or area where the gaze S1 of the served object SerUser crosses the display device 110 when viewing the target object TarObj. Therefore, the processing device 140 can display the virtual information Vinfo at the intersection position CP using the display object frame Vf. More specifically, based on various needs or different applications, the processing device 140 can determine the actual display position of the virtual information Vinfo so that the served object SerUser can see the virtual information Vinfo superimposed on the target object TarObj through the display device 110. The virtual information Vinfo can be considered as augmented reality content augmented based on the target object TarObj.

[0055] Additionally, the processing device 140 also determines whether the virtual information Vinfo corresponding to the target object TarObj is overlaid at the intersection position CP of the display device 110. If the processing device 140 determines that the virtual information Vinfo is not overlaid at the intersection position CP of the display device 110, the processing device 140 performs offset correction on the display position of the virtual information Vinfo. For example, the processing device 140 can use an information offset correction equation to perform offset correction on the position of the virtual information Vinfo, thereby optimizing the actual display position of the virtual information Vinfo.

[0056] As described in the preceding paragraphs, after the processing device 140 identifies the user (User) in the user image and selects the service recipient (SerUser), it collects the facial features of the service recipient (SerUser), determines whether the facial features match multiple facial feature points, calculates the facial position and gaze direction (S1) of the service recipient (SerUser) using the facial feature points, and generates a number (ID) corresponding to the service recipient (SerUser) and three-dimensional coordinates (x, y, y) of the facial position (SerUser). u , y u , z u ).

[0057] When multiple users are in the service area 3, the processing device 140 identifies at least one user in the user image and selects the service recipient SerUser from the multiple users in the service area 3 through a user filtering mechanism. Figure 4 This is a schematic diagram illustrating the selection of a service recipient (SerUser) by an active interactive tour guide system according to an embodiment of the present invention. Please also refer to... Figure 2 and Figure 4The processing device 140 can filter out users outside the service area 3 and select the service recipients (SerUsers) from the users in the service area 3. In one embodiment, users closer to the user image acquisition device 130 can be selected as service recipients (SerUsers) based on their location. In another embodiment, users closer to the center of the user image acquisition device 130 can be selected as service recipients (SerUsers) based on their location. In yet another embodiment, it can also be as follows... Figure 4 As shown in the figure, based on the left-right relationship of the user User, the user User who is relatively in the middle is selected as the service object SerUser.

[0058] Once the processing device 140 identifies the user (User) from the user image (Img) and selects the service target (SerUser), the service area range (Ser_Range) will be displayed at the bottom of the user image (Img). The face of the service target (SerUser) on the user image (Img) will be marked with a focus point P1, and the distance between the service target (SerUser) and the user image acquisition device 130 will be displayed (e.g., 873.3 mm). At this time, the user image acquisition device 130 will first filter out other users (User) to more accurately focus on the service target (SerUser).

[0059] After the processing device 140 selects the SerUser to be served in the user image Img, it acquires the facial features of the SerUser, calculates the facial position and gaze direction of the SerUser using the facial feature points, and generates a number (ID) corresponding to the SerUser and three-dimensional coordinates (x, y, z) of the facial position. u , y u , z u ), where the focal point P1 can be located at the three-dimensional coordinates (x, y) of the face of the serviced object SerUser. u , y u , z u Additionally, the processing device 140 also generates facial depth information (h) based on the distance between the service recipient SerUser and the user image acquisition device 130. o ).

[0060] When the SerUser moves left or right within the service area Area3, the processing device 140 uses the three-dimensional coordinates (x, y, z) of the SerUser's face position. u , y u , z u The horizontal coordinate x in ) uUsing the SerUser object as the center point, the service field range Ser_Range is dynamically shifted according to its location. Figure 5 This is a schematic diagram illustrating the adjustment of the service area range Ser_Range according to an embodiment of the present invention. Please refer to it. Figure 5 When the service object SerUser moves left or right within the service area Area3, the service area range Ser_Range will dynamically translate left and right with the face position (focus point P1) of the service object SerUser as the center point, but the size of the service area range Ser_Range can remain unchanged.

[0061] The service area range, Ser_Range, can have an initial size (e.g., 60cm) or a variable size. When the served object, SerUser, moves back and forth within the service area 3, the size of the service area range, Ser_Range, can be adjusted appropriately according to the varying distance between the served object, SerUser, and the user image acquisition device 130. For example... Figure 5 As shown, the processing device 140 uses the face position (focus point P1) of the serviced object SerUser as the center point, and calculates the facial depth information (h) of the serviced object SerUser. o Adjust the left and right dimensions of the service field range Ser_Range, that is, adjust the left range Ser_Range_L and the right range Ser_Range_R of the service field range Ser_Range.

[0062] In one embodiment, the processing device 140 can process facial depth information (h) o The left range Ser_Range_L and right range Ser_Range_R of the service field range Ser_Range are calculated as follows:

[0063]

[0064] Here, width refers to the width value of the camera resolution. For example, if the camera resolution is 1280x720, then the width is 1280; similarly, if the camera resolution is 1920x1080, then the width is 1920. FOV W The field of view of the user image acquisition device 130.

[0065] Once the SerUser being served leaves the service area Area3, the processing device 140 can no longer detect the SerUser within the service area range Ser_Range. In one embodiment, the user image acquisition device 130 resets the size of the service area range Ser_Range and moves it to its initial position, such as the bottom center. The service area range Ser_Range can be moved to its initial position gradually or immediately. In another embodiment, the processing device 140 may not move the service area range Ser_Range to its initial position, but instead select the next SerUser being served from multiple Users in the service area Area3 using a user filtering mechanism. After selecting the next SerUser, the processing device 140 then uses the three-dimensional coordinates (x, y, x) of the next SerUser's face. u , y u , z u The horizontal coordinate x in ) u Using the center point, the service field range Ser_Range is dynamically shifted according to the position of the next serviced object SerUser.

[0066] In one embodiment, the present invention also provides an active interactive tour guide system, which can select service recipients from multiple users in a service area through a user filtering mechanism, identify the target object being looked at by the service recipient based on the service recipient's gaze, and display virtual information corresponding to the target object at the intersection position of the display device. Please refer to [further details omitted]. Figure 1 , 2 The active interactive tour guide system 1 includes a light-transmitting display device 110, a target object image acquisition device 120, a user image acquisition device 130, and a processing unit 140. The light-transmitting display device 110 is positioned between at least one user (User) and multiple dynamic objects (Obj). The target object image acquisition device 120 is coupled to the display device 110 to acquire dynamic object images of the dynamic objects (Obj). The user image acquisition device 130 is coupled to the display device 110 to acquire user images of the user (User).

[0067] Processing device 140 is coupled to display device 110. Processing device 140 is used to identify a dynamic object Obj in a dynamic object image and to track the dynamic object Obj. The processing device is further used to identify at least one user User in a user image, and to select a serviced object SerUser based on the range of a service area Area 3, and to detect the gaze S1 of the serviced object SerUser. The service area Area 3 has an initial size, and the gaze S1 of the serviced object SerUser traverses the display device 110 to gaze at the target object TarObj of the dynamic object Obj. The processing device 140 is further used to identify the target object TarObj gazed at by the serviced object SerUser based on the gaze S1 of the serviced object SerUser, and to generate three-dimensional coordinates (x, y, y) corresponding to the facial position of the serviced object SerUser. u , y u , z u ) and the corresponding three-dimensional coordinates of the target object TarObj and the depth and width information of the target object (h o , w o The system calculates the intersection point CP of the line of sight S1 of the served object SerUser across the display device 110, and displays the virtual information Vinfo corresponding to the target object TarObj at the intersection point CP of the display device 110. The detailed method has been described in the previous paragraphs and will not be repeated here.

[0068] In one embodiment, when the serviced object SerUser moves, the processing device 140 uses the three-dimensional coordinates (x, y, x) of the SerUser's face position. u , y u , z u The left and right dimensions of the service area Area3 are dynamically adjusted around the center point.

[0069] In one embodiment, when the processing device 140 fails to identify the service object SerUser within the service area 3 of the user image, the service area 3 is reset to its initial size.

[0070] The target image acquisition device 120, user image acquisition device 130, and processing device 140 described in this invention are written in a way that includes parallel operation of program code, and are equipped with a multi-core central processing unit to perform parallel processing in multiple threads.

[0071] Figure 6 This is a flowchart illustrating an active interactive tour guide method 6 according to an embodiment of the present invention. Please also refer to... Figure 1 , Figure 2 as well as Figure 6 , Figure 6The process of the proactive interactive guided tour method 6 can be... Figure 1 and Figure 2 This is achieved through an active interactive navigation system 1. Here, the user (the serviced object SerUser) can view the dynamic object Obj, the target object TarObj, and their corresponding virtual information VInfo through the display device 110 of the active interactive navigation system 1.

[0072] In step S610, the target image acquisition device 120 acquires an image of the dynamic object, identifies the dynamic object Obj in the dynamic object image, and tracks the dynamic object Obj. In step S620, the user image acquisition device 130 acquires a user image, identifies the user in the user image, and selects the serviced object SerUser. As described above, both the target image acquisition device 120 and the user image acquisition device 130 may include an RGB image sensing module, a depth sensing module, an inertial sensing module, and a GPS positioning sensing module to locate the positions of the user User, the serviced object SerUser, the dynamic object Obj, and the target object TarObj.

[0073] In step S630, the facial features of the serviced object SerUser are acquired, and it is determined whether the facial features match multiple facial feature points. If the facial features match multiple facial feature points, then in step S640, the gaze S1 of the serviced object SerUser is detected. If the facial features do not match multiple facial feature points, then in step S650, image segmentation is performed to segment the user image into multiple images to be recognized. User recognition is performed on each of the multiple images to be recognized until the facial features of the serviced object SerUser in at least one of the multiple images to be recognized match multiple facial feature points. Then, in step S640, the gaze S1 of the serviced object SerUser is detected. The gaze S1 passes through the display device 110 to gaze at the target object TarObj of the dynamic object Obj.

[0074] After detecting the gaze S1 of the service recipient SerUser, in step S660, the target object TarObj that the service recipient SerUser is looking at is identified based on the gaze S1 of the service recipient SerUser, and the three-dimensional coordinates (x, y, y) corresponding to the face position of the service recipient SerUser are generated. u , y u , z u ) and the corresponding three-dimensional coordinates (x, y) of the target object TarObj. o , y o ,z o ) and the depth and width information of the target object TarObj (h o , w oIn step S670, based on the three-dimensional coordinates (x, y) of the face of the serviced object SerUser... u , y u , z u ) and the three-dimensional coordinates (x, y) of the target object TarObj. o , y o , z o ), depth and width information (h o , w o ) Calculate the intersection point CP of the line of sight S1 of the served object SerUser across the display device 110. In step S680, display the virtual information Vinfo corresponding to the target object TarObj at the intersection point CP of the display device 110.

[0075] Figure 7 This is a flowchart illustrating an active interactive tour guide method 7 according to an embodiment of the present invention, mainly for further explanation. Figure 6 Steps S610 to S660 in the active interactive guided tour method 6 shown below. Please refer to... Figure 2 , 7 In step S711, the dynamic object image is acquired using the target object image acquisition device 120. In step S712, the target object TarObj being viewed by the service object SerUser is identified based on the gaze S1 of the service object SerUser. In step S713, the pixel features of the target object TarObj are acquired. In step S714, the pixel features are compared with multiple object feature points of each corresponding dynamic object Obj stored in the database 150. If the pixel features do not match the object feature points stored in the database 150, the process returns to step S711 to continue acquiring dynamic object images. If the pixel features match the object feature points, in step S715, a number corresponding to the target object TarObj and three-dimensional coordinates (x, y, y) corresponding to the position of the target object TarObj are generated. o , y o , z o ) and the depth and width information of the target object TarObj (w o , h o ).

[0076] On the other hand, in step S721, a user image is acquired using the user image acquisition device 130. In step S722, the user (User) is identified and the service recipient (SerUser) is selected. In step S723, the facial features of the service recipient (SerUser) are acquired. In step S724, it is determined whether the facial features of the service recipient (SerUser) match multiple facial feature points. If the facial features of the service recipient (SerUser) match the facial feature points stored in the database 150, then in step S725, the gaze (S1) of the service recipient (SerUser) is detected.

[0077] If the facial features of the service recipient SerUser do not match the facial feature points stored in database 150, in step S726a, image segmentation is performed to divide the user image into multiple images to be identified. User identification is performed on each of the multiple images to be identified until the facial features of the service recipient SerUser in at least one of the multiple images to be identified match multiple facial feature points. Then, in step S725, the gaze S1 of the service recipient SerUser is detected. On the other hand, in step S726b, a supplementary lighting mechanism is implemented on the user image acquisition device 130 to improve the clarity of the user image.

[0078] After detecting the gaze S1 of the service recipient SerUser, in step S727, the facial position of the service recipient SerUser and the gaze direction of gaze S1 are calculated using facial feature points. In step S728, a number (ID) corresponding to the service recipient SerUser and the three-dimensional coordinates (x, y, y) of the facial position are generated. u , y u , z u ).

[0079] When the target object TarObj's number, corresponding to the target object TarObj's three-dimensional coordinates (x, y, y) o , y o , z o ), the depth and width information of the target object TarObj (w o , h o ), corresponding to the SerUser's ID and the three-dimensional coordinates of the face position (x, y). u , y u , z u After all the data has been generated, in step S740, the three-dimensional coordinates (x, y, z) of the face of the serviced object SerUser are used to determine the location. u , y u , z u ) and the three-dimensional coordinates (x, y) of the target object TarObj. o , y o, z o ), depth and width information (h o , w o ) Calculate the intersection point CP of the line of sight S1 of the served object SerUser across the display device 110. In step S750, display the virtual information Vinfo corresponding to the target object TarObj at the intersection point CP of the display device 110.

[0080] In one embodiment, the active interactive navigation method of the present invention can determine whether the virtual information Vinfo corresponding to the target object TarObj is overlaid and displayed at the intersection position CP of the display device 110; if it is determined that the virtual information Vinfo is not overlaid and displayed at the intersection position CP of the display device 110, the position of the virtual information Vinfo can be offset and corrected by the information offset correction equation.

[0081] If the proportion of the SerUser in the user image is too small, making it impossible to collect the SerUser's facial features and calculate the SerUser's facial position and gaze direction S1 using facial feature points, the active interactive navigation method of this invention can first segment the user image into multiple images to be recognized. These images to be recognized include a central image to be recognized and multiple peripheral images to be recognized, wherein one of these images to be recognized has an overlapping area with an adjacent image, and the one of these images to be recognized and the adjacent image can be vertically adjacent, horizontally adjacent, or diagonally adjacent. The detailed method has been described in the previous paragraphs and will not be repeated here.

[0082] When multiple users are within the service area Area 3, the proactive interactive navigation method of this invention can identify at least one user from the user image using the processing device 140, and select the service target SerUser from the multiple users in the service area Area 3 through a user filtering mechanism. Once the user is identified from the user image Img and the service target SerUser is selected, the service area range Ser_Range will be displayed at the bottom of the user image Img to more accurately focus on the service target SerUser. The service area range Ser_Range may have an initial size or a variable size.

[0083] Once the SerUser is selected in the user image Img, its facial features are captured. The facial feature points are used to calculate the SerUser's facial position and gaze direction, generating a unique ID and three-dimensional coordinates (x, y) for the SerUser's face. u , y u , z u), where the focal point P1 can be located at the three-dimensional coordinates (x, y) of the face of the serviced object SerUser. u , y u , z u Additionally, facial depth information (h) will be generated based on the distance between the service recipient SerUser and the user image acquisition device 130. o ).

[0084] When the SerUser moves left or right within the service area Area3, the active interactive navigation method described in this invention will use the three-dimensional coordinates (x, y, z) of the SerUser's face position. u , y u , z u The horizontal coordinate x in ) u Centered on the location of the SerUser, the service area range Ser_Range is dynamically shifted based on the position of the SerUser being served. When the SerUser moves left or right within the service area Area3, the service area range Ser_Range will dynamically shift left or right with the SerUser's face position (focus point P1) as the center point, but the size of the service area range Ser_Range can remain unchanged.

[0085] When the serviced object SerUser moves back and forth within the service area Area3, the width of the service area Ser_Range can be adjusted appropriately depending on the distance between SerUser and the user image acquisition device 130. The detailed implementation has been described in previous paragraphs and will not be repeated here.

[0086] In summary, the active interactive tour guide system and method described in the embodiments of the present invention can track the viewing user's gaze direction in real time, stably track moving target objects, and actively display virtual information corresponding to the target objects, providing highly accurate augmented reality information and a comfortable non-contact interactive experience. The embodiments of the present invention can also integrate internal and external perception recognition, virtual-real fusion, and system virtual-real fusion pairing calculation cores, actively using internal perception to determine the viewing angle of the visitor and then matching it with external perception AI target object recognition to realize augmented reality applications. Furthermore, the embodiments of the present invention also optimize the virtual-real fusion display position correction algorithm to perform offset correction methods, improve facial recognition of users at a distance, and prioritize the selection of service objects, which can greatly solve the problem of insufficient manpower and create an interactive experience where knowledge and information are conveyed at zero distance.

[0087] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.

Claims

1. An active interactive tour guide system, characterized in that, include: A light-transmitting display device positioned between at least one user and multiple dynamic objects; An image acquisition device for a target object, coupled to the display device, is used to acquire images of dynamic objects; A user image acquisition device, coupled to the display device, is used to acquire user images; as well as A processing device, coupled to the display device, is configured to identify and track dynamic objects in the dynamic object image, and to identify at least one user and select a service object in the user image, acquire facial features of the service object and determine whether the facial features match multiple facial feature points. If the facial features match the facial feature points, the processing device detects the gaze of the service object. If the facial features do not match the facial feature points, the processing device performs image segmentation to segment the user image into multiple images to be identified, and performs user identification on each of the images to be identified to detect the gaze of the service object, wherein the gaze passes through the display device to gaze at a target object of the dynamic objects. The processing device is further used to identify the target object being looked at by the service recipient based on the gaze, generate three-dimensional coordinates corresponding to the face position of the service recipient and three-dimensional coordinates corresponding to the position of the target object and the depth and width information of the target object, calculate the intersection position of the gaze across the display device, and display the virtual information corresponding to the target object at the intersection position of the display device.

2. The active interactive tour guide system according to claim 1, characterized in that, These images to be identified include a central image to be identified and multiple surrounding images to be identified.

3. The active interactive tour guide system according to claim 1, characterized in that, One of the images to be identified has an overlapping area with an adjacent image.

4. The active interactive tour guide system according to claim 3, characterized in that, One of the images to be identified can be adjacent to the other in a vertical, horizontal, or diagonal manner.

5. The active interactive tour guide system according to claim 1, characterized in that, The target image acquisition device, the user image acquisition device, and the processing device are all written in a way that includes parallel computation, and are paired with a multi-core central processing unit to perform parallel processing using multiple threads.

6. The active interactive tour guide system according to claim 1, characterized in that, The user image acquisition device includes an RGB image sensing module, a depth sensing module, an inertial sensing module, and a GPS positioning sensing module.

7. The active interactive tour guide system according to claim 1, characterized in that, If the facial feature matches the facial feature points, the processing device uses the facial feature points to calculate the facial position and the direction of the gaze of the service recipient, and generates a number corresponding to the service recipient and the three-dimensional coordinates of the facial position.

8. The active interactive tour guide system according to claim 7, characterized in that, The processing device identifies at least one user in the user image and selects the service object according to a service area range, wherein the service area range has an initial size, and when the service object moves, the processing device dynamically adjusts the left and right size of the service area range with the three-dimensional coordinates of the face position of the service object as the center point.

9. The active interactive tour guide system according to claim 8, characterized in that, When the processing device does not identify the served object within the service area of ​​the user image, it resets the service area to the initial size.

10. The active interactive tour guide system according to claim 1, characterized in that, Including: A database, coupled to the processing device, is used to store multiple object feature points corresponding to each of these dynamic objects; When the processing device identifies the target object being looked at by the service recipient, the processing device collects the pixel features of the target object and compares the pixel features with the object feature points; If the pixel feature matches the object feature points, the processing device generates a number corresponding to the target object, the three-dimensional coordinates of the position of the target object, and the depth and width information of the target object.

11. The active interactive tour guide system according to claim 1, characterized in that, The processing device determines whether the virtual information corresponding to the target object is overlaid and displayed at the intersection position of the display device; If the virtual information is not overlaid on the intersection of the display device, the processing device performs offset correction on the position of the virtual information.

12. An active interactive tour guide method, applicable to an active interactive tour guide system having a light-transmitting display device, a target object image acquisition device, a user image acquisition device, and a processing device, characterized in that, The display device is positioned between at least one user and multiple dynamic objects, and the processing device is used to execute the proactive interactive navigation method, which includes: The target object image acquisition device acquires images of dynamic objects, identifies the dynamic objects in the images, and tracks the dynamic objects. The system acquires a user image using the user image acquisition device, identifies at least one user in the user image, selects a service recipient, acquires the facial features of the service recipient, and determines whether the facial features match multiple facial feature points. If the facial features match the facial feature points, the system detects the service recipient's gaze. If the facial features do not match the facial feature points, the system performs image segmentation to divide the user image into multiple images to be identified. For each of these images to be identified, the system performs user identification to detect the service recipient's gaze, wherein the gaze passes through the display device to focus on a target object among the dynamic objects; and Based on the gaze, the target object being looked at by the service recipient is identified, and three-dimensional coordinates of the face position of the service recipient, three-dimensional coordinates of the position of the target object, and depth and width information of the target object are generated. The intersection point of the gaze across the display device is calculated, and virtual information corresponding to the target object is displayed at the intersection point of the display device.

13. The active interactive tour guide method according to claim 12, characterized in that, These images to be identified include a central image to be identified and multiple surrounding images to be identified.

14. The active interactive tour guide method according to claim 12, characterized in that, One of the images to be identified has an overlapping area with an adjacent image.

15. The active interactive tour guide method according to claim 14, characterized in that, One of the images to be identified can be adjacent to the other in a vertical, horizontal, or diagonal manner.

16. The active interactive tour guide method according to claim 12, characterized in that, Including: If the facial feature matches the facial feature points, then the facial feature points are used to calculate the facial position of the service recipient and the direction of the gaze. as well as Generate a number corresponding to the serviced object and the three-dimensional coordinates of the face position.

17. The active interactive tour guide method according to claim 16, characterized in that, The steps of acquiring a user image using the user image acquisition device, identifying at least one user from the user image, and selecting the service recipient further include: The system identifies at least one user in the user image and selects the service object based on a service area range, wherein the service area range has an initial size, and when the service object moves, the left and right sizes of the service area range are dynamically adjusted with the three-dimensional coordinates of the face position of the service object as the center point.

18. The active interactive tour guide method according to claim 17, characterized in that, Including: If the service object is not identified within the service area of ​​the user image, the service area is reset to the initial size.

19. The active interactive tour guide method according to claim 12, characterized in that, Including: Once the target object being viewed by the service recipient is identified, the pixel features of the target object are collected, and these pixel features are compared with the object's feature points; and If the pixel feature matches the object feature points, a number corresponding to the target object, the three-dimensional coordinates of the target object's position, and the depth and width information of the target object are generated.

20. The active interactive tour guide method according to claim 12, characterized in that, Including: Determine whether the virtual information corresponding to the target object is overlaid and displayed at the intersection position of the display device; and If the virtual information is not overlaid on the intersection of the display device, the position of the virtual information is offset and corrected.

21. An active interactive tour guide system, characterized in that, include: A light-transmitting display device positioned between at least one user and multiple dynamic objects; An image acquisition device for a target object, coupled to the display device, is used to acquire images of dynamic objects; A user image acquisition device, coupled to the display device, is used to acquire user images; as well as A processing device, coupled to the display device, is used to identify and track the dynamic objects in the dynamic object image. The processing device is further used to identify at least one user in the user image and select a service object according to a service field range, and detect the gaze of the service object, wherein the service field range has an initial size, the gaze passes through the display device to focus on a target object of the dynamic objects, and when the service object moves, the processing device dynamically adjusts the left and right size of the service field range with the three-dimensional coordinates of the face position of the service object as the center point. The processing device is further used to identify the target object being looked at by the service recipient based on the gaze, generate three-dimensional coordinates corresponding to the face position of the service recipient and three-dimensional coordinates corresponding to the position of the target object and the depth and width information of the target object, calculate the intersection position of the gaze across the display device, and display the virtual information corresponding to the target object at the intersection position of the display device.

22. The active interactive tour guide system according to claim 21, characterized in that, When the processing device does not identify the served object within the service area of ​​the user image, it resets the service area to the initial size.