AI digital human-based way director interaction system and method

By dividing semantic zones and performing consistency checks in the wayfinding interaction system, the problems of path deviation recognition and interactive feedback self-correction in dynamic environments are solved, realizing stable and intelligent adaptive navigation for AI digital humans.

CN121387084AActive Publication Date: 2026-01-23SHANGHAI ZEMSO ELECTRONICS TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511935251.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-01-23
Estimated Expiration
2045-12-22

AI Technical Summary

Technical Problem

Existing guidance interaction methods struggle to identify semantic path deviations and achieve interactive feedback self-correction in dynamic environments. They lack self-learning capabilities and self-consistent update mechanisms, leading to unstable guidance from AI digital humans.

Method used

By collecting on-screen input events, establishing reference lines, dynamically determining the user's physical location and line of sight, dividing the path into semantic zones and arranging them into an ordered sequence of semantic zones, and combining this with interaction logs for consistency verification, the AI ​​digital human can achieve accurate broadcasting and interactive control.

Benefits of technology

It achieves semantic modeling and structured expression of paths, ensuring the continuity and logical consistency of the navigation process, and has the ability to identify and automatically correct misalignments in real time, thereby improving the accuracy of human-computer interaction and continuous guidance capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121387084A_ABST
    Figure CN121387084A_ABST
Patent Text Reader

Abstract

The invention discloses a way director interaction system and method based on an AI digital human, and relates to the technical field of artificial intelligence man-machine interaction, and the method comprises the steps: collecting a screen front input event, normalizing the event into an event stream, building a reference line, and dynamically judging the physical position and the sight line area of a user under environment perception; based on a reference line, a user physical position and a sight line area, dividing semantic zones from a starting point to a target position, arranging the semantic zones into an ordered semantic zone sequence, binding an on-screen instruction and a plurality of on-screen candidate target areas for each semantic zone, dynamically sorting the candidate target areas, and establishing an interaction log; and the AI digital person broadcasts according to the ordered semantic zone sequence, highlights the on-screen indication and the on-screen candidate target area when broadcasting to the semantic zone name, and judges whether the sight of the user resides in the first candidate target area in a fixed confirmation time period. The continuity and logic consistency of path expression are kept in a complex space, and the accuracy of man-machine interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence human-computer interaction, in particular to a guiding machine interaction system and method based on an AI digital person. BACKGROUND

[0002] With the rapid development of artificial intelligence and human-computer interaction technology, interactive devices based on speech recognition, gesture recognition and visual perception have been widely used in public navigation, exhibition tour and commercial guidance scenarios. Traditional guiding machines usually use a mode based on touch input or fixed voice broadcast to provide direction prompts for users through a pre-set path and a map interface. In recent years, guiding interaction methods combining computer vision and depth perception have emerged, which can collect spatial position and line-of-sight information of users through a camera to achieve a more immersive interaction experience. AI-driven AI digital person technology further promotes this development, enabling guiding processes to have semantic understanding and situational response capabilities. AI digital persons can adjust the content of the speech, visual pointing and feedback actions according to user input and state to complete path broadcasting, semantic prompting and multi-round confirmation interaction in complex environments, significantly improving the naturalness and continuity of human-computer communication.

[0003] However, conventional guiding interaction methods still have limitations in maintaining path consistency and utilizing user behavior feedback in dynamic environments. Although traditional methods can record interaction events, they mostly stay at the input response level and lack the ability to reconstruct real semantic paths from behavior logs, making it impossible to judge the deviation relationship between the actual guiding path and the pre-set path. After semantic zone misplacement, candidate target area misselection or user rollback operation, existing methods cannot achieve path self-correction and instruction self-adaptive adjustment, resulting in a lack of self-learning ability and self-consistency update mechanism for AI digital person guidance. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides a guiding machine interaction method based on an AI digital person to solve the problems of unidentifiable semantic path deviation and uncorrectable interaction feedback in the guiding process.

[0006] To solve the above technical problems, the present application provides the following technical solutions:

[0007] In a first aspect, the present application provides a guiding machine interaction method based on an AI digital person, which includes,

[0008] Collecting screen input events and normalizing them into an event stream, establishing a reference line, dynamically judging the user's physical position and line-of-sight area under environmental perception;

[0009] Based on the reference line, the user's physical location and the line of sight area, the starting point to the target position is divided into semantic zones and arranged into an ordered semantic zone sequence, the screen indication and several candidate target areas are bound for each semantic zone, the candidate target areas are dynamically sorted, and an interaction log is established;

[0010] The AI digital person reports according to the ordered semantic zone sequence, when reporting to the semantic zone name, highlights the screen indication and the candidate target area on the screen, and judges whether the user's line of sight stays in the candidate target area at the top of the sorting within a fixed confirmation period;

[0011] Based on the interaction log, the observed semantic zone sequence is reconstructed and checked for consistency with the ordered semantic zone sequence.

[0012] As a preferred scheme of the AI digital person-based navigation machine interaction method, the collected screen front input events are normalized into an event stream, and the reference line is established, and the specific steps are,

[0013] The navigation machine uniformly records the atomic inputs generated by the screen side area and the screen front side area, each atomic input is written as an event record, the event records are strictly sorted by time, and an event stream is formed;

[0014] The geometric mapping relationship from the screen plane to the ground reference plane is obtained through the paired point positions composed of the calibration points laid on the ground and the screen pixel points, the upper end and the lower end pixel points of the screen vertical center line are taken, the upper end and the lower pixel points are converted into actual point positions on the ground reference plane, the line connecting the actual point positions is taken as the reference line, and the direction connected by the actual point positions is taken as the reference line direction.

[0015] As a preferred scheme of the AI digital person-based navigation machine interaction method, the user's physical location and the line of sight area are dynamically judged under environmental perception, and the specific steps are,

[0016] The user's physical location is obtained through the depth imaging camera of the navigation machine, the user's head orientation is obtained according to the user's head posture, and the user's head orientation is converted to the ground reference coordinate system consistent with the reference line;

[0017] The included angle between the user's head orientation and the reference line direction is compared, when the included angle does not exceed the forward angle threshold, it is considered that the user's line of sight is oriented to the front of the reference line, when the included angle falls within the left allowable interval, it is considered that the user's line of sight is oriented to the left, when it falls within the right allowable interval, it is considered that the user's line of sight is oriented to the right, and when the included angle exceeds the left allowable interval and the right allowable interval and no interaction is performed within a fixed time, it is considered that the interaction is ended;

[0018] Measuring the lateral nearest distance of the user's physical position to the reference line, when the lateral nearest distance does not exceed the forward bandwidth threshold, it is considered that the user is in the effective interaction area of the navigation machine, when the lateral nearest distance exceeds the forward bandwidth threshold and no interaction is performed within the fixed time, it is considered that the user has deviated from the effective interaction area of the navigation machine;

[0019] When the included angle does not exceed the forward angle threshold and the lateral nearest distance does not exceed the forward bandwidth threshold, the current line of sight area is classified as the reference line forward dominant area, when the included angle is in the left allowed interval and the lateral nearest distance does not exceed the forward bandwidth threshold, the current line of sight area is classified as the left adjacent area, and when the included angle is in the right allowed interval and the lateral nearest distance does not exceed the forward bandwidth threshold, the current line of sight area is classified as the right adjacent area.

[0020] As a preferred scheme of the navigation machine interaction method based on the AI digital person of the application, wherein: the starting point to the target position is divided into semantic zones and arranged into an ordered semantic zone sequence, and the specific steps are,

[0021] According to the user's physical position, the user's current position is determined as the starting point, the preset path is read in the navigation machine built-in site diagram, the reference line direction is taken as the reference, the topological nodes are identified in turn along the preset path, and the continuous road section between adjacent topological nodes is taken as a semantic zone.

[0022] From the semantic zone where the starting point is located to the semantic zone where the target position is located, an ordered semantic zone sequence is obtained.

[0023] As a preferred scheme of the navigation machine interaction method based on the AI digital person of the application, wherein: the screen indication and a plurality of candidate target areas are bound for each semantic zone, and the candidate target areas are dynamically sorted, and the specific steps are,

[0024] The direction description is recorded on each semantic zone, and the arrow, the annotation line and the AI digital person speech anchor point are aligned with the direction description of the semantic zone.

[0025] The target areas located in the geometric range of the current semantic zone and the target areas which are adjacent to the next semantic zone and intersect or tangent with the border of the current semantic zone are merged into the candidate target area set of the semantic zone.

[0026] Whenever an effective record is added to the event stream, the candidate target area set of the current semantic zone and the next semantic zone is rearranged in the order of line of sight consistency priority, zone consistency priority, event nearest association priority, nearest priority along the reference line and turning complexity priority.

[0027] As a preferred scheme of the navigation machine interaction method based on the AI digital person, when the semantic zone name is broadcasted, the on-screen indication and the on-screen candidate target area are highlighted, and the specific steps are,

[0028] When the semantic zone name field is broadcasted, an arrow is immediately lit on the screen to describe the direction of the current semantic zone, and the highlight is performed on the icon of the top candidate target area.

[0029] As a preferred scheme of the navigation machine interaction method based on the AI digital person, when the semantic zone name is broadcasted, the on-screen indication and the on-screen candidate target area are highlighted, and the specific steps are,

[0030] In the fixed confirmation period, the user's gaze landing point is sampled at equal intervals and attributed to the on-screen labeled candidate target area, and the cumulative residence time and continuity of each candidate target area covered by the gaze landing point are continuously counted. When the top candidate target area always maintains the first cumulative residence time and is not interrupted in the fixed confirmation period, it is considered that the gaze confirmation is passed, otherwise it is considered that the gaze confirmation is not passed.

[0031] If the gaze confirmation is not passed, a point selection confirmation control is presented around the top candidate target area, and a back control is presented on the side of the screen.

[0032] If the gaze confirmation is passed, the current semantic zone is marked as confirmed, and the AI digital person continues to broadcast the name and action statement of the next semantic zone, and the cycle continues until the semantic zone where the target position is located.

[0033] As a preferred scheme of the navigation machine interaction method based on the AI digital person, the observed semantic zone sequence is reconstructed based on the interaction log, and the specific steps are,

[0034] The semantic zone identifier is extracted by reading the operation category record in the interaction log, and is spliced into the observed semantic zone sequence in time sequence.

[0035] As a preferred scheme of the navigation machine interaction method based on the AI digital person, the consistency check is performed with the ordered semantic zone sequence, and the specific steps are,

[0036] The observed semantic zone sequence and the ordered semantic zone sequence are aligned through the minimum step dynamic programming table, the minimum difference step number is calculated, the earliest misaligned semantic zone is located, and one-key correction and re-direction are provided on the site diagram.

[0037] In a second aspect, the application provides a navigation machine interaction system based on an AI digital person, comprising,

[0038] A collection judgment module collects pre-screen input events and normalizes them into an event stream, establishes a reference line, and dynamically judges the user's physical location and line-of-sight area under environmental perception;

[0039] A semantic zone module, based on the reference line, the user's physical location and line-of-sight area, divides the starting point to the target location into semantic zones and arranges them into an ordered semantic zone sequence, binds a screen indicator and several candidate target areas on the screen to each semantic zone, dynamically sorts the candidate target areas, and establishes an interaction log;

[0040] An AI digital person broadcasting module, the AI digital person broadcasts according to the ordered semantic zone sequence, when broadcasting to the semantic zone name, highlights the screen indicator and the candidate target area on the screen, and judges whether the user's line of sight is located in the top candidate target area within a fixed confirmation period;

[0041] A reconstruction checking module, based on the interaction log, reconstructs the observed semantic zone sequence and checks its consistency with the ordered semantic zone sequence.

[0042] The present application has the following advantages: by dividing semantic zones and arranging them into an ordered semantic zone sequence, semantic modeling and structured expression of the path are achieved, enabling the AI digital person to make accurate broadcasts and interactive control, thereby maintaining the continuity and logical consistency of path expression in complex spaces, through consistency checking, real-time misalignment recognition and automatic correction during navigation are achieved, the AI digital person performs self-correction and self-learning, ensuring the stability and intelligent adaptability of the route guidance process, and improving the accuracy and continuous guidance ability of human-computer interaction. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0044] Fig. 1 The flowchart of the route guidance machine interaction method based on AI digital person.

[0045] Fig. 2 The schematic diagram of the route guidance machine interaction system based on AI digital person.

[0046] Fig. 3 The flowchart of the reference line establishment.

[0047] Fig. 4 The flowchart of dynamically judging the user's line-of-sight area. DETAILED DESCRIPTION

[0048] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0049] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details set forth in this description. In other instances, well-known methods have not been described in detail in order to avoid obscuring the present application.

[0050] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an embodiment that is independent or alternative to other embodiments.

[0051] Reference Figs. 1-4 For one embodiment of the present application, the embodiment provides a guide machine interaction method based on an AI digital person, comprising the following steps:

[0052] S1, collect the on-screen input events and normalize them into an event stream, establish a reference line, and dynamically determine the user's physical location and line-of-sight area under environmental perception.

[0053] The guide machine collects the atomic inputs generated in the screen side area and the on-screen side area at the start of the interaction, and records them uniformly.

[0054] Atomic input refers to the smallest and non-disassembled input event in the human-computer interaction of the guide machine, including touch down, touch up, click confirmation, back trigger, etc.

[0055] Each atomic input is written as an event record, and the event record content includes event time, event type (screen side or on-screen side), screen coordinates or spatial position associated with atomic input, and a measure representing strength or importance (such as touch pressure or dwell time). All event records are strictly sorted by time to form an event stream, and are continuously appended during the entire interaction process.

[0056] When the guide machine is first deployed, a one-time calibration is performed, and the guide machine screen pixel points corresponding to the four calibration points laid on the ground form a pair of point positions by the guide machine screen pixel points and the ground calibration points, and the geometric mapping relationship from the screen plane to the ground reference plane is obtained. Take the upper end and the lower end of the screen vertical center line as two pixel points, convert the upper end and the lower end of the two pixel points to two actual point positions on the ground reference plane, and the line connecting the two actual point positions is the reference line, and the direction connected by the two actual point positions is the reference line direction.

[0057] The depth imaging camera above the guidance machine will collect data of the area in front of the screen in real time, locate the user in the area in front of the screen, obtain the three-dimensional positions of the center of the user's head and the center of the user's torso, orthogonally project the center of the user's head to the ground reference plane in the vertical direction, and obtain the physical position of the user.

[0058] After obtaining the physical position of the user, the user's head orientation is obtained according to the user's head posture, and the user's head orientation is converted to the ground reference coordinate system consistent with the reference line.

[0059] The included angle between the user's head orientation and the direction of the reference line is compared, if the included angle does not exceed the forward angle threshold, it is considered that the user's line of sight is approximately forward to the front of the reference line, if the included angle falls into the left allowable interval, it is considered that the user's line of sight is leftward, if it falls into the right allowable interval, it is considered that the user's line of sight is rightward, and if the included angle exceeds the left allowable interval and the right allowable interval and no interaction (such as touch, click, entering the left allowable interval and the right allowable interval, etc.) is performed within a fixed time, it is considered that the interaction is ended.

[0060] The forward angle threshold is preset to 20°, under typical interaction geometry, the half visual angle from the center line of the screen to the screen edge is about 14° to 27° (determined by the ratio of screen width to interaction distance), and 20° is located in the upper part of the interval, which covers the central fixation band of common screen viewing, field viewing, and screen returning, and leaves a clear boundary for the left and right allowable intervals. Moving the boundary to 19° will be biased to a too narrow forward direction, triggering the left and right allowable intervals too early, and moving to 21° will be biased to a too wide forward direction, delaying the triggering of the left and right allowable intervals. The head orientation estimation is commonly in the order of 3° to 5° error in public scenarios, and 20° provides a multiple error margin for the forward fixation, reducing the probability of misjudging the forward as lateral, 19° reduces the margin, and edge samples are more likely to be misjudged as lateral, and 21° retains part of the obvious lateral samples in the forward direction, reducing the direction differentiation.

[0061] The left allowable interval and the right allowable interval are both greater than 20° and do not exceed 45°, the central fixation is superimposed with slight head movement, and within 30°, it is generally still able to be stably recognized, and when exceeding 40°, the user usually needs to rotate the torso together, the screen falls into the far peripheral field of view, and it is difficult to simultaneously see and understand the same thing and the same direction. The upper boundary of the allowable interval is set to 45°, which covers the true lateral viewing and avoids considering large lateral or back as effective interaction. If the upper limit of the allowable interval exceeds 45°, it will consider the posture that is not suitable for the screen as an effective lateral, which is easy to cause orientation jitter and misclassification.

[0062] The transverse closest distance from the user's physical position to the reference line is measured, if the transverse closest distance does not exceed the forward bandwidth threshold, it is considered that the user is in the effective interaction area of the guidance machine, and if the transverse closest distance exceeds the forward bandwidth threshold and no interaction is performed within a fixed time, it is considered that the user has deviated from the effective interaction area of the guidance machine.

[0063] The forward bandwidth threshold is preset as 0.8 meters, the shoulder width of a typical adult is about 0.45 meters, there is natural left and right swing of about 0.1 to 0.15 meters in static residence, and about 0.2 meters of safety margin is left, totaling about 0.8 meters, which can cover most normal standing positions and slight sideways, and will not misjudge natural small displacement as leaving the effective area. Common interaction distance is between 0.8 meters and 1.5 meters, and even if the user deviates laterally to 0.8 meters, as long as the head is still within the forward 20°, the on-screen arrow, highlight and speech anchor point are still within the stable reading range. When the lateral deviation continues to increase, the user often needs to twist or move his body, and the forward determination and confirmation are easy to be unstable. The effective field of view of the on-screen depth imaging camera in a public scene can usually cover the bandwidth within 0.8 meters, and the head and trunk key points are less blocked by the screen edge and others. If it is further relaxed, such as more than 0.9 meters, it is easier to include the person on the side into the same bandwidth, increasing the misassociation and ordering jitter. 0.8 meters as the upper limit of the single-person effective bandwidth can form a distinguishable band with the spacing of the people flow in the common corridor or lobby, reducing the probability of misjudging adjacent users into the same effective band, and ensuring that the line-of-sight consistency priority is not disturbed by the side temporary stay. 0.7 meters is too harsh for users with a large shoulder width or those carrying items, and normal micro-movement is easy to be judged as deviating from the bandwidth, resulting in unnecessary backtracking and reconfirmation. 0.9 meters still includes the obvious side displacement scene into the effective bandwidth, and is more likely to treat the person on the side or the cross-displacement standing position as an effective user, increasing the risk of misadvancement.

[0064] When the included angle does not exceed the forward angle threshold and the lateral closest distance does not exceed the forward bandwidth threshold, the current line-of-sight area is classified as the reference line forward dominant area. When the included angle is in the left allowed interval and the lateral closest distance does not exceed the forward bandwidth threshold, the current line-of-sight area is classified as the left adjacent area. When the included angle is in the right allowed interval and the lateral closest distance does not exceed the forward bandwidth threshold, the current line-of-sight area is classified as the right adjacent area.

[0065] S2, based on the reference line, the user's physical position and the line-of-sight area, divide the start point to the target position into semantic zones and arrange them into an ordered semantic zone sequence, bind on-screen indicators and a plurality of candidate target areas on the screen to each semantic zone, dynamically sort the candidate target areas, and establish an interaction log.

[0066] According to the user's physical location, the user's current location is determined as the starting point, the preset path is read in the built-in site map of the navigation machine, the reference line direction is taken as the reference, the topological nodes (such as entrances, corridor corners, stair openings, doorways, escalator drop points, etc. which can cause changes in direction or floor changes, in open spaces without obvious geometric transitions, function points are used as temporary nodes) are identified along the preset path in sequence, and the continuous road segment between adjacent topological nodes is defined as a semantic zone. Starting from the semantic zone where the starting point is located, ending with the semantic zone where the target location is located, an ordered semantic zone sequence is obtained.

[0067] The preset path refers to the fixed path from the starting point to the target location pre-set in the site map embedded in the navigation machine.

[0068] A direction description is recorded on each semantic zone, and the direction description is derived from the turning relationship of the preset path at the starting position of the semantic zone. If the route remains in the same direction as the reference line, it is marked as keeping forward. If the route has a left turning trend relative to the reference line, it is marked as turning left. If the route has a right turning trend relative to the reference line, it is marked as turning right.

[0069] Align the arrow, annotation line and AI digital human dialogue anchor point with the direction description of the semantic zone. If the direction description of the semantic zone is to keep forward, the arrow points outward along the reference line. If the direction description of the semantic zone is to turn left or right, a turning or entering prompt will be given in advance at the node where the current semantic zone enters the next semantic zone.

[0070] The arrow refers to the direction symbol on the screen, which is consistent with the direction of travel of the semantic zone. When encountering a turning semantic zone, it is turned to the next direction at the inflection point.

[0071] The annotation line refers to the thin line or path line on the screen that connects the current location to the candidate target area icon.

[0072] The AI digital human dialogue anchor point refers to the on-screen drop point or object identification that the AI digital human looks at, points to, and speaks.

[0073] The target area within the semantic zone geometry range and the target area in the next semantic zone adjacent to the border of the semantic zone are merged into the candidate target area set of the semantic zone. To avoid the candidate target area set being too wide, the remote target area not in the current semantic zone and not in the next semantic zone adjacent to the semantic zone is not included, and a lateral attribute is added to each candidate target area in the candidate target area set. Taking the representative point (the default is the geometric center of the candidate target area, and the nearest point of the passable path is used for irregular shapes) of the candidate target area on the site diagram as the reference line, if the representative point falls on the left half plane relative to the reference line, it is marked as left side, if the representative point falls on the right half plane relative to the reference line, it is marked as right side, if the representative point falls on the reference line or points to the target position in the same direction as the reference line, it is marked as forward, if the candidate target area crosses the left and right half planes at the same time, the half plane where the representative point is located is used as the reference.

[0074] Whenever the event stream adds a valid record or the current view area changes, the candidate target area set of the current semantic zone and the next semantic zone adjacent to the semantic zone is rearranged in the order of view consistency priority, zone consistency priority, event closest association priority, nearest along the reference line priority, and turning complexity priority.

[0075] View consistency priority means that if the current view area is the forward dominant area of the reference line, the candidate target area with the forward lateral attribute is placed in front of all candidate target areas, if the current view area is the left adjacent area, the candidate target area with the left lateral attribute is placed in front, if it is the right adjacent area, the candidate target area with the right lateral attribute is placed in front.

[0076] Zone consistency priority means that the candidate target area located in the current semantic zone is placed in front of the candidate target area located in the next semantic zone adjacent to the semantic zone.

[0077] Event closest association priority means that if the last valid touch or click of the event stream is associated with the on-screen indication of any candidate target area (icon hit, dialogue anchor hit, annotation line hit, highlight overlay hit, and target surface layer hit), the current candidate target area is moved up in the candidate target area set.

[0078] Nearest along the reference line priority means that from the perspective of the reference line direction, the user's physical position to the representative point of the candidate target area on the site diagram along the reference line is closer in priority.

[0079] Turning complexity priority means that the fewer turns required from the current semantic zone to the candidate target area are preferred, and if they are still in parallel, the fewer consecutive turns are preferred.

[0080] After dynamic ranking, if still completely parallel, keep the original relative order of parallel items to reduce screen jitter, when there is no any candidate target area in the line of sight area, preferentially enable the candidate target area in the next semantic zone adjacent to the current semantic zone direction description consistent with the current semantic zone direction description, if there is still no candidate target area, keep the last ranking and on-screen indication unchanged, wait for the event stream to appear new effective records or get stable line of sight area again.

[0081] Build an interaction log, which is an append-only structure, write event time and event type, trigger writing reason, current semantic zone identifier and direction description, current line of sight area, candidate target area identifier and order before and after rearrangement, and projection of user physical position on site diagram, writing opportunity includes event stream appending effective record, current line of sight area change and semantic zone switching.

[0082] S3, AI digital person reports according to the ordered semantic zone sequence, when reporting to the semantic zone name, highlight the on-screen indication and on-screen candidate target area, and judge whether the user's line of sight is in the first candidate target area within a fixed confirmation period.

[0083] Take out the current semantic zone and candidate target area set from the ordered semantic zone sequence, read the direction description and arrow of the current semantic zone, label line and AI digital person speech anchor point, and read the dynamic ranking of the candidate target area.

[0084] The AI digital person starts to report the current semantic zone name and action sentence (for example, straight ahead to the corridor turning point, and then left into the service area), when reporting to the semantic zone name field, immediately light up the arrow on the screen along the direction description of the current semantic zone, and mark the turning prompt at the end of the current semantic zone in advance, while executing highlighting on the first candidate target area icon to mark the line from the current position to the candidate target area icon, and drop the AI digital person speech anchor point on the candidate target area icon or the side, forming a visual and voice synchronization effect of the same saying and pointing. On the screen, only the first candidate target area is highlighted, and the rest of the candidate target areas are presented in a secondary style with constant light without flickering, avoiding screen jitter.

[0085] After entering the fixed confirmation period, sample the user's line of sight drop point at equal intervals and attribute it to the on-screen labeled candidate target area. In the entire fixed confirmation period, continuously count the cumulative residence time and continuity of each candidate target area covered by the line of sight drop point, if the first candidate target area always keeps the cumulative residence time first and is not interrupted by the continuous residence of other candidate target areas in the fixed confirmation period, it is considered that the line of sight confirmation is passed, otherwise it is considered that the line of sight confirmation is not passed.

[0086] The fixed confirmation period is set to 2 seconds, and the front-view line collection and display refresh of the main stream screen is 30 to 60 times per second. In the 2-second period, 60 to 120 sample points can be obtained, the statistical fluctuation is smaller, the result is stable, and enough samples can be maintained even at a lower frequency.

[0087] If the line-of-sight confirmation fails, the on-screen indication and highlight state remain unchanged, and the semantic zone is not advanced. A point selection confirmation control is presented around the highlight of the candidate target region at the top of the sequence, and a back control is presented on one side of the screen for the user to back up to the previous semantic zone or retrigger the display. When any candidate target region is selected, the selected candidate target region is immediately moved to the top of the sequence in the current semantic zone and is considered as an explicit confirmation pass. If the back control is triggered, the back operation is recorded and the previous semantic zone is returned according to the ordered semantic zone sequence, and the announcement and line-of-sight residence determination are re-performed.

[0088] If the line-of-sight confirmation passes or the explicit confirmation passes, the current semantic zone is marked as confirmed, the AI digital person continues to announce the name and action statement of the next semantic zone, the on-screen indication synchronously advances the arrow to the starting position of the next semantic zone, and presents the next transition prompt in advance. According to the dynamic sequence of the candidate target regions, the candidate target region set in the next semantic zone is reordered, and the candidate target region at the top of the sequence is highlighted on the screen. The fixed confirmation period is entered again, and the cycle continues until the target position is reached.

[0089] During the entire announcement process, the time mark, operation category (such as starting announcement, entering confirmation observation, explicit point selection confirmation, and back), current semantic zone identification and direction description, on-screen indication element identification associated with the current semantic zone (such as arrow, annotation line, AI digital person speech anchor point), candidate target region identification, pre-reordering sequence and post-reordering sequence, reordering trigger reason (event stream addition effective record or current line-of-sight region change), residence cumulative duration of the candidate target region at the top of the sequence in the confirmation and whether the continuity is interrupted, point selection confirmation and back operation, user physical position point and current line-of-sight region value at the writing time, and fixed confirmation period start and end time and whether it is terminated in advance due to explicit operation, are written into the interaction log one by one.

[0090] S4, based on the interaction log, the observed semantic zone sequence is reconstructed and checked for consistency with the ordered semantic zone sequence.

[0091] During the on-screen indication session, the records with operation category marks of confirmed, advanced, and back in the interaction log are read in chronological order, the semantic zone identification is taken out one by one, and the observed semantic zone sequence is spliced. The candidate target region actually used for each semantic zone of the observed semantic zone sequence is saved.

[0092] Confirmed is the moment when the fixed confirmation period ends and the line-of-sight confirmation is passed or the user clicks to confirm any candidate target region.

[0093] Advance is the moment when the AI digital person starts to announce the name of the next semantic zone, the on-screen arrow advances to the start of the next semantic zone, and the newly ranked first candidate target region is highlighted.

[0094] Back is the moment when the user clicks the back control, the current semantic zone is aborted and the previous semantic zone is resumed.

[0095] Observed semantic zone sequence is the actual confirmed path of the current on-screen instruction session.

[0096] Align the ordered semantic zone sequence and the observed semantic zone sequence with the minimum number of steps, construct a table, the number of rows is one more than the length of the ordered semantic zone sequence, the number of columns is one more than the length of the observed semantic zone sequence, fill in each cell from the top left corner to the bottom right corner, each cell represents the minimum number of steps required to align the first n semantic zones of the ordered semantic zone sequence with the first n semantic zones of the observed semantic zone sequence, the first row and the first column start from zero and increase by one, for any cell in the table, take the cumulative number of steps of the three arrival methods and select the minimum to write, align the previous semantic zone of the ordered semantic zone sequence with the current prefix of the observed semantic zone sequence, add one to the cumulative number of steps in the cell above, align the current prefix of the ordered semantic zone sequence with the previous semantic zone of the observed semantic zone sequence, add one to the cumulative number of steps in the cell to the left, if the two semantic zone identifiers are the same, the cumulative number of steps is carried down from the top left cell, if the two semantic zone identifiers are different, add one to the cumulative number of steps based on the top left cell.

[0097] Current prefix is the subsequence from the beginning of the semantic zone sequence to the position represented by the current cell.

[0098] The value of the cell in the bottom right corner is the minimum number of difference steps between the ordered semantic zone sequence and the observed semantic zone sequence, trace back from the bottom right corner to the top left corner, the first insertion, deletion, or replacement on the trace path is the earliest misaligned semantic zone, if the bottom right corner is zero, it means that the ordered semantic zone sequence and the observed semantic zone sequence are completely consistent.

[0099] On the venue schematic diagram, use two different line types to depict the earliest misaligned semantic zone of the ordered semantic zone sequence and the earliest misaligned semantic zone of the observed semantic zone sequence, provide a key correction button at the start of the earliest misaligned semantic zone of the ordered semantic zone sequence, click to rebind the on-screen instruction and the candidate target region, provide a re-orientation button at the start of the earliest misaligned semantic zone of the observed semantic zone sequence, take the earliest misaligned semantic zone of the observed semantic zone sequence as the current semantic zone, and re-advance.

[0100] In the semantic zone identification and direction description related to the earliest misalignment semantic zone, a constraint is recorded. The candidate target area actually used this time should be placed before the first candidate target area in this round of sorting.

[0101] Next time the same semantic zone identification and direction description enter the dynamic sorting, first check if there is a corresponding constraint. If there is, directly put the recorded candidate target area at the forefront of the candidate target area set, and the remaining candidate target areas remain in the original order. Then, dynamic sorting is performed.

[0102] If multiple constraints conflict, the most recent one takes effect. If there are still situations that cannot be satisfied at the same time, the constraint with the latest time is prioritized. The remaining candidate target areas that are not satisfied remain in the original relative order.

[0103] When the ordered semantic zone sequence and the observed semantic zone sequence are consistent, or consistent again after one-key correction and re-direction, the screen will display the semantic zone with a faded track, highlight the next semantic zone, and push the arrow to the start point of the next semantic zone. The anchor point of the AI digital human dialogue is updated to the new first candidate target area in the order. Write in the interaction log that the check is complete. If the next semantic zone exists, continue the cycle. If the current semantic zone is the semantic zone where the target position is located, display the arrival prompt and end the on-screen instruction session.

[0104] The embodiment also provides a guide machine interaction system based on an AI digital human, which comprises:

[0105] A collection and judgment module collects screen input events and normalizes them into an event stream, establishes a reference line, and dynamically judges the user's physical location and line-of-sight area under environmental perception.

[0106] A semantic zone module divides the start point to the target position into semantic zones and arranges them into an ordered semantic zone sequence based on the reference line, the user's physical location and line-of-sight area. It binds on-screen instructions and a number of candidate target areas on the screen for each semantic zone, dynamically sorts the candidate target areas, and establishes an interaction log.

[0107] An AI digital human broadcasting module broadcasts the AI digital human according to the ordered semantic zone sequence. When the semantic zone name is broadcast, highlight the on-screen instructions and the candidate target areas on the screen, and judge whether the user's line of sight is in the first candidate target area in the sorting period.

[0108] A reconstruction and check module reconstructs the observed semantic zone sequence based on the interaction log and checks the consistency with the ordered semantic zone sequence.

[0109] To sum up, the application realizes semantic modeling and structured expression of the path by dividing semantic zones and arranging them into an ordered semantic zone sequence, so that the AI digital person can conduct accurate broadcasting and interactive control, thereby maintaining the continuity and logical consistency of path expression in a complex space, realizing real-time misalignment recognition and automatic correction in the navigation process through consistency checking, and ensuring the stability and intelligent adaptability of the route guidance process through self-correction and self-learning of the AI digital person, thereby improving the accuracy and continuous guidance ability of human-computer interaction.

[0110] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application, and they should be covered in the scope of the claims of the present application.

Claims

1. A wayfinding interaction method based on AI digital human, characterized in that: include, Collect on-screen input events and normalize them into an event stream, establish reference lines, and dynamically determine the user's physical location and line of sight under environmental awareness; Based on reference lines, user physical location, and line of sight area, the starting point to the target location is divided into semantic zones and arranged into an ordered semantic zone sequence. Each semantic zone is bound to an on-screen indicator and several candidate target areas on the screen. The candidate target areas are dynamically sorted and an interaction log is established. The AI ​​digital human broadcasts according to an ordered sequence of semantic zones. When the name of a semantic zone is broadcast, the indicator and candidate target area on the screen are highlighted. The AI ​​digital human also determines whether the user’s gaze is fixed on the candidate target area at the top of the list within a fixed confirmation period. The observed semantic zone sequence is reconstructed based on the interactive logs and then checked for consistency with the ordered semantic zone sequence.

2. The wayfinding interaction method based on AI digital human as described in claim 1, characterized in that: The specific steps for collecting and normalizing input events from the screen into an event stream and establishing reference lines are as follows: The navigation device uniformly records the atomic inputs generated in the screen side area and the screen front side area. Each atomic input is written as an event record, and the event records are strictly ordered by time to form an event stream. By using the pair of points formed by the calibration points laid on the ground and the pixels on the guide camera screen, the geometric mapping relationship between the screen plane and the ground reference plane is obtained. The upper and lower pixels of the vertical center line of the screen are taken and converted into actual points on the ground reference plane. The line connecting the actual points is used as the reference line, and the direction of the line connecting the actual points is used as the direction of the reference line.

3. The wayfinding interaction method based on AI digital human as described in claim 2, characterized in that: The specific steps for dynamically determining the user's physical location and line-of-sight area under environmental perception are as follows: The system obtains the user's physical location using the depth imaging camera of the wayfinding device, acquires the user's head orientation based on the user's head posture, and converts the user's head orientation to a ground reference coordinate system consistent with the reference line. Compare the angle between the user's head orientation and the reference line direction. When the angle does not exceed the forward angle threshold, the user's gaze is considered to be facing forward of the reference line. When the angle falls into the left allowable range, the user's gaze is considered to be pointing to the left. When it falls into the right allowable range, the user's gaze is considered to be pointing to the right. When the angle exceeds both the left and right allowable ranges and no interaction occurs within a fixed time, the interaction is considered to have ended. The system measures the nearest lateral distance from the user's physical location to the reference line. When the nearest lateral distance does not exceed the forward bandwidth threshold, the user is considered to be within the effective interaction area of ​​the navigation system. When the nearest lateral distance exceeds the forward bandwidth threshold and no interaction is performed within a fixed time, the user is considered to have deviated from the effective interaction area of ​​the navigation system. When the included angle does not exceed the forward angle threshold and the lateral nearest distance does not exceed the forward bandwidth threshold, the current line of sight area is classified as the reference line's forward dominant area. When the included angle is within the left allowable range and the lateral nearest distance does not exceed the forward bandwidth threshold, the current line of sight area is classified as the left adjacent area. When the included angle is within the right allowable range and the lateral nearest distance does not exceed the forward bandwidth threshold, the current line of sight area is classified as the right adjacent area.

4. The wayfinding interaction method based on AI digital human as described in claim 3, characterized in that: The specific steps for dividing the area from the starting point to the target location into semantic zones and arranging them into an ordered semantic zone sequence are as follows: Based on the user's physical location, the user's current location is determined as the starting point. The preset path is read from the site diagram built into the navigation device. The reference line direction is used as the benchmark. The topology nodes are identified sequentially along the preset path, and the continuous road segments between adjacent topology nodes are used as semantic zones. Starting from the semantic zone of the starting point and ending at the semantic zone of the target location, an ordered sequence of semantic zones is obtained.

5. The wayfinding interaction method based on AI digital human as described in claim 4, characterized in that: The step involves binding an on-screen indicator and several candidate target regions to each semantic zone, and dynamically sorting the candidate target regions. The specific steps are as follows: Record directional descriptions in each semantic zone and align arrows, annotation lines, and AI digital human speech anchors with the directional descriptions of the semantic zones. The target regions located within the geometric range of the current semantic zone, as well as the target regions in the next semantic zone that intersect or are tangent to the boundary of the current semantic zone, are merged into a set of candidate target regions for the semantic zone. Whenever a valid record is added to the event stream, the set of candidate target regions for the current semantic zone and the next adjacent semantic zone is rearranged in the following order: line-of-sight consistency priority, zone consistency priority, event proximity priority, proximity along the reference line priority, and turning complexity priority.

6. The wayfinding interaction method based on AI digital human as described in claim 5, characterized in that: When the semantic zone name is broadcast, the indicator and candidate target areas on the screen are highlighted. The specific steps are as follows: When the broadcast enters the semantic zone name field, the arrow describing the direction of the current semantic zone is immediately highlighted on the screen, and the first candidate target area icon in the sorting is also highlighted.

7. The wayfinding interaction method based on AI digital human as described in claim 6, characterized in that: The specific steps for determining whether the user's gaze remains on the top-ranked candidate target area within a fixed confirmation period are as follows: During the fixed confirmation period, the user's gaze points are sampled at equal intervals and assigned to the candidate target areas marked on the screen. The cumulative dwell time and continuity of each candidate target area covered by the gaze points are continuously counted. If the first-ranked candidate target area maintains the first cumulative dwell time and is not interrupted during the fixed confirmation period, it is considered that the gaze confirmation is passed; otherwise, it is considered that the gaze confirmation is not passed. If the visual confirmation fails, a confirmation control will appear around the first candidate target area in the sorting, and a back control will appear on the side of the screen. If the line of sight is confirmed to be clear, the current semantic zone is marked as confirmed. The AI ​​digital human continues to announce the name and action statements of the next semantic zone, repeating until the target location is reached.

8. The wayfinding interaction method based on AI digital human as described in claim 7, characterized in that: The specific steps for reconstructing the observed semantic zone sequence based on interactive logs are as follows: By reading the operation category records in the interaction log, semantic zone identifiers are extracted and concatenated into an observation semantic zone sequence in chronological order.

9. The wayfinding interaction method based on AI digital human as described in claim 8, characterized in that: The specific steps for performing consistency verification with the ordered semantic zone sequence are as follows: By aligning the observed semantic zone sequence with the ordered semantic zone sequence using a table with the minimum number of steps in dynamic programming, the system calculates the minimum number of difference steps and locates the earliest misaligned semantic zone, and provides one-click correction and re-guidance on the site map.

10. A wayfinding interaction system based on AI digital human, based on the wayfinding interaction method based on AI digital human as described in any one of claims 1 to 9, characterized in that: include, The data acquisition and judgment module collects input events in front of the screen and normalizes them into an event stream, establishes reference lines, and dynamically judges the user's physical position and line of sight area under environmental awareness. The semantic zone module divides the area from the starting point to the target location into semantic zones based on reference lines, user physical location, and line of sight. These semantic zones are arranged into an ordered sequence. Each semantic zone is bound to an on-screen indicator and several candidate target areas on the screen. The candidate target areas are dynamically sorted, and an interaction log is established. The AI ​​digital human broadcasting module broadcasts information according to an ordered semantic zone sequence. When the semantic zone name is broadcast, the indicator and candidate target area on the screen are highlighted. Within a fixed confirmation period, it is determined whether the user's gaze is fixed on the candidate target area ranked first. The reconstruction and verification module reconstructs the observed semantic zone sequence based on the interaction log and performs a consistency check with the ordered semantic zone sequence.

Citation Information

Patent Citations

  • Robot guiding method and device based on semantic map

    CN116952250A

  • Robot navigation method and system based on visual identification

    CN120760734A

  • Multi-modal data search method and system based on AI large model

    CN120804368A

  • Indoor semantic mapping and navigation method and system based on multi-modal model

    CN120890460A

  • Systems and methods for detecting, analyzing, and evaluating interaction paths

    US20200089592A1