Methods of using a wayfinding robot and wayfinding robots
By collecting feature data of candidate users and converting it into confidence scores to filter target users, and combining voice and map data to generate recommended routes, the wayfinding robot can proactively identify and dynamically adjust its guidance strategy, solving the problem that existing wayfinding robots cannot proactively identify and dynamically optimize, and realizing intelligent navigation that accompanies the entire process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DIGITAL HUAXIA (SHENZHEN) TECHNOLOGY CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-26
AI Technical Summary
Existing wayfinding robots lack the ability to proactively identify potential users, leading to incorrect guidance and a lack of dynamic optimization, thus failing to achieve intelligent navigation that accompanies users throughout the journey.
Feature data of candidate users is collected by camera, radar and microphone components, converted into feature confidence scores, and target users are filtered based on confidence score thresholds. Recommended routes are generated by combining voice data and map data to determine guidance strategies of leading, accompanying or combined modes, and guidance is executed using mobile components.
It enables proactive identification of potential seekers, improves the adaptability of guidance and user experience, and ensures efficient guidance in open areas and safe accompaniment in congested sections.
Smart Images

Figure CN121612322B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and in particular to a method for a robot to guide users and a robot for asking for directions. Background Technology
[0002] In large transportation hubs such as high-speed rail stations and subway stations, passengers mainly rely on static signs, manual service counters, or fixed electronic terminals for directions. These methods suffer from problems such as inflexible information, limited service coverage, and inconvenient interaction, and are particularly difficult to handle temporary route changes or personalized navigation needs. Service robots that have emerged in recent years mostly remain at the "question and answer" level, requiring users to actively seek help, lacking the ability to proactively identify potential users, and their service model is passive and has low coverage.
[0003] Current robots primarily rely on voice or screen displays for route guidance, which is not intuitive and can easily lead travelers astray. Route planning often ignores real-time congestion and facility status, lacking dynamic optimization capabilities. Furthermore, robots cannot adaptively select leading or accompanying modes based on the environment and lack continuous tracking of travelers' progress, failing to achieve closed-loop guidance from departure to arrival. Therefore, there is an urgent need for a more intelligent, proactive, and continuous route guidance method. Summary of the Invention
[0004] In view of this, embodiments of this application provide a wayfinding robot guidance method and a wayfinding robot, which can effectively solve the technical problem that traditional robots often lead to guidance errors due to their lack of proactive identification of potential users.
[0005] In a first aspect, embodiments of this application provide a wayfinding robot guidance method, including:
[0006] The feature data of each candidate user in the target location is converted into feature confidence scores, and the target user is selected from each candidate user based on the feature confidence scores and confidence thresholds.
[0007] Based on the voice data of the target user and the map data of the target location, a recommended route within the target location is generated;
[0008] Based on the recommended path and the location of each candidate user, the guidance mode of the wayfinding robot is determined, including the leading mode, the accompanying mode, and the combined mode.
[0009] The mobile component is triggered to execute the guidance mode, guiding the target user to the destination location of the recommended path.
[0010] Secondly, embodiments of this application provide a wayfinding robot, comprising:
[0011] The camera component is used to collect handheld feature data and confused expression feature data of each candidate user in the target location;
[0012] A radar component is used to collect the stationing characteristic data of each of the candidate users and the map data of the target location;
[0013] Microphone assembly, used to collect voice data from the target user;
[0014] A control component is configured to convert feature data of each candidate user within the target location into feature confidence scores, and based on the feature confidence scores and confidence thresholds, filter out the target user from among the candidate users; generate a recommended path within the target location based on the voice data of the target user and the map data of the target location; and determine the guidance mode of the directions-asking robot based on the recommended path and the location of each candidate user, wherein the guidance mode includes a leading mode, an accompanying mode, and a combination mode.
[0015] A mobile component is used to execute the guidance mode, guiding the target user to the destination location of the recommended path.
[0016] Thirdly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0017] The feature data of each candidate user in the target location is converted into feature confidence scores, and the target user is selected from each candidate user based on the feature confidence scores and confidence thresholds.
[0018] Based on the voice data of the target user and the map data of the target location, a recommended route within the target location is generated;
[0019] Based on the recommended path and the location of each candidate user, the guidance mode of the wayfinding robot is determined, including the leading mode, the accompanying mode, and the combined mode.
[0020] The mobile component is triggered to execute the guidance mode, guiding the target user to the destination location of the recommended path.
[0021] The embodiments of this application have the following beneficial effects:
[0022] First, by converting the multimodal feature data of candidate users into quantitative feature confidence scores and combining them with preset thresholds to automatically filter target users, the system achieves proactive identification of potential help seekers, breaking through the limitation of traditional robots that can only respond passively.
[0023] Secondly, the robot intelligently determines the guidance mode (leading, accompanying, or combined mode) based on the recommended route and the distribution of people in the surrounding area, enabling the robot to guide efficiently in open areas and accompany safely in congested sections, thus enhancing the adaptability of the guidance process and the user experience. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This paper shows a schematic diagram of the structure of a wayfinding robot according to an embodiment of this application;
[0026] Figure 2 A flowchart illustrating a wayfinding robot guidance method according to an embodiment of this application is shown.
[0027] Component symbol explanation:
[0028] 11-Radar component, 12-Camera component, 13-Microphone component, 14-Control component, 15-Motion component. Detailed Implementation
[0029] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0030] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0031] In the following text, the terms "comprising," "having," and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more combinations thereof. Furthermore, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0032] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0033] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0034] The following describes the wayfinding robot using some specific examples.
[0035] Figure 1 A schematic diagram of a wayfinding robot according to an embodiment of this application is shown. Exemplarily, the wayfinding robot includes:
[0036] Camera component 12 is used to collect handheld feature data and confused expression feature data of each candidate user in the target location;
[0037] Radar component 11 is used to collect the stationing characteristic data of each candidate user and the map data of the target location;
[0038] Microphone assembly 13 is used to collect voice data of the target user;
[0039] The control component 14 is used to convert the feature data of each candidate user in the target location into feature confidence scores, and to filter out the target user from among the candidate users based on the feature confidence scores and confidence score thresholds; to generate a recommended path in the target location based on the voice data of the target user and the map data of the target location; to determine the guidance mode of the wayfinding robot based on the recommended path and the location of each candidate user, the guidance mode including the leading mode, the accompanying mode and the combined mode; to convert the guidance mode into guidance instructions and send the guidance instructions to the mobile component 15.
[0040] Mobile component 15 is used to execute the guidance mode, guiding the target user to the destination location of the recommended path.
[0041] The camera component 12 refers to a sensing device used to collect visual information of candidate users in the target location, including at least one multi-view depth camera or a visual sensor with RGB imaging capability. The camera component 12 is configured to acquire RGB images and depth images of candidate users to support the recognition and analysis of facial key points, expression features, the outline of the handheld object and its three-dimensional spatial distribution. The RGB image is used for instance segmentation to extract the two-dimensional outline of the handheld object, and the depth image is used to acquire the three-dimensional point cloud data corresponding to the outline, thereby calculating the spatial volume of the handheld object, and combining it with a continuous video frame sequence to realize the dynamic detection of facial movements (such as frowning, squinting, and drooping corners of the mouth) and the recognition of confused emotions.
[0042] Radar component 11 refers to a ranging sensing device used to collect environmental spatial information and pedestrian movement status, including lidar or millimeter-wave radar. Radar component 11 is configured to scan and construct a local point cloud map of the target location in real time, and acquire the location, movement trajectory, speed, and dwelling behavior data of candidate users. In particular, based on the point cloud data, the system can determine whether the candidate user is in a low-speed wandering state within a preset time period, and analyze its dwelling characteristics by combining the radius of the smallest geometric shape (such as the smallest circumscribed circle), thereby providing a confidence basis for identifying potential help seekers. At the same time, radar component 11 is also used to assist in the generation and updating of high-precision site maps, supporting the robot's autonomous navigation and obstacle avoidance.
[0043] The microphone assembly 13 refers to an array-type audio acquisition device composed of multiple microphone units, deployed on the robot body for directional pickup of sound signals from a specific direction in the environment. The microphone assembly 13 is configured to capture the voice commands of the target user and enhance the effective voice signal-to-noise ratio through beamforming and noise suppression technology, thereby improving the accuracy of voice recognition in high background noise environments such as high-speed rail stations and subway stations. In addition, the microphone assembly 13 can also assist in sensing whether the user utters a questioning statement (such as "How do I get here?"), serving as an auxiliary input feature in multimodal fusion judgment.
[0044] The control component 14 refers to the central processing and decision-making unit integrated within the wayfinding robot, including at least one processor and a non-volatile memory storing computer program instructions. When the program instructions are executed by the processor, the control component 14 performs the wayfinding robot guidance method.
[0045] The mobile component 15 refers to the autonomous mobile actuator installed on the bottom of the wheeled robot, including a drive motor, wheel system, inertial measurement unit, and motion controller. The mobile component 15 is configured to receive guidance commands from the control component 14 and autonomously travel on a recommended path according to a set guidance mode (e.g., leading 3–5 meters ahead or accompanying 1–2 meters to the side). The mobile component 15 has dynamic obstacle avoidance, path tracking, and speed adjustment capabilities. During the guidance process, it continuously monitors the relative positional relationship with the target user and can respond when the user deviates or pauses (e.g., wait, remind, or replan), ensuring the safe and stable completion of the escort mission to the destination.
[0046] A wayfinding robot is a wheeled service robot with autonomous movement and multimodal perception capabilities, deployed in large public places such as high-speed rail stations and subway stations to provide intelligent wayfinding guidance services to passengers. The wayfinding robot integrates a camera component 12, a radar component 11, a microphone component 13, a control component 14, and a motion component 15. It can proactively identify passengers who may need assistance through environmental perception, understand their destination needs, plan dynamically recommended routes, and provide physical guidance in lead, accompany, or combined modes until the destination is confirmed, achieving a closed-loop guidance process from "passive response" to "proactive escort."
[0047] Figure 2 A schematic flowchart of a wayfinding robot guidance method according to an embodiment of this application is shown. Exemplarily, the wayfinding robot guidance method includes the following steps:
[0048] Step S202: Convert the feature data of each candidate user in the target location into feature confidence scores, and select the target user from each candidate user based on the feature confidence scores and confidence score thresholds.
[0049] The feature data includes standing feature data, hand-held feature data, and confused expression feature data; the feature confidence includes standing feature confidence, hand-held feature confidence, and confused expression feature confidence.
[0050] In one embodiment, the confidence scores of the standing feature are obtained based on the standing feature data, the confidence scores of the handheld feature are obtained based on the handheld feature data, and the confidence scores of the confused expression feature are obtained based on the confused expression feature data. Based on preset movement weight coefficients, posture weight coefficients, and expression weight coefficients, the confidence scores of the standing feature, handheld feature, and confused expression feature are weighted and summed to obtain the fused confidence score. Target users with a fused confidence score greater than a confidence score threshold are determined from among the candidate users.
[0051] Among them, the lingering characteristic data refers to the data collected by the radar component 11 and the visual sensor that characterizes the behavior of the candidate user in the target location, including its movement trajectory, displacement distance, speed and spatial distribution range within a preset time period, which is used to determine whether the user is in a state of wandering, stagnation or looking around.
[0052] The confidence level of the dwell feature refers to a quantitative indicator calculated based on dwell feature data, reflecting the degree of probability that a candidate user will "dwell for a long time". Specifically, it is evaluated by constructing the minimum geometric figure of its movement path (such as the minimum circumscribed circle) and combining the radius of the figure with the duration. The higher the value, the more significant the dwell behavior.
[0053] Handheld feature data refers to multimodal image data related to the object carried by the candidate user's hand, acquired by the camera component 12. It includes two-dimensional contour information of the handheld object in the RGB image and the corresponding three-dimensional point cloud data in the depth image, which is used to identify whether there are large luggage, backpacks or other burdens.
[0054] The confidence score of handheld features refers to the quantitative value calculated based on the spatial volume of the handheld object, reflecting the possibility that the candidate user may need help due to the items they are carrying. It is obtained by denoising the point cloud of the handheld region, fitting the minimum bounding cube, and calculating the product of its length, width, and height. The larger the volume, the higher the confidence score.
[0055] Confused facial expression feature data refers to the facial dynamic information of candidate users collected through a continuous video frame sequence, including the geometric deformation features of key areas such as eyebrows, eyes, and corners of the mouth in RGB and depth images, which is used to identify micro-expression movements that express confusion, such as frowning, squinting, and drooping corners of the mouth.
[0056] The confidence score of the confused expression feature refers to the probability value output by a pre-trained emotion classification model after time series modeling based on the frequency, duration and intensity of facial movements. It represents the credibility of a candidate user being in a "confused" emotional state and the value range is usually [0, 1].
[0057] The movement weight coefficient refers to the preset weighting coefficient assigned to the "standing feature confidence" during the fusion of multi-feature confidence. It reflects the relative importance of abnormal walking behavior in the overall judgment and is determined by historical data analysis or machine learning training. It is used to calculate the fusion confidence by weighted summation.
[0058] The posture weight coefficient refers to the preset weighting coefficient assigned to the "handheld feature confidence" during the process of fusing multiple feature confidences. It reflects the weight of the impact of the inconvenience of movement caused by physical burden (such as luggage or children) on the need for help. It is usually less than or equal to the movement weight coefficient.
[0059] The facial expression weighting coefficient refers to the preset weighting coefficient assigned to the "confidence of confused facial expression features" during the process of fusing multiple feature confidence scores. It reflects the contribution of facial emotional cues in identifying potential help seekers and can be dynamically adjusted according to ambient lighting and occlusion.
[0060] Fusion confidence score is a comprehensive score obtained by weighting and summing the confidence scores of the three features of standing, holding, and expression according to their respective weights. It is used to comprehensively assess whether a candidate user needs help. The formula is: Fusion confidence score = movement weight coefficient × standing confidence score + posture weight coefficient × holding confidence score + expression weight coefficient × expression confidence score.
[0061] The confidence threshold is a preset judgment threshold used to filter target users from all candidate users. When the fusion confidence of a candidate user is greater than or equal to the threshold, it is identified as a target user who needs to be actively served, triggering the subsequent interaction process.
[0062] The target user refers to the passenger whose fusion confidence level exceeds the confidence threshold after multimodal feature analysis among multiple candidate users, and whose inquiry and guidance process has been initiated by the robot. In other words, the target is currently receiving directions and guidance services.
[0063] Candidate users refer to pedestrians who are within the perception range of the wayfinding robot, are continuously monitored, but have not yet been identified as needing services. Their behavioral, posture, facial expression and other characteristic data are being collected and analyzed in real time as a set of potential help seekers to be evaluated.
[0064] In one example, the calculation method for the confidence of the standby feature includes:
[0065] Obtain the target user's movement path and map data of the target location within a preset time period; construct the minimum geometric shape containing the movement path based on the map data; determine the confidence level of the dwelling feature based on the radius of the minimum geometric shape and the preset time period.
[0066] The preset duration refers to the time threshold used to determine whether a user is in a "long-term dwell" state, which is usually 10 to 30 seconds, preferably 20 seconds; during this time period, the mobile behavior data of candidate users are continuously collected as a time benchmark for identifying their wandering or confused state.
[0067] The movement path refers to the spatial movement trajectory of the candidate user within a preset time period, obtained by continuous tracking through radar component 11 and visual odometry. It consists of a series of discrete location coordinate points, reflecting the user's walking route and activity range within the station hall.
[0068] The target location refers to large indoor transportation hubs where the wayfinding robot is deployed and operated, including public areas with dense crowds, complex structures, and multiple path guidance needs, such as high-speed rail stations, subway stations, and airport terminals.
[0069] Map data refers to high-precision digital maps acquired by the navigation robot through radar component 11 or network connection. It includes fixed structural information (such as passages, turnstiles, elevators, and shops) and dynamic information (such as temporary closed areas and pedestrian density) within the target location, and is used for positioning, path planning, and environmental understanding.
[0070] The minimum geometric shape refers to the smallest closed spatial contour that covers all movement path points of the candidate user within a preset time period. It is usually the smallest circumcircle or the smallest convex hull. Its radius or area is used to quantify the user's activity range and help determine whether the user is wandering in place.
[0071] Specifically, after entering the subway station, the target user paced back and forth near the ticket gates, seemingly searching for the correct passage. The directions robot then began monitoring the passenger's behavior.
[0072] First, the navigation robot continuously collects the target user's pose information over a preset time period, such as 20 seconds, using radar component 11, forming a movement path composed of multiple coordinate points. It finds that the user never leaves a circular area with a diameter of approximately 3 meters. Next, based on map data of the target location, the navigation robot performs spatial analysis on all points along the movement path and constructs the smallest geometric figure containing the path, i.e., the smallest circumcircle. The radius of this small circumcircle is calculated to be, for example, 1.4 meters.
[0073] Finally, the confidence score for the lingering feature is calculated using the formula: Lingering Feature Confidence Score = f(Total Distance Moved, Path Complexity, Minimum Circumscribed Circle Radius, Duration). For example, if the target user moves less than 1.5 meters in 20 seconds, has a speed less than 0.2 meters per second, a small minimum circumscribed circle radius, and a circular path (high complexity), it is determined that the user exhibits significant lingering behavior. Therefore, the output lingering feature confidence score is 0.9 (out of 1.0), serving as an important input for subsequent multi-feature fusion.
[0074] The above embodiments quantify user pausing behavior by introducing a "minimum geometric shape" combined with a "preset duration," transforming vague human observation experience into calculable and reproducible technical indicators, effectively improving the accuracy of identifying potential seekers. Compared to methods that rely solely on a single speed or dwell time, this application is better able to distinguish between "short pauses" and "genuine confusion," avoiding false triggers.
[0075] In one example, the confidence level of handheld features is calculated as follows:
[0076] Acquire depth and RGB images of the target user; segment the RGB images to identify the outline of the object held by the target user; based on the depth image, acquire 3D point cloud data of the object outline; perform clustering processing on the 3D point cloud data of the object outline to remove outliers; fit a minimum bounding box in the 3D point cloud data after removing outliers; obtain the length, width, and height of the minimum bounding box; determine the spatial volume of the object held by the user based on the length, width, and height, and determine the confidence level of the handheld features based on the spatial volume.
[0077] Depth images refer to image data collected by camera devices with depth perception capabilities (such as multi-view depth cameras). Each pixel contains distance information from the robot's perspective to the surface of the corresponding object in the scene, which is used to construct the three-dimensional spatial distribution of the target object.
[0078] RGB images refer to standard color images, which consist of three color channels: red, green, and blue. They record the visual appearance information of candidate users and their surrounding environment and are used for object detection, semantic segmentation, and contour extraction.
[0079] The outline of a handheld object refers to the boundary of an independent object located near the user's hand area in an RGB image, identified by an image segmentation algorithm. It represents the two-dimensional shape range of items such as suitcases and backpacks held in the hand, serving as the basis for subsequent three-dimensional volume reconstruction.
[0080] 3D point cloud data refers to a set of three-dimensional spatial points obtained by converting a depth image. Each point has (x, y, z) coordinates and accurately describes the spatial shape and size distribution of the outline of a handheld object in the real world.
[0081] Outliers are abnormal noise points in 3D point cloud data that deviate from the main structure. They are usually caused by sensor errors, background interference, or uneven reflection. They need to be removed by cluster analysis before volume calculation to improve measurement accuracy.
[0082] The minimum bounding cuboid refers to the smallest axis-aligned or oriented cuboid that encloses the denoised 3D point cloud data. Its side lengths correspond to the maximum extension dimensions of the object in the length, width, and height directions, respectively, and are used to quickly estimate the overall space occupation of the object.
[0083] Spatial volume refers to the geometric volume value (V = length × width × height) calculated based on the length, width, and height of the smallest circumscribed cuboid, representing the actual size of the handheld object; this value is used to determine whether it exceeds a preset threshold (such as 50cm × 30cm × 20cm), thereby determining whether the user has a burden of weight.
[0084] Specifically, the target user enters the high-speed rail station hall carrying a large carry-on suitcase and stops near the ticket gate to check the signs. The wayfinding robot deployed in the hall initiates its perception process:
[0085] The navigation robot simultaneously acquires the RGB and depth images of the target user through the integrated camera component 12; it then processes the RGB images using an instance segmentation model to identify objects within the handheld area and extract the outline of the handheld object.
[0086] The outline of the handheld object is mapped onto a depth image, and all depth values of the corresponding region are extracted to generate the 3D point cloud data of the suitcase. Optionally, a clustering algorithm is used to analyze the 3D point cloud data to automatically identify and remove floating outliers.
[0087] A minimum bounding cuboid is fitted to the denoised point cloud, with dimensions measured as follows: length 60cm, width 40cm, height 25cm. The spatial volume V = 60×40×25 = 60000 cm³ is calculated. Since this spatial volume significantly exceeds a preset burden threshold (e.g., 30000 cm³), based on the proportion of the spatial volume exceeding the preset burden threshold, a handheld feature confidence score is output, for example, 0.82, indicating that the user is highly likely to need assistance due to carrying large luggage. Understandably, this handheld feature confidence score is then used in the fusion judgment, combined with the user's pausing and facial expression behavior, to ultimately determine whether to identify the user as the target user and proactively initiate an inquiry.
[0088] Through the above embodiments, by fusing RGB images and depth images, accurate estimation of the spatial volume of handheld objects is achieved, overcoming the shortcomings of traditional methods that rely solely on appearance and are prone to misjudgment. The introduction of point cloud denoising and minimum bounding box fitting mechanisms improves the robustness and practicality of volume calculation.
[0089] In one example, the confidence level of confused facial features is calculated as follows:
[0090] Acquire a continuous video frame sequence of the target user, including RGB images and depth images; perform facial keypoint detection on each RGB and depth image frame, and extract geometric deformation features of the eyebrow, eye, and mouth regions; based on the geometric deformation features, determine at least one facial action, which includes at least one or more combinations of frowning, squinting, and drooping corners of the mouth; perform temporal modeling on the frequency, duration, and intensity changes of facial actions in the video frame sequence to generate a facial dynamic feature sequence; input the facial dynamic feature sequence into a pre-trained emotion classification model, and output the confidence score of confused expression features representing the target user's confused emotion.
[0091] The video frame sequence refers to a set of image frames continuously acquired by the camera component 12 and arranged in chronological order. It contains dynamic facial images of the target user over a period of time. Each frame includes an RGB image and a corresponding depth image, which are used to capture the process of facial expression changes.
[0092] Geometric deformation features refer to the spatial positional changes extracted after detecting key facial points (such as eyebrows, corners of the eyes, and corners of the mouth), which characterize the stretching, compression, or tilting of the facial area. For example, a reduced distance between the eyebrows indicates frowning, and a reduced eyelid opening indicates squinting.
[0093] Facial movements refer to specific facial muscle movements identified based on geometric deformation features, including at least one or more combinations of frowning, squinting, and downturned corners of the mouth, used to determine emotional state.
[0094] Facial dynamic feature sequence refers to a structured data sequence generated by temporal modeling of the frequency, duration and intensity changes of facial movements detected at each moment in a video frame sequence, reflecting the temporal evolution of the user's emotional expression within the observation period.
[0095] An emotion classification model is a pre-trained machine learning model that takes a sequence of facial dynamic features as input and outputs the probability value of a user being in a specific emotional state (such as confusion, anxiety, or calmness). This emotion classification model is trained on a large number of labeled facial expression datasets and has high generalization ability.
[0096] Specifically, the target user stops in the subway transfer passage, looking hesitant. The directions robot then initiates its emotion recognition process:
[0097] The navigation robot invokes camera component 12 to continuously capture a 5-second video frame sequence (e.g., 75 frames, 30fps) of the target user, with each frame containing an RGB image and a depth image. A deep learning-based facial landmark detection algorithm is used to process each frame, accurately locating multiple key points such as the eyebrow arch, inner and outer corners of the eyes, upper and lower eyelid edges, and corners of the mouth.
[0098] The vertical distance between two adjacent eyebrows is calculated. If it shrinks significantly in multiple frames, it is identified as "frowning". The distance ratio between the upper and lower eyelids is analyzed. If the average opening is more than 20% lower than the baseline value, it is identified as "squinting". The vertical offset of the corners of the mouth relative to the midline of the nose and lips is detected. If the corners of the mouth are obviously bent downwards on both sides, it is identified as "drooping corners of the mouth".
[0099] The cumulative frequency of the above three facial movements within 5 seconds is as follows: for example, frowning 3 times (total duration 1.8 seconds), squinting 4 times (total duration 2.1 seconds), and drooping corners of the mouth 2 times (total duration 1.2 seconds), and the intensity shows a fluctuating upward trend;
[0100] This information is organized into a facial dynamic feature sequence, which is then input as a time series into the deployed emotion classification model. The emotion classification model (based on an LSTM network architecture) processes the facial dynamic feature sequence and outputs the confidence score of the confused expression feature of the user in the current "confused" emotional state, which is, for example, 0.87 (close to the full score).
[0101] Through the above embodiments, compared with traditional methods that rely solely on voice or behavior for judgment, the introduction of a facial dynamic feature modeling mechanism based on video frame sequences can capture users' potential intention to seek help earlier and more accurately, which is especially suitable for people with language barriers, foreigners, or groups who are unwilling to speak up.
[0102] Step S204: Based on the voice data of the target user and the map data of the target location, generate a recommended route within the target location.
[0103] Among them, voice data refers to the sound signal emitted by the target user in a natural dialogue state, which is collected by the microphone component 13 and converted into text information after noise reduction and enhancement processing; the voice data is used to express the user's inquiry intention or destination needs (such as "I want to go to the A12 ticket gate"), and is parsed into structured instructions through speech recognition and natural language processing technology, serving as the key input basis for generating personalized recommendation paths.
[0104] The recommended route refers to the optimal guidance route calculated by the control component 14 through a path planning algorithm after a comprehensive evaluation based on the destination determined by the target user's voice data, combined with high-precision map data of the target location, real-time operational status (such as temporary closed areas and elevator operation status), and environmental dynamic information (such as pedestrian density and traffic efficiency).
[0105] Specifically, the system first acquires the target user's destination information and current location, along with map data. This map data includes spatial layout information of fixed facilities, which at least include elevators, escalators, stairs, turnstiles, and accessible pathways. Then, it acquires real-time environmental status data, including: real-time pedestrian density in each area, temporary closure zone markers, operational status of key facilities, estimated queue lengths, and at least one of the following: air quality or noise level.
[0106] Then, based on the destination, current location, high-precision station map, and environmental status data, a weighted directed graph model is constructed. Each traversable node's connection edge is assigned a comprehensive cost, which consists of multiple weighted sub-costs, including:
[0107] Geometric distance cost, which is the base movement time based on the actual walking distance.
[0108] Facility waiting / operation costs introduce estimated waiting time and operation time as additional costs for facilities that require waiting or use (such as elevators and security checkpoints).
[0109] The cost of congestion delays is determined by dynamically adjusting the time penalty coefficient for passing through a specific road segment based on the current pedestrian density.
[0110] The cost of safe accessibility includes setting up "no access" signs or indefinite costs for temporary closures, faulty equipment, or routes that violate access rules.
[0111] The cost of comfort adjustment is determined by taking into account factors such as slope, presence of steps, wheelchair accessibility, noise level, and air quality, and applying differentiated weights to different user types.
[0112] The walking intuition matching cost model is based on the analysis and modeling of human natural walking preferences through historical behavioral data, which closely matches the decision-making habits of real users in terms of intersection selection, landmark utilization and turning frequency.
[0113] Finally, an incremental path planning algorithm is employed to search for the path with the minimum total cost from the starting point to the destination in the weighted directed graph, outputting the initial optimal path. During navigation, environmental change events are continuously monitored. When a critical state update affecting the path cost is detected, the cost is recalculated only for the affected local graph structure, and a priority queue management mechanism is used to trigger local path correction, generating an updated optimal path adapted to the new environment.
[0114] The above embodiments, by integrating high-precision in-station maps, real-time environmental status data and multi-dimensional cost models, construct a weighted directed graph for path planning. This makes the recommended path no longer limited to the shortest geometric distance, but comprehensively considers traffic efficiency, safety, comfort and human walking intuition, significantly improving the practicality of navigation results and user experience.
[0115] Step S206: Based on the recommended path and the location of each candidate user, determine the guidance mode of the wayfinding robot. The guidance modes include the lead mode, the companion mode, and the combination mode.
[0116] The guidance mode refers to the movement and interaction strategies adopted by the wayfinding robot in guiding the target user to the destination. It is dynamically determined by the control component 14 based on the road segment characteristics of the recommended route and the surrounding environment, aiming to achieve a safe, efficient and humanized guidance service. The guidance mode includes the lead mode, the companion mode, and a combination of the two.
[0117] The navigation mode refers to a guidance strategy in which the navigation robot travels 3 to 5 meters in front of the target user as a visual follow target, and periodically reminds the user of the direction of travel through screen arrow prompts, voice broadcasts and light signals; it is suitable for road sections with open passages, moderate traffic flow and good visibility, highlighting guidance efficiency and path visibility.
[0118] The accompanying mode refers to a guidance strategy in which the wayfinding robot stays 1 to 2 meters to the side of the target user, moves in parallel, reduces its speed and increases the frequency of voice or screen interaction (such as reminding users to watch their step); it is suitable for narrow passages, densely populated areas or complex intersections to improve safety and the user's sense of security.
[0119] The combined mode refers to a segmented guidance strategy that divides the entire recommended route into multiple sub-paths and applies either a leading mode or an accompanying mode based on the environmental characteristics of different road segments (such as population density and space width). For example, the leading mode is used for fast passage on main roads, while the accompanying mode is switched at escalator entrances to assist passengers going up and down. Ultimately, this forms an optimal guidance behavior sequence that adapts to changes throughout the entire journey.
[0120] In one embodiment, each turning position on the recommended path is determined; based on each turning position, the recommended path is divided into multiple sub-paths; for each sub-path, the sub-path is matched with each candidate user to determine the number of candidate users on the sub-path; based on the number of candidate users on the sub-path, the guidance pattern for the sub-path is determined; and the combination of guidance patterns for each sub-path is used as the guidance pattern for the navigation robot.
[0121] Among them, the turning position refers to the key node on the recommended path where the direction of travel needs to be changed, including but not limited to intersections, right-angle turns, escalator / elevator entrances and exits, and passageway junctions. At the turning position, users need to make a choice to turn left, right, go straight, or go up or down according to the guidance. It is an important basis for judging the path segmentation and guidance mode switching.
[0122] A sub-path is a continuous road segment unit that divides the entire recommended route into sections based on each turning point. Each sub-path connects two adjacent turning points (or from the starting point to the first turning point, and from the last turning point to the destination). Environmental characteristics (such as space width and pedestrian density) are analyzed independently for each sub-path, and the optimal guidance mode (leading, accompanying, or a combination) is determined accordingly to achieve refined and dynamic guidance.
[0123] Specifically, the target user with luggage has confirmed their destination as "Exit B3," and the navigation robot has completed route planning, generating a recommended route containing multiple key nodes. The following process will then be executed:
[0124] Control component 14 analyzes the recommended path and identifies all key points requiring directional changes, including: the path start point (current location); the first fork in the road A (choose to turn left into the main passage); the escalator up point B; the interchange in the transfer hall C (turn right into the exit passage); point D in front of the turnstile group; and the end point E (exit B3). These key points are marked as turning positions and used as the basis for path segmentation.
[0125] Based on the aforementioned turning positions, the entire route is divided into several continuous sub-paths: Sub-path 1: Starting point ~ A (straight open area); Sub-path 2: A ~ B ~ C (main road + escalator section); Sub-path 3: C ~ D (narrow exit passage); Sub-path 4: D ~ E (short-distance guidance after passing through the turnstiles).
[0126] The wayfinding robot uses radar component 11 and camera component 12 to collect real-time statistics on the distribution of people on each sub-path: Sub-path 1: low traffic, good visibility, no dense crowds; Sub-path 2: a small number of passengers near the escalator, with an average spacing of more than 1.5 meters; Sub-path 3: narrow passage, high passenger flow at the current time, and pedestrian spacing of less than 0.8 meters; Sub-path 4: the area after passing through the turnstile is open, with only occasional pedestrians passing by.
[0127] According to the preset rules, the system makes a judgment based on the pedestrian density and spatial characteristics of each segment: Sub-path ①: If the space is open and the pedestrian flow is sparse, the system adopts the guidance mode. The robot will move steadily forward 4 meters ahead, and the screen will display the direction of the arrow and periodically remind you with the voice "Please follow me".
[0128] Sub-path 2: If it includes escalators and requires both up and down movement, it will switch to the companion mode. The robot will approach the target user 1.2 meters to the side, slowly ascend, and give a voice prompt: "We are going up, please hold the handrail."
[0129] Sub-path 3: If the passage is narrow and the crowd is dense, continue to maintain the escort mode, increase the frequency of interaction, and flash the lights to indicate the passage status to prevent getting lost.
[0130] Sub-path 4: If the destination is near and the environment is relaxed, then switch back to the guide mode, speed up the movement, and guide the way to the final exit.
[0131] The independent modes of each sub-path are chained together to form a combined mode of "leading ~ accompanying ~ accompanying ~ leading", achieving fully adaptive guidance. Understandably, throughout the process, the control component 14 continuously monitors the user's relative position. If it detects that the user has fallen behind or deviated from the route, it pauses progress and issues a recall prompt, ensuring a reliable service loop.
[0132] The above embodiments achieve refined and dynamic decision-making for guidance modes by dividing the recommended path into multiple sub-paths based on turning positions and independently analyzing the environmental conditions and crowd distribution for each segment. Compared to traditional fixed single guidance methods, it automatically switches to companion mode at narrow, crowded, or complex nodes (such as escalators and turnstiles), reducing collision risks and ensuring stable guidance in high-density environments. Furthermore, it avoids energy waste caused by high-frequency interactions throughout the entire process, only activating high-cost behaviors (such as close companionship) in necessary sections, thus extending the robot's endurance.
[0133] Step S208: Convert the guidance mode into a guidance instruction and send the guidance instruction to the mobile component 15 to trigger the mobile component 15 to execute the guidance mode and guide the target user to the destination location of the recommended path.
[0134] The guidance instruction refers to the structured control command generated by the control component 14 according to the current guidance mode, which is used to drive the mobile component 15 to perform specific navigation behaviors. The instruction includes target motion parameters (such as travel speed, turning angle, and holding distance), interactive trigger signals (such as voice broadcast content and screen animation prompts), and mode status indicators (such as "leading the way" or "accompanying") to ensure that the robot guides the user in a way that meets the needs of the current environment.
[0135] The destination location refers to the geographic coordinates of the end point of the recommended path, corresponding to the specific facility or area that the target user needs to reach, such as a ticket gate, exit gate, elevator lobby, service desk, or entrance to a specific shop. This location is based on high-precision maps and confirmed through multi-sensor fusion technology to ensure that the robot can accurately determine the task completion status.
[0136] Specifically, the control component 14 translates each guidance mode into specific guidance instructions. These instructions are sent to the mobile component 15 via an internal communication bus. Upon receiving the instructions, the mobile component 15 drives the wheeled chassis along a set trajectory while continuously tracking the relative position of the target user using a LiDAR and a depth camera. When the target user briefly pauses to check their phone, the robot automatically stops moving forward and announces in voice, "I'm waiting for you here."
[0137] When the navigation robot confirms, based on radar positioning and visual recognition, that the target user has entered, for example, within 1.5 meters of the A12 ticket gate, it determines that it has arrived at its destination, and then plays a closing message: "You have arrived at your destination. Have a pleasant journey." After waiting for 10 seconds without receiving any new requests, it autonomously returns to its standby point to recharge.
[0138] Through the above embodiments, the abstract guidance model is accurately transformed into executable guidance instructions, achieving a seamless connection from intelligent decision-making to physical actions. The mobile component 15 dynamically adjusts its travel strategy according to different road sections, enabling both efficient guidance and safe accompaniment, significantly improving guidance reliability and user experience.
[0139] This application also provides a computer-readable storage medium for storing computer programs used in the aforementioned terminal devices. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0140] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that, as an alternative implementation, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0141] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0142] If a function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application.
[0143] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for a robot to guide users, characterized in that, The wayfinding robot includes a moving component and a control component. The control component executes the following guidance method, including: The system acquires map data of the movement paths and target locations of candidate users within a preset time period. Based on the map data, it constructs a minimum geometric shape containing the movement paths. Based on the radius of the minimum geometric shape and the preset time period, it determines the confidence level of the standing feature, obtains the confidence level of the handheld feature based on the handheld feature data, obtains the confidence level of the confused expression feature based on the confused expression feature data, and performs a weighted sum of the confidence levels of the standing feature, the handheld feature, and the confused expression feature based on preset movement weight coefficients, posture weight coefficients, and expression weight coefficients to obtain a fusion confidence level. Among the candidate users, it identifies target users whose fusion confidence level is greater than a confidence level threshold. Based on the voice data of the target user and the map data of the target location, a recommended route within the target location is generated; Based on the recommended path and the location of each candidate user, the guidance mode of the wayfinding robot is determined, including the leading mode, the accompanying mode, and the combined mode. The mobile component is triggered to execute the guidance mode, guiding the target user to the destination location of the recommended path.
2. The method according to claim 1, characterized in that, The process of obtaining handheld feature confidence based on the handheld feature data includes: Obtain the depth image and RGB image of the candidate user; The RGB image is segmented to identify the outline of the object held by the candidate user; Based on the depth image, obtain the three-dimensional point cloud data of the outline of the handheld object; Based on the 3D point cloud data of the handheld object's outline, the spatial volume of the handheld object is calculated, and based on the spatial volume, the confidence level of the handheld feature is determined.
3. The method according to claim 2, characterized in that, The calculation of the spatial volume of the handheld object based on the 3D point cloud data of the handheld object's outline includes: Clustering processing is performed on the 3D point cloud data of the handheld object outline to remove outliers in the 3D point cloud data; Fit the minimum bounding box in the 3D point cloud data after removing the outliers; Obtain the length, width, and height of the minimum bounding cuboid; The spatial volume of the handheld object is determined based on the length, width, and height.
4. The method according to claim 1, characterized in that, The process of obtaining the confidence level of the confused expression features based on the confused expression feature data includes: Obtain a continuous video frame sequence of the candidate user, the video frame sequence including RGB images and depth images; Facial keypoint detection is performed on each frame of RGB and depth images to extract geometric deformation features of the eyebrow, eye, and mouth regions; Based on the geometric deformation features, at least one facial movement is determined, and the facial movement includes at least one or more combinations of frowning, squinting, and drooping corners of the mouth. The confidence level of the confused expression feature is determined based on the frequency, duration, and intensity changes of the facial movements in the video frame sequence.
5. The method according to claim 4, characterized in that, The determination of the confidence level of the confused expression features based on the frequency, duration, and intensity changes of the facial movements in the video frame sequence includes: Temporal modeling is performed on the frequency, duration, and intensity changes of facial actions in the video frame sequence to generate a facial dynamic feature sequence; The facial dynamic feature sequence is input into a pre-trained emotion classification model, which outputs the confidence score of the confused expression feature representing the candidate user's confused emotion.
6. The method according to claim 1, characterized in that, The step of determining the guidance mode of the navigation robot based on the recommended path and the location of each candidate user includes: Determine each turning position on the recommended path; Based on each of the aforementioned turning positions, the recommended path is divided into multiple sub-paths; For each sub-path, the sub-path and each candidate user are matched to determine the number of candidate users on the sub-path. The guidance mode for the targeted sub-path is determined based on the number of candidate users on the targeted sub-path. The combination of guidance patterns for each of the sub-paths is used as the guidance pattern for the wayfinding robot.
7. A wayfinding robot, characterized in that, include: The camera component is used to collect handheld feature data and confused expression feature data of each candidate user in the target location; A radar component is used to collect the stationing characteristic data of each of the candidate users and the map data of the target location; Microphone assembly, used to collect voice data from the target user; A control component is used to acquire the movement path of the candidate user within a preset time period and map data of the target location; based on the map data, construct a minimum geometric figure containing the movement path; based on the radius of the minimum geometric figure and the preset time period, determine the confidence level of the standing feature; obtain the confidence level of the handheld feature based on the handheld feature data; obtain the confidence level of the confused expression feature based on the confused expression feature data; and perform a weighted sum of the confidence levels of the standing feature, the handheld feature, and the confused expression feature based on preset movement weight coefficients, posture weight coefficients, and expression weight coefficients to obtain a fusion confidence level; and determine the target user whose fusion confidence level is greater than a confidence level threshold among the candidate users. Based on the voice data of the target user and the map data of the target location, a recommended route within the target location is generated; Based on the recommended path and the location of each candidate user, the guidance mode of the wayfinding robot is determined, including the leading mode, the accompanying mode, and the combined mode. A mobile component is used to execute the guidance mode, guiding the target user to the destination location of the recommended path.
8. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed on a processor, implements the wayfinding robot guidance method according to any one of claims 1-6.
Citation Information
Patent Citations
Leader robot control method, control equipment and computer readable storage medium
CN115454075A