Method for guiding a robot having a body and related device
By optimizing information acquisition through communication between the robot controller and the server and the sensor module, synchronous linkage between the humanoid robot and multimedia equipment was achieved, solving the problem of insufficient equipment linkage and improving the scheduling efficiency and user experience for receiving multiple groups of people.
Patent Information
- Application Number
- CN202511745787.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-11-26
AI Technical Summary
Currently, humanoid robots lack sufficient coordination with other devices during tours, resulting in low scheduling efficiency and poor user experience when handling multiple groups of people.
By communicating with the server, the robot controller obtains the intention information of the crowd, detects and schedules the linkage of multimedia devices in real time, and optimizes information acquisition by utilizing sensor modules and intention behavior databases to achieve synchronous linkage between the robot and multimedia devices.
It improves the interconnectivity of equipment, enhancing scheduling efficiency and user experience when handling multiple groups of people.
Smart Images

Figure CN121179478B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, and more particularly to a guided tour method and related apparatus using an embodied robot. Background Technology
[0002] Humanoid robots have already been deployed as service devices in some business scenarios, such as guiding tours, teaching, companionship, and equipment operation.
[0003] However, when interacting with users to complete some business processes, it mainly relies on its own capabilities and lacks sufficient linkage with other devices in the business scenario. During the tour, the scheduling efficiency and user experience are not high enough when receiving multiple groups of people. Summary of the Invention
[0004] This application provides a guided tour method and related apparatus using a robot, which can improve equipment connectivity, increase scheduling efficiency when receiving multiple groups of people, and enhance user experience.
[0005] In a first aspect, embodiments of this application provide a guided tour method using an embodied robot, applied to a robot controller in a guided tour system. The robot controller is used to control a single robot. The guided tour system further includes a server, which is communicatively connected to the robot controller. The method includes:
[0006] The first intention information of the first group of people is obtained based on the first reception instruction. The first group of people includes the tourists currently within the crowd detection range. The first intention information includes the tour guide needs of the first group of people.
[0007] The crowd detection operation is continuously performed. If a second group of people is detected within the crowd detection range, the second intention information of the second group of people is determined. The second intention information includes the tour guide needs of the second group of people.
[0008] Send the second intent information to the server and receive response information from the server;
[0009] Based on the response information, the second group of people are informed of the tour information, which includes at least one of the following: tour robot information and estimated waiting time.
[0010] If the first group of people requires a guided tour, then the first group of people will be guided through the guided tour operation based on the first intent information. The guided tour operation includes guided tour explanation operation and / or multimedia change operation. The multimedia change operation includes changing the display content of the display device.
[0011] In one possible embodiment, the continuous execution of the crowd detection operation, wherein determining the second intent information of the second crowd if a second crowd is detected within the crowd detection range, includes:
[0012] If the population is within the detection range, the population detection operation will continue to be performed;
[0013] If the second group of people is detected, the characteristics of the first group of people are identified, and the first distance between the first group of people and the second group of people is obtained;
[0014] The acquisition method of the second intent information is determined based on the first distance and a preset distance threshold, and the acquisition method is used to constrain the information source of the second intent information;
[0015] The audio information and / or behavioral information of the second group of people are obtained based on the acquisition method described above;
[0016] Intent feature information is extracted based on the audio information and / or the behavioral information;
[0017] The second intent information is determined based on the intent feature information and the intent behavior database.
[0018] In one possible embodiment, the single robot is equipped with a first sensor module, the navigation system further includes a second sensor module, the information source includes the first sensor module and / or the second sensor module, and the acquisition method includes direct acquisition or indirect acquisition;
[0019] The method for determining the second intent information based on the first distance and a preset distance threshold includes:
[0020] If the first distance is greater than the preset distance threshold, the method of obtaining the second intent information is determined to be the indirect acquisition.
[0021] If the first distance is not greater than the preset distance threshold, the method for obtaining the second intent information is determined to be the direct acquisition.
[0022] The acquisition of audio and behavioral information of the second group of people based on the acquisition method includes:
[0023] If the acquisition method is direct acquisition, the audio information and / or behavioral information of the second group of people are acquired based on the first sensor module;
[0024] If the acquisition method is indirect acquisition, the audio information and / or behavioral information of the second group of people received from the second sensor module are acquired.
[0025] In one possible embodiment, guiding the first group of people through a guided tour based on the first intent information includes:
[0026] Obtain the tour guidance intention information of the first group of people, the tour guidance intention information including the tour preferences of the first group of people;
[0027] Determine the tour route based on the tour intent information;
[0028] Based on the guided tour route and the guided tour skill library, the robot guides the first group of people through the target exhibition hall. The guided tour skill library is used by the robot to match specific behaviors based on external information.
[0029] In one possible embodiment, multiple interconnected devices are installed within the target exhibition hall. These devices are used for displaying text, sound, or images. The guided tour operation, based on the guided route and the guide skill library, leading the first group of people through the tour includes:
[0030] The location information of the linkage device and the explanation area is detected in real time on the guided tour route. The location information of the explanation area is used to constrain the explanation content and behavior of the individual robot.
[0031] If the linkage device is detected, establish a communication connection with the linkage device based on the current explanation content;
[0032] The explanation actions are determined based on the location information of the explanation area and the guide skill library;
[0033] The control command is determined based on the current explanation content and / or the explanation action;
[0034] Send control commands to the linkage device to control the linkage device to perform linkage display operations.
[0035] In one possible embodiment, the linkage device includes multiple display screens and / or lighting devices and / or projection devices, and the step of determining the control command based on the current narration content and / or the narration action includes:
[0036] If the linkage device includes the lighting device and / or the projection device, the positional relationship of the first group of people is detected, including the relative positional relationship with the individual robot; based on the positional relationship, the technical parameters of the lighting device and / or the projection device are determined, the technical parameters constraining at least one of the angle, brightness, color, and shape of the lighting device and / or the projection device; based on the technical parameters and the current explanatory content, the control command of the lighting device and / or the projection device is determined.
[0037] If the linkage device includes multiple display screens, the control commands for the lighting device and / or the multiple display screens are determined based on the current explanatory content.
[0038] In one possible embodiment, determining the control commands for the lighting device and / or the projection device based on the technical parameters and the current narration content includes:
[0039] Based on the current content being explained, semantic decomposition is performed to identify multiple semantic keywords;
[0040] Based on the multiple semantic keywords, a pre-stored device linkage mapping table is queried to match multimedia action combinations associated with topic tags. The device linkage mapping table includes the topic tags and the multimedia action combinations corresponding to the topic tags. The multimedia action combinations are used to control the robot's actions according to the explanation content. The device linkage mapping table is pre-stored in the server and / or the robot controller.
[0041] Based on the multimedia action combination, a control instruction containing spatiotemporal synchronization parameters is determined. The control instruction containing spatiotemporal synchronization parameters is used to uniformly trigger the actions of multiple devices through a timecode synchronizer, so that the display screen and / or the lighting device and / or the projection device are superimposed on the narration text and / or the narration action and start synchronously when they reach the anchor point.
[0042] Secondly, embodiments of this application provide a guided tour device for an embodied robot, a robot controller applied to a guided tour system, the robot controller being used to control a single robot, the guided tour system further including a server, the server being communicatively connected to the robot controller, and the device comprising:
[0043] The intent acquisition module is used to acquire the first intent information of the first group of people based on the first reception instruction. The first group of people includes the tourists currently within the crowd detection range, and the first intent information includes the tour guide needs of the first group of people.
[0044] The crowd detection module is used to continuously perform crowd detection operations. If a second group of people is detected within the crowd detection range, the second intention information of the second group of people is determined. The second intention information includes the navigation needs of the second group of people.
[0045] An intent sending module is used to send the second intent information to the server and receive response information from the server;
[0046] The status sending module is used to inform the second group of people of the tour status information based on the response information, wherein the tour status information includes at least one of the following: tour robot information and estimated waiting time;
[0047] A navigation module, configured to, if the navigation requirement of the first group of people is to require navigation, lead the first group of people to perform navigation operations based on the first intent information. The navigation operations include navigation explanation operations and / or multimedia change operations, and the multimedia change operations include changing the display content of a display device.
[0048] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute some or all of the steps described in the first aspect.
[0049] In a fourth aspect, an embodiment of the present application provides a robot, including a processor, a memory, a communication interface, and one or more programs. Among them, the above one or more programs are stored in the above memory and are configured to be executed by the above processor. The above programs include instructions for performing some or all of the steps described in the first aspect of the embodiments of the present application.
[0050] In a fifth aspect, an embodiment of the present application provides a computer program product. Among them, the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps described in the first aspect of the embodiments of the present application. The computer program product may be a software installation package.
[0051] By implementing the embodiments of the present application, first intent information of a first group of people is obtained based on a first reception instruction. The first group of people includes a group of tourists currently within a people detection range, and the first intent information includes the navigation requirement of the first group of people. A people detection operation is continuously executed. If a second group of people is detected within the people detection range, second intent information of the second group of people is determined. The second intent information includes the navigation requirement of the second group of people. The second intent information is sent to the server, and a response information is received from the server. Based on the response information, navigation situation information is informed to the second group of people. The navigation situation information includes at least one of navigation robot information and an estimated waiting time. If the navigation requirement of the first group of people is to require navigation, the first group of people is led to perform navigation operations based on the first intent information. The navigation operations include navigation explanation operations and / or multimedia change operations, and the multimedia change operations include changing the display content of a display device. In this way, it is possible to more quickly dock with tourists to be served, and联动 with other devices during the docking and explanation processes, so as to improve the device linkage, improve the scheduling efficiency when receiving multiple groups of people, and improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the background art, the accompanying drawings used in the embodiments of the present invention or the background art will be described below.
[0053] Figure 1a This is a schematic diagram of the architecture of the first type of tour guide system provided in the embodiments of this application;
[0054] Figure 1b This is a schematic diagram of the architecture of the second type of tour guide system provided in the embodiments of this application;
[0055] Figure 1c This is a schematic diagram of the architecture of the third type of tour guide system provided in the embodiments of this application;
[0056] Figure 2 This is a flowchart illustrating a guided tour method using an embodied robot, as provided in an embodiment of this application.
[0057] Figure 3 This is a schematic diagram of the robot scheduling process for a guided tour method using an embodied robot provided in an embodiment of this application;
[0058] Figure 4 This is a schematic diagram of a guided tour method using an embodied robot, as provided in an embodiment of this application.
[0059] Figure 5 This is a schematic diagram of the structure of a guide device for an embodied robot according to an embodiment of this application;
[0060] Figure 6 This is a schematic diagram of the structure of another guide device for an embodied robot provided in an embodiment of this application;
[0061] Figure 7 This is a schematic diagram of the structure of a robot provided in an embodiment of this application. Detailed Implementation
[0062] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0063] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or electronic device that includes a series of steps or units is not limited to the listed steps or units, but in an alternative example includes steps or units not listed, or in an alternative example includes other steps or units inherent to these processes, methods, products, or electronic devices.
[0064] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0065] Humanoid robots have already been deployed as service devices in some business scenarios, such as guiding tours, teaching, companionship, and equipment operation.
[0066] However, when interacting with users to complete some business processes, it mainly relies on its own capabilities and lacks sufficient linkage with other devices in the business scenario. During the tour, the scheduling efficiency and user experience are not high enough when receiving multiple groups of people.
[0067] To address the aforementioned issues, this application provides a guided tour method and related apparatus using a robot avatar, which can improve equipment connectivity, increase scheduling efficiency when receiving multiple groups of people, and enhance user experience.
[0068] The guided tour method for embodied robots provided in this application embodiment can be applied to, for example... Figure 1a , Figure 1b or Figure 1c Please refer to the guide system shown. Figure 1a , Figure 1a This is a schematic diagram of the architecture of the first type of tour guide system provided in this application embodiment. The tour guide system 100 includes a server 110, a robot controller 121, a robot 120, and a linkage device 130.
[0069] The robot controller 121 is located inside the robot 120 and is used to control the robot 120. The robot controller 121 is communicatively connected to the server 110 and can communicate with the server 110 via a network. The linkage device 130 can communicate with at least one of the server 110 and the robot controller 121. Where possible, communication is based on the Multi-Agent Communication Protocol (MCP). The robot controller 121 includes a Modular Component System (MCS), which encapsulates robot functional modules (such as motion control and sensor actuation) through standardized interfaces.
[0070] In this solution, server 110 refers to a remote computer used for processing large amounts of computing tasks and storing data. A device linkage mapping table is deployed on server 110. Server 110 is used for matching robot and device actions, as well as for generating instructions and facilitating communication. Server 110 can communicate directly or indirectly with linkage device 130 to control the linkage between linkage device 130 and robot 120.
[0071] Based on this, this application provides a guided tour method and related apparatus for an embodied robot, which will be described in detail below with reference to the accompanying drawings.
[0072] Please see Figure 1b , Figure 1b This is a schematic diagram of the architecture of the second type of tour guide system provided in the embodiments of this application. The tour guide system 100 includes a server 110, a robot controller 121, a robot 120, and a linkage device 130. The robot 120 is equipped with a first sensor module 122.
[0073] The first sensor module 122 can be mounted on the head to simulate the robot's line of sight and is used to acquire image information from the outside world. The image information may include, but is not limited to, the intention information of the crowd. The first sensor module 122 can be a single camera or a combination of sensors.
[0074] Please see Figure 1c , Figure 1c This is a schematic diagram of the architecture of the third type of tour guide system provided in the embodiments of this application. The tour guide system 100 includes a server 110, a robot controller 121, a robot 120, a linkage device 130, and a second sensor module 140. The robot 120 is equipped with a first sensor module 122.
[0075] The second sensor module 140 is used to acquire image information from the outside world, which may include, but is not limited to, the intent information of the crowd. The second sensor module 140 may be a single camera or a combination of sensors. Where possible, at least two of the server 110, robot controller 121, robot 120, linkage device 130, second sensor module 140, and first sensor module 122 communicate with each other based on the Multi-Agent Communication Protocol (MCP).
[0076] Please see Figure 2 , Figure 2 This is a flowchart illustrating a guided tour method using an embodied robot, as provided in an embodiment of this application. The method is applied to a robot controller in a guided tour system. The robot controller controls a single robot. The guided tour system also includes a server, which is communicatively connected to the robot controller. Figure 2 As shown, the method includes the following steps:
[0077] S210, based on the first reception instruction, obtain the first intent information of the first group of people, the first group of people includes the tourists currently within the group detection range, and the first intent information includes the tour guide needs of the first group of people.
[0078] The first reception instruction can be a reception instruction issued by the server or generated spontaneously.
[0079] Specifically, if a single robot is equipped with a first sensor module, the first reception instruction can be determined based on the first crowd image information acquired by the first sensor module. If the tour guide system is equipped with a second sensor module, the first reception instruction can be determined based on the second crowd image information acquired by the second sensor module. Specifically, if a second sensor module is installed, the second sensor acquires a second crowd image within the crowd detection range and transmits this image to the server. The server generates a first reception instruction based on this image, and the single robot, upon receiving the first reception instruction, moves to the location of the first crowd to provide reception. Alternatively, if a first sensor module acquires a first crowd image within the crowd detection range and transmits it to the robot controller, the robot controller can either automatically generate a first reception instruction and report it to the server, or the robot controller can transmit the first crowd image, the server generates a first reception instruction based on this image, and the single robot, upon receiving the first reception instruction, moves to the location of the first crowd to provide reception.
[0080] When the robot moves near the first group of people, its controller acquires the first intention information of that group. This first intention information can be determined based on acquired images, audio, or other content. The first intention information represents the group's guidance needs; for example, whether the group needs guidance or not.
[0081] The crowd detection range can be a preset area within which people can be identified and detected; for example, the crowd detection range can be an area at the entrance or an area in the lobby.
[0082] It should be noted that the tour guide system includes multiple robots used for guiding tours. A single robot can be any one of these robots; here, the single robot facing the first group of people is the primary focus. When other single robots, besides the current one, face the second group, third group, etc., the single robots facing these groups become the primary agents for applying this method.
[0083] S220, continue to perform crowd detection operation. If a second group of people is detected within the crowd detection range, determine the second intention information of the second group of people. The second intention information includes the navigation needs of the second group of people.
[0084] In this process, after a single robot arrives in the crowd detection range based on the first reception instruction, it continuously performs crowd detection operations while receiving the first group of people. If a second group of people is detected in the crowd detection range, the second intention information of the second group of people is determined, and the second group of people is a group of people who are not received by robots.
[0085] The method for determining whether the second group of people is served by a robot includes: initiating a first query request to the server to query the central task scheduling table, which records the status of all robots in real time, including their location and reception status; the server determines whether the second group of people has been assigned a reception task and generates first feedback information; the robot controller receives the first feedback information, wherein the reception task includes tasks that other service robots are currently performing or about to perform to serve the second group of people, and the feedback information is used to indicate whether the second group of people has been served; if the second group of people has not been assigned any robot, then the second group of people is determined to be a group without robot service.
[0086] The method for determining whether the second group of people is being received by a robot includes: (a) obtaining the real-time location information of all other robots currently within the detection range of the group of people, and determining whether the location of the second group of people matches the real-time location information of any other robot; (b) and / or querying the central task scheduling table to determine whether the second group of people has been assigned a reception task; if the result of step (a) is a mismatch, and / or the result of step (b) is that no reception task has been assigned, then the second group of people is determined to be a group without robot reception.
[0087] In one possible embodiment, the continuous execution of the crowd detection operation, and determining the second intent information of the second crowd if a second crowd is detected within the crowd detection range, includes: continuously executing the crowd detection operation if the crowd is within the crowd detection range; if the second crowd is detected, locking the crowd characteristics of the first crowd and obtaining a first distance from the second crowd; determining the acquisition method of the second intent information based on the first distance and a preset distance threshold, wherein the acquisition method is used to constrain the information source of the second intent information; acquiring the audio information and / or behavioral information of the second crowd based on the acquisition method; extracting intent feature information based on the audio information and / or the behavioral information; and determining the second intent information based on the intent feature information and an intent behavior database.
[0088] The process of identifying the characteristics of the first group of people is crucial to prevent the robot from losing track of the original service recipients (the first group) when processing the second group. The first distance to the second group is obtained by calculating the centroid of the detected point cloud of the second group using a crowd center point meter. Based on this first distance, the method for acquiring the second intention information of the second group is determined. Since the accuracy of sensor acquisition is limited by distance, this distance constraint ensures the reliability of the acquired second intention information. The first distance and a preset distance threshold determine the information source, which can be sensors at different locations, such as the sensors of the current single robot or sensors from other locations. At least one of the audio and behavioral information of the second group is acquired from sensors at different locations. The behavioral information is determined based on the actions, expressions, and mannerisms of the second group, while the audio information is the voice emitted by the second group, such as "I need robot guidance." Intention feature information is extracted based on at least one of the audio and behavioral information of the second group, and then the second intention information is determined by matching the intention feature information with an intention-behavior database.
[0089] The extraction process for the aforementioned audio features can be as follows: extracting Mel-frequency cepstral coefficients to represent speech content; calculating the fundamental frequency (F0) variance to detect emotional state (e.g., variance > 50Hz when anxious). The extraction process for the aforementioned behavioral features can be as follows: determining spatial features, including crowd density and relative azimuth to exhibits; determining temporal features, including sustained staring or looking around, and repeatedly approaching a specific area.
[0090] Specifically, the audio feature extraction process may include: first, preprocessing the acquired original audio signal, including noise reduction and sound source enhancement, to obtain a standard audio signal; then, performing time-frequency domain transformation on the standard audio signal to obtain a spectrogram; extracting Mel-frequency cepstral coefficients (MFCC) based on the spectrogram; then performing dynamic feature expansion to obtain a multidimensional feature vector; and finally, performing feature fusion based on the feature vector to obtain the audio features.
[0091] Specifically, the behavioral feature extraction process may include: first, performing human posture analysis on the acquired image to determine the coordinates of human key points; then, determining spatial relationship features and motion dynamics features based on the coordinates of human key points, including relative direction angle, movement speed, angle, trajectory entropy value, amplitude, movement time, etc.; and finally, generating feature vectors based on the spatial relationship features and motion dynamics features to obtain behavioral features.
[0092] The obtained audio and / or behavioral feature information is multimodally aligned, synchronized via timestamps, and then weighted and fused. In some cases, the weights of audio and behavior can be constrained by a first distance. For example, when the first distance is greater than a preset distance threshold, the audio weight is 'a' and the behavior weight is 'b'; when the first distance is not greater than the preset distance threshold, the audio weight is 'c' and the behavior weight is 'd'. The weighted fused features are then matched against an intent-behavior database; specifically, matching can be performed through similarity calculations.
[0093] As can be seen, in this embodiment, by constraining the information source through the first distance, resource utilization can be optimized, the accuracy and efficiency of information analysis can be improved, and robustness can be enhanced. Furthermore, this can improve scheduling efficiency when receiving multiple groups of people and enhance the user experience.
[0094] In one possible embodiment, the single robot is equipped with a first sensor module, and the navigation system further includes a second sensor module. The information source includes the first sensor module and / or the second sensor module, and the acquisition method includes direct acquisition or indirect acquisition. Determining the acquisition method of the second intention information based on the first distance and a preset distance threshold includes: if the first distance is greater than the preset distance threshold, determining the acquisition method of the second intention information as indirect acquisition; if the first distance is not greater than the preset distance threshold, determining the acquisition method of the second intention information as direct acquisition. Acquiring the audio and behavioral information of the second group of people based on the acquisition method includes: if the acquisition method is direct acquisition, acquiring the audio and / or behavioral information of the second group of people based on the first sensor module; if the acquisition method is indirect acquisition, acquiring the audio and / or behavioral information of the second group of people received from the second sensor module.
[0095] The first sensor module can be a visual sensor and an audio sensor, such as a camera and microphone array, installed on the robot's head. The second sensor module can be a visual sensor and an audio sensor, such as a camera and microphone array, installed in the environment.
[0096] When the first distance is not greater than a preset distance threshold, the acquisition method is direct acquisition mode. The data source is the first sensor module mounted on the robot body, which captures images of the crowd through a high-definition camera, analyzes key points of the human skeleton (such as the OpenPose algorithm), and analyzes dynamic features such as limb movements (such as waving amplitude > 30°) and movement trajectory (wandering path entropy value > 1.2). A directional microphone array implements sound source beamforming to suppress environmental noise and focuses on collecting clear voice commands (such as "Where are the exhibits?") or emotional tone (urgent / calm).
[0097] When the initial distance exceeds a preset distance threshold, the data acquisition method is indirect, with the data source being the second sensor module mounted on the scene. This module captures images of the crowd using a high-definition camera, analyzes key points of the human skeleton (e.g., using the OpenPose algorithm), and analyzes dynamic features such as limb movements (e.g., waving amplitude > 30°) and movement trajectories (wandering path entropy > 1.2). A directional microphone array performs beamforming to suppress environmental noise, focusing on capturing clear voice commands (e.g., "Where are the exhibits?") or emotional tone (urgent / calm). The robot controller sends data requests to the central scheduling system via a wireless network (e.g., the ROS 2 DDS protocol), and the system retrieves data from the second sensor module closest to the second group of people.
[0098] As can be seen, in this embodiment, the hard switching strategy based on distance threshold takes into account both accuracy and energy consumption. By extending the robot's effective perception range through fixed environmental nodes, an integrated perception of "robot-environment" is constructed, which can improve equipment linkage, improve scheduling efficiency when receiving multiple groups of people, and improve user experience.
[0099] S230, send the second intent information to the server and receive response information from the server.
[0100] The second intent information is a structured intent judgment result generated through local processing, which primarily represents whether the second group of people needs the guided tour service.
[0101] In one possible embodiment, please refer to Figure 3 , Figure 3 This is a schematic diagram of the robot scheduling process for a guided tour method using an embodied robot provided in an embodiment of this application, as shown in the following example. Figure 3 As shown, the server first obtains the second intent information to determine whether the second group of people needs a guided tour. If the second intent information indicates that the second group of people needs a guided tour, the server makes a decision based on the central task scheduling table to see if there are any idle robots available. If there are no idle robots, the server determines the task information of robots whose tasks are about to end, calculates the estimated waiting time, and sends a response message to the robot controller of the individual robot, including the estimated waiting time. If there are idle robots, the server dispatches an idle robot in real time and sends a response message to the robot controller of the individual robot, including a notification that the robot has been dispatched. If the second intent information indicates that the second group of people does not need a guided tour, the request is discarded.
[0102] As can be seen, in this embodiment, the robot's direct questioning and information dissemination reduce the dissatisfaction of the second group and improve the user experience.
[0103] S240, based on the response information, inform the second group of the tour information, the tour information including at least one of the tour robot information and the estimated waiting time.
[0104] The information regarding the guided tour can be provided through at least one of voice broadcasting and large-screen display. Voice content generation can be achieved by applying a structured template to the aforementioned guided tour information to generate complete sentences, or by using artificial intelligence to generate complete sentences. Specifically, the robots can have numbers or codes, and the guided robot information can be these numbers or codes, with estimated waiting times accurate to minutes or seconds. The large screen used for the aforementioned display can be a fixed screen used to show the robot status, displaying the location of each robot, its estimated task completion time, the estimated waiting time for the current waiting crowd, and the corresponding robot information in a fixed area.
[0105] In this process, while providing guidance, a single robot performs action tag matching based on audio or behavior. At least one intention action mapping library is set in the robot controller or server. The intention action mapping library includes multiple skill mapping tags corresponding to specific actions or audio. When audio or behavior that matches the tag in the intention action mapping library is detected, the action corresponding to the tag is retrieved and executed.
[0106] It should be noted that the aforementioned intention-action mapping library is defined as a library, or capability shelf, encompassing various atomic skills of the robot, including action libraries such as flat-ground walking, navigation, satellite positioning, grasping, teleoperation, and custom functions. Through MCP / functionCalling or other interactive methods, smaller models (such as vision) are cascaded, and the skill library is structured and loaded into VLA / LLM. Users can then interact with the robot or issue tasks through multimodal user interaction.
[0107] For example, in scenarios where long waits cause anxiety, crowd behavior might include repeatedly checking the time or pacing; crowd sounds might include phrases like "How much longer?" or "Why is it so slow?" Based on the input audio or visual information, semantic analysis or behavior recognition is performed to determine standard semantics and standard behaviors. Based on the standard semantics and standard behaviors, they are matched with action tags in the intention-action mapping library to trigger at least one of the specific matched output actions and output speech, where the output actions include body movements and facial expressions.
[0108] S250, if the first group of people requires a guided tour, then the first group of people are led to perform a guided tour operation based on the first intent information. The guided tour operation includes guided tour explanation operation and / or multimedia change operation. The multimedia change operation includes changing the display content of the display device.
[0109] If the first group's guidance request is for a guided tour, after informing the second group of the aforementioned response information, a single robot continues to guide the locked first group based on the first intent information. The guided tour operations include providing explanations, changing actions to accompany the explanations, and changing multimedia settings to accompany the explanations. Multimedia changes involve communicating with associated multimedia devices.
[0110] In one possible embodiment, guiding the first group of people through the guided tour based on the first intent information includes: obtaining the guided tour intent information of the first group of people, the guided tour intent information including the first group of people's tour preferences; determining the guided tour route based on the guided tour intent information; and guiding the first group of people through the guided tour in the target exhibition hall based on the guided tour route and a guided tour skill library, wherein the guided tour skill library is used by the individual robot to match specific behaviors based on external information.
[0111] During the guided tour, the system can receive the proactive needs of the first group of people, i.e., the tour intention information. Then, based on the tour intention information, key semantic words are determined. In some possible cases, the BERT model can be used to fill semantic slots (e.g., "want to see bronzes" → preference = "bronzes"). Based on this need, the tour route is determined to include this need.
[0112] When determining the guided tour route, decisions can be based on multiple factors, including but not limited to the tour's intended purpose, the planned route, and booth availability. Booth availability can refer to visitor traffic or crowd density at each booth. The specific tour route can be determined based on these multiple factors and their respective weights.
[0113] The guide skill library is used to match specific explanation behaviors for the robot, such as actions and commands. Specifically, when the robot arrives near a booth, it can identify which booth it is currently in, and then provide guidance based on the actions or commands corresponding to the guide skill library for that booth. The commands are used to coordinate with multiple schedulable linkage devices near the booth.
[0114] As can be seen, in this embodiment, guiding the tour based on the proactive needs of the first group of people can improve the flexibility of the tour and thus increase user satisfaction.
[0115] In one possible embodiment, multiple interconnected devices are installed within the target exhibition hall. These interconnected devices are used for displaying text, sound, or images. The guided tour operation, based on the guided tour route and a guide skill library, involves: real-time detection of the interconnected devices and the location information of the explanation area along the guided tour route; the location information of the explanation area constrains the explanation content and behavior of the individual robot; if an interconnected device is detected, establishing a communication connection with the interconnected device based on the current explanation content; determining explanation actions based on the explanation area location information and the guide skill library; determining control instructions based on the current explanation content and / or the explanation actions; and sending control instructions to the interconnected devices to control them to perform interconnected display operations.
[0116] The linkage device includes a text display device for displaying text, an audio device for playing sound, and a display device for displaying images. The text display device can be an electronic screen or a projection device, the audio device can be a speaker, and the display device can be a projection device, an electronic screen, AR glasses, a holographic device, etc. The aforementioned linkage device is communicatively connected to at least one of the robot controller or a server.
[0117] The explanation area can be multiple closed areas pre-divided based on maps, etc. In some cases, it can be a centimeter-accurate map of the exhibition hall (error ±3cm) built by laser SLAM. The explanation area may be equipped with sensors to identify and detect the position information of the robot and the explanation area. The sensors can also communicate with the robot controller, which can obtain the positional relationship between itself and the explanation area.
[0118] When a linkage device is detected, the control command for controlling the linkage device is determined based on the guide skill library matched with the current explanation content. For example, the skill library matching logic could be: when the explanation content describes an oil painting, the matching action is to point an arm to the details of the painting, and the control command is to focus the projector on a specific area; when the explanation content introduces the bronze casting process, the matching action is to simulate the pouring action, and the control command is to display a 3D casting animation on the AR glasses; when the explanation content reaches the war history area, the control command is to play battlefield sound effects and dim the lights.
[0119] As can be seen, in this embodiment, the content and space are strongly coupled by matching the explanation area, and the devices are activated on demand by the collaboration between the robot and multiple linkage devices. The system efficiency is significantly optimized, thereby improving the tour guide effect and enhancing the user experience.
[0120] In one possible embodiment, the linkage device includes multiple display screens and / or lighting devices and / or projection devices. The step of determining control commands based on the current narration content and / or the narration actions includes: if the linkage device includes the lighting device and / or the projection device, detecting the positional relationship of the first group of people, the positional relationship including the relative positional relationship with the individual robot; determining the technical parameters of the lighting device and / or the projection device based on the positional relationship, the technical parameters constraining at least one of the angle, brightness, color, and shape of the lighting device and / or the projection device; determining the control commands of the lighting device and / or the projection device based on the technical parameters and the current narration content; if the linkage device includes multiple display screens, determining the control commands of the lighting device and / or the multiple display screens based on the current narration content.
[0121] The display screen is used to show visual information (such as text, images, and videos) related to the content being explained. Lighting devices adjust parameters such as angle, brightness, color, and shape to create atmosphere or highlight key points. Projection devices project dynamic or static content (such as animations and charts) onto specific areas. These devices are dynamically adjusted based on the content being explained and actions taken, enhancing interactivity and immersion.
[0122] If the linkage device includes at least one of lighting or projection devices, the positional relationship of the crowd can be detected using sensors (such as cameras, infrared radar, and ultrasound) to capture the crowd's coordinates in real time and calculate the spatial relationship between the crowd and the robot (such as distance and azimuth). The process of determining the technical parameters can be as follows: adjusting the illumination direction of the lights / projection based on the crowd's position (e.g., when the crowd is concentrated on the left, the lights deflect to the left); dynamically adjusting based on ambient light or crowd distance (e.g., reducing brightness to avoid glare when people are close); combining with the theme of the content being explained (e.g., switching to a blue tone when explaining "the ocean"); and forming specific patterns through projection or light arrays (e.g., focusing a circular light spot). The process of generating instructions based on the content being explained can be as follows: triggering parameter changes with keywords (e.g., mentioning "volcano" causes the projection to display a volcano animation, and the lights to turn orange-red); and adjusting parameters according to the chapter's progress (e.g., when entering the "summary stage," the lights gradually dim, and the projection displays a summary chart).
[0123] If the linkage device includes multiple display screens, the current explanatory content is analyzed, broken down into multiple dimensions, and distributed to different screens. Specifically, Natural Language Processing (NLP) can be used to extract keywords and semantics. Based on the matching of keywords and semantics, the corresponding image or text content is displayed on the display screens.
[0124] Specifically, the content capture process involves: acquiring a real-time text stream via microphone and speech recognition; capturing the presenter's gestures and orientation using motion sensors (such as skeletal tracking); continuously scanning the audience's coordinates; calculating the optimal coverage angle and brightness threshold based on the audience's coordinates; calling a content keyword matching library to map color and shape templates; and dynamically allocating display screen permissions based on the robot's or the first group's orientation; packaging technical parameters and content links into standard protocols (such as MQTT messages and HTTP APIs); adjusting lighting / projection parameters via the DMX protocol or a custom API; or pushing HTML / video content to a designated screen and calling system APIs to adjust screen brightness. It should be noted that the above communication process is only an example; in actual use, the communication path can be selected according to the specific circumstances.
[0125] As can be seen, in this embodiment, the technical parameters are coupled with the explanation scene through the closed-loop control of "environmental perception - content analysis - device linkage". This improves the efficiency of information transmission while taking into account interactivity, comfort and energy efficiency, thereby enhancing device linkage and improving user experience.
[0126] In one possible embodiment, determining the control instructions for the lighting device and / or the projection device based on the technical parameters and the current narration content includes: semantically decomposing the current narration content to determine multiple semantic keywords; querying a pre-stored device linkage mapping table based on the multiple semantic keywords to match multimedia action combinations associated with theme tags, wherein the device linkage mapping table includes the theme tags and the multimedia action combinations corresponding to the theme tags, the multimedia action combinations being used for the robot to perform action control according to the narration content, and the device linkage mapping table being pre-stored in the server and / or the robot controller; and determining control instructions containing spatiotemporal synchronization parameters based on the multimedia action combinations, wherein the control instructions containing spatiotemporal synchronization parameters are used to uniformly trigger the actions of multiple devices through a timecode synchronizer, so that the display screen and / or the lighting device and / or the projection device are superimposed on the narration words and / or the narration actions and start synchronously when they reach the anchor point.
[0127] In this process, semantic decomposition determines multiple semantic keywords. This can be achieved by calculating word weights using a pre-set dictionary or models such as BERT and TF-IDF, and then extracting the Top N keywords.
[0128] The device linkage mapping table can be a library or capability shelf encompassing various atomic skills of the robot, including action libraries such as satellite positioning capabilities, action libraries, teleoperation, and custom functions. Smaller models (e.g., vision-based) are cascaded through MCP / functionCalling or other interactive methods, while the skill library is structured and loaded into VLA / LLM. Users interact with the robot or issue tasks through multimodal interaction. The device linkage mapping table maps semantic keywords to predefined device action combinations, achieving a "content-device" association. The topic tags in the device linkage mapping table are used for matching and can correspond to specific multimedia action combinations. Each topic tag corresponds to a set of device actions (which can include combinations of lights, projections, and screens). Action parameters can include angle (e.g., light pointing at a crowd), brightness (e.g., 70%), color (e.g., #FF0000), and shape (e.g., circular light spot).
[0129] Specifically, control commands containing spatiotemporal synchronization parameters are determined based on the multimedia action combinations to ensure precise temporal and spatial synchronization of multiple device actions, consistent with the narration rhythm. Automatic recognition determines the trigger time points of the narration words / actions (e.g., the timecode of the word "eruption"). Device action combinations are bound to anchor point timecodes, and multiple device actions are coordinated through a timecode synchronizer (e.g., Adobe SpeedGrade, a self-developed synchronization engine) to ensure that at the anchor point moment, at least one action superimposed on the narration words and actions—from the display screen, lighting devices, or projection devices—is executed simultaneously.
[0130] Specifically, the semantic keyword extraction process can be as follows: First, input the real-time explanation content, load a pre-trained word segmentation model (such as Jieba, HanLP) for word segmentation. Filter out stop words (such as "de", "shi"), and retain nouns and verbs. Extract the 3-5 keywords with the highest weights through the TF-IDF or TextRank algorithm. Output the keyword list. Traverse the keywords and query the pre-stored device linkage mapping table. If multiple theme tags are matched (such as "volcanic eruption" + "energy release"), merge the action combinations (such as lights + projector + screen), and then output the multimedia action combination (such as {"light": {"color": "red"}, "projector":{"content": "lava.mp4"}}). If based on a semantic anchor point (such as the keyword "eruption"), locate the time code of this word through speech recognition. If based on an action anchor point (such as the presenter pointing to the left), detect the action occurrence time through the skeleton tracking algorithm. Bind the action combination to the time code and add spatial parameters (such as the light angle being consistent with the presenter's orientation). Generate a unified instruction format (JSON / XML) and push it to the device controller. Output the control instruction with spatio-temporal parameters.
[0131] It can be seen that in this embodiment, by semantic parsing, the device actions are automatically matched to adapt to different explanation contents. The time code synchronizer ensures that the multi-device actions are strictly aligned with the explanation words / actions, enhancing the viewing experience. The device linkage mapping table supports dynamic updates. When adding new theme tags, only the device action combinations need to be configured, without modifying the core logic. Furthermore, it improves the device linkage and the user experience.
[0132] For example, please refer to Figure 4 , Figure 4 which is a schematic diagram of the scenario of a navigation method for an embodied robot provided by an embodiment of the present application. As Figure 4As shown, when the first group of people enters the detection range, a single robot A greets them based on their initial intent. During this process, a second group of people also enters the detection range. At this point, the robot closest to the second group is likely the single robot A currently greeting the first group. After locking onto the first group, robot A prepares to greet the second group. Robot A inquires or directly detects whether the second group requires a guided tour, generating second intent feature information based on the acquired audio and / or behavioral information of the second group. If the second intent feature information indicates that the second group requires a tour, the server's response informs the second group that a guiding robot B has been assigned, along with the estimated waiting time. If a screen is available to display robot status, the information on the screen is updated. If the second group is detected to be impatient, anxious, angry, or dissatisfied, comforting or explanatory actions or audio are triggered. After completing the preparations for greeting the second group, robot A leads the locked first group on a guided tour.
[0133] It should be noted that the aforementioned device linkage mapping table and intent behavior database can be updated during actual use.
[0134] As can be seen, by implementing the embodiments of this application, the robot controller obtains the first intent information of a first group of people based on a first reception instruction. The first group of people includes tourists currently within the crowd detection range, and the first intent information includes the guidance needs of the first group of people. It continuously performs crowd detection operations; if a second group of people is detected within the crowd detection range, it determines the second intent information of the second group of people, which includes the guidance needs of the second group of people. It sends the second intent information to the server and receives response information from the server. Based on the response information, it informs the second group of the guidance status information, which includes at least one of the guidance robot information and the estimated waiting time. If the guidance needs of the first group of people are that they require guidance, it leads the first group of people through a guidance operation based on the first intent information. The guidance operation includes guidance explanation operations and / or multimedia change operations, whereby the multimedia change operations include changing the display content of the display device. In this way, it can more quickly connect with tourists waiting to be served and link with other devices during the connection and explanation process, thereby improving device linkage, scheduling efficiency when receiving multiple groups of people, and enhancing the user experience.
[0135] Please see Figure 5 , Figure 5This is a schematic diagram of the structure of a guide device for an embodied robot according to an embodiment of this application. The guide device 500 for the embodied robot includes: an intent acquisition module 510, a crowd detection module 520, an intent sending module 530, a situation sending module 540, and a guide module 550. The system includes the following modules: an intent acquisition module 510, which acquires first intent information of a first group of people based on a first reception instruction. The first group includes visitors currently within the crowd detection range, and the first intent information includes the guide request of the first group. A crowd detection module 520 continuously performs crowd detection operations. If a second group is detected within the crowd detection range, it determines the second intent information of the second group, which includes the guide request of the second group. An intent sending module 530 sends the second intent information to the server and receives response information from the server. A status sending module 540 informs the second group of the guide status information based on the response information. The guide status information includes at least one of the guide robot information and the estimated waiting time. A guide module 550, which, if the guide request of the first group is to require a guide, leads the first group through a guide operation based on the first intent information. The guide operation includes a guide explanation operation and / or a multimedia change operation, whereby the multimedia change operation includes changing the display content of the display device.
[0136] In one possible embodiment, the crowd detection module 520, while continuously performing crowd detection operations, determines the second intent information of the second crowd if a second crowd is detected within the crowd detection range. Specifically, this is used for:
[0137] If the population is within the detection range, the population detection operation will continue to be performed;
[0138] If the second group of people is detected, the characteristics of the first group of people are identified, and the first distance between the first group of people and the second group of people is obtained;
[0139] The acquisition method of the second intent information is determined based on the first distance and a preset distance threshold, and the acquisition method is used to constrain the information source of the second intent information;
[0140] The audio information and / or behavioral information of the second group of people are obtained based on the acquisition method described above;
[0141] Intent feature information is extracted based on the audio information and / or the behavioral information;
[0142] The second intent information is determined based on the intent feature information and the intent behavior database.
[0143] In one possible embodiment, the single robot is equipped with a first sensor module, and the navigation system further includes a second sensor module. The information source includes the first sensor module and / or the second sensor module, and the acquisition method includes direct acquisition or indirect acquisition. The crowd detection module 520, in determining the acquisition method of the second intent information based on the first distance and a preset distance threshold, is specifically used for:
[0144] If the first distance is greater than the preset distance threshold, the method of obtaining the second intent information is determined to be the indirect acquisition.
[0145] If the first distance is not greater than the preset distance threshold, the method for obtaining the second intent information is determined to be the direct acquisition.
[0146] In acquiring the audio and behavioral information of the second crowd based on the acquisition method, the crowd detection module 520 is specifically used for:
[0147] If the acquisition method is direct acquisition, the audio information and / or behavioral information of the second group of people are acquired based on the first sensor module;
[0148] If the acquisition method is indirect acquisition, the audio information and / or behavioral information of the second group of people received from the second sensor module are acquired.
[0149] In one possible embodiment, the tour guide module 550, in guiding the first group of people through the tour based on the tour route and the tour guide skill library, is specifically used for:
[0150] Obtain the tour guidance intention information of the first group of people, the tour guidance intention information including the tour preferences of the first group of people;
[0151] Determine the tour route based on the tour intent information;
[0152] Based on the guided tour route and the guided tour skill library, the robot guides the first group of people through the target exhibition hall. The guided tour skill library is used by the robot to match specific behaviors based on external information.
[0153] In one possible embodiment, multiple linkage devices are installed in the target exhibition hall. These linkage devices are used for displaying text, sound, or images. The guide module 550, in guiding the first group of people through the tour based on the first intent information, is specifically used for:
[0154] The location information of the linkage device and the explanation area is detected in real time on the guided tour route. The location information of the explanation area is used to constrain the explanation content and behavior of the individual robot.
[0155] If the linkage device is detected, establish a communication connection with the linkage device based on the current explanation content;
[0156] The explanation actions are determined based on the location information of the explanation area and the guide skill library;
[0157] The control command is determined based on the current explanation content and / or the explanation action;
[0158] Send control commands to the linkage device to control the linkage device to perform linkage display operations.
[0159] In one possible embodiment, the linkage device includes multiple display screens and / or lighting devices and / or projection devices, and the guide module 550, in determining control commands based on the current narration content and / or the narration actions, is specifically used for:
[0160] If the linkage device includes the lighting device and / or the projection device, the positional relationship of the first group of people is detected, including the relative positional relationship with the individual robot; based on the positional relationship, the technical parameters of the lighting device and / or the projection device are determined, the technical parameters constraining at least one of the angle, brightness, color, and shape of the lighting device and / or the projection device; based on the technical parameters and the current explanatory content, the control command of the lighting device and / or the projection device is determined.
[0161] If the linkage device includes multiple display screens, the control commands for the lighting device and / or the multiple display screens are determined based on the current explanatory content.
[0162] In one possible embodiment, the tour guide module 550, in determining the control instructions for the lighting device and / or the projection device based on the technical parameters and the current narration content, is specifically configured to:
[0163] Based on the current content being explained, semantic decomposition is performed to identify multiple semantic keywords;
[0164] Based on the multiple semantic keywords, a pre-stored device linkage mapping table is queried to match multimedia action combinations associated with topic tags. The device linkage mapping table includes the topic tags and the multimedia action combinations corresponding to the topic tags. The multimedia action combinations are used to control the robot's actions according to the explanation content. The device linkage mapping table is pre-stored in the server and / or the robot controller.
[0165] Based on the multimedia action combination, a control instruction including spatio-temporal synchronization parameters is determined. The control instruction including spatio-temporal synchronization parameters is used to uniformly trigger multi-device actions through a time code synchronizer, so that the display screen and / or the lighting device and / or the projection device are superimposed and start synchronously when the explanatory words and / or the explanatory actions reach the anchor point.
[0166] It is worth noting that, among them, for the specific functional implementation method of the guiding device 500 of the embodied robot, refer to the description of the guiding method of the embodied robot shown above. Figure 2 For example, the intention acquisition module 510 is used to implement the relevant content of executing S210, the crowd detection module 520 is used to implement the relevant content of executing S220, the intention sending module 530 is used to implement the relevant content of executing S230, the situation sending module 540 is used to implement the relevant content of executing S240, and the guiding module 550 is used to implement the relevant content of executing S250. Each unit or module in the guiding device 500 of the embodied robot can be separately or all combined into one or several other units or modules to form, or some of the units or modules can be further split into multiple smaller units or modules in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above units or modules are divided based on logical functions. In practical applications, the function of one unit (or module) is implemented by multiple units (or modules), or the functions of multiple units (or modules) are implemented by one unit (or module).
[0167] It can be seen that the guiding device of the embodied robot described in the embodiments of the present application obtains the first intention information of the first crowd based on the first reception instruction. The first crowd includes the visiting crowd currently within the crowd detection range, and the first intention information includes the guiding requirements of the first crowd; continuously performs crowd detection operations. If a second crowd is detected within the crowd detection range, then determines the second intention information of the second crowd, and the second intention information includes the guiding requirements of the second crowd; sends the second intention information to the server and receives a response information from the server; based on the response information, informs the second crowd of the guiding situation information, and the guiding situation information includes at least one of guiding robot information and estimated waiting time; if the guiding requirement of the first crowd is to require guiding, then based on the first intention information, leads the first crowd to perform guiding operations, and the guiding operations include guiding explanation operations and / or multimedia change operations, and the multimedia change operations include changing the display content of the display device. It can achieve a faster connection with the visiting users to be served, and联动 with other devices during the connection and explanation processes, so as to improve the device linkage, improve the scheduling efficiency when receiving multiple batches of crowds, and improve the user experience.
[0168] In the case of using integrated units, please refer to Figure 6 , Figure 6 This is a schematic diagram of the structure of another embodied robot guide device provided in the embodiments of this application, as shown below. Figure 6 As shown, the guide device 500 for the embodied robot includes a processing module 502 and a communication module 501. The processing module 502 controls and manages the actions of the guide device 500, for example, executing the steps of the intent acquisition module 510, the crowd detection module 520, the intent sending module 530, the situation sending module 540, and the guide module 550, and / or performing other processes of the technology described herein. The communication module 501 is used for interaction between the guide device 500 and other devices. Figure 6 As shown, the guide device 500 of the embodied robot may also include a storage module 503, which is used to store the program code and data of the guide device 500 of the embodied robot.
[0169] The processing module 502 can be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication module 501 can be a transceiver, RF circuitry, or a communication interface, etc. The storage module 503 can be a memory.
[0170] All relevant content in each scenario involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here. The guide device 500 of the above-mentioned embodied robot can perform the above-mentioned... Figure 2 The guided tour method shown is for embodied robots.
[0171] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a robot proposed in an embodiment of this application, as shown below. Figure 7As shown, the robot 120 includes a processor 1210, a memory 1220, a communication interface 1230, and one or more programs 1221, which are stored in the memory 1220 and configured to be executed by the processor 1210.
[0172] The processor 1210, memory 1220, and communication interface 1230 are interconnected and perform communication between them.
[0173] The memory 1220 can be a volatile memory such as dynamic random access memory (DRAM) or a non-volatile memory such as a hard disk drive (HDD). The memory 1220 stores a set of executable program code, and the processor 1210 calls one or more programs 1221 stored in the memory 1220 to execute some or all of the steps of any of the guided tour methods for embodied robots described in the above embodiments.
[0174] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes a robot.
[0175] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include a robot.
[0176] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0177] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0178] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0179] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0180] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0181] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer robot (which may be a personal computer, a robot, or a network robot, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0182] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0183] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for guiding a somatic robot, characterized by, A robot controller applied to a tour guide system, the robot controller being used to control a single robot, the tour guide system further comprising a server, the server being communicatively connected with the robot controller, the method comprising: obtaining first intention information of a first crowd based on a first reception instruction, the first crowd comprising a tour crowd currently located in a crowd detection range, the first intention information comprising tour guide requirements of the first crowd; continuously performing a crowd detection operation, and determining second intention information of a second crowd if the second crowd is detected to be located in the crowd detection range, the second intention information comprising tour guide requirements of the second crowd; sending the second intention information to the server, and receiving response information from the server; informing the second crowd of tour guide situation information based on the response information, the tour guide situation information comprising at least one of tour guide robot information and an expected waiting time; if the tour guide requirements of the first crowd are to be guided, leading the first crowd to perform a tour guide operation based on the first intention information, the tour guide operation comprising a tour guide explanation operation and / or a multimedia changing operation, the multimedia changing operation comprising changing display content of a display device.
2. The method of claim 1, wherein, The continuously performing a crowd detection operation, and determining second intention information of a second crowd if the second crowd is detected to be located in the crowd detection range, comprises: continuously performing the crowd detection operation if the second crowd is detected to be located in the crowd detection range; locking crowd features of the first crowd and obtaining a first distance to the second crowd if the second crowd is detected; determining an obtaining manner of the second intention information based on the first distance and a preset distance threshold, the obtaining manner being used to constrain information sources of the second intention information; obtaining audio information and / or behavior information of the second crowd based on the obtaining manner; extracting intention feature information based on the audio information and / or the behavior information; determining the second intention information based on the intention feature information and an intention behavior database.
3. The method of claim 2, wherein, The single robot is provided with a first sensor module, the tour guide system further comprises a second sensor module, the information sources comprise the first sensor module and / or the second sensor module, and the obtaining manner comprises direct obtaining or indirect obtaining. The determining an obtaining manner of the second intention information based on the first distance and a preset distance threshold comprises: if the first distance is greater than the preset distance threshold, determining that the obtaining manner of the second intention information is the indirect obtaining; if the first distance is not greater than the preset distance threshold, determining that the obtaining manner of the second intention information is the direct obtaining. The obtaining audio information and behavior information of the second crowd based on the obtaining manner comprises: if the obtaining manner is the direct obtaining, obtaining the audio information and / or the behavior information of the second crowd based on the first sensor module; if the obtaining manner is the indirect obtaining, obtaining the audio information and / or the behavior information of the second crowd received from the second sensor module.
4. The method of claim 1, wherein, The leading the first crowd to perform a tour guide operation based on the first intention information comprises: Obtaining tour intention information of a first crowd, the tour intention information including tour preferences of the first crowd; Determining a tour route based on the tour intention information; Leading the first crowd in the target exhibition hall to perform the tour operation based on the tour route and a tour skill library, the tour skill library being used for the single robot to match specific behaviors based on external information.
5. The method of claim 4, wherein, The target exhibition hall is provided with a plurality of linkage devices for performing text, sound or picture display, and leading the first crowd to perform the tour operation based on the tour route and the tour skill library includes: Real-time detecting the linkage devices and explanation area position information on the tour route, the explanation area position information being used to constrain explanation content and behaviors of the single robot; If the linkage device is detected, establishing a communication connection between the current explanation content and the linkage device based on the current explanation content; Determining an explanation action based on the explanation area position information and the tour skill library; Determining a control instruction based on the current explanation content and / or the explanation action; Sending the control instruction to the linkage device to control the linkage device to perform linkage display operation.
6. The method of claim 5, wherein, The linkage device includes a plurality of display screens and / or light devices and / or projection devices, and determining the control instruction based on the current explanation content and / or the explanation action includes: If the linkage device includes the light device and / or the projection device, detecting a position relationship of the first crowd, the position relationship including a relative position relationship with the single robot; determining technical parameters of the light device and / or the projection device based on the position relationship, the technical parameters being used to constrain at least one of an angle, a brightness, a color and a shape of the light device and / or the projection device; determining the control instruction of the light device and / or the projection device based on the technical parameters and the current explanation content; If the linkage device includes a plurality of the display screens, determining the control instruction of the light device and / or the plurality of the display screens based on the current explanation content.
7. The method of claim 6, wherein, Determining the control instruction of the light device and / or the projection device based on the technical parameters and the current explanation content includes: Performing semantic decomposition based on the current explanation content to determine a plurality of semantic keywords; Querying a pre-stored device linkage mapping table based on the plurality of semantic keywords to match a multimedia action combination associated with a theme tag, the device linkage mapping table including the theme tag and the multimedia action combination corresponding to the theme tag, the multimedia action combination being used for the robot to perform action control with the explanation content, the device linkage mapping table being pre-stored in the server and / or the robot controller; Determining a control instruction with space-time synchronization parameters based on the multimedia action combination, the control instruction with space-time synchronization parameters being used to trigger multiple device actions uniformly through a time code synchronizer to make the display screens and / or the light devices and / or the projection devices start synchronously when explanation words and / or the explanation action reach an anchor point.
8. A guide device for a somatic robot, characterized by comprising: A robot controller applied to a guide system, the robot controller being used to control a single robot, the guide system further comprising a server, the server being communicatively connected with the robot controller, the apparatus comprising: an intention obtaining module, configured to obtain first intention information of a first crowd based on a first reception instruction, the first crowd comprising a visiting crowd currently located in a crowd detection range, the first intention information comprising guide demand of the first crowd; a crowd detection module, configured to continuously perform a crowd detection operation, and determine second intention information of a second crowd if the second crowd is detected to be located in the crowd detection range, the second intention information comprising guide demand of the second crowd; an intention sending module, configured to send the second intention information to the server, and receive response information from the server; a situation sending module, configured to inform guide situation information to the second crowd based on the response information, the guide situation information comprising at least one of guide robot information and expected waiting time; a guide module, configured to lead the first crowd to perform a guide operation based on the first intention information if the guide demand of the first crowd is to be guided, the guide operation comprising a guide explanation operation and / or a multimedia changing operation, the multimedia changing operation comprising changing display content of a display device.
9. A computer-readable storage medium, characterized in that, A guide program of a body-equipped robot, comprising execution instructions, when a processor of an electronic device executes the execution instructions, the processor executes the method in any one of claims 1 to 7.
10. A robot, characterized in that An electronic device comprising a processor and a memory storing execution instructions, the memory storing one or more programs; when the processor executes the execution instructions stored in the memory, the processor executes the method in any one of claims 1 to 7.
Citation Information
Patent Citations
Guiding method for machine room robot
CN109822581A
Navigation robot control method and device, electronic equipment and storage medium
CN112506197A