Multi-robot cooperative service control method driven by tourist behavior recognition
By transmitting visual continuity information between service robots in the scenic area, the problem of tourist identification in cross-field and occluded environments was solved, enabling continuous service control of multi-robot systems and improving the accuracy and timeliness of scenic area services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING KUNZHIHAOXIANG AVIATION TECHNOLOGY CO LTD
- Filing Date
- 2026-05-20
- Publication Date
- 2026-06-26
Smart Images

Figure CN122275006A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent robots, computer vision recognition, and smart scenic area service control technology, specifically a multi-robot linkage service control method driven by tourist behavior recognition. Background Technology
[0002] With the development of smart scenic areas, cultural tourism service robots, and mobile visual perception technology, scenic area services are gradually evolving from traditional manual patrols, fixed guide equipment, and single service terminals to robot patrols, video perception, visitor status analysis, and intelligent service scheduling. Among existing publicly available technologies, some solutions attempt to complete functions such as guidance, early warning, inquiry, and crowd control through scenic area visitor flow analysis, visitor behavior perception, or service robots. For example, patent application CN117592638A discloses a method and system for analyzing and managing visitor flow in scenic areas. Addressing the issues of visitor flow prediction and service resource allocation in scenic areas, it proposes analyzing and managing visitor flow based on the carrying capacity of scenic area service resources. It points out that traditional prediction methods based on historical data and experience have poor accuracy, while methods for predicting visitor numbers based on convolutional neural networks suffer from problems such as strong reliance on daily traffic and queuing data, high computational costs, and short response time for scenic areas. Furthermore, this type of visitor flow management often uses the carrying capacity of the scenic area or the number of visitors as the basis for judgment, easily leading to macro-level management focused on the group size, while failing to continuously identify and provide follow-up services for fine-grained service needs of individual tourists during their movement, such as seeking help, getting lost, lingering, deviating from the route, or approaching risky areas.
[0003] Furthermore, patent application CN115797873A discloses a crowd density detection method, system, device, storage medium, and robot. This method applies deep learning algorithms to a scenic area robot, using an image processing chip to detect crowd density and executing inspection or monitoring modes based on whether the flow of people is sparse or crowded. This is used for scenic area crowd monitoring, early warning, and guidance for evacuation. While this approach improves the robot's ability to perceive crowd density and provides decision-making support for scenic area safety management, it primarily focuses on crowd density, congestion levels, and warning limits, thus addressing the detection and early warning of regional crowd flow. For situations where the same target tourist is obstructed, leaves the field of view, moves across areas, or enters another robot's service area from another, existing solutions typically lack a visual continuity mechanism that encapsulates and transmits the tourist's visual appearance features, image location, spatial location, movement direction, behavioral sequence, and service behavior status. This can easily lead to problems such as unstable target identity continuity, cross-robot tracking interruptions, and inability to continuously determine service status.
[0004] Existing scenic area service robots or intelligent management systems still have the following shortcomings in practical applications: On the one hand, single-robot services are limited by their own visual acquisition range, occlusion environment, tourist density, and movement path. Once a target tourist is obscured by crowds, facilities, or buildings, or leaves the current robot's field of vision, the robot struggles to continuously grasp the tourist's subsequent behavioral state, resulting in the inability to reliably transmit previously identified service behaviors such as seeking help, waiting, lingering, and risky approach to subsequent service stages. On the other hand, even if multiple robots are set up in the scenic area, if the robots only share locations, assign tasks, or coordinate paths, without a succession matching rule based on the target tourist's visual behavior sequence and service behavior state, the second robot is prone to mismatch when faced with multiple candidate tourists who look similar, move in similar directions, or enter the field of vision simultaneously, leading to incorrect service targets, delayed service actions, or duplicate services. Furthermore, in existing technologies, after a service robot performs a service action, the service is usually terminated based on the issuance of an instruction or the completion of the action, lacking a closed-loop mechanism where a visual verification robot collects post-service video images and judges whether the service has truly been completed based on the tourist's post-service visual behavior. When tourists fail to respond, deviate from the service path, remain in risk areas, or the service behavior status changes to a higher risk status, the system struggles to update the linkage service control instructions in a timely manner. Summary of the Invention
[0005] The purpose of this invention is to provide a multi-robot linkage service control method driven by tourist behavior recognition, thereby solving some of the drawbacks and shortcomings pointed out in the background art.
[0006] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows: A multi-robot linkage service control method driven by tourist behavior recognition, comprising: a first robot acquiring a video image sequence of a target tourist, recognizing its visual behavior sequence and determining the service behavior state; and generating visual continuity information when the target tourist meets a preset visual continuity triggering condition.
[0007] The visual continuity information includes the visual appearance features of the target tourist, image location, target spatial location, direction of movement, visual behavior sequence, and service behavior status. The target spatial location is obtained by mapping the image location to coordinates. The visual continuity information is sent to the second robot, which uses this information to identify the corresponding target tourist in the video images it has collected, and continues to identify the target tourist's visual behavior sequence and service behavior status.
[0008] Based on the identified service behavior status, multi-robot linkage service control instructions are generated to control the participating robots to perform service actions.
[0009] Furthermore, when the second robot determines the corresponding target tourist, it matches the visual appearance features, image position, target spatial position, movement direction and visual behavior sequence of the candidate tourist with the corresponding information in the visual continuity information, and judges whether the service behavior status of the candidate tourist conforms to the preset state transition rules based on the service behavior status in the visual continuity information.
[0010] If the matching result meets the preset matching threshold and the state transition result conforms to the preset state transition rule, then the candidate tourist is determined as the target tourist.
[0011] Furthermore, when the preset visual continuity triggering condition includes the target tourist being obscured or leaving the field of vision of the first robot, the first robot extracts the last visual behavior segment of the target tourist before being obscured or leaving the field of vision and adds it to the visual continuity information; the second robot extracts the first visual behavior segment of the candidate tourist after entering its field of vision, and when the last visual behavior segment and the first visual behavior segment meet the preset behavior connection condition, it continues to identify the service behavior status of the candidate tourist.
[0012] Furthermore, when generating multi-robot linkage service control instructions, a service robot and a visual verification robot are determined from the first robot, the second robot, or the collaborative robot; the service robot is controlled to perform service actions, the visual verification robot is controlled to collect post-service video images, and when the post-service visual behavior of the target tourist is not found to meet the service completion conditions based on the post-service video images, the multi-robot linkage service control instructions are updated.
[0013] Furthermore, the preset state transition rules include a set of allowed transition states corresponding to the service behavior states; after the second robot identifies the service behavior state of a candidate tourist, it compares the service behavior state with the set of allowed transition states, and excludes the candidate tourist if the service behavior state does not belong to the set of allowed transition states.
[0014] Furthermore, the visual continuity information also includes the final visual behavior segment before the target tourist leaves the field of vision of the first robot; the second robot performs a behavior connection judgment between the initial visual behavior segment after the candidate tourist enters its field of vision and the final visual behavior segment; when the two meet the preset behavior connection conditions in terms of changes in body orientation, changes in movement direction and changes in action posture, it is determined that the candidate tourist meets the visual behavior matching conditions.
[0015] Furthermore, when multiple candidate tourists all meet the preset conditions of the matching results and state transition results, the second robot determines the candidate tourist ranked first in the comprehensive ranking as the target tourist based on the matching confidence level of each candidate tourist and the preset transition priority between service behavior states.
[0016] Furthermore, the post-service visual behavior is compared with the pre-service triggering behavior corresponding to the subsequent identified service behavior state; if the post-service visual behavior still includes the pre-service triggering behavior, or if no target completion behavior matching the service action is identified, it is determined that the post-service visual behavior does not meet the service completion conditions.
[0017] The target completion behaviors include one or more of the following: following the service robot, leaving the risk area, meeting up with fellow personnel, stopping wandering, stopping seeking help, and entering the target service area.
[0018] Furthermore, the reasons for service non-completion are determined based on the post-service video images. These reasons include one or more of the following: target tourist not responding, deviation from the service path, continued service behavior state, entry into a visually obstructed area, and service behavior state transitioning to a high-risk state. Based on the reasons for service non-completion, the role assignments, spatial positions, or service actions of the service robot and the visual verification robot are updated.
[0019] Furthermore, within a preset review period, the post-service video images are subjected to frame sequence recognition, and the number of reproduced frames of the pre-service triggered behavior and the number of continuous frames of the target completion behavior are counted respectively. When the number of reproduced frames reaches a preset reproduction threshold, or the number of continuous frames does not reach a preset continuous threshold, it is determined that the post-service visual behavior does not meet the service completion conditions.
[0020] This invention utilizes a first robot to continuously acquire video image sequences of target tourists and perform visual behavior recognition to determine the tourist's service behavior status. It generates visual continuity information when the tourist is obstructed, leaves the view, or enters a subsequent service area, enabling the transfer of the tourist's visual appearance features, image location, target spatial location, movement direction, visual behavior sequence, and service behavior status between different robots. Consequently, a second robot can accurately identify the corresponding target tourist from its acquired video images based on this multi-dimensional continuity information, avoiding misidentification issues caused by relying solely on single appearance or location features, and improving the stability of tourist continuity recognition across different fields of view, regions, and obstructed environments.
[0021] This invention combines visual behavior sequences with service behavior states for continuation judgment, enabling the second robot not only to confirm the target tourist's identity but also to continuously assess changes in their state, such as seeking help, lingering and observing, deviating from the path, approaching risks, or waiting for service. This avoids the loss of service status or interruption of the service process due to the limited field of vision of the first robot. By jointly matching the target's spatial location, direction of movement, and behavior sequence, this invention enhances the information collaboration capabilities among multiple robots and improves the accuracy of service object identification and service demand judgment in complex scenic environments.
[0022] This invention further generates multi-robot collaborative service control instructions based on the service behavior status after successive identification, controlling the participating robots to execute corresponding service actions. This transforms robot service from a single-machine independent response to a continuous collaborative response based on the target tourist's behavior status. This method can reduce duplicate, missed, and erroneous services, improve the timeliness and reliability of scenic area guidance, assistance response, risk alerts, and service follow-up, thereby enhancing the proactive service capabilities and closed-loop control effect of multi-robot systems in smart scenic area service scenarios. Attached Figure Description
[0023] Figure 1 This is a flowchart of the multi-robot linkage service control method driven by tourist behavior recognition according to the present invention.
[0024] Figure 2 This is a schematic diagram of the corridor deployment and visual continuity triggering in Embodiment 1 of the present invention.
[0025] Figure 3 This is a bar chart showing the continuity of candidate tourist behavior in Embodiment 1 of the present invention.
[0026] Figure 4 This is a comprehensive matching and ranking diagram of candidate tourists in Embodiment 1 of the present invention.
[0027] Figure 5 This is a schematic diagram of the deployment of three robots and two-wheeled linkage service in Embodiment 2 of the present invention.
[0028] Figure 6 This is a rating chart for the service and visual review roles in Embodiment 2 of the present invention.
[0029] Figure 7 This is the verification frame statistics and service completion determination diagram in Embodiment 2 of the present invention. Detailed Implementation
[0030] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0031] Combined with appendix Figure 1This invention discloses a multi-robot collaborative service control method driven by tourist behavior recognition. The first robot continuously performs visual tracking and behavior recognition processing on a target tourist. Specifically, the first robot acquires a video image sequence of the target tourist through its own visual acquisition device, and extracts visual behavioral features of the target tourist in consecutive frames according to time sequence, such as human orientation, head rotation, gait changes, dwelling state, limb movements, and relative relationship with the scene, thereby generating a corresponding visual behavior sequence. After obtaining the visual behavior sequence, the current service behavior state of the target tourist is determined by combining it with pre-established service behavior judgment rules. The service behavior judgment rules can be set based on the amplitude of changes in human orientation, changes in movement speed, dwelling time, changes in distance from risk areas or service areas, and the frequency of limb movements within several consecutive frames. For example, when the target tourist continuously faces the service robot within a preset time and lingers or waves, it can be determined as a state of seeking help or waiting for service. When the target tourist continuously approaches the boundary of a risk area without showing a tendency to leave, it can be determined as a state of approaching a risky area. When the target tourist repeatedly turns in multiple directions and moves back and forth over short distances, it can be determined as a state of lingering or observing. Service behavior status can be used to characterize whether the target tourist is currently in a state of seeking help, observing, forming an intention to follow, deviating from the path, approaching a risky area, waiting for service, or other behavioral states related to the robot's service response. When the target tourist meets the preset visual continuity trigger conditions, the first robot generates visual continuity information. The preset visual continuity trigger conditions may include situations such as the target tourist about to leave the effective field of view of the first robot, the target tourist entering an occluded area, an adverse change in the visual observation angle between the first robot and the target tourist, or the target tourist moving towards the service area of another robot. Leaving the effective field of view can be determined based on the continuous movement direction of the target tourist in the image edge area, the distance between the target box and the image boundary, and the proportion of the visible area. Entering an occluded area can be determined based on the continuous decrease in the visible area of the target box, the increase in the occlusion ratio of key parts, or the continuous decrease in tracking confidence. By generating visual continuity information at the above times, the behavior recognition process of the target tourist can be continuously transmitted between different robots, thereby reducing recognition interruptions caused by field of view switching.
[0032] Visual connection information includes the visual appearance features, image position, target spatial position, moving direction, visual behavior sequence, and service behavior status of the target tourist. Among them, the visual appearance features can include the clothing color distribution, human body contour features, local texture features, body shape proportion features, and gait characterization information of the target tourist, which are used to support the re-identification of the same target tourist among different robots. The image position is the position area of the target tourist in the current video image of the first robot, and the target spatial position is obtained by conversion according to the image position in combination with the pre-calibrated scene coordinate mapping relationship, so as to convert the target position in the image coordinates into the spatial position in the actual service scene. The scene coordinate mapping relationship can be established by the calibration parameters of the robot camera, the installation pose parameters, and the scene reference plane mapping relationship. When necessary, the target spatial position can also be corrected by combining depth information or binocular ranging results. The moving direction is used to characterize the current traveling trend of the target tourist, the visual behavior sequence is used to reflect the continuous behavior change process of the target tourist in a period of time before connection, and the service behavior status is used to characterize the recognition result of the current service demand or behavior risk of the first robot for the target tourist. After the first robot sends the visual connection information to the second robot, the second robot detects and compares the tourist objects that appear within its field of view based on the video images it collects, and determines the tourist object corresponding to the target tourist according to the received visual appearance features, spatial position, moving direction, visual behavior sequence, and service behavior status. After the second robot completes the determination of the corresponding target tourist, it continues to obtain the subsequent video image sequence of the target tourist, continues to identify its visual behavior sequence, and updates its service behavior status. Subsequently, multi-robot collaborative service control instructions are generated according to the service behavior status after connection recognition, and the robots participating in the collaboration are controlled to execute corresponding service actions. The service actions can include one or more of approaching and guiding, voice prompting, path leading, risk dissuasion, accompanying and following, information inquiry response, collaborative containment, or visual review. Through the above processing method, the visual connection recognition of the target tourist among different robots can be realized, and the multi-robot collaborative service control can be completed based on the connection recognition result.
[0033] After receiving the visual continuity information from the first robot, the second robot performs visitor detection on the currently captured video images, extracting candidate visitors within its field of vision. For each candidate visitor, it extracts visual appearance features, image location, target spatial location, movement direction, and visual behavior sequence. Subsequently, the second robot matches this information of each candidate visitor with the corresponding information in the visual continuity information. Specifically, matching visual appearance features determines the similarity between the candidate visitor and the target visitor in terms of clothing, body shape, and local texture; matching image location with target spatial location determines whether the candidate visitor is located within the area where the target visitor might appear after the continuity; matching movement direction determines whether the candidate visitor's movement trend is consistent with the target visitor's movement trend when leaving the first robot's field of vision; and matching visual behavior sequence determines whether the candidate visitor's action changes during the continuity period are connected to the target visitor's previous actions. After completing the above matching, the second robot further performs state transition judgment on the currently identified service behavior state of the candidate visitor based on the service behavior state in the visual continuity information to determine whether the candidate visitor meets the preset state transition rules. The matching process can be performed using a sub-scoring and weighted summarization method, where visual appearance features, spatial location, movement direction, and visual behavior sequence each correspond to a matching score. A comprehensive matching result is then generated based on preset weights. The preset matching threshold can be set to a value that the comprehensive matching result reaches, or a value that ranks highly among all candidate tourists and whose difference from the next best candidate tourist reaches a preset discrimination threshold. If a candidate tourist's matching result meets the preset matching threshold and its state transition result conforms to preset state transition rules, then the candidate tourist is identified as the target tourist, thereby achieving continuous and stable recognition of the same tourist object among different robots. By incorporating the above matching criteria, those skilled in the art can determine the target tourist based on appearance, location, direction, and behavioral continuity.
[0034] Furthermore, the preset state transition rules include a set of permissible transition states corresponding to the service behavior states. That is, the service behavior states determined by the first robot after completing the previous stage of identification correspond to a range of permissible service behavior states in subsequent time periods. After identifying the service behavior states of candidate tourists, the second robot compares these states with the set of permissible transition states. If a service behavior state does not belong to the set of permissible transition states, it indicates that although the candidate tourist may have some similarity to the target tourist in appearance or location, its behavioral state evolution process is inconsistent with the expected behavioral continuity of the target tourist, and therefore the candidate tourist is excluded. The set of permissible transition states can be pre-established according to the reasonableness of behavioral continuity. For example, when the previous stage service behavior state is a request for help state, the set of permissible transition states may include waiting for service, following guidance, short-term observation, or continuing to request help states; when the previous stage service behavior state is a risk approach state, the set of permissible transition states may include being dissuaded from staying, turning away, short-term observation, or continuing to approach at risk states; when the previous stage service behavior state is a path deviation state, the set of permissible transition states may include returning to the correct path, waiting for guidance, or continuing to deviate. By introducing transition constraints for service behavior states, we can avoid the misidentification problem that occurs when relying solely on appearance features or spatial location for target confirmation, while also providing a clear basis for state transition judgment.
[0035] Furthermore, the visual continuity information also includes the final visual behavior segment before the target tourist leaves the first robot's field of vision. This final visual behavior segment characterizes the target tourist's behavioral changes during the last period before the continuity, such as the trend of body orientation rotation, the trend of change in movement direction, and the continuity of action posture. After detecting a candidate tourist entering its field of vision, the second robot extracts the initial visual behavior segment after the candidate tourist enters its field of vision and performs a behavior continuity judgment between the initial and final visual behavior segments. When both satisfy preset behavior continuity conditions in terms of body orientation change, movement direction change, and action posture change, the candidate tourist is determined to meet the visual behavior matching conditions. Both the final and initial visual behavior segments can be composed of several consecutive frames before and after the continuity. The preset behavior continuity conditions may include one or more of the following: consistent or nearly consistent body orientation change trends, movement direction angle less than a preset angle threshold, consistent or continuous transition in action posture category, and gait rhythm changes not exceeding a preset fluctuation range. The above processing method can correlate the continuity of the target tourist's behavior before leaving with the continuity of the candidate tourist's behavior after entering, thereby further enhancing the credibility of identifying the same target tourist and reducing continuity ambiguity caused by short-term occlusion, local similar appearance, or multiple people overlapping.
[0036] In some implementations, when multiple candidate tourists simultaneously meet the preset conditions for matching results and state transition results, the second robot further sorts the candidate tourists. Specifically, the second robot calculates the comprehensive ranking result of each candidate tourist based on the matching confidence level and the preset transition priority between service behavior states. The matching confidence level comprehensively reflects the degree of matching in visual appearance features, spatial location, movement direction, and visual behavior. The preset transition priority characterizes the reasonableness and credibility of the evolution of different service behavior states from one stage to the next. The comprehensive ranking result can be generated by weighting the matching confidence level and the transition priority. When the comprehensive ranking results are the same or the difference is less than a preset ranking distinction threshold, the distance deviation between the candidate tourist and the predicted target spatial location point, the order in which they enter the second robot's field of vision, and the stability of continuous tracking can be further compared to determine the final result. After comprehensively considering the above factors, the second robot determines the candidate tourist ranked first in the comprehensive ranking as the target tourist. By using the above multi-factor joint ranking method, even in complex scenarios with multiple similar candidate objects, the uniqueness and reliability of the target tourist determination result can be improved, providing an accurate target basis for subsequent multi-robot linkage service control.
[0037] When the preset visual continuity triggering conditions include the target tourist being occluded or leaving the first robot's field of vision, the first robot reads and analyzes the continuous images of the target tourist before the occlusion or before leaving the field of vision, extracts the final visual behavior segment, and adds the final visual behavior segment to the visual continuity information and sends it to the second robot. The final visual behavior segment is used to characterize the continuous behavioral changes of the target tourist in the final stage before continuity, which may include the trend of changes in body orientation, the continuity of movement direction, the changes in gait rhythm, and the evolution of limb posture. By incorporating this final visual behavior segment into the visual continuity information, the second robot can not only match based on the target tourist's visual appearance features and spatial position information in the subsequent recognition process, but also make continuity judgments based on the continuity features of the behavior before leaving the first robot's field of vision, thereby improving the stability and accuracy of cross-robot visual continuity recognition. When extracting the final visual behavior segment, the first robot can jointly analyze the trajectory of human key points, motion vectors, and target box center trajectory in the continuous images, and retain the behavioral change data in the most recent period before the continuity trigger, so that the second robot can perform connection comparison.
[0038] After receiving visual continuity information, the second robot detects candidate tourists in the currently acquired video images and extracts initial visual behavior segments for candidate tourists entering its field of vision. These initial visual behavior segments characterize the continuous movement changes of the candidate tourist during the initial period after entering the second robot's field of vision. Subsequently, the second robot performs a behavior connection judgment between the initial and final visual behavior segments. When the two segments meet preset behavior connection conditions, the second robot continues to identify the service behavior status of the candidate tourist. These preset behavior connection conditions limit the consistency or continuity between the final and initial visual behavior segments in terms of changes in body orientation, movement direction, continuity of action posture, and changes in behavioral rhythm. If the two segments have a continuous connection in these aspects, it indicates a high probability that the candidate tourist corresponds to the target tourist previously identified by the first robot; therefore, the second robot continues to perform subsequent visual behavior recognition and service behavior status judgment. Through this processing method, even if the target tourist temporarily leaves the first robot's field of vision due to short-term occlusion or cross-regional movement, the system can still continuously identify the target tourist within the second robot's field of vision by utilizing the connection relationship between the preceding and following behavior segments. Furthermore, if the initial visual behavior segment after a candidate tourist enters the field of vision is not long enough to complete a stable judgment, several subsequent frames can be added to form an extended initial visual behavior segment before a connection judgment is made, thereby avoiding misjudgment due to incomplete target display in the initial stage of entering the field of vision.
[0039] When generating multi-robot collaborative service control instructions, the system determines service robots and visual verification robots from the first robot, the second robot, and collaborative robots based on the current position of the target tourist, the service behavior state after consecutive recognition, the distribution positions of each robot, the movement ability, the field of view coverage range, and the current task status. The service robot is used to directly perform service actions for the target tourist, and the visual verification robot is used to continuously obtain the post-service video images of the target tourist after the service actions are executed to visually confirm the service results. After completing the role determination, control the service robot to execute the corresponding service actions according to the multi-robot collaborative service control instructions. The service actions may include one or more of approaching and guiding, leading away along a path, voice prompting, risk dissuasion, accompanying and following, information guidance, or collaborative reminder. At the same time, control the visual verification robot to continuously collect video of the target tourist after the service and identify the post-service visual behavior of the target tourist based on the post-service video images. When the recognition result indicates that the post-service visual behavior does not meet the service completion conditions, update the multi-robot collaborative service control instructions to enable the robots participating in the collaboration to readjust the service execution method and improve the service closed-loop control ability. The determination of the service robot and the visual verification robot can be carried out by using a role scoring method. Among them, the service robot is preferably selected as the one that is closer to the target tourist, has a shorter time to reach the target position, has voice or interaction capabilities adapted to the current service behavior state, and the current task load is lower than the preset upper limit. The visual verification robot is preferably selected as the one that has continuous visual coverage of the area where the target tourist is located, has a large visual angle difference from the service robot, and does not affect the movement path of the service robot. When the same robot meets both the service and verification conditions at the same time, the robot can first execute the service actions and another collaborative robot can take over the verification, or when the robot has the ability to perform verification while providing service, it can assume both roles simultaneously. When updating the multi-robot collaborative service control instructions, at least update one of the service robot role, visual verification robot role, target standing position, service path, prompt content, or service action combination to form a closed-loop adjustment plan.
[0040] Furthermore, to determine whether the service is completed, the post-service visual behavior is compared with the pre-service triggering behavior corresponding to the subsequently identified service behavior state. The pre-service triggering behavior characterizes the original behavioral basis that triggers the robot to perform the service action, such as continuous lingering, obvious request for help, approaching a risk area, deviating from the predetermined travel direction, or stopping to wait for guidance. If the post-service visual behavior still includes the pre-service triggering behavior after the service action is performed, it indicates that the target tourist's original service needs or risk state have not been eliminated. If no target completion behavior matching the service action is identified, it is also determined that the current service result has not met the expected requirements. In this embodiment, the target completion behavior includes one or more of the following actions: following the service robot, leaving the risk area, meeting up with fellow travelers, stopping lingering, stopping requesting help, and entering the target service area. That is, when the target tourist exhibits completion behavior consistent with the performed service action after the service, it can be determined that the service has received a valid response. Conversely, when the target tourist maintains the original triggering behavior after the service, or does not exhibit the corresponding completion behavior, it is determined that the post-service visual behavior does not meet the service completion conditions. A pre-established correspondence can be established between service actions and target completion behaviors. For example, when a service action is to lead away from a risky area or dissuade someone from doing so, the corresponding target completion behavior could be leaving the risky area or turning to a safe passage. When a service action is to approach and guide, accompany, or provide information guidance, the corresponding target completion behavior could be following the service robot, entering the target service area, or stopping lingering. When a service action is to respond to an inquiry or provide a collaborative reminder, the corresponding target completion behavior could be stopping seeking help, meeting up with fellow travelers, or ending the waiting period. By establishing this correspondence between service actions and target completion behaviors, the determination of service completion conditions can have a clear basis for identification.
[0041] In some implementations, the system also determines the reason for service incompleteness based on post-service video images. Specifically, the reason for service incompleteness can be identified by combining the target tourist's movement trajectory, changes in body orientation, dwelling state, interaction response, and environmental location after the service action is performed. Reasons for service incompleteness include one or more of the following: target tourist not responding, deviating from the service path, continued service behavior, entering a visually obstructed area, and the service behavior state transitioning to a high-risk state. Specifically, a target tourist not responding can be characterized by not following, turning, or avoiding after the service robot issues a prompt or guidance; deviating from the service path can be characterized by the target tourist turning to another area instead of moving in the direction guided by the service robot; continued service behavior can be characterized by behaviors such as seeking help, lingering, or dangerous approach continuing after the service; entering a visually obstructed area can be characterized by the target tourist entering an obstructed environment, causing the verification and identification to be interrupted; and the service behavior state transitioning to a high-risk state can be characterized by the target tourist further approaching a dangerous area or exhibiting a higher level of abnormal behavior during the service process. After determining the reason for service incompleteness, the system updates the role assignments, spatial positions, or service actions of the service robot and the visual verification robot. For example, when a target tourist does not respond, a robot that is closer and has stronger interactive capabilities can be selected as a new service robot, with added voice prompts or repeated guidance actions. When a target tourist deviates from the service path, the positioning relationship between the service robot and the visual verification robot can be replanned to form a new guidance path and verification perspective. When a target tourist enters a visually obstructed area, other collaborative robots can be dispatched to supplement the visual coverage area. When the service behavior status is continuous, the verification time can be extended, and the service robot can be kept in front of or slightly to the side of the target tourist to provide continuous guidance. When the service behavior status shifts to a high-risk status, the number of collaborative robots involved can be increased, and the service actions can be switched to higher-priority risk dissuasion, perimeter control reminders, or safe removal actions. Through the above methods, the linkage service control strategy can be dynamically updated according to different reasons for non-compliance, thereby improving the service success rate and continuous control capability in complex scenarios.
[0042] Furthermore, within a preset review period, frame sequence recognition is performed on the post-service video images to count the number of recurring frames of the pre-service triggered behavior and the number of sustained frames of the target completion behavior. The number of recurring frames reflects the degree to which the pre-service triggered behavior reappears or persists after the service action is executed, while the number of sustained frames reflects the retention of the target completion behavior in consecutive image frames. When the number of recurring frames reaches a preset recurrence threshold, it indicates that the target tourist continues to exhibit the original triggered behavior after the service, suggesting insufficient service effectiveness. When the number of sustained frames does not reach a preset duration threshold, it indicates that although the target tourist may briefly exhibit the target completion behavior, this behavior has not formed a stable and continuous state, and the service cannot be considered effectively completed. Therefore, when the number of recurring frames reaches the preset recurrence threshold, or the number of sustained frames does not reach the preset duration threshold, the post-service visual behavior is determined not to meet the service completion conditions. The preset review period can be set according to the service action type and scene travel distance, and the preset recurrence threshold and preset duration threshold can be set according to the number of consecutive frames, cumulative frames, or percentage thresholds. For example, if a pre-service triggered behavior occurs continuously for a preset number of frames or accumulates to a preset proportion during the review period, it can be determined that the number of reproduced frames has reached a preset reproduction threshold. Similarly, if the target completion behavior continues to reach a preset number of frames or accumulates to a preset proportion during the review period, it can be determined that the number of continuous frames has reached a preset duration threshold. In cases where there is short-term occlusion or local recognition interruption in the post-service video image, fault tolerance processing can be performed on continuous frame statistics within a preset interruption tolerance range to avoid distortion of completion judgment due to momentary occlusion. By introducing a frame sequence statistics mechanism within the preset review duration, completion judgments can be avoided based solely on a single frame image or instantaneous action, thereby improving the stability, objectivity, and reliability of service result judgments and providing an accurate basis for subsequent updates to multi-robot collaborative service control instructions.
[0043] Example 1:
[0044] A company deployed two mobile robots on a connecting corridor outside the visitor service center of a scenic area. The first robot was positioned near the western entrance of the corridor, and the second robot was positioned near the eastern corner. There was a continuous coverage area of approximately 6 meters between the two robots. This coverage area contained landscape pillars and guide signs, which could cause temporary obstruction by tourists. Both robots were equipped with visible light cameras, positioning modules, and onboard processing units. The corridor floor was pre-calibrated to reflect planar coordinates, allowing image coordinates to be mapped to scene coordinates. For clarity, the target object identified by the first robot is designated as Visitor 1, and the three subsequent candidate objects detected by the second robot are designated as Visitor 2, Visitor 3, and Visitor 4. Figure 2 The diagram illustrates the corridor deployment, the continuous coverage area, the obstruction position of the landscape pillars, the initial movement trajectory of visitor 1, the relative arrangement of the first and second robots, and the spatial positional relationship when visual continuity is triggered.
[0045] The first robot continuously acquires video image sequences of the connecting corridor outside the visitor service center, with a frame rate of 20 frames per second. When visitor 1 enters approximately 7.2 meters in front of the first robot, the robot continuously extracts their orientation, head rotation, dwell time, gait changes, and limb movements. The visitor lingers near the service center's sign for 6.2 seconds, raising their hand twice to signal, accompanied by head rotations and short, stationary movements. Based on this, the first robot identifies the visitor's behavior as a request for assistance. Subsequently, visitor 1 moves along the corridor from west to east, with the target frame center gradually approaching the right edge of the image. Using pre-defined scene coordinate mapping, the target spatial position before leaving the first robot's main field of view is calculated to be 8.6 meters X-coordinate and 2.4 meters Y-coordinate, with a movement direction of 28° east-northeast relative to the corridor's main axis. When visitor 1 approaches the area obscured by a landscape pillar, the first robot detects that the visible area of the target decreases from 82% to 38%, and the target frame center continues to approach the image edge, thus triggering visual continuation processing. Figure 2 It can be seen that after Tourist 1 approaches the landscape pillar and enters the continuous coverage area, the first robot's ability to continuously observe Tourist 1 decreases significantly, thus meeting the visual continuity triggering condition.
[0046] Upon triggering visual continuity, the first robot writes the color distribution of tourist 1's clothing, body contour features, local texture features, gait characteristics, image position before departure, target spatial position, direction of movement, preceding visual behavior sequence, and request for help status into the visual continuity information. It also extracts 12 consecutive frames before tourist 1 leaves the field of vision to form a final visual behavior segment. In this final visual behavior segment, tourist 1's body orientation changes from 12° to 26°, and the direction of movement changes from 22° to 29°. The limb movement is characterized by raising and lowering the right hand while maintaining a posture facing the service area entrance. After receiving the visual continuity information, the second robot detects tourists 2, 3, and 4 in the current video image and extracts their visual appearance features, image position, target spatial position, direction of movement, initial visual behavior segment, and current service behavior status. Figure 2 The diagram also shows the process of visual continuity information being sent from the first robot to the second robot, as well as the distribution of candidate objects within the second robot's field of vision.
[0047] To determine whether the candidate object and Visitor 1 exhibit behavioral continuity before and after occlusion, the second robot first calculates the behavioral continuity for each candidate object. Since the most stable indicators of continuity for the same object in cross-robot transition scenarios are changes in body orientation, changes in movement direction, and continuation of action posture, these three factors are normalized and weighted summed. The behavioral continuity is defined as follows:
[0048]
[0049] In the formula, Indicates the first One candidate object, Indicates the first The degree of behavioral coherence among candidate objects This indicates the difference in body orientation change between the initial visual behavior segment of the candidate and the final visual behavior segment of tourist 1. This represents the difference in the direction of movement between two action segments. Indicates the consistency of action posture, with a value ranging from 0 to 1. This indicates the allowed body orientation difference threshold. This represents the threshold for the allowed difference in movement directions. , , These represent the weights of the three factors, and their sum is 1. In this embodiment, considering that the body orientation and movement direction before and after occlusion are more directly related to subsequent identification, we take... , , ,Pick , The reason for adopting this calculation method is that changes in body orientation and direction of movement can reflect the continuous movement trend of tourist 1 before and after leaving the field of vision of the first robot, and the consistency of action posture can reflect the continuity of local actions such as raising hands, standing, and looking to the side. The three together constitute the main basis for the connection of behavior before and after occlusion. Figure 3 The comparison results of tourist 2, tourist 3 and tourist 4 in this embodiment in terms of orientation continuity component, direction continuity component, consistency of action posture and behavior continuity are shown.
[0050] The second robot substitutes the values of tourists 2, 3, and 4 into the above calculation formula. In the 10 consecutive frames after entering the second robot's field of vision, tourist 2's body orientation change is 8°, its movement direction change is 10°, and its action posture consistency is 0.92. Therefore, its behavioral continuity is...
[0051]
[0052] Tourist 3's body orientation change is 10°, movement direction change is 12°, and action posture consistency is 0.87. Therefore, their behavioral coherence is...
[0053]
[0054] Tourist 4's body orientation change is 30°, movement direction change is 35°, and action posture consistency is 0.63. Therefore, their behavioral coherence is...
[0055]
[0056] Therefore, it can be seen that tourists 2 and 3 are quite similar to tourist 1 in terms of continuity of body orientation, direction of movement, and posture, while tourist 4's behavioral continuity is significantly weaker. Combined with... Figure 3 Further analysis reveals that Visitor 2 exhibits the highest overall behavioral continuity, with an orientation continuity component of 0.822, a directional continuity component of 0.833, and a motion posture consistency of 0.920. This indicates a stronger continuous correspondence between Visitor 2's initial visual behavioral segment and Visitor 1's final visual behavioral segment before leaving the first robot. While Visitor 3 also maintains high behavioral continuity, it remains lower than Visitor 2 overall. Visitor 4 shows significant deviations in body orientation and movement direction changes, resulting in a significantly lower behavioral continuity.
[0057] After calculating the behavioral continuity, the second robot further incorporates visual appearance features, spatial position, movement direction, and behavioral continuity into the comprehensive matching calculation. Since cross-robot continuity recognition does not rely solely on single appearance information, and positional continuity and behavioral continuity after occlusion are equally important, the four factors are uniformly normalized before calculating the comprehensive matching score. The comprehensive matching score is defined as follows:
[0058]
[0059] In the formula, Indicates the first The overall matching score of each candidate. Indicates visual appearance matching degree. Indicates the degree of matching of the target spatial location. Indicates the degree of matching in the direction of movement. This indicates the degree of coherence of the aforementioned behaviors. , , , These represent the weights of the corresponding items, and the sum of the four is 1. In this embodiment, considering the high recognizability of clothing appearance in the corridor scene, and the equally crucial connection between actions before and after occlusion, we take... , , , The reason for using this calculation method is that... It directly reflects the degree of similarity among the same tourists in terms of clothing color distribution, body shape, and local texture. This is used to constrain whether the candidate object falls within a reasonable area where visitor 1 appears after leaving the first robot's field of vision. Used to constrain whether the trend of movement continues. This supplements the continuity of behavior before and after occlusion, thereby making the confirmation of succession more stable. Figure 4The overall matching score ranking results of tourists 2, 3 and 4 in this embodiment are shown.
[0060] In this embodiment, the second robot obtained the following normalized results for the three candidate objects: Visual appearance matching degree of visitor 2. The target spatial location matching degree is 0.86. The matching degree in the direction of movement is 0.81. It is 0.90, combined with the above. Calculations yielded
[0061]
[0062] Visual appearance matching degree of tourist 3 The target spatial location matching degree is 0.81. The matching degree in the direction of movement is 0.85. It is 0.87, combined with the above. Calculations yielded
[0063]
[0064] Visual appearance matching degree of tourist 4 The target spatial location matching degree is 0.83. The matching degree in the direction of movement is 0.67. It is 0.71, combined with the above. Calculations yielded
[0065]
[0066] In this embodiment, the overall matching threshold is set to 0.75. Based on this, tourist 2 and tourist 3 meet the matching requirements, while tourist 4 is below the threshold and is initially excluded. Combined with... Figure 4 It can be seen that the comprehensive matching scores of Tourist 2, Tourist 3, and Tourist 4 are 0.852, 0.830, and 0.676, respectively. The ranking relationship between the candidates is relatively clear, with Tourist 2 ranking first, Tourist 3 ranking second, and Tourist 4 significantly below the matching threshold.
[0067] To avoid misidentification due to similar appearance and location, the second robot also performs service behavior state transition judgment on candidate objects that meet the matching threshold. Since the first robot identified tourist 1 as being in a request for help state before the connection, the set of allowed transition states corresponding to this state was pre-set as continuing to request help, waiting for service, and forming an intention to follow guidance. The second robot identified that tourist 2 was currently briefly pausing towards the service area entrance and observing the directional signs again, and determined its service behavior state as waiting for service. The second robot also identified that tourist 3, after entering the field of vision, still briefly raised his hand to signal and turned slightly in place, and determined its service behavior state as continuing to request help. Tourist 4, on the other hand, quickly crossed the corridor without any intention to linger, and determined its service behavior state as quickly leaving. Therefore, tourist 2 and tourist 3 both belong to the set of allowed transition states, while tourist 4 does not, and is therefore excluded.
[0068] When multiple candidate objects simultaneously meet the comprehensive matching threshold and all belong to the set of allowed transition states, the second robot continues to perform comprehensive sorting. In this embodiment, the transition priority from the help-seeking state to the waiting-for-service state is set to 0.92, the transition priority from the help-seeking state to the continuing-help-seeking state is set to 0.88, and the comprehensive matching score difference threshold is set to 0.03. Since the comprehensive matching scores of tourist 2 and tourist 3 are 0.852 and 0.830 respectively, and their difference is 0.022, which is less than 0.03, it is determined that they have a high probability of corresponding in terms of appearance, location, direction, and behavioral continuity. Therefore, a further state transition priority is introduced for sorting. The waiting-for-service state transition priority corresponding to tourist 2 is higher than the continuing-help-seeking state transition priority corresponding to tourist 3. Therefore, the second robot ultimately identifies tourist 2 as the target tourist corresponding to tourist 1 and continues to acquire its subsequent video image sequences to continuously identify its subsequent visual behavior. Figure 4 The overall matching score ranking results shown indicate that Visitor 2 not only has the highest overall matching score, but also has greater rationality in the subsequent state transition priority comparison.
[0069] In this embodiment, after the second robot completes target confirmation, it records the current position of tourist 2 as X coordinate 14.1m, Y coordinate 2.7m, and its movement direction remains approximately 31° east of north, maintaining continuity with the movement trend of tourist 1 before leaving the first robot's field of vision. Therefore, even under conditions where the landscape pillar causes temporary occlusion and multiple similar candidate objects appear simultaneously within the second robot's field of vision, by jointly constraining visual appearance features, the target spatial position corresponding to the image location, the movement direction, the connection relationship between the final and initial behavioral segments, and the service behavior state transition relationship, stable visual continuity recognition of the same tourist object between the first and second robots can be achieved, providing an accurate target basis for subsequent coordinated service control.
[0070] Example 2:
[0071] A company deploys three mobile robots at the fork formed by the intersection of the water-facing walkway and the main tourist route in the scenic area. Robot 1 is arranged at the west entrance of the main tourist route, Robot 2 is arranged in the central area of the fork, and Robot 3 is arranged on the side of the water-facing walkway close to the guardrail. All three robots are equipped with video acquisition devices, positioning modules, voice interaction modules, and vehicle-mounted processing units. The plane coordinate calibration of the scene ground has been completed in advance, and the target tourist status information and control instructions can be exchanged in real time between the robots. In this embodiment, after the previous sequential recognition, the target tourist has been continuously tracked by Robot 2. Robot 2 determines the current service behavior status of this tourist as approaching risk and deviating from the path. Specifically, Tourist 1 continuously stays about 1.4 m away from the water-facing guardrail and turns towards the water-facing walkway twice at the fork, without continuing to move forward along the main tourist route. Based on this status, the system generates the first-round multi-robot linkage service control instructions. The behavior set before service is set as continuously approaching the guardrail and deviating from the main tourist route, and the target completion behavior is set as leaving the risk area and following the service robot into the main tourist route. Figure 5 It shows the deployment positions of the three robots in this embodiment, the relative positions of Tourist 1 in the first and second rounds, the service paths in the first and second rounds, the visual verification paths, and the path relationship of Tourist 1 returning to the main tourist route.
[0072] When generating the first-round linkage control instructions, the system needs to determine the service robot and the visual verification robot from Robot 1, Robot 2, and Robot 3. Since the service robot needs to have a shorter arrival distance, a shorter arrival time, stronger interaction capabilities, and lower task occupancy priority, and the visual verification robot needs to have a higher field of view coverage and more stable tracking ability priority, the system uses a unified scoring structure to evaluate the roles of each robot. The role score is defined as:
[0073]
[0074] In the formula, represents the role score of the th robot, taking 1, 2, 3, represents the straight-line distance from the th robot to the current position of Tourist 1, represents the distance normalization reference value, represents the estimated time for the th robot to reach near Tourist 1 according to the current traffic conditions, represents the time normalization reference value, represents the field of view coverage of this robot for the area where Tourist 1 is located, with a value range of 0 to 1, This represents the robot's interaction capability coefficient, with a value ranging from 0 to 1. This represents the current task load factor, with a value ranging from 0 to 1. to These are the weighting coefficients, and the sum of the five factors is 1. The reason for using this calculation formula is that the determination of service robots and visual verification robots is related to distance, time, field of view, interaction, and load, but the importance of each factor is different. Therefore, the same scoring structure can be used and the different roles can be selected by adjusting the weights.
[0075] In this embodiment, the distance normalized baseline value is used during the first round of character selection. Time normalized baseline The distances of visitor 1 to robots 1, 2, and 3 are 6.4m, 3.8m, and 2.9m, respectively, with estimated arrival times of 5.3s, 3.1s, and 2.7s. Their field-of-view coverage is 0.46, 0.68, and 0.84, respectively; their interaction coefficients are 0.72, 0.91, and 0.78, respectively; and their task load coefficients are 0.20, 0.35, and 0.62, respectively. When selecting service robots, considering the need to balance proximity efficiency and interaction capability, [the robot is selected]. , , , , Substituting the values, we get:
[0076]
[0077] Therefore, Robot 2 was selected as the first service robot. When selecting a robot for visual verification, considering that the verification process places greater emphasis on field of view coverage and payload capacity, [the following was chosen]. , , , , Substituting the values, we obtain the review scores for Robot 1 as 0.544, Robot 2 as 0.680, and Robot 3 as 0.682. Therefore, Robot 3 is determined to be the first-round visual review robot. Based on these results, the system generates the first round of linkage control commands, controlling Robot 2 to execute voice prompts and guide the main tour route, controlling Robot 3 to continuously collect post-service video images on the side of the guardrail, and Robot 1 to remain on standby at the entrance of the main tour route and supplement the distant field of vision. Figure 6 The comparison results of the first round of service scores, the first round of review scores, the second round of service scores, and the second round of review scores in this embodiment are shown. The first round of service scores, from highest to lowest, are Robot 2, Robot 3, and Robot 1. The first round of review scores, from highest to lowest, are Robot 3, Robot 2, and Robot 1, which is consistent with the above role determination results.
[0078] After receiving the control command, Robot 2 approaches Visitor 1 at a speed of 0.8 m / s and plays a path prompt voice at a distance of approximately 1.2 m, guiding Visitor 1 to the main tour route. Robot 3, positioned beside the guardrail, continuously collects post-service video images from an oblique side view to verify the behavioral changes of Visitor 1 after the service action. The first round of verification is set to 8 seconds with a video frame rate of 20 frames per second, resulting in a total of 160 frames for the first round of verification. Within this verification period, the system identifies the post-service visual behavior frame by frame, counting the number of recurring frames of the pre-service triggered behavior and the number of sustained frames of the target completion behavior. To ensure a direct correspondence between the completion determination and the frame sequence statistics, the service completion determination is defined as:
[0079]
[0080] In the formula, This indicates the service completion determination result. This indicates that the service is complete. This indicates that the service was not completed. This indicates the number of frames in which the behavior triggered before the service was executed was reproduced. Indicates the number of frames that the target action lasts. Indicates the threshold number of frames to reproduce. This represents the threshold for the number of consecutive frames. The reason for using this criterion is that service completion requires not only a significant fading of the original risk-triggered behavior, but also that the completion behavior forms a stable and continuous state; therefore, both conditions must be met simultaneously. Below the reproducibility threshold and The conditions are: not lower than the sustained threshold. In this embodiment, Set to 30 frames per second. Set to 50 frames. Figure 7 This embodiment shows the number of frames for reproducing the pre-service triggering behavior, the number of frames for the target completion behavior, and the service completion determination result corresponding to the first and second rounds of review.
[0081] During the first round of review, the system detected that although Visitor 1 briefly turned their head after receiving the voice prompt from Robot 2, they did not continue moving in the guided direction. Furthermore, Visitor 1 turned towards the waterfront walkway twice and lingered near the railing. Therefore, the number of frames reproducing the pre-service triggered behavior was [not specified]. The statistics show 44 frames. Meanwhile, Tourist 1 only briefly moved towards the main tour route, failing to establish a stable, continuous state of following Robot 2. The number of frames for the completion of the target behavior is [not specified]. The statistics show 22 frames. Substituting the above data into the decision formula, since... and Therefore, we get , it is determined that the first-round service is not completed. Further analysis of the post-service video image shows that although Visitor 1 had a short reaction to the voice prompt, no effective response was formed, and their movement trajectory deviated towards the guardrail again. Therefore, the reason for the failure to complete this round is identified as the lack of an effective response and deviation from the service path. At the same time, the side-view image of Robot 3 shows that the distance between Visitor 1 and the guardrail once shortened to 1.1 m, indicating that the state of approaching risk continues. Combining Figure 7 It can be further seen that in the first-round review, the number of reproduced frames is higher than the reproduction threshold, and the continuous frames are lower than the continuous threshold. Therefore, the result of the first-round determination is that the service is not completed.
[0082] After the first-round service is not completed, the system updates the multi-robot collaborative service control instructions according to the reason for the failure to complete. Since the lack of an effective response and deviation from the service path indicate that simple voice guidance is insufficient, and Visitor 1 has further approached the guardrail, the system recalculates the service scores and review scores of each robot at the current moment. At this time, the distances between Visitor 1 and Robot 1, Robot 2, and Robot 3 are updated to 5.8 m, 3.4 m, and 1.6 m respectively, and the estimated arrival times are updated to 4.9 s, 2.8 s, and 1.5 s respectively. The remaining parameters remain unchanged corresponding to the current capabilities of the robots. After recalculating according to the weight of the first-round service robot selection, it can be obtained that
[0083]
[0084] Therefore, the service robot is switched from Robot 2 to Robot 3. Considering that Robot 3 should not be responsible for the main review task after switching to perform the close-range path leading action, after recalculating according to the review weight, the review score of Robot 2 is higher than that of Robot 1. Therefore, Robot 2 is changed to the new visual review robot, and Robot 1 moves forward about 1.5 m to the main tour route entrance to supplement the long-distance tracking view after Visitor 1 returns to the main route. The updated control instructions include that Robot 3 implements a close-range path leading from the guardrail side, Robot 2 moves to a position slightly to the left and behind Visitor 1 to continuously collect the post-service video image, and Robot 1 keeps the front route entrance area unobstructed and provides long-distance vision support. Combining Figure 6 As can be seen from the shown results, in the second-round service scores, Robot 1, Robot 2, and Robot 3 are 0.550, 0.722, and 0.751 respectively. Therefore, the service robot is switched to Robot 3; when comparing the second-round visual review scores between Robot 1 and Robot 2, Robot 1 and Robot 2 are 0.558 and 0.690 respectively. Therefore, the visual review robot is switched to Robot 2.
[0085] When Robot 3 performs the updated service action, it first cuts along the edge of the railing to a position approximately 0.9m in front of Visitor 1, then guides Visitor 1 away from the railing using a combination of voice prompts and slow movement. Robot 2 continuously collects video images after the service from a position slightly to the left of the rear, focusing on identifying whether Visitor 1 still lingers facing the railing, repeatedly turns towards the waterfront walkway, or deviates from the guided path. The second round of review is also set to 8 seconds, with the video frame rate remaining at 20 frames per second, for a total of 160 frames. During this review period, the system detects that the distance between Visitor 1 and the railing gradually increases to over 2.8m, and Visitor 1 continuously follows Robot 3 towards the main tour route. Pre-service triggered behaviors only occur sporadically, and the number of reproduced frames is [not specified]. The statistics show that the target completion behavior is manifested in leaving the risk area and stably following Robot 3 into the main tour route, with a duration of 12 frames. The statistics show 88 frames. Substituting this set of data into the aforementioned judgment formula, since... and Therefore, we get The service is deemed complete upon completion. (Combined with...) Figure 7 As can be seen, the number of reproduced frames in the second round of review was lower than the reproduction threshold, and the number of sustained frames was higher than the sustained threshold. Therefore, the result of the second round of review was that the service was completed.
[0086] After the second round of verification, the system updates Tourist 1's current status to indicate that they have left the risk area and resumed passage on the main route, ending the closed loop of this round of multi-robot collaborative service. Thus, based on the risk approach and path deviation status obtained after subsequent identification, the system first determines the first-round service robot and the visual verification robot through role scoring, executes service actions, and counts the number of frames of pre-service triggered behavior reproduction and the number of frames of target completion behavior duration. Then, based on the reasons for service incompleteness, the system reassigns roles, adjusts positions, and updates service actions, ultimately completing the closed-loop collaborative service control for Tourist 1.
Claims
1. A multi-robot linkage service control method driven by tourist behavior recognition, characterized by include: The first robot acquires video image sequences of the target tourists, identifies their visual behavior sequences, and determines the service behavior status. When the target tourist meets the preset visual continuity trigger conditions, visual continuity information is generated; The visual continuity information includes the visual appearance characteristics of the target tourist, image location, target spatial location, direction of movement, visual behavior sequence, and service behavior status. The target spatial location is obtained by mapping the image location to coordinates. The visual continuity information is sent to the second robot, which then identifies the corresponding target tourist in the video images it has collected, and continues to identify the target tourist's visual behavior sequence and service behavior status. Based on the identified service behavior status, multi-robot linkage service control instructions are generated to control the participating robots to perform service actions. 2.The tourist behavior recognition driven multi-robot coordinated service control method according to claim 1, characterized in that: When the second robot determines the corresponding target tourist, it matches the candidate tourist's visual appearance features, image position, target spatial position, movement direction and visual behavior sequence with the corresponding information in the visual continuity information, and judges whether the candidate tourist's service behavior status conforms to the preset state transition rules based on the service behavior status in the visual continuity information. If the matching result meets the preset matching threshold and the state transition result conforms to the preset state transition rule, then the candidate tourist is determined as the target tourist. 3.The tourist behavior recognition driven multi-robot coordinated service control method according to claim 1, characterized in that: When the preset visual continuity triggering condition includes the target tourist being obscured or leaving the field of vision of the first robot, the first robot extracts the last visual behavior segment of the target tourist before being obscured or leaving the field of vision and adds it to the visual continuity information; the second robot extracts the first visual behavior segment of the candidate tourist after entering its field of vision, and when the last visual behavior segment and the first visual behavior segment meet the preset behavior connection condition, it continues to identify the service behavior status of the candidate tourist.
4. The multi-robot linkage service control method driven by tourist behavior recognition according to claim 1, characterized in that: When generating multi-robot linkage service control instructions, a service robot and a visual verification robot are determined from the first robot, the second robot, or the collaborative robot; the service robot is controlled to perform service actions, the visual verification robot is controlled to collect post-service video images, and when the post-service visual behavior of the target tourist is not found to meet the service completion conditions based on the post-service video images, the multi-robot linkage service control instructions are updated.
5. The multi-robot linkage service control method driven by tourist behavior recognition according to claim 2, characterized in that: The preset state transition rules include a set of allowed transition states corresponding to the service behavior state; After the second robot identifies the service behavior status of a candidate tourist, it compares the service behavior status with the set of allowed transfer statuses. If the service behavior status does not belong to the set of allowed transfer statuses, the candidate tourist is excluded.
6. The multi-robot linkage service control method driven by tourist behavior recognition according to claim 2, characterized in that: The visual continuity information also includes the final visual behavior segment before the target tourist leaves the field of vision of the first robot; the second robot performs a behavior connection judgment between the initial visual behavior segment after the candidate tourist enters its field of vision and the final visual behavior segment; when the two meet the preset behavior connection conditions in terms of changes in body orientation, changes in movement direction and changes in action posture, it is determined that the candidate tourist meets the visual behavior matching conditions. 7.The tourist behavior recognition driven multi-robot coordinated service control method according to claim 2, characterized in that: When multiple candidate tourists meet the preset conditions of the matching results and state transition results, the second robot determines the candidate tourist ranked first as the target tourist based on the matching confidence of each candidate tourist and the preset transition priority between service behavior states.
8. The multi-robot linkage service control method driven by tourist behavior recognition according to claim 4, characterized in that: Compare the visual behavior after the service with the pre-service triggering behavior corresponding to the service behavior status after subsequent identification; If the post-service visual behavior still includes the pre-service triggering behavior, or if no target completion behavior matching the service action is identified, it is determined that the post-service visual behavior does not meet the service completion conditions. The target completion behaviors include one or more of the following: following the service robot, leaving the risk area, meeting up with fellow personnel, stopping wandering, stopping seeking help, and entering the target service area.
9. The multi-robot linkage service control method driven by tourist behavior recognition according to claim 4, characterized in that: The reasons for service non-completion are determined based on the post-service video images. These reasons include one or more of the following: target tourist does not respond, deviates from the service path, service behavior status remains unchanged, enters a visually obstructed area, or service behavior status changes to a high-risk state. Based on the reasons for service non-completion, the role assignments, spatial positions, or service actions of the service robot and the visual verification robot are updated.
10. The multi-robot linkage service control method driven by tourist behavior recognition according to claim 8, characterized in that: Within a preset review period, the video images after the service are subjected to frame sequence recognition, and the number of recurring frames of the pre-service triggered behavior and the number of continuous frames of the target completion behavior are counted respectively. When the number of reproduced frames reaches the preset reproduction threshold, or the number of continuous frames does not reach the preset continuous threshold, it is determined that the visual behavior after the service does not meet the service completion conditions.