A people flow tracking and intent prediction system and method for embodied intelligent camera networks

CN122715084APending Publication Date: 2026-09-08KUNG FU ROBOT (JIANGSU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610875902.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

[0014]本发明的目的在于提供一种具身智能相机网络的人流跟踪与意图预测方法、系统、设备及存储介质,以解决现有技术中机器人对动态人员缺乏前瞻性感知、相机之间难以形成全局连续轨迹、预测结果难以直接约束机器人规划以及关键观测资源难以动态调配的问题

Benefits of technology

第一,本发明通过构建相机拓扑图、全局坐标系和轨迹生命周期管理机制,将多个相机的局部观测结果融合为全局连续轨迹,因此能够在大范围作业空间内持续感知人员运动状态,提高对遮挡、离场和再出现情形的目标连续建模能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122715084A_ABST
    Figure CN122715084A_ABST
Patent Text Reader

Abstract

This invention discloses a system and method for pedestrian tracking and intent prediction using an embodied intelligent camera network. Distributed intelligent camera nodes are deployed within the work area. At the edge, pedestrian and robot target detection, feature extraction, and trajectory fragment generation are performed. The central processing unit combines camera topology relationships to perform cross-camera association and global trajectory fusion. Based on historical trajectories, scene semantics, and the robot's future motion state, it outputs the future location distribution of pedestrians, intent categories, and prediction uncertainties. This further forms a dynamic risk field for human-robot interaction safety, driving the robot to perform global path correction, local obstacle avoidance, speed adjustment, and coordinated early warning. This solution can improve human-robot collaborative safety, continuous perception capabilities, and planning foresight in dynamic crowd environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of embodied intelligent perception, human-machine collaborative control, and intelligent robot safety planning, and particularly to a method, system, device, and storage medium for pedestrian tracking and intent prediction using an embodied intelligent camera network.

[0002] More specifically, the present invention relates to a system-level technical solution that couples distributed fixed cameras, edge visual perception, cross-camera target continuous tracking, human motion intention prediction under semantic constraints, group behavior analysis, and mobile robot prediction-driven planning. Background Technology

[0003] With the rapid development of intelligent manufacturing, warehousing and logistics, park delivery, and semi-open service scenarios, it has become commonplace for mobile robots to move in the same space as on-site personnel. Robots not only need to complete tasks such as handling, inspection, resupply, and retrieval, but also need to continuously ensure passage efficiency and interaction safety in dynamically changing crowd environments.

[0004] Existing mobile robots typically rely on vehicle-mounted LiDAR, depth cameras, ultrasonic or millimeter-wave sensors to achieve obstacle detection and local obstacle avoidance. This type of solution is applicable to static obstacles at close range or dynamic obstacles with low speed, but its ability to perceive corners, the back of shelves, aisle intersections, and large areas where people gather is still limited due to limitations in installation height, field of view, occlusion, and changes in vehicle posture.

[0005] When personnel suddenly move through shelves, workstations, access control channels, or corner areas, such as walking in the wrong direction, stopping, meeting, or dispersing, relying solely on the robot's own sensors often only allows for reactive obstacle avoidance after the personnel have already entered the local perception range. This type of reactive obstacle avoidance is prone to problems such as frequent sudden stops, redundant detours, large speed fluctuations, unstable task rhythm, and poor human-machine interaction experience.

[0006] To enhance robots' global understanding of complex environments, existing technologies include solutions that utilize fixed cameras to monitor workshops or warehouses. These solutions are typically used for security recording, personnel counting, boundary detection, or area monitoring. Most cameras operate independently, and even with centralized video management, the video streams are often only used for manual review, lacking the ability to maintain long-term tracking of the same person's continuous trajectories across cameras.

[0007] Another approach attempts to assist robots in task control through multi-view visual information. For example, multiple cameras can be arranged around the control panel, workstation, or target equipment to help the robot identify the object to be grasped, understand task instructions, or estimate the next action. This type of approach focuses on the robot's operational accuracy or the generation of task action sequences. Its core object is usually the robot and its task target, rather than multi-person behavior modeling in an open and dynamic environment.

[0008] The prior art document CN121245791A cited in the search report discloses a robot control method based on large model combinations. This method predicts the robot's next action and generates control commands through multi-view images, task instructions, and reference calibration objects. From a technical perspective, this document focuses on solving the motion reasoning and control problems during the execution of a single robot task, and does not establish a continuous cross-view relay tracking mechanism for multiple personnel targets within a workshop.

[0009] Furthermore, the aforementioned comparative documents do not disclose the use of historical trajectory sequences, scene semantic relationships, and robot future paths as joint inputs to infer the future location distribution and behavioral intentions of personnel, nor do they disclose the generation of group behavior analysis results, dynamic risk fields, and prediction-driven path planning constraints based on multi-person target prediction results.

[0010] Even when conventional target detection, single-camera tracking, or simple area counting are mechanically combined, they can usually only provide the location and status of personnel at a certain moment. It is difficult to obtain a consistent trajectory representation across regions and time periods, and even more difficult to provide quantifiable future risk information before the planning time. Therefore, existing technologies still have significant gaps in the complete chain of "continuous perception - future prediction - proactive planning".

[0011] Furthermore, workshop and warehouse environments are generally characterized by narrow passageways, numerous intersections, frequent temporary enclosed areas, and personnel work behavior being significantly constrained by workstation rules. If the inherent semantic, rule-based, and temporary scheduling information of the scenario is not incorporated into the pedestrian flow modeling process, the intention prediction results are prone to inconsistencies with the scenario constraints, thereby weakening the guiding role for robot planning.

[0012] Furthermore, due to the large number of cameras on-site and the limitations of edge computing power and network bandwidth, it is difficult to control the overall system cost while ensuring the perception quality of key areas if key camera resources cannot be dynamically allocated based on robot task priorities and regional risk status. Existing solutions typically lack a dynamic viewpoint scheduling mechanism driven by robot tasks.

[0013] Therefore, a new technical solution is urgently needed to simultaneously achieve continuous tracking of human targets, reliable prediction of future intentions and human flow trends, timely identification of abnormal group behavior, and coordinated control of robot planning and viewpoint resources in dynamic mixed environments, thereby improving the safety and operational efficiency of human-machine collaboration. Summary of the Invention

[0014] The purpose of this invention is to provide a method, system, device and storage medium for people tracking and intention prediction using an embodied intelligent camera network, in order to solve the problems in the prior art such as the lack of forward-looking perception of dynamic people by robots, the difficulty in forming a global continuous trajectory between cameras, the difficulty in directly constraining robot planning with prediction results and the difficulty in dynamically allocating key observation resources.

[0015] To achieve the above objectives, this invention proposes an embodied intelligent camera network scheme for dynamic crowd work areas. Multiple distributed intelligent camera nodes are deployed in the work area. Each intelligent camera node detects human targets and mobile robot targets at the edge, performs pose estimation, feature extraction, and local trajectory segment generation, and only uploads compact target metadata instead of continuously uploading the original video stream, thus balancing timeliness and communication cost.

[0016] Furthermore, this invention maintains a global coordinate system, a camera topology map, and a scene semantic map at the central processing end. The camera topology map is used to describe the adjacency relationship, relay direction relationship, and coverage overlap relationship between cameras; the scene semantic map is used to describe semantic attributes such as passage areas, workstation areas, waiting areas, intersection areas, door restricted areas, no-entry areas, danger zones, and temporary closed areas.

[0017] During the global trajectory construction phase, the central processing unit performs cross-camera association based on the temporal, spatial, motion, and appearance information in the target metadata. Unlike matching based solely on appearance similarity, this invention incorporates topological reachability constraints, velocity continuity constraints, regional semantic consistency constraints, and trajectory lifecycle management into the matching process simultaneously to improve trajectory continuity in occlusion, short-term disappearance, and dense intersection scenarios.

[0018] In the prediction phase, this invention constructs a joint input for each person target, consisting of historical trajectory sequences, local scene semantics, future trajectories of neighboring robots, and group context. The intention prediction model then outputs the position distribution, intention category, and prediction uncertainty for multiple future time steps. The intention categories may include categories such as moving straight along a passage, approaching a workstation, crossing a passage, temporarily stopping, avoiding robots, following group movement, and leaving the current area.

[0019] At the group level, this invention jointly analyzes the future location distribution and historical continuous trajectories of multiple targets to generate time-varying pedestrian density maps, mainstream flow direction estimation results, clustering areas, congestion formation trends, and abnormal event labels. Abnormal event labels may include events such as reverse movement, abnormal lingering, rapid crossing, crowding, or obstruction of the robot's path, to further enhance the risk sensitivity of the planning.

[0020] During the planning phase, this invention constructs a dynamic risk field based on individual intention prediction results, group behavior analysis results, and prediction uncertainties, and applies this risk field simultaneously to the robot's global planning and local control. In the global planning phase, the robot adjusts its path cost based on the comprehensive risk of different regions within a future time window; in the local control phase, the robot makes real-time corrections to its speed, acceleration, turning angular velocity, or obstacle avoidance constraints based on short-term dynamic risks.

[0021] During the resource scheduling phase, this invention also provides a dynamic viewpoint scheduling mechanism. The central processing unit dynamically selects key observation cameras and reallocates sampling frequency, image quality, or re-identification resources based on robot task priority, current location, future path, camera load, network bandwidth, and pedestrian density distribution, thereby ensuring that critical path areas have higher quality perception and prediction support.

[0022] During the closed-loop update phase, the robot path planning module feeds back the planned trajectory, expected speed, and behavioral patterns for several future time points to the central processing unit. The central processing unit uses this feedback as additional input for predicting human intentions, forecasting the potential response of the human to the robot's movements, thereby constructing a human-environment-robot closed-loop collaborative model.

[0023] Optionally, the present invention can integrate audio-visual prompting devices, ground projection devices, on-site display terminals, or wearable prompting terminals into the system. When the prediction result indicates that the robot is about to enter a high-risk area, the system can not only adjust the robot's trajectory but also proactively output prompts to nearby personnel, thereby forming a two-way safety intervention mechanism.

[0024] Optionally, the present invention also supports an online correction mechanism. The system fine-tunes the matching threshold, risk weight, model parameters, or regional semantic weights online based on the deviation between the actual trajectory of the person and the predicted trajectory, the deviation between the actual trajectory of the robot and the planned trajectory, and the stability of the cross-camera association results.

[0025] Compared with the prior art, the present invention has at least the following beneficial effects: First, by constructing a camera topology map, a global coordinate system, and a trajectory lifecycle management mechanism, this invention integrates the local observation results of multiple cameras into a global continuous trajectory. Therefore, it can continuously sense the movement status of personnel in a large-scale work space and improve the ability to continuously model targets in situations of occlusion, departure, and reappearance.

[0026] Second, this invention integrates historical trajectory, scene semantics, the future motion state of neighboring robots, and group context for prediction. Therefore, it not only obtains the position information in the current state, but also the future position distribution, intention category, and uncertainty. Based on this, the robot can make planning corrections before the risk actually occurs, realizing the transformation from passive obstacle avoidance to prediction-driven active obstacle avoidance.

[0027] Third, this invention incorporates the results of group behavior analysis into the construction of a dynamic risk field. Therefore, the system not only focuses on the local risks of individual targets, but also identifies the impact of congestion trends, flow conflicts and abnormal events on the overall traffic environment, thereby more accurately constraining the behavior of robots in high-risk areas such as intersections and narrow passages.

[0028] Fourth, this invention provides a dynamic viewpoint scheduling mechanism, which can still prioritize the perception quality of key areas and key time periods in real-world environments where edge computing power and network bandwidth are limited, thereby improving the engineering feasibility and resource utilization efficiency of system deployment.

[0029] Fifth, by feeding back the robot's future trajectory to the central processing unit to predict the response of personnel, this invention can, to a certain extent, simulate the future movement trend of personnel after being influenced by the robot's behavior, thereby further improving the realism of human-computer interaction prediction and the executability of planning results. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of the overall architecture of the intelligent camera network-embedded people tracking and intention prediction system of the present invention.

[0031] Figure 2 This is a schematic diagram of the edge sensing, metadata generation and uploading process of the present invention.

[0032] Figure 3 This is a schematic diagram of the cross-camera association and global continuous trajectory fusion process of the present invention.

[0033] Figure 4 This is a schematic diagram of the personnel intention prediction model under semantic constraints of the present invention.

[0034] Figure 5 This is a schematic diagram of the group behavior analysis and abnormal event identification process of the present invention.

[0035] Figure 6 This is a schematic diagram of the predictive-driven hierarchical path planning for robots according to the present invention.

[0036] Figure 7 This is a schematic diagram of the viewpoint dynamic scheduling and resource reallocation process of the present invention.

[0037] Figure 8 This is a schematic diagram of the closed-loop collaborative timing of the inventor, environment, and robot. Detailed Implementation

[0038] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. Without departing from the concept of the present invention, those skilled in the art can make equivalent substitutions or modifications to the relevant structures, parameters, and implementation steps, and such substitutions or modifications should all fall within the scope of protection of the present invention.

[0039] In this specification, the term "central processing terminal" can be implemented by a central server, an edge server cluster, an industrial control computing platform, or a cloud-edge collaborative computing platform; the term "intelligent camera node" can be a fixed camera with a built-in AI processor, a gimbal camera, or a visual acquisition terminal with local inference capabilities; and the term "mobile robot" can be an autonomous mobile robot, an unmanned forklift, a delivery robot, an inspection robot, or other mobile devices that perform path planning in a dynamic environment.

[0040] For ease of explanation, some functional modules of the system are described using reference numerals in the following embodiments. For example, 100 represents a distributed intelligent camera network, 110 represents an intelligent camera node, 120 represents an edge perception module, 130 represents a target metadata encapsulation module, 200 represents a central processing unit, 210 represents a cross-camera association module, 220 represents a global trajectory fusion module, 230 represents a scene semantic management module, 240 represents an intent prediction module, 250 represents a group behavior analysis module, 260 represents a dynamic risk field generation module, 270 represents a viewpoint dynamic scheduling module, 300 represents a mobile robot, 310 represents a global path planning module, 320 represents a local obstacle avoidance control module, 330 represents an execution feedback module, and 340 represents a field prompting module.

[0041] Example 1: System overall architecture and deployment relationship.

[0042] like Figure 1 As shown, this embodiment deploys multiple smart camera nodes 110 in a smart manufacturing workshop, logistics warehouse, or industrial park distribution transfer area. Each smart camera node 110 is installed on the ceiling, pillars, workstation side, or at the corner of a passageway to form continuous coverage of key passages, intersections, access control points, workstation entrances, and robot operation paths.

[0043] Each smart camera node 110 includes an image acquisition unit, an edge perception module 120, a target metadata encapsulation module 130, and a communication unit. The edge perception module 120 preferably deploys a lightweight target detection network and a re-identification feature extraction network to output detection boxes, target categories, appearance feature vectors, and local trajectory fragments; the target metadata encapsulation module 130 is used to convert the recognition results into structured messages and upload them to the central processing terminal 200.

[0044] The central processing unit 200 communicates with the scene semantic management module 230, the cross-camera association module 210, the global trajectory fusion module 220, the intent prediction module 240, the group behavior analysis module 250, the dynamic risk field generation module 260, and the viewpoint dynamic scheduling module 270. The mobile robot 300 exchanges planning status, risk constraints, and execution feedback with the central processing unit 200 via wired or wireless networks. The on-site prompting module 340 receives early warning triggering information from the central processing unit 200 and uses it to output audio-visual prompts, projection prompts, or display terminal prompts to on-site personnel.

[0045] In this embodiment, the scene semantic management module 230 maintains an electronic map, which includes not only geometric boundaries and obstacle information, but also semantic attributes such as channel direction, workstation function, passable time window, robot priority area, restricted area, temporary construction area, and personnel hotspot. By unifying the semantic map with the global coordinate system, the system can directly reference regional semantic features during the cross-camera fusion and prediction stages.

[0046] Example 2: Edge perception and target metadata generation.

[0047] like Figure 2 As shown, each smart camera node 110 acquires image data at a preset frame rate. The edge perception module 120 first performs human and mobile robot target detection. The detection network can adopt YOLO series networks, Transformer-based detection networks, or other lightweight detection models suitable for edge deployment.

[0048] After detection, the system extracts appearance feature vectors, 2D skeleton key points, head and shoulder orientation, gait rhythm, or posture category information for each target. For mobile robot targets, robot orientation, loading status, or operational status labels can also be extracted to subsequently determine the interaction relationship between the robot and the human.

[0049] The edge-side tracking unit generates local trajectory segments based on adjacent detection results within a short time window and calculates the stability score for each trajectory segment. If a target is continuously observed within the current camera's field of view, the local trajectory segments are continuously updated; if the target is about to leave the field, the departure position, departure velocity, and departure time are recorded.

[0050] The target metadata encapsulation module 130 encapsulates the timestamp, camera identifier, target category, bounding box coordinates, keypoint coordinates, position estimate mapped to the global coordinate system, velocity estimate, orientation estimate, appearance feature vector, detection confidence, and local trajectory identifier before sending them to the central processing unit 200. By transmitting only metadata instead of continuously transmitting the original video stream, bandwidth consumption can be significantly reduced and deployment scalability enhanced.

[0051] In another embodiment, when the edge perception module 120 detects a decrease in the confidence level of a target in an intersection area, a robot's inevitable passage area, or a high-density area, it can temporarily upload low-bitrate image fragments or cropped area images for the central processing unit 200 to review, thereby improving the perception reliability in key scenarios while ensuring communication efficiency.

[0052] Example 3: Cross-camera association and global trajectory fusion.

[0053] like Figure 3 As shown, after receiving the target metadata uploaded by each camera node, the cross-camera association module 210 first filters candidate association relationships based on the camera topology map. If two cameras have no spatial adjacency, no overlapping field of view, or no reasonable crossing path, the system does not directly perform relay matching on their detection results to reduce false associations.

[0054] For each target entering the candidate set, the system calculates an appearance similarity score, a motion continuity score, a displacement reachability score, and a region semantic consistency score. The appearance similarity score reflects the distance relationship between two detected targets in the re-identification feature space; the motion continuity score comprehensively characterizes the consistency between the target's departure velocity, arrival velocity, and motion direction; the displacement reachability score reflects whether the target can move from one position to the next within a given time interval; and the region semantic consistency score reflects whether the target's motion direction matches the channel direction or region connectivity.

[0055] In an example implementation, the system can construct a comprehensive matching score S, satisfying S = w1 × Sa + w2 × Sm + w3 × St + w4 × Ss, where Sa represents the appearance similarity score, Sm represents the motion continuity score, St represents the displacement reachability score, Ss represents the region semantic consistency score, and w1 to w4 represent the corresponding weights. When S is greater than a preset threshold, the two local trajectory segments are merged into the same global trajectory.

[0056] The global trajectory fusion module 220 maintains a global trajectory list. Each global trajectory corresponds to a unique global target identifier, trajectory lifecycle status, position sequence, velocity sequence, and historical semantic region sequence for the most recent time steps. When a target disappears from the field of view of all cameras within a short period of time, the global trajectory does not terminate immediately but enters a pending confirmation state. If the target reappears in a topologically adjacent camera within an adjacent time window, the original global identifier is continued to be used to avoid trajectory interruption caused by short-term occlusion.

[0057] To improve the stability of associations in complex scenarios, the system can also perform special processing for target intersections, multiple people walking in close parallel proximity, and targets entering and exiting through doorways or shelves. For example, when two people have high similarity in appearance but different speeds and directions of travel, the system corrects the association results by using historical trajectory continuity relationships and regional semantic constraints.

[0058] Example 4: Scene semantic modeling and semantic map maintenance.

[0059] In this invention, the scene semantic map is not merely a regular grid map or topological map, but a multi-layered map encompassing spatial semantics, rule semantics, and task semantics. The spatial semantic layer describes the locations of passageways, workstations, shelves, access control systems, zebra crossings, buffer zones, and temporary obstacles; the rule semantic layer describes rules such as one-way traffic, no-entry, deceleration, priority, and personnel waiting areas; and the task semantic layer describes key work areas, temporarily closed areas, robot task hotspots, and personnel work hotspots for the current shift.

[0060] Scene semantic maps can be constructed through manual annotation or converted from BIM models, digital twin maps of the factory area, or existing navigation maps for robots. For temporary semantic information, such as temporary construction areas, short-term storage areas, equipment maintenance areas, or manually fenced areas, it can be written in real time by the upper-level management system or automatically identified and written by the camera network.

[0061] When the trajectories of human and robot targets are projected onto the scene semantic map, the system can obtain the semantic relationship between the target and the area, such as descriptions like "located in the main passage," "approaching the workstation entrance," "about to cross the intersection," or "entering near the boundary of the robot-dedicated passage." This type of relationship directly participates in the construction of the input for the intent prediction model.

[0062] Example 5: Predicting People's Intent under Semantic Constraints.

[0063] like Figure 4As shown, the intent prediction module 240 constructs a prediction sample for each person target to be predicted. A prediction sample includes at least: the position sequence, velocity sequence, orientation sequence, gait change state within the historical time window, relative distance to neighboring robots, semantic region, set of semantic regions reachable ahead, and contextual information of the neighboring people group.

[0064] In a preferred implementation, the trajectory temporal coding branch uses a Transformer encoder or a gated recurrent network to encode the target's historical trajectory sequence, the scene semantic coding branch uses a graph neural network, convolutional network, or embedded lookup table method to encode the region's functional attributes, topological connectivity, and rule constraints, and the robot interaction coding branch encodes the future paths, speed change trends, and encounter times of neighboring robots.

[0065] After fusing the aforementioned multi-branch features, the output includes a location probability heatmap, an intent category probability vector, and an uncertainty estimate for multiple future time steps. The location probability heatmap represents the probability of the target appearing in different spatial locations in the future; the intent category probability vector represents the type of behavior the target may take; and the uncertainty estimate reflects the model's confidence in the current prediction results.

[0066] In this embodiment, the preferred intent categories include continuing along the passage, slowing down and stopping, approaching the workstation, crossing the passage, turning into an adjacent branch, avoiding robots, following the flow of the group, and leaving the current area. For different scenarios, the set of intent categories can also be redefined according to industry needs, such as defining specific behavior categories like queuing, waiting for elevators, and entering service counters in airport, hospital, or shopping mall environments.

[0067] During the model training phase, location regression loss, heatmap reconstruction loss, intent classification loss, and uncertainty regularization loss can be used in combination. During the online inference phase, the system updates the prediction results at a fixed frequency and only sends personnel prediction information related to the robot's future planning range to the robot to control the communication burden.

[0068] Considering that human behavior may be influenced by robot movement, this invention preferably incorporates the robot's future planned trajectory as additional input. For example, when the robot will pass in front of a person within the next two seconds, the person's future trajectory is usually different from the future trajectory when the robot is stationary. Incorporating the robot's future motion state into the model input can significantly improve the predictive accuracy in close human-robot interaction scenarios.

[0069] Example 6: Group behavior analysis and abnormal event identification.

[0070] like Figure 5As shown, the group behavior analysis module 250 takes the global trajectory of multiple targets, their future location distribution, and the probability of their intention categories as input, and periodically generates a time-varying density map and a direction consistency map. The time-varying density map is obtained by superimposing the future location probabilities of each target; the direction consistency map is used to reflect whether the mainstream movement directions of multiple targets in the same area are consistent.

[0071] The system can perform cluster analysis on time-varying density maps to identify clustered areas, potential congestion areas, and high-flow intersection areas. If an area maintains high density and low directional consistency over multiple future time steps, it can be determined that the area has flow direction conflicts or congestion risks.

[0072] For anomaly event identification, the system can identify at least one or more of the following behaviors: moving in the opposite direction in a one-way channel, lingering in a high-frequency robot passageway for an extended period, suddenly accelerating across the robot path, multiple people crowding and blocking workstation entrances, and loitering near restricted areas. The anomaly event identification results may include estimates of the event level, scope of impact, and duration.

[0073] The results of group behavior analysis are not only used to construct risk fields, but can also be fed back to the management system for on-site safety supervision. For example, if the system continuously detects that a high density of intersections will occur in the next few seconds, it can issue manual control suggestions in advance or adjust the robot task allocation strategy to reduce local conflicts.

[0074] Example 7: Construction of dynamic risk field.

[0075] The dynamic risk field generation module 260 generates a time-varying risk field based on the future location distribution of each person's target, the risk level of their intention category, the trend of group density, the labels of abnormal events, and the uncertainty of prediction. The risk field can be in raster form, continuous function form, or graph structure node weight form.

[0076] In one example implementation, the risk value R of a spatial location at a predicted time can be determined by the following factors: the probability of a single person's future location, the risk level of a single person's intention, the population density, the anomaly event, and the uncertainty correction. The higher the probability of a single person's future location, the more likely the intention is to lead to an interaction with a robot, the higher the population density in the area, the higher the level of the anomaly event, or the greater the prediction uncertainty, the higher the corresponding risk value.

[0077] By introducing uncertainty into the risk field, the system can avoid adopting overly aggressive path-crossing strategies when model predictions are unstable. In other words, when the system lacks sufficient confidence in a person's future behavior, the robot can automatically switch to a more conservative planning mode to enhance safety redundancy.

[0078] Example 8: Prediction-driven hierarchical path planning for robots.

[0079] like Figure 1 and Figure 6 As shown, Figure 1 The mobile robot 300 includes a global path planning module 310, a local obstacle avoidance control module 320, and an execution feedback module 330. Figure 6 The prediction-driven hierarchical path planning process implemented based on the aforementioned module is illustrated. The global path planning module 310 generates candidate global paths based on the map static cost, semantic rule cost, and dynamic risk field cost; the local obstacle avoidance control module 320 generates control commands based on the short-term risk field, robot dynamics constraints, and the current state of surrounding obstacles.

[0080] During the global planning phase, this invention can employ the A* algorithm, Dijkstra's algorithm, Hybrid A* algorithm, search-based time-extended planning algorithm, or other graph search planning methods. Unlike existing solutions, this invention maps the pedestrian flow risk field within the future time window into path cost terms, causing the planning results to tend to bypass areas of impending congestion, rather than simply avoiding currently existing pedestrian obstacles.

[0081] In the local control phase, this invention can employ dynamic windowing, model predictive control, velocity obstacle avoidance, or a combination thereof. The local controller adjusts the velocity, angular velocity, acceleration limit, and yielding strategy based on short-term prediction results. For example, if there is a high probability that a person will cross in front of the robot within two seconds, the robot can decelerate in advance instead of stopping abruptly after the person enters the emergency braking distance.

[0082] When the dynamic risk field indicates that the current global path will pass through a high-risk area within a future time window, the robot can trigger a local or global replanning. If the risk stems from temporary congestion and is short-lived, the robot can choose a low-speed waiting strategy; if the risk stems from continuous abnormal clustering, the robot can switch to an alternative route. By distinguishing between risk type and persistence, the system can reduce unnecessary large-scale detours.

[0083] Example 9: Dynamic Viewpoint Scheduling and Resource Reallocation.

[0084] like Figure 7 As shown, the viewpoint dynamic scheduling module 270 acquires the robot's task priority, future path, camera load, bandwidth utilization, and pedestrian risk distribution, and then performs hierarchical management of camera nodes. Camera nodes located in the robot's key future passage areas, high-risk interaction areas, or complex intersection areas are marked as key observation cameras.

[0085] For key observation cameras, the system can increase the sampling frame rate, improve the coding bit rate, enable a higher precision re-identification model, or increase the local buffer length; for cameras far from the robot's task path and with sparse crowds, the system can appropriately reduce the sampling frequency or use a lighter edge model to concentrate resources on areas that truly affect the planning quality.

[0086] In another embodiment, if the key observation camera malfunctions, is blocked, or experiences network congestion, the viewpoint dynamic scheduling module 270 can quickly select a replacement camera based on the camera topology map and appropriately increase the resource allocation of the replacement camera, thereby maintaining the continuous observation capability of the key area.

[0087] Example 10: Human-Environment-Robot Closed-Loop Collaboration.

[0088] like Figure 1 and Figure 8 As shown, the mobile robot 300 feeds back its future planned trajectory, target speed, and behavior pattern to the central processing unit 200 in each planning cycle. Based on this, the central processing unit 200 updates its predictions of human responses to the robot's behavior and recalculates the risk field. The updated risk field and key observation camera strategy are then fed back to the mobile robot 300 and the intelligent camera node 110. When the predicted risk exceeds a preset threshold, the central processing unit 200 can also trigger the on-site prompting module 340 to output prompt information to surrounding personnel. Simultaneously, the central processing unit 200 can perform online corrections based on execution feedback, thus forming a collaborative closed loop of perception, prediction, planning, execution, prompting, and correction.

[0089] The significance of this closed-loop mechanism lies in the fact that human behavior is not fixed but reacts to robot movements. For example, after observing a robot slowing down to yield, a person might speed up to cross the intersection; upon observing a robot approaching at high speed, they might temporarily stop and wait. By incorporating the robot's future trajectory into human predictions, the system can more closely resemble real-world interactions.

[0090] Example 11: On-site coordinated early warning and multi-level safety response.

[0091] When the risk level output by the dynamic risk field generation module 260 exceeds the preset threshold, in addition to sending commands to the robot to slow down, detour, or stop, the system can also use the on-site prompt module 340 to link with sound and light devices, ground projection devices, voice broadcasting equipment, or signboard terminals to remind nearby personnel to pay attention to the robot's passage or to warn of the risk of congestion ahead.

[0092] Different risk levels can be addressed with different handling mechanisms. For example, medium-risk events can be handled with speed reduction and voice prompts; high-risk events can be handled with stopping, detours, and projected warnings; and extremely high-risk events can be simultaneously reported to the workshop management system and trigger manual intervention requests. This multi-level handling mechanism avoids a one-size-fits-all approach to all risks, improving the balance between production pace and safety.

[0093] Example 12: Online calibration and continuous model optimization.

[0094] During system operation, the actual trajectory of personnel, cross-camera correlation results, actual robot trajectory, and risk handling results can be recorded as online samples. Based on these online samples, the system periodically corrects the appearance matching threshold, topology reachability threshold, risk weight, and intent prediction model parameters.

[0095] In a preferred embodiment, when a certain type of error continues to increase, the system adjusts the corresponding weights accordingly. For example, if prediction deviations occur frequently in intersection areas, the weights of group context and robot interaction features in the model can be increased; if false associations occur frequently in severely occluded areas, the weights of topology consistency constraints and trajectory lifecycle verification can be increased.

[0096] Example 13: Typical Application Scenario 1.

[0097] In logistics and warehousing scenarios, robots need to move toy boxes back and forth along the main aisle, while personnel frequently move between side aisles and workstations. Because the entrances to side aisles are obstructed by shelves, robots relying solely on their own sensors struggle to detect personnel about to cross. With this invention, a camera node above the side aisle entrance can observe personnel exit trends in advance. Based on this, the system predicts that the person will enter the main aisle within two seconds and writes this prediction into the robot's risk field. The robot then reduces its speed and adjusts its yielding strategy accordingly.

[0098] Example 14: Typical Application Scenario 2.

[0099] In a smart manufacturing workshop, multiple workstations are distributed on both sides of the same main aisle. Personnel often move materials in groups and briefly stop and communicate at intersections. The group behavior analysis module 250 can identify the aggregation trend and flow conflict trend at the intersection in advance. Based on this, the global planning module 310 increases the passage cost of the intersection in the next few seconds, causing the robot to choose to detour through the secondary aisle, thereby avoiding passive queuing or frequent emergency stops.

[0100] Example 15: Typical Application Scenario 3.

[0101] In park delivery or hospital logistics delivery environments, service recipients may randomly stop, walk in groups, or meet temporarily in passageways. Since the semantics of these behaviors are not entirely equivalent to industrial scenarios, this invention can redefine semantic regions and intent categories to incorporate service counters, elevator entrances, waiting areas, or access control points into a semantic map, enabling the same technical framework to be extended to a wider range of human-machine mixed environments.

[0102] Example 16: A preferred time configuration for the method flow.

[0103] In a non-limiting example, the edge perception module 120 outputs target metadata at a frequency of 10 to 25 frames per second, cross-camera association and trajectory fusion are updated at a frequency of 5 to 10 Hz, the intent prediction module 240 outputs prediction results for a time window of 2 to 5 seconds at a frequency of 5 Hz, and the robot local control module 320 executes control command updates at a frequency of 10 Hz or higher. Of course, the specific frequency can be flexibly configured according to the scene size, robot speed, and hardware capabilities.

[0104] Example 17: A preferred data structure for the method flow.

[0105] In a preferred data structure, each global trajectory record includes a global target identifier, the position, velocity, direction, semantic region, neighboring robot identifiers, predicted intent category, prediction confidence, and state update time for the most recent N time points. Through this unified data structure, the central processing unit 200, the group behavior analysis module 250, and the robot planning module can efficiently exchange information, facilitating engineering implementation.

[0106] Example 18: System calibration and coordinate unification.

[0107] Before the system goes live, intrinsic and extrinsic parameter calibrations can be performed on each smart camera node 110, and a mapping relationship between each camera's field of view and the global coordinate system can be established. For environments with ground elevation differences, ramps, or local platforms, a 3D coordinate correction model can be used to improve the projection accuracy of the target position. If the on-site layout changes, such as moving shelves, modifying workstations, or adding isolation barriers, the scene semantic management module 230 will synchronously update the geometric boundaries and semantic labels of the corresponding areas to ensure that cross-camera association and intent prediction are still based on the latest environment.

[0108] Example 19: Reliability Degradation and Fault-Tolerant Operation.

[0109] When the system detects that some cameras are offline, image quality has significantly degraded, network links are congested, or the computing power of some modules in the central processing unit is insufficient, it can automatically enter a degraded operation mode. In degraded operation mode, the system prioritizes camera perception and risk calculation within the robot's future path coverage area, reducing the refresh rate for non-critical areas; at the same time, it increases the safety margin and conservative weights in local planning, enabling the robot to maintain acceptable safe operation capabilities under limited perception conditions. Therefore, this invention does not rely on all cameras always operating at the highest quality, but rather possesses the fault tolerance required for engineering deployment.

[0110] Example 20: Data closed loop and continuous iteration.

[0111] This invention also establishes a continuous optimization closed loop from on-site operation to model iteration. Specifically, the system uses indicators such as cross-camera association success rate, prediction error, abnormal event handling results, robot emergency stops, detours, and travel time as operational evaluation indicators, and periodically generates model evaluation reports. Researchers or maintenance personnel can use the evaluation results to retrain and redeploy the detection model, re-identification model, intent prediction model, risk field weights, and semantic rules offline. Through this continuous iteration mechanism, the system can gradually improve prediction accuracy and planning stability as the scene changes and data accumulates.

[0112] Example 21: Privacy Protection and Industrial Implementation.

[0113] When implemented in industrial settings, an edge-side anonymization strategy can be adopted, uploading only structured target metadata, anonymized target identifiers, and necessary cropped image segments, rather than storing the complete video long-term, thereby reducing privacy and data compliance risks. For video evidence that must be retained, controlled caching can be implemented based on event level and time window. This implementation demonstrates that the present invention not only features algorithmic innovation but also addresses the practical needs of factories, warehouses, and public service spaces regarding network load, operational costs, and data compliance.

[0114] Example 22: Further comparison with the prior art.

[0115] Compared with multi-view control schemes that focus on individual robot movements, the technical focus of this invention is not on generating single-step movements for the robot, but on establishing a complete technical chain around dynamic crowd environments, which includes "continuous perception across cameras - semantically constrained future prediction - group risk modeling - prediction-driven planning - dynamic viewpoint scheduling".

[0116] First, this invention focuses on human targets and human-robot interaction risks in its modeling, rather than simply using cameras as auxiliary sensors for robot object identification. Second, the camera network in this invention establishes a truly global continuous trajectory management mechanism through topology graphs, global coordinates, and trajectory lifecycles, rather than simply using multi-view parallel sampling. Third, this invention outputs position distribution, intent category, uncertainty, and group risk, rather than just the robot's next action control input.

[0117] Furthermore, this invention directly writes the prediction results into the dynamic risk field and applies them to hierarchical planning, addressing the issues of early avoidance and overall traffic efficiency in complex dynamic environments. Simultaneously, this invention also achieves resource-level optimization through dynamic viewpoint scheduling. These multiple features support each other, jointly achieving system-level synergistic effects, and are not simply independent stacks of each other.

[0118] In summary, this invention not only enhances the ability to continuously perceive dynamic crowds, but also enables robots to proactively avoid obstacles, facilitate smooth passage, and provide coordinated early warnings using future risk information. It has clear engineering application value and a basis for invention patent protection.

[0119] The above are merely preferred embodiments of the present invention. Those skilled in the art can make various substitutions, deletions, or modifications to the module settings, algorithm combinations, parameter configurations, and communication methods of the present invention without departing from the essential spirit and technical concept of the present invention. All such substitutions, deletions, or modifications should be included within the protection scope of the present invention.

[0120] Explanation of reference numerals in the attached figures 100 Distributed Smart Camera Network; 110 Smart Camera Nodes; 120 Edge Sensing Modules; 130 Target Metadata Encapsulation Modules; 200 Central processing unit; 210 Cross-camera association module; 220 Global trajectory fusion module; 230 Scene semantic management module; 240 Intent Prediction Module; 250 Group Behavior Analysis Module; 260 Dynamic Risk Field Generation Module; 270 Viewpoint Dynamic Scheduling Module; 300 Mobile robot; 310 Global path planning module; 320 Local obstacle avoidance control module; 330 Execution feedback module; 340 On-site prompting module.

Claims

1. A method for pedestrian tracking and intent prediction using an embodied intelligent camera network, characterized in that, The steps include the following: S1. Deploy multiple distributed intelligent camera nodes within the work area, establish a camera topology map with known coverage relationships, and construct a global coordinate system and scene semantic map corresponding to the work area; each intelligent camera node collects image data within its field of view in real time. S2. At the edge side of each of the smart camera nodes, perform human target and mobile robot target detection, target re-identification feature extraction, pose estimation and local trajectory segment generation on the image data, and form target metadata including timestamp, camera identifier, target category, two-dimensional or three-dimensional position, velocity, pose, appearance feature vector and detection quality score; S3. The central processing unit receives the target metadata uploaded by each of the intelligent camera nodes. Based on the camera topology map, time continuity constraints, displacement reachability constraints, motion continuity constraints, appearance similarity constraints, and regional semantic consistency constraints, it performs association matching on the same target detected across cameras, generates a global target identifier, and merges multiple local trajectory segments into a global continuous trajectory. S4. For each person target, extract the historical time window trajectory sequence from the global continuous trajectory, and combine it with the regional functional attributes, passage rule attributes, obstacle constraint attributes and the future motion state of neighboring robots related to the person target in the scene semantic map, input it into the intention prediction model, and output the position distribution, intention category probability and prediction uncertainty of the person target in the future time window. S5. Based on the global continuous trajectory, future location distribution and intention category probability of multiple personnel targets, perform group behavior analysis to obtain personnel gathering areas, mainstream flow direction, congestion trend, reverse movement events, abnormal lingering events or sudden crossing events. S6. Based on the future location distribution, intention category probability, prediction uncertainty, group behavior analysis results, and the current and future planned states of the mobile robot, construct a dynamic risk field for robot path planning. S7. Input the dynamic risk field into the robot path planning module, adjust the regional cost of the candidate path in the global planning stage, adjust the obstacle avoidance constraints, speed limit and turning strategy in the local planning stage, and feed back the updated robot future planning state to the central processing terminal to form a closed-loop update for personnel intention prediction and viewpoint scheduling.

2. The method according to claim 1, characterized in that, In step S3, when performing cross-camera target association matching, candidate camera sets are first filtered based on the camera topology map. Then, time continuity thresholds and displacement reachability thresholds are established based on the target's departure time, departure position, departure speed in the previous camera and the target's entry time and entry position in the next camera. Appearance similarity score, motion continuity score, and regional semantic consistency score are calculated. Only when the comprehensive matching score is greater than the preset threshold, the corresponding local trajectory segments are merged into the same global target identifier.

3. The method according to claim 1, characterized in that, In step S4, the intent prediction model adopts a dual-branch network structure that combines a trajectory temporal coding branch and a scene semantic coding branch. The trajectory temporal coding branch is used to encode the position, speed, heading changes and stop status in the historical trajectory, while the scene semantic coding branch is used to encode the channel area, workstation area, intersection area, restricted area and danger area. After feature fusion, the model simultaneously outputs the position probability heatmap, target travel intent category probability and prediction variance or confidence interval for multiple future time moments.

4. The method according to claim 1, characterized in that, In steps S4 and S6, the scene semantic map includes at least the area boundary, the connectivity between adjacent areas, the one-way passage rule, the safe distance rule, the human-machine mixed passage priority rule, and the information of the temporary closed area; the dynamic risk field is jointly constructed by the probability of the future location of the personnel, the risk level of the personnel's intention, the group density weight, the abnormal event weight, and the prediction uncertainty weight.

5. The method according to claim 1, characterized in that, In step S5, the group behavior analysis includes: generating a time-varying density map based on the future location distribution of personnel; calculating a directional consistency index based on trajectory direction vectors; identifying clustered areas and congestion-forming areas based on clustering results; and generating abnormal event labels and writing the abnormal event labels into the dynamic risk field when backtracking, prolonged lingering, sudden crossing of the robot's planned path, or area density exceeding a threshold is detected.

6. The method according to claim 1, characterized in that, It also includes a viewpoint dynamic scheduling step: the central processing unit dynamically selects at least one key observation camera node based on the robot task priority, personnel density in the target area, camera load status, network bandwidth status, and the overlap relationship of each camera's field of view. It improves the sampling frequency, encoding quality, or re-identification computing resource allocation of the key observation camera node, and reduces the sampling frequency of non-key observation camera nodes to ensure the continuity of perception in key areas and the overall system resource utilization efficiency.

7. The method according to claim 1, characterized in that, In step S7, when the dynamic risk field characterization robot is about to enter a high-risk interaction area within a preset time window, the central processing terminal issues a deceleration, detour, or stop command to the robot, and links the audio-visual prompt device, ground projection device, or on-site display terminal to issue a prompt to nearby personnel; after the event is resolved, the associated matching parameters, intent prediction model parameters, or risk field weights are corrected online based on the robot's actual execution trajectory and the personnel's real trajectory.

8. A pedestrian tracking and intent prediction system with an integrated intelligent camera network, characterized in that, The system comprises a distributed intelligent camera network, a central processing unit, a scene semantic management module, a group behavior analysis module, a dynamic risk field generation module, a viewpoint dynamic scheduling module, and at least one mobile robot. The distributed intelligent camera network includes multiple intelligent camera nodes, each with an edge perception module and a target metadata encapsulation module configured at its edge for acquiring image data of the work area and generating target metadata. The central processing unit is configured with a cross-camera association module, a global trajectory fusion module, and an intent prediction module for performing cross-camera association, global trajectory fusion, and personnel intent prediction. The scene semantic management module maintains a semantic map and camera topology map of the work area. The group behavior analysis module outputs aggregation areas, flow directions, and abnormal event labels. The dynamic risk field generation module constructs a dynamic risk field based on prediction results. The viewpoint dynamic scheduling module adjusts camera resource allocation based on the robot's task status and camera load status. The mobile robot performs global path correction, local obstacle avoidance, and speed control based on the dynamic risk field and provides feedback on future planning status or execution feedback to the central processing unit. The system is configured to perform the method described in any one of claims 1 to 7.

9. An electronic device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Robot control method and device based on large model combination and electronic equipment

    CN121245791A