A pan-tilt fast tracking security warning robot instruction execution system
Patent Information
- Application Number
- CN202610680646.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-25
AI Technical Summary
该方式触发条件单一,不能准确反映目标朝向图像边界的运动趋势和底盘周向障碍情况,容易出现两类问题:一是底盘补偿触发过晚,目标已经脱离云台视野;二是底盘补偿触发过早或过频繁,导致机器人运动不稳定,并增加与周边障碍物或人员发生干涉的风险
本发明不是仅依据目标在图像中的当前位置生成云台控制指令,而是针对各警戒目标构建警戒意图传播图,将目标侵入警戒扇区的深度、目标朝向核心警戒位置的位移、云台剩余转角、底盘周向可通行距离以及目标脱离视野风险共同纳入传播计算,得到各警戒目标的跟踪意图值。由此能够在多目标同时出现时优先选择更需要跟踪、更可能产生警戒风险或更易丢失的当前主跟踪目标,避免云台在多个目标之间无序摆动,提高警戒响应的针对性和及时性。
Smart Images

Figure CN122632901A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of security robot technology, and more specifically to a gimbal-based rapid tracking command execution system for security surveillance robots. Background Technology
[0002] With the increasing demand for unmanned patrols and proactive surveillance in scenarios such as industrial parks, factories, warehousing and logistics facilities, substations, boundary fences, parking areas, and key entrances and exits, security surveillance robots are gradually being used to replace or assist human personnel in tasks such as mobile patrols, abnormal target detection, on-site evidence collection, and remote alarms. Existing security surveillance robots are typically equipped with a mobile chassis, camera devices, a gimbal mechanism, environmental perception sensors, and a controller. They patrol and move via the mobile chassis, and the gimbal drives the camera devices to turn towards the target area. Target detection algorithms are used to identify people, vehicles, or other abnormal moving targets.
[0003] In existing technologies, gimbal tracking typically generates gimbal rotation commands based on the deviation of the target's center from the image's center. When the target deviates from the image center, the controller controls the gimbal to rotate in the target's direction, bringing the target back to the vicinity of the image center. This method can meet basic monitoring needs in scenarios with fixed cameras or targets with low movement speeds, but it still has significant shortcomings in security and surveillance robot scenarios. Because the robot itself has a mobile chassis, the gimbal tracking process is affected not only by the target's movement but also by changes in chassis posture, gimbal mechanical limitations, surrounding obstacles, and the robot's traversable space. If gimbal control commands are still generated solely based on the target image deviation, the gimbal may repeatedly follow the target's historical position, leading to tracking lag or target loss.
[0004] Furthermore, existing multi-target tracking methods typically prioritize targets based on factors such as target category, detection confidence, or distance, failing to comprehensively reflect the degree to which targets intrude into the warning zone, the risk of targets leaving the gimbal's field of view, the remaining gimbal rotation angle, and the feasibility of chassis compensation. For example, one target, although having a high warning level, may be located in the center of the image and tracked stably; another target, although having a lower warning level, may be rapidly moving towards the image boundary and approaching the gimbal's limit. Traditional prioritization methods struggle to promptly identify the risk of losing the latter target, potentially causing the robot to miss the optimal tracking and compensation opportunity.
[0005] Furthermore, existing collaborative control of the gimbal and mobile chassis often employs a method where the chassis turning is triggered only after the gimbal angle exceeds a threshold. This method has a single trigger condition and cannot accurately reflect the target's movement trend toward the image boundary or the presence of obstacles around the chassis. This easily leads to two types of problems: first, the chassis compensation is triggered too late, by which time the target has already left the gimbal's field of view; second, the chassis compensation is triggered too early or too frequently, resulting in unstable robot movement and increasing the risk of interference with surrounding obstacles or people.
[0006] Furthermore, when a target is briefly obscured by pillars, vehicles, trees, or other objects, existing systems typically employ expanded scanning or re-detection for target re-acquisition. This method has a large search range, slow response, and is prone to mistaking other similar targets for the original target, affecting the continuous evidence collection and tracking stability of alert events. Therefore, there is an urgent need for a security alert robot command execution system that can integrate target alert intent, gimbal field of view boundary risks, chassis compensation capabilities, and short-term target memory loss, in order to improve the response speed, continuity, and reliability of gimbal tracking. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides a gimbal-based rapid tracking command execution system for a security surveillance robot. The system acquires continuous image frames, the current angle of the gimbal, and the distance to circumferential obstacles. It extracts the center coordinates, scale changes, direction of motion, and confidence level of the target. A warning intent propagation map is constructed, including the target, warning sectors, gimbal margin, chassis accessibility, obstacle suppression, and loss risk nodes, yielding a tracking intent value. Boundary pressure values are generated based on the target's distance to the image boundary, inter-frame displacement towards the boundary, remaining gimbal rotation angle, and circumferential accessible chassis distance. The current primary tracking target is determined, and gimbal advance rotation angle, segmented rotation speed curve, and chassis compensation rotation angle are generated. When the confidence level decreases, a target afterimage sector is generated and recaptured, improving tracking stability in fast-moving and short-term occlusion scenarios.
[0008] This application provides a gimbal-based rapid tracking security and surveillance robot command execution system, including a gimbal imaging device, a mobile chassis, an environmental sensing device, a gimbal angle detection device, and an edge computing controller; The gimbal imaging device is used to acquire continuous image frames, the gimbal angle detection device is used to obtain the current angle of the gimbal, and the environmental perception device is used to obtain the chassis attitude and the distance to circumferential obstacles. The edge computing controller is used to determine the remaining turning angle of the gimbal based on the current angle of the gimbal, determine the circumferential passable distance of the chassis based on the circumferential obstacle distance, and extract the center coordinates, scale changes, motion direction and confidence of each warning target from continuous image frames; The edge computing controller is also used to construct a warning intent propagation map for each warning target, including target node, warning sector node, gimbal margin node, chassis passage node, obstacle suppression node, and loss risk node, to obtain the tracking intent value; to obtain the gimbal field of view boundary pressure value based on the distance from the warning target to the image boundary, the inter-frame displacement towards the image boundary, the remaining gimbal rotation angle, and the chassis circumferential passage distance; to determine the current main tracking target based on the tracking intent value and the gimbal field of view boundary pressure value, and to generate the gimbal advance rotation angle, the gimbal segmented rotation speed curve, and the chassis compensation rotation angle; The edge computing controller is also used to control the gimbal imaging device to rotate according to the gimbal advance angle and the gimbal segment rotation speed curve, control the mobile chassis to compensate for steering according to the chassis compensation angle, and generate a target afterimage sector when the confidence of the current main tracking target is lower than the tracking threshold and control the gimbal imaging device to perform re-acquisition within the target afterimage sector.
[0009] Preferably, the gimbal imaging device includes a visible light camera and / or an infrared camera; the gimbal angle detection device includes a horizontal angle encoder and a pitch angle encoder; the environmental sensing device includes an inertial measurement unit and a circumferential ranging sensor; the mobile chassis includes a chassis controller, a drive motor and a steering actuator; and the edge computing controller is communicatively connected to the gimbal imaging device, the gimbal angle detection device, the environmental sensing device and the mobile chassis.
[0010] Preferably, the edge computing controller is used to determine the center coordinates of the warning target based on the target detection boxes in consecutive image frames; determine the scale change of the warning target based on the area change or diagonal length change of the target detection boxes in adjacent frames; determine the motion direction of the warning target based on the displacement direction of the center coordinates of adjacent frames; and fuse the target detection confidence with the target matching consistency of adjacent frames to obtain the confidence of the warning target.
[0011] Preferably, the target node is used to characterize the center coordinates, scale changes, direction of movement, and confidence level of the warning target; the warning sector node is used to characterize the warning level and core warning position of the warning area where the warning target is located; the gimbal margin node is used to characterize the remaining rotation angle from the current angle of the gimbal to the preset mechanical limit angle; the chassis passage node is used to characterize the circumferential passage distance of the mobile chassis in different steering directions; the obstacle suppression node is used to characterize the degree of restriction of obstacles on the compensated steering of the mobile chassis; and the loss risk node is used to characterize the risk of the warning target leaving the gimbal's field of vision.
[0012] Preferably, the warning intent propagation map includes: an intrusion propagation edge from the target node to the warning sector node; a boundary risk propagation edge from the target node to the lost risk node; a gimbal margin propagation edge from the lost risk node to the gimbal margin node; a compensation demand propagation edge from the gimbal margin node to the chassis passage node; and a reverse suppression propagation edge from the obstacle suppression node to the chassis passage node. The weight of the intrusion propagation edge increases with the depth of the warning target's intrusion into the warning sector; the weight of the boundary risk propagation edge increases with the inter-frame displacement of the warning target towards the image boundary; the weight of the gimbal margin propagation edge increases with the decrease of the remaining gimbal rotation angle; and the weight of the reverse suppression propagation edge increases with the decrease of the circumferential obstacle distance.
[0013] Preferably, the edge computing controller is used to initialize the node value of the target node with the alert level, confidence level and intrusion depth of the alert sector of the alert target; propagate the node value of the target node along the propagation edge in the alert intent propagation graph for 2 to 4 rounds; deduct the reverse inhibition amount from the obstacle inhibition node after each round of propagation, and normalize the node value after propagation; and fuse the target node value and the lost risk node value after propagation to obtain the tracking intent value of the alert target.
[0014] Preferably, the gimbal field of view boundary pressure value is determined by boundary distance pressure, boundary approach pressure, gimbal margin pressure, and chassis compensation release amount; wherein, the smaller the distance from the warning target to the image boundary, the greater the boundary distance pressure; the greater the inter-frame displacement of the warning target toward the image boundary, the greater the boundary approach pressure; the smaller the remaining gimbal rotation angle, the greater the gimbal margin pressure; the greater the circumferential travel distance of the moving chassis in the direction that moves the warning target away from the image boundary, the greater the chassis compensation release amount; the edge computing controller is used to superimpose the boundary distance pressure, boundary approach pressure, and gimbal margin pressure, and subtract the chassis compensation release amount to obtain the gimbal field of view boundary pressure value.
[0015] Preferably, the edge computing controller is used to weightedly fuse the tracking intent value of each alert target and the gimbal's field of view boundary pressure value to obtain the target execution value of each alert target; determine the alert target with the highest target execution value as the current primary tracking target; determine the gimbal's advance rotation angle based on the center coordinates, movement direction, current gimbal angle, and remaining gimbal rotation angle of the current primary tracking target; and generate a segmented gimbal rotation speed curve including a rapid acquisition segment, a pressure release segment, and a low-speed locking segment based on the gimbal's field of view boundary pressure value of the current primary tracking target.
[0016] Preferably, when the gimbal field of view boundary pressure value of the current main tracking target is greater than the compensation trigger threshold, and the circumferential obstacle distance of the moving chassis in the pressure release direction is greater than the safe distance, the edge computing controller generates a chassis compensation angle; the direction of the chassis compensation angle is to move the current main tracking target away from the image boundary and to return the current angle of the gimbal to the preset intermediate working angle range; when the gimbal field of view boundary pressure value of the current main tracking target is greater than the compensation trigger threshold, and the circumferential obstacle distance of the moving chassis in the pressure release direction is less than or equal to the safe distance, the edge computing controller prohibits the generation of chassis compensation angle and reduces the rotation speed of the fast interception segment in the gimbal segmented rotation speed curve.
[0017] Preferably, the edge computing controller is used to cache the center coordinates, scale changes, angle change directions, warning sectors, and current angular velocity of the current primary target in the most recent 3 to 8 frames before the confidence of the current primary target falls below the tracking threshold; when the confidence of the current primary target falls below the tracking threshold, it determines the center direction of the target afterimage sector based on the cached angle change directions and the current angular velocity of the gimbal; it determines the expansion angle of the target afterimage sector based on the inter-frame displacement velocity and scale change before the current primary target was lost; and it controls the gimbal imaging device to perform recapture according to a fan-shaped search sequence expanding outward from the center direction of the target afterimage sector.
[0018] Compared with the prior art, the technical solution of the present invention has the following beneficial effects: This invention does not generate gimbal control commands solely based on the target's current position in the image. Instead, it constructs a warning intent propagation map for each warning target, incorporating the target's depth of intrusion into the warning sector, the target's displacement towards the core warning position, the remaining gimbal rotation angle, the chassis's circumferential passable distance, and the risk of the target leaving the field of view into the propagation calculation to obtain the tracking intent value for each warning target. This allows for the priority selection of the primary tracking target when multiple targets appear simultaneously, prioritizing those that require more tracking, are more likely to pose a warning risk, or are more easily lost. This avoids the gimbal oscillating erratically among multiple targets, improving the targeting and timeliness of the warning response.
[0019] This invention determines the target's tendency to leave the field of view by measuring the pressure value at the gimbal's field of view boundary, rather than simply relying on whether the gimbal angle exceeds a threshold to trigger chassis compensation. The system comprehensively considers the distance from the target to the image boundary, the inter-frame displacement of the target towards the boundary, the remaining gimbal rotation angle, and the circumferential passable distance of the chassis to determine whether to generate a chassis compensation rotation angle. When the boundary pressure is high and the chassis has safe turning conditions, the moving chassis performs compensation turning in advance, bringing the gimbal back to the preset intermediate working angle range; when the circumferential obstacle distance is insufficient, chassis movement is suppressed and the gimbal's rotation speed during the rapid acquisition segment is reduced, thus balancing rapid tracking and motion safety.
[0020] This invention caches the center coordinates, scale changes, angle change directions, location of the alert sector, and current angular velocity of the gimbal in the most recent frames before the confidence of the current primary target decreases. When the target's confidence falls below the tracking threshold due to occlusion, edge cutout, or brief recognition failure, a target afterimage sector is generated based on the cached information, and the gimbal is controlled to re-acquire the target according to a fan-shaped search sequence expanding outward from the center of the afterimage sector. This avoids the problems of traditional global scanning, such as large search range, long search time, and easy accidental locking of other targets, improving the continuity of alert event evidence collection and the reliability of the gimbal's rapid recovery of tracking. Attached Figure Description
[0021] Figure 1This is a schematic diagram of the overall structure of a security and surveillance robot command execution system for rapid gimbal tracking, as described in this application. Figure 2 This application relates to a security and surveillance robot with gimbal-based rapid tracking; Figure 3 This is a schematic diagram of the instruction execution flow of a security and surveillance robot instruction execution system for rapid gimbal tracking, as described in this application. Figure 4 This is a schematic diagram of the nodes and propagation edges of the warning intent propagation graph in this application; Figure 5 This is a schematic diagram of the control of gimbal field of view boundary pressure, chassis collaborative compensation and target afterimage sector re-acquisition in this application. Detailed Implementation
[0022] Those skilled in the art will understand that, in order to make the above-mentioned objects, features, and beneficial effects of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Figure 1 , 2 This application illustrates a gimbal-based rapid tracking security and surveillance robot command execution system, including a gimbal imaging device, a mobile chassis, an environmental sensing device, a gimbal angle detection device, and an edge computing controller; The gimbal imaging device is used to acquire continuous image frames, the gimbal angle detection device is used to obtain the current angle of the gimbal, and the environmental perception device is used to obtain the chassis attitude and the distance to circumferential obstacles. In one specific implementation, such as Figure 2 The process method shown includes a security surveillance robot comprising a mobile chassis, a two-degree-of-freedom gimbal mounted on the mobile chassis, a gimbal imaging device mounted on the gimbal, a gimbal angle detection device mounted at the gimbal pivot, an environmental sensing device mounted on the mobile chassis, and an edge computing controller.
[0023] The pan-tilt imaging device includes a visible light camera and an infrared camera. The visible light camera is used to acquire continuous image frames of the surveillance area during the day or in well-lit environments, while the infrared camera is used to acquire continuous image frames at night, in low light, or in backlight environments. Both the visible light and infrared cameras are communicatively connected to the edge computing controller and can send image frames to the edge computing controller via Ethernet, USB, MIPI, or other image transmission interfaces. The edge computing controller receives continuous image frames at a preset frame rate, such as 25 frames per second or 30 frames per second, and appends a frame number and acquisition timestamp to each image frame, enabling subsequent determination of the displacement direction, displacement velocity, and scale changes of the surveillance target between adjacent frames.
[0024] The gimbal angle detection device includes a horizontal angle encoder and a pitch angle encoder. The horizontal angle encoder is installed on the horizontal rotation axis of the gimbal and is used to detect the horizontal angle of the gimbal relative to the forward reference direction of the robot chassis. The pitch angle encoder is installed on the pitch rotation axis of the gimbal and is used to detect the pitch angle of the gimbal relative to the horizontal reference plane. The horizontal and pitch angle encoders can be absolute encoders or incremental encoders combined with a zero-position calibration switch. The edge computing controller periodically reads the horizontal and pitch angles and uses the read angles as the current angle of the gimbal. To correspond with image frames, the edge computing controller binds the current angle of the gimbal to image frames with adjacent timestamps, thereby obtaining the actual orientation of the gimbal at the time of each image frame acquisition.
[0025] The environmental sensing device includes an inertial measurement unit (IMU) and a circumferential ranging sensor. The IMU is fixedly installed in the central area of the mobile chassis or near the chassis controller to acquire the chassis attitude. The chassis attitude includes at least the changes in chassis roll angle, pitch angle, and yaw angle, and may also include chassis angular velocity and linear acceleration. The edge computing controller determines whether the robot is accelerating, braking, turning, navigating on a slope, or experiencing an external impact based on the data output by the IMU, in order to subsequently evaluate the impact of chassis motion on the gimbal imaging stability.
[0026] The circumferential ranging sensor can include at least one of LiDAR, millimeter-wave radar, ultrasonic sensors, or depth cameras. In a preferred embodiment, ranging sensors are respectively installed at the front, front left, front right, rear left, and rear right of the mobile chassis, or a two-dimensional LiDAR is installed on the top of the mobile chassis to obtain obstacle distances in different circumferential directions of the robot. The edge computing controller divides the space around the robot into multiple directional intervals, such as the front, front left, front right, left side, right side, and rear intervals, and records the distance to the nearest obstacle in each directional interval to form the circumferential obstacle distance.
[0027] When the robot performs a surveillance task, the edge computing controller synchronously receives continuous image frames, the current angle of the gimbal, the chassis attitude, and the circumferential obstacle distance. Specifically, when the gimbal imaging device acquires the Nth frame of image, the edge computing controller reads the horizontal angle, pitch angle, chassis attitude, and circumferential obstacle distance closest to the acquisition time of that image frame, and forms a set of state data. This state data includes the image frame number, image acquisition time, gimbal horizontal angle, gimbal pitch angle, chassis roll angle, chassis pitch angle, chassis yaw change, and circumferential obstacle distance in each direction range. For example, when a security surveillance robot patrols near a factory fence, the gimbal imaging device continuously acquires images of the fenced area. If a person is detected in the image, the edge computing controller not only acquires the image frame containing the person but also simultaneously acquires the current horizontal and pitch angles of the gimbal and reads the chassis inertial attitude and the robot's circumferential obstacle distance. If there is a nearby obstacle in front of the right side of the robot, the edge computing controller can identify the restricted steering state in that direction when generating the chassis compensation steering angle later, thus avoiding blindly compensating for steering to the right front of the moving chassis.
[0028] Through the above implementation methods, such as Figure 5 As shown, the system can maintain a temporal correspondence between continuous image frames, gimbal angle, chassis attitude, and circumferential obstacle distance, providing a reliable data foundation for subsequent determination of the remaining gimbal rotation angle, the circumferential passable distance of the chassis, the movement direction of the warning target, the pressure value of the gimbal's field of view boundary, and the chassis compensation rotation angle.
[0029] The edge computing controller is used to determine the remaining turning angle of the gimbal based on the current angle of the gimbal, determine the circumferential passable distance of the chassis based on the circumferential obstacle distance, and extract the center coordinates, scale changes, motion direction and confidence of each warning target from continuous image frames; In one specific implementation, after receiving continuous image frames from the gimbal, the current angle of the gimbal, the chassis attitude, and the circumferential obstacle distance, the edge computing controller first performs time alignment processing on the aforementioned data. Specifically, the edge computing controller uses the acquisition timestamp of the image frame as a reference, selects the gimbal horizontal angle, gimbal pitch angle, chassis attitude data, and circumferential obstacle distance data closest to the acquisition time of that image frame, and combines them into a set of tracking status data. This tracking status data serves as input for subsequent target extraction, calculation of the remaining gimbal rotation angle, and calculation of the chassis circumferential passable distance.
[0030] The gimbal angle detection device outputs the current horizontal and vertical angles of the gimbal in real time. The edge computing controller pre-stores the mechanical limit angles of the gimbal, including the left horizontal limit angle, right horizontal limit angle, upper vertical limit angle, and lower vertical limit angle. Based on the current horizontal angle, the edge computing controller calculates the remaining angles from the corresponding mechanical limit angles when the gimbal rotates left and right, and calculates the remaining angles from the corresponding mechanical limit angles when the gimbal tilts up and down, based on the current vertical angle. This yields the remaining gimbal rotation angle in different rotation directions. For example, if the horizontal rotation range of the gimbal is -120° to +120° on the right, and the current horizontal angle is +80°, then the remaining rotation angle for continuing to rotate to the right is smaller, and the remaining rotation angle for rotating to the left is larger. If the target continues to move towards the right edge of the image, the edge computing controller will determine that the gimbal's rightward tracking margin is insufficient. Subsequently, when generating the gimbal's field-of-view boundary pressure value and chassis compensation angle, it will increase the boundary pressure and chassis compensation requirements in that direction. In another scenario, if the gimbal's current angle is near the middle, such as around 0° horizontally, the gimbal has a large remaining rotation angle in both left and right directions. The system prioritizes target tracking through the gimbal's own rotation rather than immediately triggering chassis compensation steering. In this way, the remaining gimbal rotation angle is not simply the current angle value, but rather represents the angle margin the gimbal can continue to rotate in the target's direction of movement, used to determine whether the gimbal possesses independent continuous tracking capability.
[0031] The environmental sensing device acquires circumferential obstacle distances in multiple directions around the robot. The edge computing controller divides the space around the robot into several directional intervals, such as the front, left front, right front, left side, right side, left rear, right rear, and rear intervals. Each directional interval corresponds to one or more circumferential ranging data points. The edge computing controller first filters the circumferential obstacle distances within each directional interval, removing obviously abnormal ranging values, such as instantaneous jump values, values exceeding the effective detection range of the sensors, and isolated values inconsistent with continuous ranging results. Then, the edge computing controller selects the nearest obstacle distance within each directional interval as the effective obstacle distance for that direction.
[0032] After determining the effective obstacle distance, the edge computing controller combines the preset safety distance, the width of the mobile chassis, the turning radius, and the current chassis posture to determine the circumferential passable distance of the chassis in the corresponding direction. Specifically, when the distance to the nearest obstacle within a certain directional interval is greater than the preset safety distance, and the direction meets the space required for chassis turning, that direction is determined to be a passable direction. When the distance to the nearest obstacle is less than or equal to the preset safety distance, or when the chassis posture indicates that the robot is in a state of significant tilt, sharp turn, or instability, the circumferential passable distance of the chassis in that direction is reduced, or even set to an impassable state. For example, when the mobile chassis needs to compensate for a right turn to release the pressure on the right boundary of the gimbal, the edge computing controller reads the circumferential obstacle distances in the right front interval and the right interval. If there is a nearby wall or pedestrian in the right-front area, the system determines that the right-side compensation steering space is insufficient, and reduces the right-side chassis circumferential clearance. If there are no nearby obstacles in the right-side and right-front areas, the system determines that there is a large right-side chassis circumferential clearance, allowing the chassis to slowly compensate for steering to the right when generating the subsequent chassis compensation steering angle. Therefore, the chassis circumferential clearance is not an obstacle distance directly measured by a single sensor, but a chassis action feasibility parameter determined by the edge computing controller based on the circumferential obstacle distance, safety distance, chassis size, steering radius, and chassis attitude.
[0033] The edge computing controller performs target detection on consecutive image frames, identifying people, vehicles, animals, or other abnormal moving objects in the images. The target detection results are output as bounding boxes, which include the coordinates of the top-left and bottom-right corners, the target category, and the detection confidence score. The edge computing controller determines the center coordinates of the target based on the position of the bounding box. Specifically, the horizontal midpoint of the bounding box is used as the horizontal coordinate of the target center, and the vertical midpoint is used as the vertical coordinate. This center coordinate represents the position of the target in the image. For example, if a person is detected in a frame, and their bounding box is located in the right-hand area of the image, then the target's center coordinates are also located on the right side of the image. The edge computing controller can then determine that the person is offset to the right relative to the image center, and the pan-tilt unit needs to rotate to the right or the chassis needs to compensate for the rightward steering.
[0034] To improve stability, the edge computing controller can also smooth the target center coordinates across multiple consecutive frames. For example, when the target detection box in a certain frame experiences a slight jump due to lighting, occlusion, or jitter, the edge computing controller can correct the center coordinates of that frame by combining the trend of center coordinate changes in previous and subsequent frames, thus avoiding frequent small jitters of the gimbal due to single-frame detection errors.
[0035] The edge computing controller determines scale changes based on the size variations of the target detection box for the same vigilant target in consecutive image frames. These scale changes can be determined through variations in the target detection box area, height, width, or diagonal length. In one implementation, the edge computing controller preferentially uses changes in the target detection box area to characterize scale changes. When the area of the detection box for the same target gradually increases in adjacent frames, it indicates that the target may be approaching the robot or approaching the gimbal's field of view; when the detection box area gradually decreases, it indicates that the target may be moving away from the robot; when the change in the detection box area is not significant, it indicates that the relative distance between the target and the robot changes little. For example, if the height and area of the detection box for the same person target continuously increase in five consecutive frames, while the target center moves towards the vigilance core area, the edge computing controller determines that the person target has a high approach trend. This scale change can subsequently be used to improve the vigilance intent value of the target node, or to assist in determining the expansion angle of the target afterimage sector when the target confidence decreases. In another implementation, when the area of the target detection box is greatly affected by changes in posture, the edge computing controller can use the change in the diagonal length of the target detection box as a scale change parameter to reduce the impact of factors such as people turning around and changes in the lateral angle of vehicles on scale judgment.
[0036] The edge computing controller performs inter-frame association on the same surveillance target in consecutive image frames. Inter-frame association can be determined based on the overlap of target detection boxes, target category consistency, center coordinate distance, and similarity of target appearance features. Through inter-frame association, the edge computing controller can determine whether a surveillance target in frame N is the same as a surveillance target in frame N+1. After completing inter-frame association, the edge computing controller determines the target's motion direction based on the direction of change of the center coordinates of the same target in adjacent frames. If the target's center coordinates continuously change to the right of the image, the target is determined to have a rightward motion trend; if the target's center coordinates continuously change to the left of the image, the target is determined to have a leftward motion trend; if the target's center coordinates simultaneously change to the top or bottom of the image, its vertical motion trend is further determined. In security surveillance scenarios, the motion direction is used not only to determine the target's movement direction in the image but also to determine whether the target is facing the image boundary, whether it is facing the core surveillance location, and whether it may leave the pan-tilt-zoom (PTZ) camera's field of view. For example, if the current main tracking target is located on the right side of the image and continues to move to the right for several consecutive frames, the edge computing controller determines that the target is moving towards the right image boundary; if the remaining rightward rotation angle of the gimbal is small at this time, a higher gimbal field of view boundary pressure value will be generated subsequently.
[0037] The edge computing controller acquires the target detection confidence score output by the target detection algorithm and combines it with the target matching consistency of adjacent frames to obtain the confidence score of the warning target. The target detection confidence score indicates the reliability of identifying a target as a person, vehicle, or abnormal object in a single frame image; the target matching consistency of adjacent frames indicates whether the target is stably present and has a continuous motion trajectory in consecutive frames. In one implementation, when the target detection confidence score is high and the target can be stably matched in multiple consecutive frames, the edge computing controller determines that the warning target has a high confidence score. When the target detection confidence score decreases, or when the target exhibits discontinuity, occlusion, detection box drift, or category change in adjacent frames, the edge computing controller lowers the confidence score of the warning target. For example, if a person target is briefly occluded after passing a pillar, the target detection box may disappear within one or two frames or the confidence score may decrease significantly. At this point, the edge computing controller does not immediately delete the target completely. Instead, it marks the target's confidence as decreasing and combines this with the target's center coordinates before loss, scale changes, angle change direction, and location within the warning sector to provide a basis for generating subsequent target afterimage sectors. By fusing target detection confidence and inter-frame matching consistency, it avoids false alarms or missed alarms caused by relying solely on single-frame detection results and provides a stable basis for the continuous tracking and re-acquisition of the current primary target.
[0038] In some embodiments, when the security patrol robot patrols near a park fence, the gimbal imaging device continuously captures consecutive image frames of the fenced area. After receiving the Nth frame, the edge computing controller identifies two human targets based on the target detection box. The first human target is located in the front left of the image and is moving slowly, while the second human target is located in the right side of the image and is continuously moving towards the right boundary of the image. At the same time, the gimbal angle detection device reports that the current horizontal angle of the gimbal is deflected to the right. Based on this angle and the right mechanical limit angle, the edge computing controller determines that the remaining right-side gimbal rotation angle is small. The environmental perception device reports that there are no nearby obstacles on the right and front right of the robot. Based on this, the edge computing controller determines that the circumferential travel distance of the right-side chassis is large.
[0039] The edge computing controller further extracts the center coordinates, scale changes, direction of motion, and confidence level of the two human targets from consecutive image frames. For the second human target, its center coordinates continuously move to the right, and the detection box area gradually increases, indicating that it may be approaching the robot and moving towards the right boundary of the image. The edge computing controller uses this information as the basis for subsequent calculations of the warning intent propagation map and the gimbal's field of view boundary pressure value, enabling the system to prioritize whether the target needs rapid tracking and whether chassis compensation steering is required.
[0040] Through the above implementation, the edge computing controller can form unified, stable, and usable basic state variables for control decisions from sensor data and continuous image frames. This provides data support for subsequent construction of a warning intent propagation map, generation of gimbal field-of-view boundary pressure values, determination of the current primary tracking target, and execution of gimbal and chassis coordinated control. Furthermore, it extracts the center coordinates, scale changes, motion direction, and confidence level of each warning target from continuous image frames. The center coordinates are used to determine the distance from the warning target to the image boundary, the motion direction is used to determine the inter-frame displacement of the warning target towards the image boundary, and the scale changes and confidence levels are used to determine the tracking stability state of the warning target. The center coordinates of the warning target extracted by the edge computing controller are used to calculate the positional relationship of the target relative to the image center and image boundary. The scale changes are used to determine the target's approach trend and the expansion range of the target's afterimage sector. The motion direction is used to determine the target's displacement towards the warning core position, the inter-frame displacement towards the image boundary, and the center direction of the target's afterimage sector. The confidence level is used to initialize the target node value in the warning intent propagation map and trigger the generation of the target's afterimage sector.
[0041] The edge computing controller is also used to construct a warning intent propagation graph for each warning target, including target nodes, warning sector nodes, gimbal margin nodes, chassis access nodes, obstacle suppression nodes, and loss risk nodes, such as... Figure 4 As shown, the tracking intention value is obtained; the gimbal field of view boundary pressure value is obtained based on the distance from the warning target to the image boundary, the inter-frame displacement towards the image boundary, the remaining gimbal rotation angle, and the circumferential passable distance of the chassis; the current main tracking target is determined based on the tracking intention value and the gimbal field of view boundary pressure value, and the gimbal advance rotation angle, the gimbal segmented rotation speed curve, and the chassis compensation rotation angle are generated; In one specific implementation, after obtaining the center coordinates, scale changes, direction of movement, confidence level, remaining gimbal rotation angle, and circumferential passable distance of each alert target, the edge computing controller constructs an alert intent propagation map for each alert target. The alert intent propagation map is used to comprehensively determine whether the alert target should be prioritized for tracking and whether there is a risk of it soon leaving the gimbal's field of view.
[0042] Before the robot performs its patrol mission, the edge computing controller divides the area around the robot into multiple warning sectors. For example, these can be divided into sectors such as the front front, left front, right front, left side, right side, rear, and key entrance sectors. Each warning sector is configured with a warning level and a core warning location. The warning level characterizes the importance of the area. For example, factory entrances / exits, fence gaps, and hazardous equipment areas can be assigned higher warning levels; ordinary roads, green belts, or non-key areas can be assigned lower warning levels. The core warning location is a location within the sector that requires special protection or monitoring, such as access control gates, fence lines, equipment entrances, or the boundary lines of prohibited areas. Once the edge computing controller identifies a warning target from consecutive image frames, it determines the corresponding warning sector based on the target's position in the image, the current angle of the gimbal, and the robot's current orientation. It further determines the target's intrusion depth and movement trend relative to the core warning location within that sector.
[0043] For each alert target, the edge computing controller constructs a corresponding alert intent propagation graph. This graph includes at least target nodes, alert sector nodes, gimbal margin nodes, chassis access nodes, obstacle suppression nodes, and loss risk nodes. Target nodes characterize the state of the alert target itself, including target center coordinates, scale changes, direction of movement, and confidence level. For example, target center coordinates indicate the target's position in the image, scale changes indicate whether the target is approaching the robot or the core alert area, direction of movement indicates whether the target is moving towards the image boundary or the core alert position, and confidence level indicates the reliability of target detection and continuous tracking. Alert sector nodes characterize the importance of the alert area where the target is located, including alert level, core alert position, and the depth of the target's intrusion into the alert sector. If the target is located in a key entry sector and is moving towards the core alert position, this node has a high propagation effect. Gimbal margin nodes characterize the remaining turning angle of the gimbal in the target's direction of movement. If the target moves towards the right edge of the image, the edge computing controller focuses on reading the remaining turning angle of the gimbal to the right; if the target moves towards the left edge of the image, it focuses on reading the remaining turning angle of the gimbal to the left. The smaller the remaining turning angle, the weaker the gimbal's ability to continue tracking the target independently. The chassis access node characterizes the circumferential travel distance of the moving chassis in different compensation directions. For example, when the target continues to move towards the right edge of the image, the system determines whether there is sufficient space for the moving chassis to compensate to the right or right front, and uses this to determine the node value of the chassis access node. The obstacle suppression node characterizes the degree to which obstacles restrict the chassis compensation action. If there are nearby obstacles, people, or walls in the pressure release direction, the obstacle suppression node has a strong reverse suppression effect, reducing the feasibility of the chassis compensation action. The loss risk node characterizes the risk of the target leaving the gimbal's field of view. If the target is close to the image boundary, continues to move towards the image boundary, the remaining turning angle of the gimbal is small, or the target confidence decreases, the node value of the loss risk node increases.
[0044] After constructing the nodes, the edge computing controller establishes propagation edges between each node. First, an intrusion propagation edge is established from the target node to the warning sector node. This edge represents the degree to which the target enters the warning area and approaches the core warning position. The deeper the target intrudes into the warning sector, or the closer the target's movement direction is to the core warning position, the greater the weight of this intrusion propagation edge. Second, a boundary risk propagation edge is established from the target node to the loss-of-view node. This edge represents the risk of the target leaving the field of view due to approaching the image boundary. The closer the target's center coordinates are to the image boundary, and the more the displacement direction of consecutive frames is towards that image boundary, the greater the weight of the boundary risk propagation edge. Third, a gimbal margin propagation edge is established from the loss-of-view node to the gimbal margin node. This edge represents the relationship between loss risk and gimbal tracking capability. When the target is leaving the field of view in a certain direction, and the remaining turning angle of the gimbal in that direction is small, the weight of the gimbal margin propagation edge increases, indicating that the ability of the gimbal to continue tracking alone is insufficient. Fourth, a compensation demand propagation edge is established from the gimbal margin node to the chassis passage node. This propagation edge indicates whether chassis compensation is needed when the gimbal has insufficient margin. The smaller the remaining gimbal angle, the more easily the chassis passage node is affected by the compensation demand propagation edge, thus increasing the likelihood of chassis compensation turning. Fifth, establish a reverse inhibition propagation edge from the obstacle inhibition node to the chassis passage node. This propagation edge represents the inhibitory effect of obstacles on chassis compensation. The smaller the circumferential obstacle distance in the pressure release direction, the greater the weight of the reverse inhibition propagation edge, and the lower the executability of the chassis passage node. Through these propagation edges, the edge computing controller can couple factors such as "whether the target is important," "whether the target is about to be lost," "whether the gimbal still has margin," and "whether the chassis can safely compensate" into the same graph structure for calculation, instead of using isolated threshold judgments separately.
[0045] In one specific implementation, the edge computing controller first initializes the node value of the target node based on the target's alert level, confidence level, and depth of intrusion into the alert sector. If the target is a key alert object such as personnel or vehicles, has a high detection confidence level, and has entered a high-level alert sector, its initial target node value is higher; if the target is located in a low-level alert area and has a low confidence level, its initial target node value is lower. Subsequently, the edge computing controller propagates the node value of the target node along the propagation edge in the alert intent propagation graph for a limited number of rounds, preferably 2 to 4 rounds. Too few propagation rounds make it difficult to fully reflect the impact of gimbal margin, chassis accessibility, and obstacle suppression on the tracking intent; too many propagation rounds easily increase computation time and reduce the speed of fast tracking response. Therefore, using 2 to 4 rounds of propagation can balance propagation sufficiency and real-time performance.
[0046] After each round of propagation, the edge computing controller deducts the reverse inhibition amount from the obstacle suppression node and normalizes the values of each node after propagation, keeping them within a preset range to avoid misjudgment caused by abnormal amplification of a certain node value. After propagation is complete, the edge computing controller fuses the propagated target node value with the lost risk node value to obtain the tracking intent value of the warning target. This tracking intent value not only reflects the warning risk of the target itself, but also reflects whether the target is likely to leave the gimbal's field of view and whether the gimbal or chassis needs to be moved in advance. For example, the first target is located in a key entrance area but is in the center of the image and moves slowly; the second target is located in a normal area but is moving rapidly towards the right boundary of the image, and the remaining rightward turning angle of the gimbal is small. Traditional methods of sorting by warning level may prioritize tracking the first target, while this implementation, through the propagation effect of lost risk nodes and gimbal margin nodes, can improve the tracking intent value of the second target, allowing the robot to prioritize the target that is about to be lost, thereby preventing the target from leaving the field of view.
[0047] After obtaining the tracking intent value of each alerted target, the edge computing controller further calculates the gimbal field-of-view boundary pressure value for each alerted target. This value characterizes the pressure exerted by the alerted target on the gimbal's field-of-view boundary, i.e., the degree to which the target is about to leave the current image field of view and the gimbal's ability to continue tracking is insufficient. In one specific implementation, the edge computing controller first calculates the distance from the target to the image boundary based on the target's center coordinates. The image boundary includes the left boundary, right boundary, top boundary, and bottom boundary. The closer the target is to a certain boundary, the easier it is for it to leave the field of view from that boundary, and the greater the corresponding boundary distance pressure. Secondly, the edge computing controller determines whether the target is moving towards the image boundary based on the changes in the target's center coordinates in consecutive frames. If the target is located on the right side of the image and moves continuously to the right, it is considered that the target has an inter-frame displacement towards the right boundary, and the boundary approach pressure increases; if the target is close to the right boundary but is moving back towards the image center, the boundary approach pressure decreases. Finally, the edge computing controller determines the gimbal margin pressure in conjunction with the remaining gimbal rotation angle. If the target moves towards the right boundary, and the remaining rightward turning angle of the gimbal is small, it indicates that the mechanical margin for the gimbal to continue tracking to the right is insufficient, increasing the gimbal margin pressure. If the gimbal still has a large remaining turning angle in that direction, the gimbal margin pressure is small. Finally, the edge computing controller determines the chassis compensation release amount based on the circumferential travel distance of the chassis. If the moving chassis has a large travel distance in the direction that allows the target to move away from the image boundary, the chassis can release the gimbal field of view boundary pressure through compensation steering, resulting in a large chassis compensation release amount. If there are obstacles or personnel in that direction, the chassis compensation release amount is small. The edge computing controller superimposes the boundary distance pressure, boundary approach pressure, and gimbal margin pressure, and subtracts the chassis compensation release amount to obtain the gimbal field of view boundary pressure value. The resulting gimbal field of view boundary pressure value can simultaneously reflect how close the target is to the boundary, whether it is moving towards the boundary, whether the gimbal still has tracking margin, and whether the chassis can safely compensate.
[0048] When multiple targets are present in an image, the edge computing controller calculates the tracking intent value and the PTZ field-of-view boundary pressure value for each target, and then weights and fuses these values to obtain the target execution value for each target. The tracking intent value primarily reflects the importance and necessity of tracking the target, while the PTZ field-of-view boundary pressure value primarily reflects the urgency of the target about to leave the field of view. The edge computing controller identifies the target with the highest execution value as the current primary tracking target. To avoid frequent switching between multiple targets, in a preferred approach, the edge computing controller also sets a target switching hysteresis condition. That is, when the execution value of a new target is only slightly higher than that of the current primary tracking target, the system maintains the current primary tracking target unchanged; only when the execution value of the new target exceeds the current primary tracking target by a preset difference, or when the confidence level of the current primary tracking target drops below the tracking threshold, is the current primary tracking target switched. This method reduces jittery switching of the PTZ between multiple near-risk targets. For example, consider the simultaneous presence of personnel and vehicle targets at the entrance of a park. The vehicle target has a high alert level but is located in the center of the image; the personnel target has a slightly lower alert level but is moving rapidly towards the edge of the image and is about to enter the fence blind spot. The edge computing controller combines the tracking intent value of both and the PTZ field of view boundary pressure value. If the personnel target has a higher target execution value, it will be identified as the current primary tracking target, and the PTZ will be prioritized to track the personnel target to prevent it from entering the monitoring blind spot.
[0049] After determining the current primary tracking target, the edge computing controller generates a gimbal advance angle based on the target's center coordinates, direction of movement, current gimbal angle, and remaining gimbal rotation angle. In one specific implementation, the edge computing controller first determines the offset direction and magnitude of the target's center coordinates relative to the image center. If the target is located on the right side of the image, the horizontal component of the advance gimbal rotation angle points to the right; if the target is located on the left side, the horizontal component points to the left. If the target is simultaneously above or below the image, the advance gimbal rotation angle also includes a pitch component. The edge computing controller further adjusts the advance gimbal rotation angle based on the target's direction of movement. For example, when the target is on the right side of the image and continuously moves to the right, the advance gimbal rotation angle not only corresponds to the target's current center position but also appropriately points to the position the target might reach in the next moment; when the target's direction of movement is towards the image center, the advance is reduced to avoid excessive gimbal rotation. Simultaneously, the edge computing controller limits the advance gimbal rotation angle based on the remaining gimbal rotation angle. When the gimbal has insufficient remaining rotation angle in a certain direction, the gimbal will rotate in advance to within the safe range allowed in that direction, and trigger the subsequent chassis compensation rotation angle generation process to release the mechanical margin of the gimbal through chassis steering.
[0050] The edge computing controller generates a segmented gimbal rotation speed curve based on the gimbal's field-of-view boundary pressure value, target center deviation, and gimbal's advance rotation angle for the current primary tracking target. This segmented rotation speed curve includes a rapid acquisition segment, a pressure release segment, and a low-speed lock segment. The rapid acquisition segment is used to quickly approach the target's predicted direction at a high rotation speed when the target deviates significantly from the image center or when the target's execution value is high. The purpose of this stage is to shorten the time it takes for the gimbal to turn from its current orientation to the current primary tracking target, preventing the target from continuing to move away from the image center. The pressure release segment is used in conjunction with chassis compensation when the gimbal's field-of-view boundary pressure value is high. If the target is close to the image boundary and the remaining gimbal rotation angle is insufficient, the edge computing controller reduces the gimbal's rapid rotation speed in the pressure release segment and, based on the chassis compensation rotation angle, slowly turns the moving chassis, gradually returning the gimbal's current angle to the preset intermediate working angle range. This stage prevents the gimbal from continuing to approach the mechanical limit and reduces image jitter caused by simultaneous rapid movements of the gimbal and chassis. The low-speed lock segment is used for minor corrections after the current primary tracking target enters the image center neighborhood. At this point, the gimbal no longer uses a high rotation speed, but instead makes low-speed fine adjustments based on the residual deviation between the target center coordinates and the image center, so that the target is stably kept near the center of the image, which is convenient for subsequent identification, evidence collection and alarm.
[0051] The edge computing controller determines whether to generate a chassis compensation angle based on the gimbal's field-of-view boundary pressure value of the current primary tracking target, the remaining gimbal rotation angle, and the circumferential passable distance of the chassis. When the gimbal's field-of-view boundary pressure value of the current primary tracking target exceeds the compensation trigger threshold, and the circumferential obstacle distance of the mobile chassis in the pressure release direction is greater than the safe distance, the edge computing controller generates a chassis compensation angle. The direction of this chassis compensation angle is to move the current primary tracking target away from the image boundary and return the current gimbal angle to the preset intermediate working angle range. For example, if the current primary tracking target continues to move towards the right boundary of the image, the gimbal has already rotated a large angle to the right and the remaining rightward rotation angle is insufficient, and there are no nearby obstacles on the robot's right front and right sides, the edge computing controller generates a rightward chassis compensation angle. After the mobile chassis performs this compensation turn, the robot adjusts towards the target direction, preventing the gimbal from continuing to approach the right mechanical limit, thus restoring the gimbal's subsequent usable rotation angle. Conversely, when the gimbal's field of view boundary pressure value of the current main tracking target is high, but there is an obstacle in the direction of pressure release on the mobile chassis, causing the circumferential obstacle distance to be less than or equal to the safe distance, the edge computing controller prohibits the generation of chassis compensation rotation angle and reduces the rotation speed of the fast interception segment in the gimbal segment rotation speed curve, so as to avoid the robot performing dangerous turns when space is insufficient, and at the same time reduce the loss of target or severe screen shaking caused by high-speed gimbal movements.
[0052] For example, when a security patrol robot patrols a fenced area in a warehouse park, the gimbal imaging device captures continuous image frames. The edge computing controller identifies three targets: the first target is a distant vehicle located in the center of the image; the second target is a person located on the right side of the image and moving towards the right boundary; the third target is an animal located in a low-alert area on the left. The edge computing controller constructs alert intent propagation maps for each of the three targets. The first target has a high alert level but a low risk of being lost; the second target, although having a slightly lower alert level, is moving towards the right boundary of the image, and the gimbal has a small remaining rightward turning angle, so the loss risk node and the gimbal margin node significantly increase its tracking intent value during propagation; the third target has a low tracking intent value due to its low alert level and low confidence. Subsequently, the edge computing controller calculates the gimbal field-of-view boundary pressure value for each of the three targets. Because the second target is close to the right boundary, continuously moving to the right, and has a small remaining rightward turning angle, the second target has the highest gimbal field-of-view boundary pressure value. After fusing the tracking intent value and the gimbal field-of-view boundary pressure value, the edge computing controller determines the second target as the current primary tracking target. Next, the edge computing controller generates a rightward gimbal pre-turn angle based on the center coordinates and movement direction of the second target. Based on its higher boundary pressure, it generates a segmented gimbal rotation curve including a rapid interception phase, a pressure release phase, and a low-speed locking phase. Simultaneously, the environmental perception device reports that there are no nearby obstacles to the robot's right front, and the circumferential travel distance of the chassis meets the compensation conditions. The edge computing controller then generates a rightward chassis compensation turn angle. During execution, the gimbal first quickly turns to the predicted position of the second target, then slowly moves the chassis to the right to compensate for the turn, gradually returning the gimbal to the middle working angle range, ultimately maintaining the second target near the center of the image during the low-speed locking phase.
[0053] Through the above implementation methods, the edge computing controller can comprehensively judge the target's alert intention and the pressure at the field of view boundary in complex scenarios such as multiple targets, targets rapidly approaching the image boundary, and gimbal approaching mechanical limits, and generate rapid tracking commands in coordination between the gimbal and the chassis. This improves the tracking stability and response speed of the security surveillance robot for fast-moving targets. The edge computing controller is also used to control the gimbal imaging device to rotate according to the gimbal advance angle and the gimbal segment rotation speed curve, control the mobile chassis to compensate for steering according to the chassis compensation angle, and generate a target afterimage sector when the confidence of the current main tracking target is lower than the tracking threshold and control the gimbal imaging device to perform re-acquisition within the target afterimage sector.
[0054] In one specific implementation, after determining the current primary tracking target, the edge computing controller generates the gimbal advance angle, the gimbal segmented rotation speed curve, and the chassis compensation angle, and sends these parameters to the gimbal control board and the chassis controller, respectively. The gimbal control board controls the rotation of the gimbal imaging device based on the gimbal advance angle and the gimbal segmented rotation speed curve, while the chassis controller controls the moving chassis to perform compensated steering based on the chassis compensation angle. During execution, the edge computing controller continuously receives gimbal angle feedback, chassis attitude feedback, and continuous image frame feedback to determine whether to continue execution, reduce speed, regenerate instructions, or enter the re-acquisition process.
[0055] After determining the current primary tracking target, the edge computing controller generates the gimbal advance angle based on the target's center coordinates, direction of movement, current gimbal angle, and remaining gimbal rotation angle. This advance gimbal angle includes a horizontal advance angle and a pitch advance angle. For example, when the target is located on the right side of the image and moves continuously to the right, the edge computing controller determines the horizontal advance angle as a rightward angle; when the target is located above the image and moves upward, the edge computing controller determines the pitch advance angle as an upward angle. The advance gimbal angle does not only correspond to the deviation of the target's current center coordinates from the image center, but also incorporates the target's direction of movement to compensate for the target's potential location, allowing the gimbal to preferentially turn towards the predicted position of the target. Upon receiving the advance gimbal angle, the gimbal control board converts it into target angles for the gimbal's horizontal and pitch rotation motors. The gimbal angle detection device provides real-time feedback of the horizontal and pitch angles, and the edge computing controller or gimbal control board determines whether the gimbal is approaching the target angle based on the feedback angles. When the error between the actual angle of the gimbal and the target angle is less than the preset angle error, it is considered that the gimbal has completed the early angle rotation or entered the low-speed locking stage.
[0056] In one specific implementation, the segmented rotation speed curve of the gimbal includes a rapid acquisition segment, a pressure release segment, and a low-speed locking segment. Instead of simply sending a single speed value to the gimbal, the edge computing controller generates a rotation speed curve that varies with the execution phase based on the target execution value of the current primary tracked target, the gimbal's field-of-view boundary pressure value, the target center deviation, and the remaining gimbal rotation angle.
[0057] The fast acquisition phase occurs when the currently tracked target is far from the image center or when the target's execution value is high. During this phase, the gimbal control board controls the gimbal to rotate at a high angular velocity in the direction corresponding to the pre-planned gimbal rotation angle, bringing the target into the vicinity of the gimbal's field of view as quickly as possible. If the inter-frame displacement of the target towards the image boundary is large, the duration of the fast acquisition phase can be appropriately extended; if the target center deviation decreases rapidly, the fast acquisition phase ends early to avoid gimbal overshoot.
[0058] In the pressure release phase, the gimbal enters a pressure release phase when the pressure value at the gimbal's field of view boundary of the currently tracked target exceeds the compensation trigger threshold, or when the current angle of the gimbal approaches the preset mechanical limit angle. During this phase, the edge computing controller reduces the gimbal's rotation speed and coordinates with the compensated steering of the moving chassis to gradually return the current angle of the gimbal to the preset intermediate working angle range. This process prevents the gimbal from continuing to rotate rapidly towards the mechanical limit direction and reduces image jitter caused by the simultaneous movement of the gimbal and chassis.
[0059] In the low-speed locking phase, once the current primary target enters the neighborhood of the image center, the gimbal enters the low-speed locking phase. The edge computing controller, based on the residual deviation between the current primary target's center coordinates and the image center, controls the gimbal to make small, low-speed corrections, keeping the primary target stably near the center of the image. If the target center deviation remains below a preset deviation threshold during the low-speed locking phase, the current gimbal angle is maintained; if the target center deviation increases again, the gimbal's advance rotation angle and segmented rotation speed curves are regenerated. Through these segmented rotation speed curves, the system can quickly intercept targets with large deviations, and when the gimbal's boundary pressure is high, it coordinates with the chassis to release pressure, ensuring stable locking after the target enters the central region, avoiding overshoot, jitter, or slow response caused by traditional single-speed control.
[0060] In one specific implementation, before generating the chassis compensation angle, the edge computing controller first determines whether the gimbal's field-of-view boundary pressure value of the current primary tracking target is greater than the compensation trigger threshold, and whether the circumferential obstacle distance of the moving chassis in the pressure release direction is greater than the safe distance. When the gimbal's field-of-view boundary pressure value is greater than the compensation trigger threshold, and the circumferential obstacle distance in the pressure release direction is greater than the safe distance, the edge computing controller sends the chassis compensation angle to the chassis controller. The chassis controller controls the drive motor and steering actuator to slowly turn the moving chassis along the pressure release direction. The pressure release direction refers to the direction that moves the current primary tracking target away from the image boundary and returns the current angle of the gimbal to a preset intermediate working angle range. For example, if the current primary tracking target continues to move towards the right boundary of the image, the gimbal has already rotated a large angle to the right, and the remaining right-side gimbal rotation angle is small, then the pressure release direction is to the right or to the right front. If the circumferential obstacle distances to the right front and right sides are both greater than the safe distance, the edge computing controller generates a right-side chassis compensation angle. After the mobile chassis performs compensated steering, the robot adjusts towards the target direction, and the horizontal angle of the gimbal relative to the robot chassis gradually decreases, thus restoring the gimbal's usable right-hand turning angle. When the pressure value at the gimbal's field of view boundary exceeds the compensation trigger threshold, but the circumferential obstacle distance in the pressure release direction is less than or equal to the safe distance, the edge computing controller prohibits chassis compensated steering. At this time, the system does not directly drive the mobile chassis closer to the obstacle, but instead reduces the rotation speed of the fast acquisition segment in the gimbal's segmented rotation speed curve, and tries to keep the target within the field of view through low-speed gimbal correction. If the target continues to approach the image boundary and the chassis still does not have the conditions for compensation, the edge computing controller can further trigger the target afterimage sector preparation process so that the target can be quickly reacquired after a brief loss.
[0061] In a preferred embodiment, the edge computing controller employs a gimbal-priority, chassis-assisted execution timing. Specifically: when the target deviates from the image center but the gimbal has sufficient remaining rotation angle, only the gimbal is controlled to execute the advance rotation angle and segmented rotation speed curve; when the target continues to move towards the image boundary, the gimbal's remaining rotation angle is insufficient, or the gimbal's field-of-view boundary pressure value exceeds the compensation trigger threshold, the moving chassis is then controlled to execute compensating steering. The gimbal execution cycle can be set shorter than the chassis execution cycle. For example, the gimbal control board receives angle and speed control commands with a shorter cycle to quickly respond to target center deviation; the chassis controller executes compensating steering with a longer cycle to avoid frequent chassis start-stop or sharp turns. Through this dual-timescale control method, the gimbal undertakes the task of rapid response, while the chassis undertakes the task of slow attitude compensation, thus balancing rapid target tracking and robot driving safety. During execution, the edge computing controller continuously compares the current angle of the gimbal with the preset intermediate working angle range. When the chassis compensating steering gradually brings the current angle of the gimbal back to the intermediate working angle range, the edge computing controller reduces or stops the chassis compensating rotation angle, retaining only the low-speed lock control of the gimbal. This can prevent the target from deviating from the frame due to excessive steering of the chassis.
[0062] During normal tracking of the current primary target, the edge computing controller continuously caches the state information of the current primary target for several recent frames, preferably 3 to 8 recent frames. The cached information includes center coordinates, scale changes, direction of angle changes, location within the warning sector, and the current angular velocity of the gimbal. When the confidence level of the current primary target falls below the tracking threshold due to occlusion, lighting changes, excessively fast movement, or proximity to the image boundary, the edge computing controller does not immediately abandon the target or initiate a global scan. Instead, it generates a target afterimage sector based on the cached information. Specifically, the edge computing controller determines the center direction of the target afterimage sector based on the direction of change of the center coordinates in consecutive frames before the target was lost and the current angular velocity of the gimbal. For example, if the target moved continuously to the right of the image before being lost, and the gimbal was rotating to the right, the center direction of the target afterimage sector is set to the gimbal direction corresponding to the right of the target loss point; if the target moved to the upper right before being lost, the center direction of the target afterimage sector is set to the horizontal and pitch combination direction corresponding to the upper right.
[0063] The edge computing controller also determines the expansion angle of the target afterimage sector based on the inter-frame displacement velocity and scale change before the target is lost. When the target moves quickly or changes scale significantly before being lost, it indicates that the target may have moved to a wider area in a short time, so the expansion angle of the target afterimage sector is larger. When the target moves slowly and changes scale little, the expansion angle of the target afterimage sector is smaller to avoid mis-locking other targets due to an excessively wide search range. For example, if the current primary tracked target is a fast-running person whose center coordinates continuously move to the right before being lost, and the detection box area gradually increases, the edge computing controller determines that the person may reappear from behind a nearby occlusion on the right, thus generating a target afterimage sector with a larger expansion angle centered on the right. If the target is a person moving slowly in the distance with little scale change before being lost, a narrower target afterimage sector is generated.
[0064] After generating the target afterimage sector, the edge computing controller controls the gimbal imaging device to perform recapture within the target afterimage sector. Recapture is not a global cruise scan, but rather a directional search within a local sector where the target is most likely to reappear. In one specific implementation, the gimbal imaging device first turns towards the center direction of the target afterimage sector, acquiring image frames and performing target detection near the center. If no candidate target matching the current primary tracked target is detected, the gimbal scans according to a fan-shaped search sequence expanding outward from the center direction of the target afterimage sector. This fan-shaped search sequence may sequentially include the center direction, a small-angle position to the left of the center direction, a small-angle position to the right of the center direction, a larger-angle position to the left of the center direction, and a larger-angle position to the right of the center direction, until the target afterimage sector is covered. The edge computing controller acquires continuous image frames in each search direction and detects candidate targets. If a candidate target is detected, it is matched against cached information based on the candidate target's category, scale change direction, appearance location, movement direction, and the vigilance sector it belongs to. The edge computing controller restores a candidate target as the original primary tracking target only if the candidate target's scale change direction, angle change direction, and location within the warning sector are consistent with those of the current primary tracking target before it was lost. If a candidate target, although of the same category, has a motion direction that is significantly inconsistent with the buffered angle change direction, or if its location is outside the target afterimage sector, the edge computing controller will not lock the candidate target as the original primary tracking target to reduce the risk of false locking. After the original primary tracking target is successfully recaptured, the edge computing controller resumes the process of generating the gimbal advance angle and segmented rotation speed curve for that target and recalculates its gimbal field-of-view boundary pressure value. If the target is not recaptured within the preset search time or preset number of scans, the edge computing controller ends the afterimage sector search and recalculates the tracking intent value and gimbal field-of-view boundary pressure value based on the warning target in the current image frame to determine a new primary tracking target.
[0065] For example, when a security patrol robot is patrolling near a factory fence, the gimbal imaging device detects a person approaching the fence from the right. The edge computing controller identifies this person as the current primary tracking target and generates a rightward gimbal pre-turn angle and a segmented gimbal rotation curve including a rapid acquisition phase, a pressure release phase, and a low-speed lock-on phase. The gimbal first rapidly rotates to the right during the rapid acquisition phase, bringing the person near the center of the image. Subsequently, the person continues to move towards the right boundary of the image, increasing the pressure value at the gimbal's field of view boundary while simultaneously decreasing the remaining rightward turn angle. The environmental perception device detects no nearby obstacles in front of the robot's right side, and the edge computing controller sends a rightward chassis compensation turn angle to the chassis controller. The moving chassis slowly compensates by turning to the right, gradually reducing the rightward turn angle of the gimbal relative to the chassis, and the gimbal returns to the preset intermediate working angle range. During the tracking process, the person briefly passes behind a pillar, causing the confidence level of the current primary tracking target to fall below the tracking threshold. The edge computing controller generates a target afterimage sector centered on the right side of the pillar based on the center coordinates, scale changes, angular change direction, location of the warning sector, and the current angular velocity of the gimbal in the most recent frames before the person disappeared. It then controls the gimbal to first search the center direction of this afterimage sector and then expand the search to both sides. After the person reappears on the right side of the pillar, the edge computing controller restores them to the original primary tracking target based on their motion direction and scale changes matching the cached information, and continues the gimbal's rapid tracking.
[0066] Through the above implementation methods, the system enables the gimbal to quickly capture the target, the chassis to release the gimbal boundary pressure in a timely manner, and to perform local re-capture based on the target afterimage sector when the target is temporarily obscured or the confidence level decreases, thereby improving the continuous tracking capability and command execution reliability of the security and surveillance robot in complex dynamic scenarios.
[0067] Preferably, the gimbal imaging device includes a visible light camera and / or an infrared camera; the gimbal angle detection device includes a horizontal angle encoder and a pitch angle encoder; the environmental sensing device includes an inertial measurement unit and a circumferential ranging sensor; the mobile chassis includes a chassis controller, a drive motor and a steering actuator; and the edge computing controller is communicatively connected to the gimbal imaging device, the gimbal angle detection device, the environmental sensing device and the mobile chassis.
[0068] Preferably, the edge computing controller is used to determine the center coordinates of the warning target based on the target detection boxes in consecutive image frames; determine the scale change of the warning target based on the area change or diagonal length change of the target detection boxes in adjacent frames; determine the motion direction of the warning target based on the displacement direction of the center coordinates of adjacent frames; and fuse the target detection confidence with the target matching consistency of adjacent frames to obtain the confidence of the warning target.
[0069] Preferably, the target node is used to characterize the center coordinates, scale changes, direction of movement, and confidence level of the warning target; the warning sector node is used to characterize the warning level and core warning position of the warning area where the warning target is located; the gimbal margin node is used to characterize the remaining rotation angle from the current angle of the gimbal to the preset mechanical limit angle; the chassis passage node is used to characterize the circumferential passage distance of the mobile chassis in different steering directions; the obstacle suppression node is used to characterize the degree of restriction of obstacles on the compensated steering of the mobile chassis; and the loss risk node is used to characterize the risk of the warning target leaving the gimbal's field of vision.
[0070] Preferably, the warning intent propagation map includes: an intrusion propagation edge from the target node to the warning sector node; a boundary risk propagation edge from the target node to the lost risk node; a gimbal margin propagation edge from the lost risk node to the gimbal margin node; a compensation demand propagation edge from the gimbal margin node to the chassis passage node; and a reverse suppression propagation edge from the obstacle suppression node to the chassis passage node. The weight of the intrusion propagation edge increases with the depth of the warning target's intrusion into the warning sector; the weight of the boundary risk propagation edge increases with the inter-frame displacement of the warning target towards the image boundary; the weight of the gimbal margin propagation edge increases with the decrease of the remaining gimbal rotation angle; and the weight of the reverse suppression propagation edge increases with the decrease of the circumferential obstacle distance.
[0071] Preferably, the edge computing controller is used to initialize the node value of the target node with the alert level, confidence level and intrusion depth of the alert sector of the alert target; propagate the node value of the target node along the propagation edge in the alert intent propagation graph for 2 to 4 rounds; deduct the reverse inhibition amount from the obstacle inhibition node after each round of propagation, and normalize the node value after propagation; and fuse the target node value and the lost risk node value after propagation to obtain the tracking intent value of the alert target.
[0072] Preferably, the gimbal field of view boundary pressure value is determined by boundary distance pressure, boundary approach pressure, gimbal margin pressure, and chassis compensation release amount; wherein, the smaller the distance from the warning target to the image boundary, the greater the boundary distance pressure; the greater the inter-frame displacement of the warning target toward the image boundary, the greater the boundary approach pressure; the smaller the remaining gimbal rotation angle, the greater the gimbal margin pressure; the greater the circumferential travel distance of the moving chassis in the direction that moves the warning target away from the image boundary, the greater the chassis compensation release amount; the edge computing controller is used to superimpose the boundary distance pressure, boundary approach pressure, and gimbal margin pressure, and subtract the chassis compensation release amount to obtain the gimbal field of view boundary pressure value.
[0073] Preferably, the edge computing controller is used to weightedly fuse the tracking intent value of each alert target and the gimbal's field of view boundary pressure value to obtain the target execution value of each alert target; determine the alert target with the highest target execution value as the current primary tracking target; determine the gimbal's advance rotation angle based on the center coordinates, movement direction, current gimbal angle, and remaining gimbal rotation angle of the current primary tracking target; and generate a segmented gimbal rotation speed curve including a rapid acquisition segment, a pressure release segment, and a low-speed locking segment based on the gimbal's field of view boundary pressure value of the current primary tracking target.
[0074] Preferably, when the gimbal field of view boundary pressure value of the current main tracking target is greater than the compensation trigger threshold, and the circumferential obstacle distance of the moving chassis in the pressure release direction is greater than the safe distance, the edge computing controller generates a chassis compensation angle; the direction of the chassis compensation angle is to move the current main tracking target away from the image boundary and to return the current angle of the gimbal to the preset intermediate working angle range; when the gimbal field of view boundary pressure value of the current main tracking target is greater than the compensation trigger threshold, and the circumferential obstacle distance of the moving chassis in the pressure release direction is less than or equal to the safe distance, the edge computing controller prohibits the generation of chassis compensation angle and reduces the rotation speed of the fast interception segment in the gimbal segmented rotation speed curve.
[0075] Preferably, the edge computing controller is used to cache the center coordinates, scale changes, angle change directions, warning sectors, and current angular velocity of the current primary target in the most recent 3 to 8 frames before the confidence of the current primary target falls below the tracking threshold; when the confidence of the current primary target falls below the tracking threshold, it determines the center direction of the target afterimage sector based on the cached angle change directions and the current angular velocity of the gimbal; it determines the expansion angle of the target afterimage sector based on the inter-frame displacement velocity and scale change before the current primary target was lost; and it controls the gimbal imaging device to perform recapture according to a fan-shaped search sequence expanding outward from the center direction of the target afterimage sector.
[0076] This application discloses a PTZ-based rapid tracking security and surveillance robot command execution system. The hardware includes: a mobile chassis, a PTZ imaging device, a PTZ angle detection device, an environmental sensing device, an edge computing controller, a PTZ control board, a chassis controller, a power management device, a communication device, and an audio-visual warning device. All hardware components are connected via power lines, data communication lines, and control lines, forming a closed-loop system from sensing and data acquisition to edge computing, PTZ execution, chassis compensation, and warning output.
[0077] The mobile chassis serves as the load-bearing and movement foundation for the security and surveillance robot, comprising the chassis body, drive wheels, drive motors, steering actuators, braking mechanisms, chassis controller, wheel speed sensors, and an odometer. The chassis controller is connected to the drive motors, steering actuators, braking mechanisms, wheel speed sensors, and odometer. It receives chassis compensation angle, steering speed, and stop commands from the edge computing controller and controls the drive motors and steering actuators to complete the chassis compensation steering. The mobile chassis also supports the gimbal imaging device, environmental sensing device, edge computing controller, power management device, and communication device.
[0078] The gimbal imaging device is mounted on a mobile chassis and includes a two-degree-of-freedom gimbal, a visible light camera, an infrared camera, and a zoom lens. The two-degree-of-freedom gimbal includes a horizontal rotation mechanism, a pitch rotation mechanism, a horizontal drive motor, and a pitch drive motor. The visible light camera and the infrared camera are fixedly mounted on the gimbal's pitch bracket and can rotate synchronously with the gimbal in the horizontal and pitch directions. The visible light camera is used to acquire continuous image frames in well-lit environments, while the infrared camera is used to acquire continuous image frames in nighttime, low-light, backlight, or smoky environments. The cameras connect to the edge computing controller via Ethernet, USB, MIPI, or LVDS interfaces, sending continuous image frames to the edge computing controller.
[0079] The gimbal angle detection device includes a horizontal angle encoder and a pitch angle encoder. The horizontal angle encoder is mounted on the horizontal rotation axis of the gimbal and is used to acquire the horizontal angle of the gimbal relative to the forward reference direction of the moving chassis. The pitch angle encoder is mounted on the pitch rotation axis of the gimbal and is used to acquire the pitch angle of the gimbal relative to the horizontal reference plane. Both the horizontal and pitch angle encoders are connected to the gimbal control board or edge computing controller and can transmit angle data via CAN bus, RS485, SPI, I2C, or a serial interface. The edge computing controller determines the remaining rotation angle of the gimbal based on the acquired current gimbal angle and the preset mechanical limit angle.
[0080] The gimbal control board is located inside the gimbal or on the mobile chassis, communicating with the edge computing controller and connecting to the horizontal drive motor, pitch drive motor, horizontal angle encoder, and pitch angle encoder. After generating the gimbal advance angle and segmented speed curves, the edge computing controller sends gimbal control commands to the gimbal control board. The gimbal control board controls the rotation of the horizontal and pitch drive motors based on the gimbal advance angle, performing different speed controls according to rapid interception, pressure release, and low-speed lock-up segments. The gimbal control board also feeds back the real-time angle, speed status, execution completion status, and abnormal status of the gimbal to the edge computing controller.
[0081] The environmental sensing device includes one or more of the following: an inertial measurement unit (IMU), a circumferential ranging sensor, a lidar, a millimeter-wave radar, an ultrasonic sensor, a depth camera, and a collision sensor. The IMU is fixedly mounted in the center of the mobile chassis or near the chassis controller to acquire chassis attitude, including roll angle, pitch angle, yaw angle change, angular velocity, and acceleration. The circumferential ranging sensor is positioned at the front, left front, right front, left side, right side, and rear of the mobile chassis to acquire the distance to circumferential obstacles in different directions around the robot. The circumferential ranging sensor can be a lidar, millimeter-wave radar, ultrasonic sensor, or depth camera. The edge computing controller determines the circumferential passable distance of the chassis based on the circumferential obstacle distance. The collision sensor is located on the chassis shell or anti-collision edge to send a collision signal to the edge computing controller when the robot encounters an obstacle. The environmental sensing device connects to the edge computing controller via a CAN bus, Ethernet, RS485, UART, or USB interface.
[0082] The edge computing controller is the core control hardware of the system and can be an embedded computing platform, an industrial computer, an AI edge computing box, or a robot main control board. The edge computing controller connects to the gimbal imaging device, gimbal control board, gimbal angle detection device, environmental perception device, chassis controller, communication device, and audio-visual warning device. The edge computing controller receives continuous image frames, the current angle of the gimbal, chassis attitude, and circumferential obstacle distances; extracts the center coordinates, scale changes, direction of motion, and confidence level of the warning target from the continuous image frames; constructs a warning intent propagation map; generates gimbal field-of-view boundary pressure values; determines the current primary tracking target; generates the gimbal advance rotation angle, gimbal segmented rotation speed curve, chassis compensation rotation angle, and target afterimage sector; and issues execution commands to the gimbal control board and chassis controller.
[0083] The chassis controller is located inside the mobile chassis, communicating with the edge computing controller and connected to the drive motor, steering actuator, braking mechanism, wheel speed sensor, and odometer. After generating the chassis compensation steering angle, the edge computing controller sends the compensation steering angle, compensation steering speed, and safety limit parameters to the chassis controller. The chassis controller then controls the mobile chassis to perform compensation steering along the pressure release direction, moving the currently tracked target away from the image boundary and returning the gimbal's current angle to the preset intermediate working angle range. The chassis controller also feeds back chassis motion status, wheel speed, mileage, steering angle, and fault information to the edge computing controller.
[0084] The power management unit includes a battery pack, a power management module, a voltage regulator module, a charging interface, and a power protection circuit. The battery pack powers the mobile chassis, gimbal imaging device, gimbal control board, edge computing controller, environmental sensing device, communication device, and audio-visual warning device. The power management module monitors battery level, voltage, current, and temperature, and provides feedback on the power status to the edge computing controller. The voltage regulator module converts the battery output voltage to the operating voltage required by each hardware component, such as providing a stable DC power supply to the edge computing controller and the corresponding voltages to the camera, radar, encoder, and control board.
[0085] The communication device includes a wireless communication module and / or a wired communication interface. The wireless communication module can be a 4G module, 5G module, Wi-Fi module, private network communication module, or LoRa communication module. The wired communication interface can be an Ethernet interface or a serial communication interface. The communication device connects to the edge computing controller and is used to send images of the monitored target, tracking status, alarm information, robot position, gimbal angle, and chassis status to the remote security platform. It can also receive patrol tasks, warning area configurations, target tracking strategies, or manual takeover commands issued by the remote security platform. The audio-visual warning device includes one or more of warning lights, speakers, buzzers, and supplementary lighting. The audio-visual warning device connects to the edge computing controller. When the edge computing controller determines that the currently tracked target is a high-risk warning target, or that the target has entered a core warning position, it can control the warning lights to flash, the speaker to play warning audio, or control the supplementary lighting to enhance the quality of nighttime evidence-gathering images.
[0086] The hardware connections are as follows: The gimbal imaging device connects to the edge computing controller, sending continuous image frames to the edge computing controller via Ethernet, USB, MIPI, or LVDS interfaces. The gimbal angle detection device connects to the gimbal control board / edge computing controller, with the horizontal and vertical angle encoders feeding back the current gimbal angle via CAN, RS485, SPI, or I2C interfaces. The edge computing controller connects to the gimbal control board and the gimbal drive motors, sending the gimbal advance angle and segmented rotation speed curves to the gimbal control board, which then controls the horizontal and vertical drive motors. The environmental sensing device connects to the edge computing controller, with the inertial measurement unit and circumferential distance sensor sending chassis attitude and circumferential obstacle distance to the edge computing controller. The edge computing controller connects to the chassis controller and the mobile chassis actuators, sending the chassis compensation angle to the chassis controller, which then controls the drive motors and steering actuators to complete the compensation steering. The chassis controller connects to the edge computing controller, providing feedback on wheel speed, mileage, steering angle, braking status, and fault status. The edge computing controller connects to the communication device and the remote security platform, uploading alarm information, target images, and robot status via the communication device. The edge computing controller also connects to the audio-visual warning device. When the edge computing controller identifies a high-risk target or a target enters the core warning location, it controls the audio-visual warning device to issue on-site warnings. The power management device connects to all power-consuming hardware. It supplies power to the mobile chassis, gimbal imaging device, gimbal control board, edge computing controller, environmental sensing device, communication device, and audio-visual warning device, and provides feedback on the power status to the edge computing controller.
[0087] In a preferred embodiment, the edge computing controller acts as the master control node, connecting to a visible light camera, an infrared camera, and a lidar via Ethernet; connecting to the chassis controller, gimbal control board, and inertial measurement unit via a CAN bus; connecting to a circumferential ultrasonic sensor via RS485; connecting to a wireless communication module via USB; and connecting to an audible and visual warning device via GPIO or a relay interface. The gimbal control board connects to the horizontal drive motor, the pitch drive motor, the horizontal angle encoder, and the pitch angle encoder. The chassis controller connects to the drive motor, the steering actuator, the braking mechanism, the wheel speed sensor, and the odometer. This connection method enables the edge computing controller to generate gimbal fast tracking commands and chassis compensation commands in real time after receiving image, angle, attitude, and obstacle information, and executes them through the gimbal control board and the chassis controller, thereby realizing gimbal fast tracking and collaborative warning of the security warning robot.
[0088] Compared with the prior art, the technical solution of the present invention has the following beneficial effects: (1) This invention does not generate gimbal control commands solely based on the target's current position in the image. Instead, it constructs a warning intent propagation map for each warning target, incorporating the target's intrusion depth into the warning sector, the target's displacement toward the core warning position, the remaining gimbal rotation angle, the chassis's circumferential passable distance, and the risk of the target leaving the field of view into the propagation calculation to obtain the tracking intent value of each warning target. This allows for the priority selection of the current primary tracking target that is more in need of tracking, more likely to pose a warning risk, or more easily lost when multiple targets appear simultaneously, avoiding disorderly oscillation of the gimbal among multiple targets and improving the targeting and timeliness of the warning response.
[0089] (2) This invention determines the trend of the target leaving the field of view by the pressure value of the gimbal's field of view boundary, instead of simply relying on whether the gimbal angle exceeds a threshold to trigger chassis compensation. The system comprehensively considers the distance from the target to the image boundary, the inter-frame displacement of the target toward the boundary, the remaining gimbal rotation angle, and the circumferential passable distance of the chassis to determine whether to generate a chassis compensation rotation angle. When the boundary pressure is high and the chassis has safe turning conditions, the moving chassis performs compensation turning in advance, so that the gimbal returns to the preset intermediate working angle range; when the circumferential obstacle distance is insufficient, the chassis movement is suppressed and the rotation speed of the gimbal during the fast acquisition segment is reduced, thereby taking into account both fast tracking and motion safety.
[0090] (3) Before the confidence of the current main target decreases, the present invention caches the center coordinates, scale changes, angle change directions, the warning sector it is located in, and the current angular velocity of the PTZ in the most recent several frames. When the confidence of the target is lower than the tracking threshold due to occlusion, edge cutting out of the screen, or brief recognition failure, the present invention generates a target afterimage sector based on the cached information and controls the PTZ to re-acquire it according to a fan-shaped search sequence that expands outward from the center of the afterimage sector. This avoids the problems of large search range, long time consumption, and easy accidental locking of other targets in traditional global scanning, and improves the continuity of evidence collection for warning events and the reliability of the PTZ to quickly resume tracking.
[0091] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products, and therefore this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.
[0092] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.
Claims
1. A gimbal-based rapid tracking command execution system for security and surveillance robots, characterized in that, It includes a gimbal imaging device, a mobile chassis, an environmental sensing device, a gimbal angle detection device, and an edge computing controller; The gimbal imaging device is used to acquire continuous image frames, the gimbal angle detection device is used to obtain the current angle of the gimbal, and the environmental perception device is used to obtain the chassis attitude and the distance to circumferential obstacles. The edge computing controller is used to determine the remaining turning angle of the gimbal based on the current angle of the gimbal, determine the circumferential passable distance of the chassis based on the circumferential obstacle distance, and extract the center coordinates, scale changes, motion direction and confidence of each warning target from continuous image frames; The edge computing controller is also used to construct a warning intent propagation map for each warning target, including target node, warning sector node, gimbal margin node, chassis passage node, obstacle suppression node, and loss risk node, to obtain the tracking intent value; to obtain the gimbal field of view boundary pressure value based on the distance from the warning target to the image boundary, the inter-frame displacement towards the image boundary, the remaining gimbal rotation angle, and the chassis circumferential passage distance; to determine the current main tracking target based on the tracking intent value and the gimbal field of view boundary pressure value, and to generate the gimbal advance rotation angle, the gimbal segmented rotation speed curve, and the chassis compensation rotation angle; The edge computing controller is also used to control the gimbal imaging device to rotate according to the gimbal advance angle and the gimbal segment rotation speed curve, control the mobile chassis to compensate for steering according to the chassis compensation angle, and generate a target afterimage sector when the confidence of the current main tracking target is lower than the tracking threshold and control the gimbal imaging device to perform re-acquisition within the target afterimage sector.
2. The PTZ-based rapid tracking security surveillance robot command execution system according to claim 1, characterized in that, The gimbal imaging device includes a visible light camera and / or an infrared camera; the gimbal angle detection device includes a horizontal angle encoder and a pitch angle encoder; the environmental sensing device includes an inertial measurement unit and a circumferential ranging sensor; the mobile chassis includes a chassis controller, a drive motor and a steering actuator; the edge computing controller is communicatively connected to the gimbal imaging device, the gimbal angle detection device, the environmental sensing device and the mobile chassis respectively.
3. The PTZ-based rapid tracking security surveillance robot command execution system according to claim 1, characterized in that, The edge computing controller is used to determine the center coordinates of the warning target based on the target detection boxes in consecutive image frames; determine the scale change of the warning target based on the area change or diagonal length change of the target detection boxes in adjacent frames; determine the motion direction of the warning target based on the displacement direction of the center coordinates of adjacent frames; and fuse the target detection confidence with the target matching consistency of adjacent frames to obtain the confidence of the warning target.
4. The PTZ-based rapid tracking security surveillance robot command execution system according to claim 1, characterized in that, The target node is used to characterize the center coordinates, scale changes, direction of movement, and confidence level of the warning target; the warning sector node is used to characterize the warning level and core warning position of the warning area where the warning target is located; the gimbal margin node is used to characterize the remaining rotation angle from the current angle of the gimbal to the preset mechanical limit angle; the chassis passage node is used to characterize the circumferential passage distance of the mobile chassis in different steering directions; the obstacle suppression node is used to characterize the degree of restriction of obstacles on the compensated steering of the mobile chassis. The lost risk node is used to characterize the risk of the warning target leaving the pan-tilt-zoom (PTZ) camera's field of view.
5. The PTZ-based rapid tracking security surveillance robot command execution system according to claim 4, characterized in that, The warning intent propagation map includes: intrusion propagation edges from the target node to the warning sector node; boundary risk propagation edges from the target node to the lost risk node; gimbal margin propagation edges from the lost risk node to the gimbal margin node; compensation demand propagation edges from the gimbal margin node to the chassis passage node; and reverse suppression propagation edges from the obstacle suppression node to the chassis passage node. The weight of the intrusion propagation edge increases with the depth of the warning target's intrusion into the warning sector; the weight of the boundary risk propagation edge increases with the inter-frame displacement of the warning target towards the image boundary; the weight of the gimbal margin propagation edge increases with the decrease of the remaining gimbal rotation angle; and the weight of the reverse suppression propagation edge increases with the decrease of the circumferential obstacle distance.
6. The PTZ-based rapid tracking command execution system for a security surveillance robot according to claim 5, characterized in that, The edge computing controller is used to initialize the node value of the target node with the alert level, confidence level and intrusion depth of the alert sector; propagate the node value of the target node along the propagation edge in the alert intent propagation graph for 2 to 4 rounds; deduct the reverse inhibition amount from the obstacle inhibition node after each round of propagation, and normalize the node value after propagation. The target node value and the lost risk node value are fused together to obtain the tracking intent value of the warning target.
7. The PTZ-based rapid tracking command execution system for a security surveillance robot according to claim 1, characterized in that, The gimbal field-of-view boundary pressure value is determined by boundary distance pressure, boundary approach pressure, gimbal margin pressure, and chassis compensation release. Specifically, the smaller the distance from the target to the image boundary, the greater the boundary distance pressure; the greater the inter-frame displacement of the target towards the image boundary, the greater the boundary approach pressure; the smaller the remaining gimbal rotation angle, the greater the gimbal margin pressure; and the greater the circumferential travel distance of the moving chassis in the direction that moves the target away from the image boundary, the greater the chassis compensation release. The edge computing controller is used to superimpose the boundary distance pressure, boundary approach pressure, and gimbal margin pressure, and subtract the chassis compensation release to obtain the gimbal field-of-view boundary pressure value.
8. The PTZ-based rapid tracking security surveillance robot command execution system according to claim 7, characterized in that, The edge computing controller is used to weightedly fuse the tracking intent value of each alert target and the gimbal's field of view boundary pressure value to obtain the target execution value of each alert target; the alert target with the highest target execution value is determined as the current main tracking target; and the gimbal advance rotation angle is determined based on the center coordinates, movement direction, current gimbal angle, and remaining gimbal rotation angle of the current main tracking target. Based on the gimbal field of view boundary pressure value of the current main tracking target, a segmented gimbal rotation speed curve is generated, including a rapid acquisition segment, a pressure release segment, and a low-speed lock segment.
9. The PTZ-based rapid tracking command execution system for a security surveillance robot according to claim 1, characterized in that, When the gimbal field of view boundary pressure value of the current main tracking target is greater than the compensation trigger threshold, and the circumferential obstacle distance of the moving chassis in the pressure release direction is greater than the safe distance, the edge computing controller generates a chassis compensation angle; the direction of the chassis compensation angle is to move the current main tracking target away from the image boundary and to return the current angle of the gimbal to the preset intermediate working angle range. When the gimbal field of view boundary pressure value of the current main tracking target is greater than the compensation trigger threshold, and the circumferential obstacle distance of the mobile chassis in the pressure release direction is less than or equal to the safe distance, the edge computing controller prohibits the generation of chassis compensation rotation angle and reduces the rotation speed of the fast interception segment in the gimbal segment rotation speed curve.
10. A PTZ-based rapid tracking command execution system for a security surveillance robot according to claim 1, characterized in that, The edge computing controller is used to cache the center coordinates, scale changes, angle change direction, warning sector, and current angular velocity of the current main tracking target in the last 3 to 8 frames before the confidence of the current main tracking target is lower than the tracking threshold. When the confidence of the current primary target is lower than the tracking threshold, the center direction of the target ghost sector is determined based on the cached angle change direction and the current angular velocity of the gimbal; the expansion angle of the target ghost sector is determined based on the inter-frame displacement velocity and scale change before the current primary target is lost. It also controls the gimbal imaging device to perform re-capture according to a fan-shaped search sequence that expands outward from the center of the target afterimage sector.