Remote video acquisition method and system based on low-delay active view angle control of cooperative mechanical arm
By establishing a coordinate mapping relationship between the terminal and the end effector of the robotic arm through collaborative robotic arm, smoothing filtering and trajectory planning are performed, solving the problems of viewpoint control and stability in remote video technology. This achieves low-latency, secure viewpoint control and stable video stream, improving the efficiency of remote observation and user experience.
Patent Information
- Application Number
- CN202511606773.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-02-10
AI Technical Summary
Existing remote video technologies have shortcomings in terms of perspective control, image stability, efficiency of multi-view switching, labor costs, security, network adaptability, and intelligent optimization, resulting in slow response speed, inaccurate positioning, significant security risks, and poor user experience.
A low-latency active view control method based on a collaborative robotic arm is adopted. By collecting remote terminal posture data in real time, a coordinate mapping relationship between the terminal and the end effector of the robotic arm is established. Smoothing filtering and trajectory planning are performed. Combined with virtual security envelope and layered coding transmission technology, active, low-latency view control and stable video stream are achieved.
It achieves efficient, stable, and secure perspective control for remote observation, reduces labor costs, improves network adaptability and user experience, and ensures operational safety.
Smart Images

Figure CN121509816A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of remote video, and particularly relates to a remote video acquisition method and system based on a cooperative mechanical arm and low-latency active view angle control. BACKGROUND
[0002] In remote observation, guidance and other scenarios, remote video technology plays an important role. However, existing remote video solutions have many problems, limiting their effectiveness and efficiency in practical applications. Most existing solutions rely on voice instructions and manual movement to adjust the camera angle, lack a low-latency, continuous "remote terminal pose → camera pose" direct control chain, resulting in passive remote personnel view angle and slow response speed. Handheld devices or traditional gimbals are prone to jitter during movement, lack effective layered filtering and prediction to suppress high-frequency disturbance mechanisms, resulting in angle jumps and unstable pictures. In the detection process, multiple observation objects and angles are involved, and the view angle needs to be frequently switched, but the existing technology requires manual repositioning when switching the view angle, which is inefficient and prone to inaccurate positioning problems. In addition, to achieve multi-view switching or stable shooting, additional photographers or assistants are usually needed to operate the equipment, increasing labor costs. Traditional gimbals or simple mechanical arms do not establish a virtual safety envelope, which may cause overreach or interference with operators, equipment, etc. during operation, posing a safety hazard. In the case of network fluctuations or high packet loss rates, the video transmission experience of existing technologies is significantly degraded, and problems such as stuttering and blurred pictures are prone to occur, affecting the real-time and accuracy of remote observation. In the existing technology, the generation and invocation of preset view angles are not intelligent enough to automatically optimize them according to user operation habits and commonly used view angles, resulting in low hit rates for commonly used view angles and poor user experience. In summary, the existing technology has obvious shortcomings in view angle control, picture stability, multi-view switching efficiency, labor costs, safety, network adaptability and intelligent optimization, and there is an urgent need for a new remote video acquisition method and system that can solve the above problems to improve the efficiency, quality and safety of remote observation. SUMMARY
[0003] The present application proposes a remote video acquisition method and system based on a cooperative mechanical arm and low-latency active view angle control to solve the above-mentioned problems of existing technology.
[0004] To achieve the above-mentioned purpose, the present application provides a remote video acquisition method based on a cooperative mechanical arm and low-latency active view angle control, comprising the following steps:
[0005] Real-time acquisition of remote terminal pose data;
[0006] Based on the attitude data of the remote terminal, the desired pose of the end effector of the collaborative robotic arm is calculated through a pre-established coordinate mapping relationship, wherein the coordinate mapping relationship relates the local coordinate system of the remote terminal and the camera coordinate system of the end effector of the collaborative robotic arm.
[0007] The desired pose sequence is smoothed by a smoothing filter to generate a smooth pose sequence.
[0008] The smooth pose sequence is sent to the robotic arm control terminal;
[0009] The robotic arm control unit performs trajectory planning based on a smooth pose sequence and under the constraints of a preset virtual safety envelope, generating the robotic arm's execution trajectory.
[0010] The collaborative robotic arm is controlled to move according to the execution trajectory of the robotic arm.
[0011] Video streams are captured by a camera mounted at the end of a collaborative robotic arm;
[0012] The video stream is encoded and transmitted to the remote terminal for display.
[0013] Optionally, the coordinate mapping relationship includes:
[0014] By identifying markers placed in the ground coordinate system using a camera installed on the remote terminal, the transformation relationship T_LG from the local coordinate system of the remote terminal to the ground coordinate system is determined.
[0015] The transformation relationship T_GW from the base coordinate system to the ground coordinate system of the collaborative robotic arm is obtained by measurement;
[0016] The transformation relationship T_WC from the end-effector coordinate system to the base coordinate system is obtained through the kinematic model of the collaborative robotic arm.
[0017] Based on the transformation relationships T_LG, T_GW, and T_WC, a coordinate mapping relationship is established from the local coordinate system of the remote terminal to the camera coordinate system at the end of the collaborative robotic arm.
[0018] Optionally, generating the smooth pose sequence includes:
[0019] High-frequency jitter determination is performed on the desired pose sequence;
[0020] Based on the jitter determination result, an adaptive smoothing window is selected to smooth the desired pose sequence;
[0021] The smoothed pose sequence is subjected to exponential moving average filtering and Kalman filtering;
[0022] The results of the exponential moving average filter and the Kalman filter are combined to output a smooth pose sequence.
[0023] Optionally, the trajectory planning under the constraints of a preset virtual safety envelope includes:
[0024] Determine whether the smooth pose is within the virtual safety envelope;
[0025] If the smooth pose is located within the virtual safety envelope, then it is taken as the target pose;
[0026] If the smooth pose is outside the virtual safe envelope, the smooth pose is projected onto the boundary of the virtual safe envelope, and the projected pose is used as the target pose.
[0027] Based on the target pose, interpolation and speed limiting are performed to generate the execution trajectory.
[0028] Optionally, encoding the video stream includes:
[0029] The video stream is determined as the main bitstream;
[0030] Based on the region of interest selected in the video stream, the region of interest is tracked, and a secondary bitstream containing the region of interest is generated based on the tracking results, wherein the encoding quality of the secondary bitstream is higher than the encoding quality of the corresponding region in the main bitstream;
[0031] The main stream and the secondary stream are transmitted synchronously to the remote terminal.
[0032] Optionally, the method also includes the generation and invocation of preset viewpoints:
[0033] When the remote terminal is in follow mode, the stationary pose of the collaborative robotic arm is recorded to form a candidate pose set;
[0034] Cluster analysis is performed on the candidate pose set to generate preset poses;
[0035] The collaborative robotic arm is controlled according to the preset pose, and the collaborative robotic arm is controlled to move to the preset pose and remain fixed.
[0036] Optionally, the K-Medoids algorithm is used to perform cluster analysis on the candidate pose set, whose distance metric function integrates the position information of the pose and the attitude information represented by quaternions.
[0037] This invention also provides a remote video acquisition system based on low-latency active viewpoint control using a collaborative robotic arm, comprising:
[0038] A remote terminal used to acquire attitude data, send pose commands, and receive and display video streams;
[0039] A collaborative robotic arm with a camera mounted at its end effector;
[0040] The robotic arm control terminal is communicatively connected to the collaborative robotic arm and the remote terminal. It is used to control the movement of the collaborative robotic arm according to the received pose instructions, and to encode and transmit the video stream captured by the camera.
[0041] Compared with the prior art, the present invention has the following advantages and technical effects:
[0042] This invention establishes a direct control chain between the terminal posture and the robotic arm's end effector pose, enabling remote observers to control their viewpoint actively and with low latency, significantly improving response speed and operational efficiency. Hierarchical filtering algorithms and robotic arm end effector easing control effectively suppress image jitter and improve image stability. K-Medoids clustering to generate preset poses and multi-segment spline interpolation transition algorithms enable smooth and rapid multi-view switching, reducing shaking and discomfort during switching. The collaborative robotic arm possesses autonomous positioning capabilities, reducing reliance on on-site filming personnel and lowering labor costs. Virtual safety envelope modeling and out-of-bounds projection constraints ensure operational safety. Hierarchical encoding of primary and secondary streams, ROI region enhancement, ABR adaptive bitrate adjustment, and dynamic FEC redundancy mechanisms improve network adaptability, ensuring smooth video even under network fluctuations or high packet loss rates. The system supports structured auditing and chained hashing, fully recording the operation process and network status, facilitating subsequent review and traceability, and improving compliance and security. Furthermore, the system can mask sensitive areas before video encoding and implement access control based on user permissions, effectively reducing the risk of privacy leaks. By dynamically adjusting cluster centers, the system improves the hit rate of commonly used viewpoints, enhancing the user experience. Attached Figure Description
[0043] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0044] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention;
[0045] Figure 2 This is a schematic diagram illustrating the security envelope setting according to an embodiment of the present invention;
[0046] Figure 3 This is a diagram illustrating the smoothing effect of an embodiment of the present invention. Detailed Implementation
[0047] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0048] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0049] Example 1
[0050] like Figure 1 As shown, this embodiment provides a remote video acquisition method based on low-latency active viewpoint control using a collaborative robotic arm, including the following steps:
[0051] Real-time acquisition of attitude data from remote terminals;
[0052] Based on the attitude data of the remote terminal, the desired pose of the end effector of the collaborative robotic arm is calculated through a pre-established coordinate mapping relationship, wherein the coordinate mapping relationship relates the local coordinate system of the remote terminal and the camera coordinate system of the end effector of the collaborative robotic arm.
[0053] The desired pose sequence is smoothed by a smoothing filter to generate a smooth pose sequence.
[0054] The smooth pose sequence is sent to the robotic arm control terminal;
[0055] The robotic arm control unit performs trajectory planning based on a smooth pose sequence and under the constraints of a preset virtual safety envelope, generating the robotic arm's execution trajectory.
[0056] The collaborative robotic arm is controlled to move according to the execution trajectory of the robotic arm.
[0057] Video streams are captured by a camera mounted at the end of a collaborative robotic arm;
[0058] The video stream is encoded and transmitted to the remote terminal for display.
[0059] The specific implementation process is as follows:
[0060] The preliminary preparation process is as follows:
[0061] S1. Using QR codes and other markers, and with base calibration, establish the mapping from the terminal L local coordinate system to the ground G.
[0062] S2. Measure the distance between the remote robotic arm end base and the ground, and establish a mapping from G to the robotic arm end base W, and then to the end camera C;
[0063] A high-contrast QR code (or ArUco, AprilTag) is printed on the terminal calibration board, and its size and the world coordinate origin are known.
[0064] During power-on initialization:
[0065] The terminal camera recognizes the QR code and analyzes the center point Ptag and the posture qtag;
[0066] Establish a local coordinate system L for the terminal by combining the mobile phone's IMU coordinate system;
[0067] The TLG is obtained by using the reference marker placed on the ground G using QR. (The process requires L->Q->G, with the intermediate step Q omitted; TGQ is obtained from the calibration plate).
[0068] The relative height h between the robotic arm base and the ground is measured using a laser rangefinder or depth camera, and the horizontal distance from the base to the center of the QR code is measured using a ruler or laser rangefinder. The TGW is then calculated.
[0069] Multiplying the two together yields the global mapping TLG×TGW.
[0070] During the actual construction, it is necessary to measure the height of the robotic arm base and calibration plate for attitude calculation T_GW and T_LG.
[0071] The robotic arm is positioned at a relatively central location, waiting for the terminal to perform the matching process.
[0072] The terminal selects a suitable posture to scan the identification code, and once the initial position is selected, the robotic arm begins to synchronize.
[0073] P_c = T_WC * T_GW * T_LG * P_l
[0074] Where T represents the 4×4 homogeneous transformation matrix that transforms the coordinates to another coordinate system, and P_l is the tooth homogeneous coordinate (4×1) of the local point at the terminal.
[0075] q_c = q_WC ⊗ q_GW ⊗ q_LG ⊗ q_l
[0076] Where ⊗ represents the quaternion product, q_l is the quaternion of the terminal in the local coordinate system, and the result q_c is the quaternion of the end camera coordinate system.
[0077] S3. Determine the safe working envelope of the robotic arm using a teaching method;
[0078] The operator manually guides the robotic arm's end effector to the boundary points of the safe workspace using a teach pendant or in Teach Mode. The system automatically records several boundary point sets Pi and calculates the safe envelope E using α-shape or convex hull algorithms.
[0079] Envelope parameters (vertices coordinates, boundary equations) are stored in the control platform configuration file, along with a version number. During runtime, the collision detection module verifies in real-time whether the pose is within the envelope.
[0080] Envelope settings such as Figure 2 The diagram shows the safety envelope settings from side and top views. Gray represents the reachable area, green represents the safe area, and brown represents the area where a collision may occur; these areas should be removed from the green section.
[0081] From a top-down perspective, the orange circle represents the operator's position, while the green area should be avoided.
[0082] The runtime process is as follows (basic functions):
[0083] S4. Acquire the real-time attitude of the terminal (position + quaternion, sampling 60–120Hz, preferably 80–100Hz).
[0084] The terminal collects data from the mobile phone's IMU (accelerometer, gyroscope, magnetometer), and calculates the quaternion attitude q using a sensor fusion algorithm (such as Madgwick or Mahony). L ;
[0085] The position p is obtained by combining the displacement information provided by the mobile phone camera or ARKit / ARCore. L .
[0086] After being timestamped, the data is converted to a ground coordinate system:
[0087] ;
[0088] Then it is passed to the robotic arm to complete the full-chain mapping.
[0089] S5. Perform hierarchical filtering on the mapped target attitude sequence: high-frequency jitter determination → adaptive window smoothing → EMA / Kalman prediction fusion to output a smooth target.
[0090] High-frequency jitter determination: Calculate the standard deviation of the samples from the short window and the long window, compare it with the set parameters, and determine whether it is high-frequency jitter.
[0091] Adaptive window smoothing: Select the window length W based on whether there is jitter. When there is jitter, use a larger window; when there is no jitter, use a smaller window and perform averaging.
[0092] EMA + Kalman: Perform EMA and Kalman operations on position and attitude, then fuse them to obtain the result; attitude needs to be interpolated between EMA attitude and KF attitude using SLERP.
[0093] The specific value can be set as follows:
[0094] High-frequency jitter determination: For short window Ws=5 and long window Wl=20, if σshort>1.5σlong and σshort<0.01 m, it is considered jitter;
[0095] Adaptive window: W=11 when jittering, W=3 when no jittering;
[0096] EMA parameters αpos=0.4, αquat=0.3;
[0097] Kalman filtering: process noise Q = 10⁻⁴I, measurement noise R = 10⁻³I;
[0098] The fusion weight αs = 0.7, biased towards the EMA output.
[0099] Smoothing effect Figure 3 As shown.
[0100] S6. Send PosePack to synchronize pose;
[0101] PosePack includes the fields shown in Table 1:
[0102] Table 1
[0103] Field Type Description header uint16 Packet header identification timestamp uint64 Timestamp (ms) position float[3] Three-dimensional position (m) quaternion float[4] Quaternion (x, y, z, w) version uint8 Protocol version signature uint32 Check signature reserve bytes[8] Reserved field
[0104] PosePack uses UDP for real-time transmission and performs packet reordering and timeout discarding at the application layer; it can switch to WebRTC DataChannel mode when necessary.
[0105] S7. After receiving PosePack, the robotic arm verifies the signature / version to prevent network transmission packets from being tampered with.
[0106] The data packet signature uses HMAC-SHA256, and the key is pre-shared between the terminal and the robotic arm.
[0107] Each version update (algorithm or calibration parameters) comes with a version number. If the version is inconsistent after the robotic arm receives the update, it will refuse to execute the update and request a synchronization update.
[0108] S8. The robotic arm performs final filtering / interpolation / speed limiting to obtain a safe execution trajectory.
[0109] After the robotic arm receives the target pose sequence:
[0110] The trajectory was smoothed using cubic spline interpolation;
[0111] Speed limiting strategy: linear velocity not exceeding 0.5 m / s, angular velocity not exceeding 60° / s;
[0112] If the posture changes abruptly, the interpolation interval is extended to ensure motion continuity.
[0113] Additionally, a low-pass filter (Butterworth first order, cutoff 10 Hz) is applied for end-point trajectory easing.
[0114] S9. Check if the pose is within the envelope. If not, project the pose onto the virtual envelope plane.
[0115] The safety envelope is represented by a convex hull or α-shape:
[0116] ;
[0117] When pose P∉E, calculate the projection of the nearest point:
[0118] ;
[0119] If the projection distance exceeds the limit (>10 mm), the system will issue a warning and stop the movement.
[0120] S10: The end effector outputs the pose; if a collision is detected, it stops.
[0121] The robotic arm is equipped with a torque sensor or a current feedback detection module at its end effector.
[0122] If an external force exceeding the threshold Fth=10N or an abnormal torque is detected, stop immediately.
[0123] Recovery process:
[0124] The system records the last safe pose;
[0125] Stop and enter standby mode;
[0126] Manual confirmation is required before the operation can be repeated.
[0127] S11: The camera components simultaneously capture video, and the combination of ABR adaptive bitrate and dynamic FEC encoding improves the smoothness of network transmission.
[0128] ABR Adjustment Formula:
[0129] ;
[0130] Where R represents fluency, L is the current network latency, and β=0.1.
[0131] Automatically reduce resolution when packet loss rate p > 5%.
[0132] FEC redundancy rate:
[0133] ;
[0134] Typically, ρmin = 0.05 and k = 3 are used. The redundancy coefficient of the ROI region is doubled.
[0135] The encoder calls the FFmpeg / LibWebRTC module, which carries the ROI mask through the RTP extension header to achieve differentiated bitrate allocation and redundancy control.
[0136] S12. Transmit via WebRTC and decode and display the screen on the terminal;
[0137] Signaling flow (simplified);
[0138] Both ends exchange SDP Offers / Answer and ICEcandidates via a signaling server (WebSocket / HTTPS).
[0139] Signaling message example (JSON):
[0140] { "type":"offer", "sdp":"v=0...","client_id":"arm001"};
[0141] { "type":"answer", "sdp":"v=0...","client_id":"termA"};
[0142] { "type":"ice", "candidate":"candidate:1 1 UDP 2122260223 ..."};
[0143] Upon receiving the answer / ICE, the client triggers setRemoteDescription / addIceCandidate to establish the P2P connection. If a direct connection fails, a TURN relay is used.
[0144] Codec configuration (using H.264 as an example);
[0145] Negotiating the encoder in SDP:
[0146] m=video ... and a=rtpmap:96 H264 / 90000, etc.
[0147] It is recommended to define the profile-level as follows: a=fmtp:96 packetization-mode=1;profile-level-id=42e01f (Baseline / Constrained).
[0148] In the edge encoder (robotic arm end), set the following parameters: FPS, GOP (Gap-to-Gap), maximum bitrate, and ROI. Example parameters: fps=30, gop=60, max_bitrate=2500kbps.
[0149] Decoding and display synchronization
[0150] Use RTP timestamps to map to NTP (RTCP SR / SDES) to ensure media time base matching.
[0151] The UI decoder sorts / reassembles frames by RTP timestamp and buffers them based on playout_delay (e.g., 50–150 ms) before playback.
[0152] If there are primary and secondary bitstreams: at the decoding end, the primary frame and the secondary frame are associated with the same time base (RTP timestamp / segment_id), the primary frame is rendered first, and then the ROI region pixels are overlaid / replaced.
[0153] Additional Feature 1 - Follow and Point Mode:
[0154] S13 and S4-S10 can be displayed in follow mode;
[0155] S14. On the terminal interaction UI, you can choose to enable the point mode, generate multiple pre-made points, and name them. The points are generated from the clustering of users at the time. Clicking on these points will cause the robotic arm to stop directly at the point coordinates, making it easier for the inspection personnel to view more stably. In point mode, the robotic arm will move directly and fix itself at the corresponding position of the point for stable observation.
[0156] UI interaction (core events);
[0157] State machine: FOLLOWING ↔ PRESET_SELECTION ↔ PRESET_MODE.
[0158] Main buttons / events:
[0159] Start / Stop Follow: Start / stop real-time following (send control signaling).
[0160] Capture Preset Candidate: Adds the current pose to the candidate pool (local);
[0161] Show Presets: Lists clustering results (thumbnail + name);
[0162] Enter Preset(p_id): Sends a PresetTrigger to the robotic arm.
[0163] Exit Preset: Resume following.
[0164] K-Medoids Implementation (Can be used on both terminal and edge devices)
[0165] Input: Set of candidate poses (Position + Attitude), Distance Measurement:
[0166] ;
[0167] in, The angle between the two quaternions (rad).
[0168] Algorithm (PAM-style K-Medoids brief steps):
[0169] Initialize by selecting K medoids (randomly or based on density).
[0170] Assign each point to the nearest medoid.
[0171] For each medoid, try swapping it with points within the cluster and calculate the total cost. If the cost is lower, replace the medoid.
[0172] Iterate until convergence or the upper limit of iterations is reached.
[0173] Parameters: K is adaptive (3–10), or the optimal K is selected using Silhouette / Davies-Bouldin.
[0174] Output: pose, hit_count, and last_hit_ts for each medoid.
[0175] Point storage and retrieval;
[0176] Storage format (JSON) example:
[0177] {
[0178] "preset_id":"p_001",
[0179] "pose_W": {"pos":[x,y,z],"quat":[w,x,y,z]},
[0180] "joint_cache":[...], "hit_count":12, "name":"Instrument A_Front",
[0181] "version":"v1", "created_by":"user_01"
[0182] };
[0183] Invocation: The terminal sends PresetTrigger{preset_id,mode}; the robotic arm verifies reachability, loads the cached joint solution, executes the pre-calculated trajectory, and is directly in place.
[0184] Additional feature 2 - ROI extraction and transfer;
[0185] S15. The user selects a polygonal region on the terminal device, and the robotic arm tracks the image region by matching feature points. The tracked region will be encoded into the secondary bitstream at a higher resolution during encoding.
[0186] Feature point matching (used for initial ROI identification);
[0187] Use KLT optical flow (Pyramidal Lucas-Kanade) for dense / sparse tracking (computationally lightweight and low latency).
[0188] If the matching quality deteriorates, revert to detection + matching based on descriptor ORB (re-detect).
[0189] Strategies for handling failures;
[0190] Confidence threshold: If the number of matches < N_min or the tracking error > ε, the decision is considered a failure.
[0191] Degradation strategy:
[0192] (a) Fast re-detection (searching around the ROI using template matching / detector);
[0193] (b) If multiple attempts fail, send a UI prompt to the terminal ("Location failed, please try again").
[0194] (c) Pause the secondary stream enhancement and revert to the main stream to avoid erroneous enhancement.
[0195] S16. Create a secondary bitstream for the tracking region, and use low QP to encode the ROI region separately to ensure that important regions are clear;
[0196] Block-based enhancements are as follows;
[0197] Cropping a ROI rectangle (or the minimum bounding rectangle of a polygon binding) from the captured frame and encoding that patch separately as a secondary stream (either a standalone encoder instance or the encoder's roi-enhancement mode).
[0198] During transmission, the auxiliary stream contains the original pixel / compressed difference of the ROI and segment_id, roi_id.
[0199] QP rules;
[0200] Define the mainstream benchmark QP: QP_base (determined by ABR).
[0201] ROI secondary stream QP is set as follows:
[0202] ;
[0203] Typical: ΔQP=6-10; QPmin=18.
[0204] If the network deteriorates, first increase QP_base (reduce the overall image quality), but keep the difference in QP_roi no less than ΔQP_min=4 (to ensure relative clarity).
[0205] The auxiliary stream synchronizes with the mainstream stream;
[0206] Using the same RTP timestamp / segment_id or carrying base_rtp_ts in the secondary stream header, the secondary stream frame is aligned with the corresponding main stream frame using the timestamp at the decoding end, and then the corresponding area of the main stream is replaced or superimposed with the pixels of the secondary stream.
[0207] If the auxiliary stream is delayed or lost, the decoding end will revert to displaying only the main stream.
[0208] Moreover, this invention employs an adaptive bitrate and dynamic forward error correction (FEC) mechanism, which ensures smooth video playback even under network fluctuations or high packet loss rates, with virtually no stuttering when the packet loss rate is below 5%.
[0209] This embodiment also provides a remote video acquisition system based on low-latency active viewpoint control using a collaborative robotic arm, including:
[0210] A remote terminal used to acquire attitude data, send pose commands, and receive and display video streams;
[0211] A collaborative robotic arm with a camera mounted at its end effector;
[0212] The robotic arm control terminal is communicatively connected to the collaborative robotic arm and the remote terminal. It is used to control the movement of the collaborative robotic arm according to the received pose instructions, and to encode and transmit the video stream captured by the camera.
[0213] The method is implemented through the system described above.
[0214] The present invention achieves the following effects through the above methods:
[0215] Low-latency active view control;
[0216] Corresponding technical means: Establish an attitude mapping link (L→G→W→C) from the terminal to the end of the robotic arm, and adopt the WebRTC real-time transmission protocol.
[0217] Implementation principle: Real-time attitude data is collected by the terminal IMU, and after being smoothed by layered filtering, it is directly mapped to the coordinate system of the robotic arm to achieve attitude synchronization of "what you point to is what you see"; WebRTC's P2P channel and SRTP encrypted transmission avoid the latency caused by intermediate server caching.
[0218] Shake reduction and improved image stability;
[0219] Corresponding technical means: layered filtering algorithm (high frequency jitter judgment → adaptive window smoothing → EMA / Kalman fusion) and robotic arm end-effector easing control.
[0220] Implementation principle: After detecting high-frequency disturbances, the smoothing window is adaptively expanded, and predictive filtering and attitude interpolation are fused to suppress hand tremors and sensor noise; the speed-limited trajectory at the end of the robotic arm further reduces abrupt changes in the field of view.
[0221] Smooth and fast multi-view switching;
[0222] Corresponding technical means: K-Medoids clustering to generate preset poses and multi-segment spline interpolation transition algorithm.
[0223] Implementation principle: In "follow mode", the terminal automatically clusters frequently used points and performs cubic spline interpolation based on the start and end postures during switching to achieve smooth trajectory transition.
[0224] High reliability and adaptive transmission in weak network conditions;
[0225] Corresponding technical measures: hierarchical coding of primary and secondary bitstreams, ROI region enhancement (low QP coding), ABR adaptive bitrate adjustment and dynamic FEC redundancy.
[0226] Implementation principle: Key detection areas (ROIs) are encoded separately and their redundancy is increased. When the network is congested, priority is given to ensuring the clarity of ROIs. The bit rate and redundancy are adjusted in real time by monitoring bandwidth and packet loss rate.
[0227] Intelligent optimization and self-evolution of commonly used poses;
[0228] Corresponding technical means: pose clustering self-evolution mechanism of terminal and robotic arm.
[0229] Implementation principle: The system dynamically adjusts the cluster center based on the frequency and duration of user operations to improve the hit rate of commonly used viewpoints; the preset point library is updated synchronously at the robotic arm end to reduce repetitive operations.
[0230] Operational safety and boundary crossing protection;
[0231] Corresponding technical means: virtual safe envelope modeling and out-of-bounds projection constraints (convex hull / α-shape model).
[0232] Implementation principle: During the teaching phase, the boundary points of the robotic arm are collected to establish an envelope model. During runtime, the pose is calculated in real time to determine whether it is within the safe zone. If it exceeds the boundary, it is automatically projected onto the boundary plane.
[0233] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A remote video acquisition method based on low-latency active viewpoint control using a collaborative robotic arm, characterized in that, Includes the following steps: Real-time acquisition of attitude data from remote terminals; Based on the attitude data of the remote terminal, the desired pose of the end effector of the collaborative robotic arm is calculated through a pre-established coordinate mapping relationship, wherein the coordinate mapping relationship relates the local coordinate system of the remote terminal and the camera coordinate system of the end effector of the collaborative robotic arm. The desired pose sequence is smoothed by a smoothing filter to generate a smooth pose sequence. The smooth pose sequence is sent to the robotic arm control terminal; The robotic arm control unit performs trajectory planning based on a smooth pose sequence and under the constraints of a preset virtual safety envelope, generating the robotic arm's execution trajectory. The collaborative robotic arm is controlled to move according to the execution trajectory of the robotic arm. Video streams are captured by a camera mounted at the end of a collaborative robotic arm; The video stream is encoded and transmitted to the remote terminal for display.
2. The method according to claim 1, characterized in that, The coordinate mapping relationship includes: By identifying markers placed in the ground coordinate system using a camera installed on the remote terminal, the transformation relationship T_LG from the local coordinate system of the remote terminal to the ground coordinate system is determined. The transformation relationship T_GW from the base coordinate system to the ground coordinate system of the collaborative robotic arm is obtained by measurement; The transformation relationship T_WC from the end-effector coordinate system to the base coordinate system is obtained through the kinematic model of the collaborative robotic arm. Based on the transformation relationships T_LG, T_GW, and T_WC, a coordinate mapping relationship is established from the local coordinate system of the remote terminal to the camera coordinate system at the end of the collaborative robotic arm.
3. The method according to claim 1, characterized in that, The generation of the smooth pose sequence includes: High-frequency jitter determination is performed on the desired pose sequence; Based on the jitter determination result, an adaptive smoothing window is selected to smooth the desired pose sequence; The smoothed pose sequence is subjected to exponential moving average filtering and Kalman filtering; The results of the exponential moving average filter and the Kalman filter are combined to output a smooth pose sequence.
4. The method according to claim 1, characterized in that, The trajectory planning under the constraints of a preset virtual safety envelope includes: Determine whether the smooth pose is within the virtual safety envelope; If the smooth pose is located within the virtual safety envelope, then it is taken as the target pose; If the smooth pose is outside the virtual safe envelope, the smooth pose is projected onto the boundary of the virtual safe envelope, and the projected pose is used as the target pose. Based on the target pose, interpolation and speed limiting are performed to generate the execution trajectory.
5. The method according to claim 1, characterized in that, Encoding the video stream includes: The video stream is determined as the main bitstream; Based on the region of interest selected in the video stream, the region of interest is tracked, and a secondary bitstream containing the region of interest is generated based on the tracking results, wherein the encoding quality of the secondary bitstream is higher than the encoding quality of the corresponding region in the main bitstream; The main stream and the secondary stream are transmitted synchronously to the remote terminal.
6. The method according to claim 1, characterized in that, The method also includes the generation and invocation of preset viewpoints: When the remote terminal is in follow mode, the stationary pose of the collaborative robotic arm is recorded to form a candidate pose set; Cluster analysis is performed on the candidate pose set to generate preset poses; The collaborative robotic arm is controlled according to the preset pose, and the collaborative robotic arm is controlled to move to the preset pose and remain fixed.
7. The method according to claim 6, characterized in that, The K-Medoids algorithm is used to perform cluster analysis on the candidate pose set. Its distance metric function integrates the position information of the pose and the attitude information represented by quaternions.
8. A remote video acquisition system based on low-latency active viewpoint control using a collaborative robotic arm, for implementing the method of any one of claims 1 to 7, characterized in that, include: A remote terminal used to acquire attitude data, send pose commands, and receive and display video streams; A collaborative robotic arm with a camera mounted at its end effector; The robotic arm control terminal is communicatively connected to the collaborative robotic arm and the remote terminal. It is used to control the movement of the collaborative robotic arm according to the received pose instructions, and to encode and transmit the video stream captured by the camera.