An image processing transmission method suitable for display control tracking and related products
By fusing sensor data with the motion state of the image acquisition device, main and redundant data streams are generated. Adaptive PID servo control and dynamic bandwidth allocation are adopted to solve the problem of display, control and tracking under unstable wireless channels, and to achieve accurate tracking and reliable transmission in complex dynamic scenarios.
Patent Information
- Application Number
- CN202511373589.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-25
AI Technical Summary
In dynamic scenarios, existing technologies struggle to maintain the smoothness of image processing and transmission and the reliability of control commands in situations where wireless channels are unstable. This is especially true when the target is moving at high speed or with high mobility, or when it is briefly obscured, which can easily lead to a decrease or loss of tracking accuracy.
By fusing high-frequency sensor data with the motion state of image acquisition equipment, the target position is predicted, and main and redundant data streams are generated. Adaptive PID servo control and dynamic bandwidth allocation strategies are adopted to ensure the integrity of critical data and the reliability of control commands.
It achieves accurate target tracking in complex and dynamic scenarios, improves tracking robustness, avoids tracking failures caused by motion coupling, and ensures the smoothness of the display terminal screen and the reliability of control commands.
Smart Images

Figure CN120856841B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing and image transmission technology, specifically to an image processing and transmission method and related products suitable for display and control tracking. Background Technology
[0002] Intelligent photoelectric tracking systems typically integrate functions such as image acquisition, target recognition, servo control, and data communication. They are widely used in various mobile or fixed platforms, such as drones, vehicle-mounted equipment, security monitoring, and robots, to achieve continuous locking, observation, and tracking of specific targets.
[0003] In dynamic scenarios, when the target moves at high speed and with high mobility or is briefly occluded, detection algorithms that rely solely on a single frame image are prone to problems such as decreased tracking accuracy or even target loss.
[0004] With the increasing mobility and networking of systems, wireless channels (such as Wi-Fi) are being used more and more for data transmission. Wireless channels are inherently susceptible to environmental interference, exhibiting instability such as bandwidth fluctuations, data packet loss, and latency jitter. This can cause stuttering, screen tearing, or interruptions in the video received from a remote location, and may also delay the transmission of control commands. Summary of the Invention
[0005] In order to solve the above problems, the present invention aims to provide an image processing and transmission method and related products suitable for display and control tracking, which can proactively ensure the integrity of key data in unstable wireless channels, while handling network congestion and ensuring the smoothness of the display terminal screen and the reliability of control commands.
[0006] This invention is achieved through the following technical solution:
[0007] An image processing and transmission method suitable for display and control tracking includes:
[0008] Acquire image data, as well as sensor data characterizing the motion state of the image acquisition device carrier;
[0009] Perform target detection on image data to determine the current location of the target;
[0010] By fusing the target's current location with sensor data, the predicted location of the target can be determined.
[0011] Based on the predicted next position, control commands are generated to drive the image acquisition device carrier to track the target;
[0012] Image data is encoded to generate a video stream, and the video stream is divided into a main data stream and a redundant data stream;
[0013] The main data stream is sent through the first transmission channel, while redundant data streams are sent in parallel through the second transmission channel.
[0014] Optionally, methods for acquiring image data and sensor data characterizing the motion state of the image acquisition device carrier include:
[0015] It receives image data generated by the image acquisition device within a defined exposure time interval in parallel, and simultaneously receives discrete angular velocity sample sequences and discrete GPS position sample sequences output by the sensor unit;
[0016] For each frame of image data, the precise total angular displacement of the image acquisition device carrier during the exposure period is calculated by multiplying each discrete angular velocity sample with its sampling time interval within the exposure time interval, and then summing all the products.
[0017] The calculated precise total angular displacement, along with the GPS location data smoothed by an adaptive filtering algorithm, is strictly timestamped with the corresponding image data.
[0018] The aligned image data, total angular displacement, and position data are encapsulated into a synchronization data packet.
[0019] Optionally, performing target detection on the image data to determine the current location of the target includes the following steps:
[0020] Based on the target state determined in the previous frame and the sensor data in the current frame, a predicted position of the target in the current image data is predicted by a Kalman filter, and a region of interest is generated around the predicted position.
[0021] Within the region of interest, the base object detection model is run to detect one or more candidate objects, and the detection confidence and preliminary feature vector of each candidate object are obtained;
[0022] For each candidate target, a comprehensive confidence score as a true target is calculated by weighted summing of three components: detection confidence component, spatial proximity component, and appearance feature similarity component.
[0023] Among all candidate targets, the candidate target with the highest comprehensive confidence score and that score is greater than the preset confirmation threshold is identified as the final target of the current frame, and its bounding box center position is taken as the current position of the target. At the same time, its depth feature vector template is updated.
[0024] Optionally, the step of fusing the target's current location with sensor data to predict the target's predicted location includes:
[0025] Determine the posterior state vector of the previous frame after filtering and updating;
[0026] Based on the total angular displacement at the current moment, determine the nonlinear transformation function used to compensate for the motion of the equipment itself;
[0027] The prior state vector at the current moment is predicted by using a nonlinear state transition function and combining the posterior state vector and total angular displacement.
[0028] The Jacobian matrix of the state transition function, the posterior error covariance matrix of the previous time step, and the process noise covariance matrix are calculated to update and predict the prior error covariance matrix of the current time step.
[0029] The position component is extracted from the prior state vector and used as the predicted position of the target.
[0030] Optionally, the step of generating control commands for driving the image acquisition device carrier to track the target based on the predicted next position includes:
[0031] Calculate the error vector between the preset center point and the predicted position of the image;
[0032] Based on the extended Kalman filter, after the measurement update of the posterior state estimate and its posterior error covariance matrix at the current time, the gain of the proportional-derivative controller is dynamically adjusted; wherein, the adjustment value of the proportional gain is proportional to the magnitude of the velocity component in the posterior state estimate, and the adjustment value of the derivative gain is proportional to the trace of the position correlation submatrix in the posterior error covariance matrix.
[0033] By applying the discrete proportional-integral-derivative (DI-DE) control law, combined with the error vector and dynamically adjusted gain, the control command vector for driving the equipment carrier is calculated. The control law includes a proportional term proportional to the current error, an integral term proportional to the cumulative sum of historical errors, and a derivative term proportional to the rate of change of the current error.
[0034] The calculated control command vector is converted into low-level instructions that can be executed by the servo motor or turntable controller.
[0035] Optionally, the steps of encoding image data to generate a video stream and dividing the video stream into a main data stream and a redundant data stream include:
[0036] Image data is encoded to obtain data packets. During the encoding process, timestamp information and GPS information are written into the supplementary enhancement information field of the bitstream data in real time.
[0037] While generating the video bitstream, the data packets are dynamically classified into critical data packets and non-critical data packets according to their type and importance to decoding. Among them, critical data packets include at least data packets containing supplemental enhancement information fields, I-frame data packets, P-frame header information and motion vector data packets.
[0038] The adaptive forward error correction redundancy rate is calculated based on the channel packet loss rate periodically fed back from the receiving terminal.
[0039] Several consecutive key data packets are grouped into a single encoding group, and a corresponding number of redundant data packets are generated for this encoding group based on the redundancy rate using a forward error correction algorithm.
[0040] The redundant data stream contains all generated redundant data packets; the main data stream contains all encoded data packets.
[0041] Optionally, the step of sending the main data stream through the first transmission channel and simultaneously sending redundant data streams in parallel through the second transmission channel includes:
[0042] At the IP layer of the network protocol, a high-priority QoS tag is set for the primary data stream, and a low-priority QoS tag is set for redundant data streams.
[0043] Based on the real-time estimation of the total available bandwidth of the transmission channel, a bandwidth allocation strategy that separates primary and secondary flows is adopted to calculate the bandwidth allocated to the primary data flow and the redundant data flow. This allocation strategy ensures that the bandwidth allocation of the primary data flow is given priority, and then the remaining available bandwidth is allocated to the redundant data flow.
[0044] Independent rate control is applied to the main data stream and redundant data stream based on dynamically allocated bandwidth, thus smoothing the data packet transmission rate.
[0045] An image processing and transmission system suitable for display and control tracking includes:
[0046] The data acquisition module is used to acquire image data and sensor data that characterizes the motion state of the image acquisition device carrier;
[0047] The target detection module is used to detect targets in image data to determine the current location of the targets;
[0048] The target prediction module is used to fuse the target's current position with sensor data to predict the target's predicted position.
[0049] The control command generation module is used to generate control commands based on the predicted position to drive the image acquisition device carrier to track the target.
[0050] The encoding and data stream partitioning module is used to encode image data to generate a video stream and partition the video stream into a main data stream and a redundant data stream.
[0051] The data transmission module is used to send the main data stream through the first transmission channel, and simultaneously send redundant data streams in parallel through the second transmission channel.
[0052] A computer-readable storage medium storing a computer program, characterized in that, when executed by a processor, the computer program implements the image processing and transmission method for display and control tracking as described in any of the preceding claims.
[0053] A computer program product includes a computer program / instructions that, when executed by a processor, implement the image processing and transmission method for display control tracking as described above.
[0054] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0055] This invention synchronizes high-frequency sensor data with camera exposure range to estimate and compensate for the platform's complex motion, and deeply fuses the platform motion with the target motion, thereby achieving high-precision prediction of the target's true motion state. It employs an adaptive PID servo control strategy to enable the control system to adjust its response characteristics in real time according to changes in the tracking state. A data transmission mechanism is constructed, and channel quality feedback is combined to dynamically adjust the redundancy of forward error correction. Furthermore, a service quality grading and dynamic bandwidth allocation strategy is used to prioritize the transmission of the core video stream.
[0056] This invention can calculate the true motion trajectory of the target and achieve accurate tracking in complex dynamic scenarios such as violent platform movement or high-speed target maneuvering, greatly improving the robustness of tracking and effectively avoiding tracking failure caused by motion coupling.
[0057] This invention constructs a proactive intelligent flow control data transmission mechanism, which can proactively ensure the integrity of critical data in unstable wireless channels, thereby ensuring the smoothness of the display terminal screen and the reliability of control commands. Attached Figure Description
[0058] The accompanying drawings illustrate exemplary embodiments of the present invention and, together with the description thereof, serve to explain the principles of the invention. These drawings are included to provide a further understanding of the invention and are incorporated in and constitute a part of this specification, but do not constitute a limitation on the embodiments of the present invention.
[0059] Figure 1 This is a flowchart illustrating the image processing and transmission method for display and control tracking according to the present invention.
[0060] Figure 2 This is a schematic diagram of the process for acquiring image data and sensor data in Embodiment 3 of the present invention.
[0061] Figure 3 This is a flowchart illustrating the process of determining the current position of a target according to Embodiment 4 of the present invention. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0063] It should also be noted that, for ease of description, only the parts relevant to the present invention are shown in the accompanying drawings.
[0064] Where there is no conflict, the embodiments and features described herein can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0065] Example 1
[0066] This embodiment provides a complete workflow for an image processing and transmission method suitable for display and control tracking. It integrates a collaborative workflow of environmental perception, intelligent prediction, closed-loop control, and reliable transmission, solving the problem of stable tracking and high-quality remote presentation of specific targets in dynamic and complex environments.
[0067] Conceptually, it can be divided into three core stages: the first is the multimodal perception and recognition stage, which is how the system acquires and understands the state of the target and itself; the second is the decision-making and control stage, which is how the system reacts based on its understanding to maintain lock on the target; and the last is the encoding and transmission stage, which is how the system efficiently and reliably transmits visual information to remote users.
[0068] like Figure 1 As shown below, the key technical steps of this method will be explained in detail:
[0069] The system acquires image data and sensor data characterizing the motion state of the image acquisition device carrier. The system acquires data in parallel from two dimensions: first, it acquires real-time visual image data from image acquisition devices such as photoelectric cameras; second, it acquires motion state data describing the attitude, angular velocity, and position of the camera-carrying motion platform (such as a drone or vehicle gimbal) from sensors such as gyroscopes and global positioning systems (GPS).
[0070] The system performs target detection on image data to determine the current location of the target. It uses built-in computer vision algorithms (usually deep learning-based intelligent recognition models) to analyze and process the acquired image data. It identifies preset tracking targets in complex image backgrounds and outputs their position information in the current image frame coordinate system, such as the center coordinates of the target's bounding box.
[0071] The target's current position is fused with sensor data to predict its predicted position. Fusion combines the target's visual position information with the platform's own motion state information, and performs calculations through prediction algorithms (such as Kalman filtering) to eliminate image changes caused by the platform's own motion, thereby more accurately predicting the target's true position in the next moment.
[0072] Based on the predicted next position, control commands are generated to drive the image acquisition device carrier to track the target. The system compares the predicted target position with the image center (or the desired tracking point) and calculates the deviation or error between the two. Subsequently, the servo control algorithm generates specific control commands based on this error, such as commands to drive the motor to rotate at a specific angle or move at a specific angular velocity. These commands ultimately drive the turntable or gimbal carrying the camera to move toward the predicted target position, so as to dynamically keep the target in the center of the field of view.
[0073] Image data is encoded to generate a video stream, which is then divided into a main data stream and a redundant data stream. A video stream refers to a data format suitable for network transmission, which is converted from consecutive image frames through compression encoding (such as the H.264 standard). In this embodiment, it is divided into two parts: one part is the main data stream containing complete video information, used for real-time playback; the other part is the redundant data stream, which typically contains error correction information or backups of critical information for data recovery, with the aim of enhancing the reliability of transmission.
[0074] The system sends the main data stream through the first transmission channel, while simultaneously sending a redundant data stream in parallel through the second transmission channel. Utilizing the network interface, the system sends out both data streams generated in the previous step in parallel. Two different transmission channels are implemented using Quality of Service (QoS) grading or different network ports. The main data stream prioritizes low latency and real-time performance, while the redundant data stream serves as a safeguard. When data loss occurs in the main data stream due to network fluctuations during transmission, the receiving end can use information from the redundant data stream to repair or supplement the data, thereby ensuring the integrity and smoothness of the remote video feed.
[0075] Example 2
[0076] Based on the method framework described in Embodiment 1, this embodiment provides a more specific and detailed technical implementation scheme based on a specific hardware platform and software process. The core hardware architecture of this embodiment mainly consists of a front-end photoelectric camera, a field-programmable gate array (FPGA) for data preprocessing and interface conversion, and a HiSilicon AI chip SD3403 as the main processing unit.
[0077] The overall workflow of this embodiment aims to illustrate the complete technical chain from the initial acquisition of image signals to the final intelligent tracking control and high-reliability transmission. The key technical aspects of this process will be described in detail step by step below.
[0078] I. Image Data Reception and Preprocessing Process
[0079] This process describes the complete path of image data from physical signals to digital information that can be processed by AI chips.
[0080] First, the image sensor inside the photoelectric camera captures optical image signals and converts them into electrical signals. After processing inside the camera, the image data is converted into LVDS format.
[0081] Subsequently, the LVDS differential signal is transmitted to the FPGA chip. The FPGA first acquires and parses the input LVDS signal, and then temporarily buffers the raw image data in DDR. During buffering, the FPGA can execute some hardware-level preprocessing algorithms, such as edge enhancement operations on the image, thereby reducing the overall system processing latency.
[0082] Finally, the FPGA transmits the preprocessed image and a copy of the original image to the core AI chip SD3403 via MIPI. Upon receiving the MIPI signal, the SD3403 chip parses the synchronization code through its interface module and converts it into internally usable image data, preparing for the next step of intelligent analysis.
[0083] II. Image Intelligent Processing and Encoding Process
[0084] This process is mainly completed on the SD3403 chip.
[0085] The SD3403 chip integrates an NPU that runs a YOLOv5 object detection model. The system scales the input 1080P resolution image to 640x640 before sending it to the NPU for inference calculations, and the processing latency for a single frame can be controlled within 8ms.
[0086] Once the NPU identifies and locates the target, the VGS hardware acceleration module inside the SD3403 chip immediately intervenes to perform OSD overlay operations. The target's coordinates are converted into a red rectangle and overlaid onto the original video frame, thus clearly identifying the target visually. Simultaneously, the GPS location information acquired by the system is also overlaid onto the image in character form.
[0087] The processed image (i.e., with OSD information overlaid) is sent to the SD3403's MPP hardware encoder. The encoder uses the mainstream H.264 video compression standard for efficient compression. The system writes the current timestamp and GPS information into the SEI field of the H.264 bitstream. The encoding bitrate can be dynamically configured via mobile devices, for example, adjusted based on the strength of the wireless signal, such as setting it to 4Mbps.
[0088] III. Target Tracking, Motor Control, and Reliable Transmission Process
[0089] This process describes how the system translates the results of intelligent analysis into physical actions and ensures stable and reliable data transmission.
[0090] In automatic tracking mode, once a target is initially identified, the SD3403 chip extracts a feature template of that target. In subsequent video frames, the system predicts the target's motion trajectory and performs template matching and confidence assessment near the predicted location to update the target's precise position. Based on the deviation between the updated target position and the image center, the system calculates control commands and drives a servo motor via PWM signals to control the turntable's rotation, thus achieving automatic target tracking. Furthermore, the system also supports a variety of manual control modes, allowing operators to send control commands via a mobile phone touchscreen, computer mouse, or head gestures from AR glasses.
[0091] Mobile phone: The mobile phone software can collect the sliding direction and distance, or the direction button click (the distance of a single click is fixed) data by sliding and clicking on the touch screen. After conversion, the data is packaged according to the communication protocol and sent to the camera via WIFI. The camera receives and parses the data and uses PWM to control the servo motor to make rotation.
[0092] Computer: By pressing and holding the left mouse button, sliding on the screen, and clicking the direction keys, the computer software collects the sliding direction and distance, or the clicked direction key (the distance of a single click is fixed), and packages the data according to the communication protocol. Then, it sends the data to the camera via WIFI / wired network. The camera receives and parses the data, and controls the servo motor to make rotation movements through the serial port.
[0093] AR Glasses: The AR glasses are equipped with an IMU, gyroscope, and accelerometer. The gyroscope collects high-frequency dynamic data based on the wearer's head rotation, while the accelerometer collects low-frequency static data. It calculates the head's orientation in 3D space in real time, uses the gyroscope to detect the direction, and the accelerometer assists in calculating and correcting the direction of gravity, outputting a precise rotation angle. Low-pass filtering is used for basic image stabilization. Kalman filtering is employed for dynamic prediction, reducing latency. Finally, the collected data is sent to the camera via Wi-Fi. The camera receives and parses the data, and uses a serial port to control a servo motor to perform the rotation movement.
[0094] In terms of data transmission and playback, this embodiment employs a retransmission mechanism combined with local storage to combat network fluctuations. While sending the video stream to the remote end via RTSP (Real-Time Streaming Protocol), the SD3403 chip simultaneously stores this stream, containing SEI timestamp information, as a file on its local eMMC hard drive. When the remote mobile device detects a video interruption due to poor network conditions, it can record the last received SEI timestamp and send a retransmission request containing that timestamp to the SD3403. Upon receiving the request, the SD3403 quickly locates the corresponding data position in the locally stored file based on the timestamp and retransmits the video stream from that position, thus achieving seamless video data recovery and ensuring data integrity.
[0095] Example 3
[0096] like Figure 2 As shown, this embodiment provides a detailed description of "acquiring image data and sensor data characterizing the motion state of the image acquisition device carrier".
[0097] In complex tracking tasks, image acquisition devices (such as cameras) and motion sensors (such as gyroscopes) typically operate at different frequencies. The former acquires images at a relatively low frame rate (e.g., 30 frames per second), while the latter outputs data at an extremely high sampling rate (e.g., 1000 times per second). This embodiment addresses the problem of data misalignment on the time axis caused by the mismatch in operating frequencies. The detailed implementation process is as follows:
[0098] Parallel reception by the image acquisition device within the exposure time interval The image data generated internally is clearly associated with the start and end timestamps of each frame of the exposure. and end timestamp .
[0099] It also synchronously receives discrete angular velocity sample sequences output by the sensor unit. and Discrete Global Positioning System (GPS) location sample sequences .
[0100] For each frame of the image data, by adjusting the exposure time interval The discrete angular velocity sample sequence within the image is integrated to calculate the precise total angular displacement of the image acquisition device carrier during the exposure period. , ;in:
[0101] This is the total angular displacement vector along three axes during the exposure of a single frame image;
[0102] For discrete time The collected angular velocity vector samples;
[0103] The sampling time interval of the angular velocity sensor;
[0104] The index range for summation.
[0105] The calculated precise total angular displacement The GPS location data, after being smoothed by an adaptive filtering algorithm, is strictly timestamped with the corresponding image data; the alignment reference is the exposure time interval of the image frame.
[0106] The aligned image data, total angular displacement, and position data are encapsulated into a synchronization data packet.
[0107] Example 4
[0108] like Figure 3 As shown, this embodiment provides a detailed explanation of "target detection of image data to determine the current position of the target". It uses historical motion information to guide the detection of the current frame and uses a multi-dimensional information comprehensive evaluation model to select the unique real target from all possible candidate targets. It can effectively deal with challenging scenarios such as the target being temporarily occluded, the presence of similar interference objects in the background, and changes in the appearance of the target itself.
[0109] Prediction and search region generation:
[0110] Based on the target state determined in the previous frame and the sensor data in the current frame, a predicted position of the target in the current image data is predicted using a Kalman filter. The Kalman filter is an algorithm used for state estimation of dynamic systems in the presence of uncertainty. By processing the target's determined state (such as position and velocity) from the previous frame and the sensor data from the current frame, the Kalman filter predicts the most likely location of the target in the current image frame. .
[0111] Candidate target detection and feature extraction:
[0112] Within the region of interest, a base object detection model (such as the YOLOv5 model) is run to detect one or more candidate objects. and obtain each candidate target Detection confidence and preliminary feature vectors The confidence level of the detection model represents the probability that the candidate is the target, and the preliminary feature vector represents the appearance characteristics of the candidate target.
[0113] Overall confidence level assessment:
[0114] For each candidate target Calculate its overall confidence score as a true target. : The three components are: detection confidence component, spatial proximity component, and appearance feature similarity component.
[0115] These are preset weighting coefficients;
[0116] Candidate targets output by the basic target detection model The detection confidence level;
[0117] Candidate targets The center coordinates of the bounding box;
[0118] The coordinates of the target center position predicted by the Kalman filter;
[0119] The Euclidean distance between the candidate target location and the predicted location;
[0120] The standard deviation of the Gaussian function is used to control the degree to which positional deviations penalize the scores.
[0121] Candidate targets in the current frame The depth feature vector;
[0122] This is the depth feature vector template for targets already identified in the previous frame;
[0123] This represents the cosine similarity between two feature vectors.
[0124] Target confirmation and template update:
[0125] The overall confidence score among all candidate targets The highest score and the score is greater than the preset confirmation threshold. The candidate target is confirmed as the final target of the current frame; that is, only when the highest score exceeds the threshold will the candidate target be finally confirmed as the real target of the current frame; its bounding box center position is... The current position of the target is used to update its deep feature vector template. .
[0126] Example 5
[0127] This embodiment provides a detailed explanation of "fusing the target's current location with sensor data to predict the target's predicted location" as described in Embodiment 1, specifically including the following steps:
[0128] Based on the previous moment The updated posterior state vector after filtering The posterior state vector refers to the optimal estimate after correction of the actual observations. This state vector is a 4×1 vector that contains multi-dimensional motion information of the target, specifically divided into position and velocity components.
[0129] According to the current time Total angular displacement The nonlinear transformation function used to compensate for the motion of the equipment itself was determined. ;
[0130] Through nonlinear state transition function To predict the current moment Prior state vector , ;
[0131] Update and predict the current moment Prior error covariance matrix : ;in:
[0132] For the current moment The predicted error covariance matrix;
[0133] For the previous moment Updated error covariance matrix;
[0134] State transition function Regarding the state vector exist The Jacobian matrix obtained at the given location;
[0135] This is the process noise covariance matrix, used to describe the uncertainty of the state transition model itself;
[0136] From the calculated prior state vector Extract the position component as the predicted position of the target. .
[0137] Nonlinear transformation function The mathematical expression is: ;in:
[0138] State vector Positional components in;
[0139] To apply an affine transformation matrix to the position vector of the previous time step;
[0140] State transition function The mathematical expression is:
[0141] ;in:
[0142] for The state vector.
[0143] State vector The velocity component in;
[0144] The displacement of the target itself based on a uniform motion model;
[0145] From time arrive The time interval.
[0146] Example 6
[0147] This embodiment details the process of "generating control commands for driving the image acquisition device carrier to track the target based on the predicted next position." It deeply couples the output of the state estimation algorithm (extended Kalman filter) with the PID control algorithm to achieve real-time optimization of the control strategy. The implementation process is detailed below:
[0148] Calculate the preset center point of the image With the predicted next position Error vector between , The magnitude and direction of this vector intuitively represent the range and direction of rotation required by the servo system to realign the target with the center.
[0149] Based on the extended Kalman filter at the current time Posterior state estimation after measurement update and its posterior error covariance matrix Dynamically adjust the proportional gain in the PID controller and differential gain , ;in:
[0150] Each represents the current time. Proportional gain and differential gain;
[0151] This is the preset base gain value.
[0152] For the posterior state vector The Euclidean norm of the medium velocity component;
[0153] The posterior error covariance matrix The trace of the position-related submatrix;
[0154] This is a preset scaling factor used to adjust the speed and the degree of uncertainty impact;
[0155] Applying the discrete PID control law, combined with the aforementioned error vector And the dynamically adjusted gain, calculate the current time. Control command vectors for driving the device carrier , ;in:
[0156] This is the integral gain;
[0157] To control the time interval of the cycle;
[0158] This is the cumulative integral term of the error;
[0159] The rate of change of the error;
[0160] The calculated control command vector This vector is then converted into low-level instructions executable by the servo motor or turntable controller. The system ultimately converts this vector into the low-level instruction format required by the servo motor or turntable controller, such as a pulse-width modulation (PWM) signal with a specific duty cycle or a serial communication data packet conforming to a specific protocol.
[0161] Example 7
[0162] This embodiment details the process of "encoding image data to generate a video stream and dividing the video stream into a main data stream and a redundant data stream." By pre-generating and sending redundant error correction data at the sending end, the receiving end can proactively and instantly recover lost information when data packet loss occurs. This improves the smoothness and integrity of video transmission in unstable network environments without increasing additional round-trip latency. The specific implementation process is detailed below:
[0163] The image data is encoded to obtain a data packet. During the encoding process, timestamp information and GPS information are written into the supplemental enhancement information (SEI) field of the bitstream data in real time. Important metadata (specifically timestamp information and GPS information) is directly written into the SEI field of the video bitstream. SEI is a channel in the video coding standard specifically used to carry additional information. By embedding this metadata, the tight binding of key synchronization information with the video frame is ensured.
[0164] While generating the video stream, data packets are dynamically classified into critical and non-critical data packets. The classification is based on the importance of the data packet to the receiving end's ability to successfully decode and reconstruct the image. The critical data packets include at least: data packets containing the SEI field, I-frame data packets, and P-frame header information and motion vector data packets.
[0165] Based on the channel quality parameters periodically fed back from the receiving terminal, the adaptive forward error correction redundancy rate is dynamically calculated. , ;in:
[0166] This is the base redundancy rate under good channel conditions;
[0167] It is a preset gain coefficient;
[0168] This refers to the current channel packet loss rate fed back from the receiver.
[0169] Continuous Each key data packet is treated as a coding group, and a forward error correction algorithm is used to generate a coding group. One redundant data packet;
[0170] The redundant data stream contains all the generated redundant data packets; the main data stream contains all the encoded data packets.
[0171] Example 8
[0172] This embodiment provides a specific implementation method for parallel transmission of the generated "main data stream" and "redundant data stream". Through a dual mechanism combining "priority marking" and "dynamic bandwidth allocation", a dynamically adaptable transmission order is established. The specific implementation process is detailed below:
[0173] At the IP layer of the network protocol, high-priority differential service code point (DSC) markers are set for all packets in the main data stream; simultaneously, low-priority DSC markers are set for all packets in the redundant data stream. DSCP allows embedding a marker in the IP header to inform network devices (such as routers) along the path of the expected service level of the packet. This reduces the probability of packets in the main data stream being dropped during network congestion, thus granting them priority forwarding.
[0174] Based on the total available bandwidth of the transmission channel Real-time estimation is achieved by dynamically calculating the bandwidth allocated to the main data stream using a primary-secondary stream separation bandwidth allocation strategy. and bandwidth allocated to redundant data streams : Only after the bandwidth requirements of the main data stream are met will the remaining channel capacity be allocated to the redundant data stream; if the channel is so congested that the target bit rate of the main data stream cannot be met, the bandwidth allocation of the redundant data stream will be reduced to zero, i.e., transmission will be suspended.
[0175] in: and Each represents the current time. The actual transmission bandwidth allocated to the main data stream and redundant data streams;
[0176] The target bitrate of the main data stream is set by the video encoder.
[0177] The total available bandwidth of the current channel is estimated through real-time transmission control protocol feedback or other detection methods;
[0178] Based on the dynamically allocated bandwidth respectively and The token bucket algorithm is used to independently control the rate of the main data stream and the redundant data stream, smoothing the data packet transmission rate, avoiding additional network congestion caused by excessively high instantaneous data transmission rate, and ensuring the final execution of the bandwidth allocation strategy.
[0179] Example 9
[0180] An image processing and transmission system suitable for display and control tracking includes:
[0181] The data acquisition module is used to acquire image data and sensor data that characterizes the motion state of the image acquisition device carrier;
[0182] The target detection module is used to detect targets in image data to determine the current location of the targets;
[0183] The target prediction module is used to fuse the target's current position with sensor data to predict the target's predicted position.
[0184] The control command generation module is used to generate control commands based on the predicted position to drive the image acquisition device carrier to track the target.
[0185] The encoding and data stream partitioning module is used to encode image data to generate a video stream and partition the video stream into a main data stream and a redundant data stream.
[0186] The data transmission module is used to send the main data stream through the first transmission channel, and simultaneously send redundant data streams in parallel through the second transmission channel.
[0187] A computer-readable storage medium storing a computer program, characterized in that, when executed by a processor, the computer program implements the image processing and transmission method for display and control tracking as described in any of the preceding claims.
[0188] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instruction data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM, EEPROM, flash memory or other solid-state storage technologies, CD-ROM, DVD or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The aforementioned system memories and mass storage devices can be collectively referred to as memory.
[0189] A computer program product includes a computer program / instructions that, when executed by a processor, implement the image processing and transmission method for display control tracking as described above.
[0190] Computer program products include computer programs or instruction sets used to perform specific tasks or achieve specific functions. These programs or instructions are designed to be executed by a processor to implement a series of predefined steps or operations. The program product may be stored in various forms of computer storage media, such as memory, hard disks, solid-state drives, optical discs, or other forms of digital storage devices. It may exist in the form of compiled binary code or in the form of scripts or bytecode that can be executed by an interpreter. Through carefully designed algorithms and logical instructions, the program product enables the processor to process data in a specific order and manner, performing various functions such as data analysis, user interaction, and device control.
[0191] In the description of this specification, the references to terms such as "one embodiment / mode," "some embodiments / modes," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment / mode or example is included in at least one embodiment / mode or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / mode or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments / modes or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments / modes or examples described in this specification, as well as the features of different embodiments / modes or examples.
[0192] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0193] Those skilled in the art should understand that the above embodiments are merely for illustrating the present invention and are not intended to limit the scope of the invention. Those skilled in the art can make other changes or modifications based on the above invention, and these changes or modifications still fall within the scope of the present invention.
Claims
1. An image processing and transmission method suitable for display and control tracking, characterized in that, include: Acquire image data, as well as sensor data characterizing the motion state of the image acquisition device carrier; Perform target detection on image data to determine the current location of the target; By fusing the target's current location with sensor data, the predicted location of the target can be determined. Based on the predicted next position, control commands are generated to drive the image acquisition device carrier to track the target; Image data is encoded to generate a video stream, and the video stream is divided into a main data stream and a redundant data stream; The main data stream is sent through the first transmission channel, while redundant data streams are sent in parallel through the second transmission channel. The step of fusing the target's current location with sensor data to predict the target's predicted location includes: Determine the posterior state vector of the previous frame after filtering and updating; Based on the total angular displacement at the current moment, determine the nonlinear transformation function used to compensate for the motion of the equipment itself; The prior state vector at the current moment is predicted by using a nonlinear state transition function and combining the posterior state vector and total angular displacement. The Jacobian matrix of the state transition function, the posterior error covariance matrix of the previous time step, and the process noise covariance matrix are calculated to update and predict the prior error covariance matrix of the current time step. The position component is extracted from the prior state vector and used as the predicted position of the target.
2. The image processing and transmission method for display and control tracking according to claim 1, characterized in that, Methods for acquiring image data and sensor data characterizing the motion state of the image acquisition device carrier include: It receives image data generated by the image acquisition device within a defined exposure time interval in parallel, and simultaneously receives discrete angular velocity sample sequences and discrete GPS position sample sequences output by the sensor unit; For each frame of image data, the precise total angular displacement of the image acquisition device carrier during the exposure period is calculated by multiplying each discrete angular velocity sample with its sampling time interval within the exposure time interval, and then summing all the products. The calculated precise total angular displacement, along with the GPS location data smoothed by an adaptive filtering algorithm, is strictly timestamped with the corresponding image data. The aligned image data, total angular displacement, and position data are encapsulated into a synchronization data packet.
3. The image processing and transmission method for display and control tracking according to claim 1, characterized in that, Performing object detection on image data to determine the current location of the object includes the following steps: Based on the target state determined in the previous frame and the sensor data in the current frame, a predicted position of the target in the current image data is predicted by a Kalman filter, and a region of interest is generated around the predicted position. Within the region of interest, the base object detection model is run to detect one or more candidate objects, and the detection confidence and preliminary feature vector of each candidate object are obtained; For each candidate target, a comprehensive confidence score as a true target is calculated by weighted summing of three components: detection confidence component, spatial proximity component, and appearance feature similarity component. Among all candidate targets, the candidate target with the highest comprehensive confidence score and that score is greater than the preset confirmation threshold is identified as the final target of the current frame, and its bounding box center position is taken as the current position of the target. At the same time, its depth feature vector template is updated.
4. The image processing and transmission method suitable for display and control tracking according to claim 1, characterized in that, The steps for generating control commands to drive the image acquisition device carrier to track the target based on the predicted next position include: Calculate the error vector between the preset center point and the predicted position of the image; Based on the extended Kalman filter, after the measurement update of the posterior state estimate and its posterior error covariance matrix at the current time, the gain of the proportional-derivative controller is dynamically adjusted; wherein, the adjustment value of the proportional gain is proportional to the magnitude of the velocity component in the posterior state estimate, and the adjustment value of the derivative gain is proportional to the trace of the position correlation submatrix in the posterior error covariance matrix. By applying the discrete proportional-integral-derivative (DI-DE) control law, combined with the error vector and dynamically adjusted gain, the control command vector for driving the equipment carrier is calculated. The control law includes a proportional term proportional to the current error, an integral term proportional to the cumulative sum of historical errors, and a derivative term proportional to the rate of change of the current error. The calculated control command vector is converted into low-level instructions that can be executed by the servo motor or turntable controller.
5. The image processing and transmission method for display and control tracking according to claim 1, characterized in that, The steps of encoding image data to generate a video stream and dividing the video stream into a main data stream and a redundant data stream include: Image data is encoded to obtain data packets. During the encoding process, timestamp information and GPS information are written into the supplementary enhancement information field of the bitstream data in real time. While generating the video bitstream, the data packets are dynamically classified into critical data packets and non-critical data packets according to their type and importance to decoding. Among them, critical data packets include at least data packets containing supplemental enhancement information fields, I-frame data packets, P-frame header information and motion vector data packets. The adaptive forward error correction redundancy rate is calculated based on the channel packet loss rate periodically fed back from the receiving terminal. Several consecutive key data packets are grouped into a single encoding group, and a corresponding number of redundant data packets are generated for this encoding group based on the redundancy rate using a forward error correction algorithm. The redundant data stream contains all generated redundant data packets; the main data stream contains all encoded data packets.
6. The image processing and transmission method for display and control tracking according to claim 5, characterized in that, The steps of sending the main data stream through the first transmission channel and simultaneously sending redundant data streams in parallel through the second transmission channel include: At the IP layer of the network protocol, a high-priority QoS tag is set for the primary data stream, and a low-priority QoS tag is set for redundant data streams. Based on the real-time estimation of the total available bandwidth of the transmission channel, a bandwidth allocation strategy that separates primary and secondary flows is adopted to calculate the bandwidth allocated to the primary data flow and the redundant data flow. This allocation strategy ensures that the bandwidth allocation of the primary data flow is given priority, and then the remaining available bandwidth is allocated to the redundant data flow. Independent rate control is applied to the main data stream and redundant data stream based on dynamically allocated bandwidth, thus smoothing the data packet transmission rate.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the image processing and transmission method for display and control tracking as described in any one of claims 1-6.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the image processing and transmission method for display and control tracking as described in any one of claims 1-6.
Citation Information
Patent Citations
Multi-mode competitive collaborative fusion sensing method for unmanned vehicle in mining area
CN119594962A
Light wave tracking device and method
JP2003099127A