An autonomous recognition and tracking system for unmanned equipment targets based on visual processing technology

By using the combination of Voronoi graph division, Sobel operator edge detection and deep learning model in unmanned target autonomous recognition and tracking systems, the problem that traditional technology is difficult to stably identify and track targets in high dynamic environments is solved, and more efficient target recognition and tracking performance is achieved.

CN119151990BActive Publication Date: 2025-05-16SHANDONG LANJIAN INTELLIGENT EQUIP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411188845.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-05-16
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

In a highly dynamic environment, traditional image processing technology is difficult to identify and track targets stably, and is prone to missing targets or false positives, and there is a lack of effective strategies to manage priority and tracking order between multiple targets in multi-objective scenarios.

Method used

An unmanned equipment target autonomous recognition and tracking system based on visual processing technology is adopted. The system includes an image acquisition module, a grid division module, an area division module, a marking area analysis module, a target classification module, a target tracking module and a remote communication module. Through the combination of Voronoi graph division method, edge detection of Sobel operator and deep learning model, adaptive processing of image features and precise recognition and tracking of targets are achieved.

Benefits of technology

It improves the system's adaptability and recognition efficiency in complex dynamic environments, enhances the processing ability of multi-target scenarios, reduces false positives and missing targets, and improves the recognition and tracking performance of unmanned equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119151990B_ABST
    Figure CN119151990B_ABST
Patent Text Reader

Abstract

This invention discloses an autonomous target recognition and tracking system for unmanned equipment based on vision processing technology, belonging to the field of visual technology. It includes an image acquisition module, which acquires image data from an acquisition device, performs frame-synchronized acquisition based on a clock signal, and preprocesses the image data; and a grid division module, which, based on the Voronoi diagram partitioning method, determines the number of seed points by the number of pixels and performs preliminary edge detection to determine the grid side length. The method of this invention further filters out grid regions requiring additional subdivision using the Sobel operator edge detection method. Combined with the Voronoi diagram partitioning method, the system can adaptively adjust the size and shape of local image windows, achieving better adaptability to image features. By calculating pixel intensity differences and SAD values ​​based on Voronoi regions, it can effectively identify targets in the image that exhibit both subtle changes and overall displacement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual processing technology, in particular to an unmanned equipment target autonomous recognition and tracking system based on visual processing technology. Background Art

[0002] Unmanned equipment usually refers to some unmanned automatic equipment, usually referring to some automated equipment such as drones and unmanned vehicles. The target recognition and tracking system based on visual processing technology obtains environmental information through visual sensors such as cameras. Such equipment can already realize unmanned autonomous recognition and tracking functions in some simple environments, and with the improvement of software frameworks and applications, it has a better user experience in intelligent tracking.

[0003] Although intelligent tracking related technologies have made significant progress, in high-dynamic environments, due to the rapid movement of targets and drastic changes in lighting conditions, traditional image processing technologies are difficult to stably identify and track targets, and are prone to losing targets or false alarms. In multi-target scenarios, there is a lack of effective strategies to manage the priority and tracking order between multiple targets, resulting in a poor user experience for unmanned equipment in complex environments and reducing the recognition efficiency of unmanned equipment. Summary of the invention

[0004] In view of the above problems existing in the existing unmanned equipment target autonomous recognition and tracking system based on visual processing technology, the present invention is proposed.

[0005] Therefore, the problem to be solved by the present invention is that although significant progress has been made in intelligent tracking related technologies, in high dynamic environments, due to the rapid movement of targets and drastic changes in lighting conditions, traditional image processing technologies are difficult to stably identify and track targets, and are prone to losing targets or false alarms. In addition, in multi-target scenarios, there is a lack of effective strategies to manage the priority and tracking order between multiple targets, resulting in a poor user experience of unmanned equipment in complex environments and reducing the recognition efficiency of unmanned equipment.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: an unmanned equipment target autonomous identification and tracking system based on visual processing technology, which comprises:

[0007] The image acquisition module acquires image data according to the acquisition device, performs frame synchronization acquisition based on the clock signal, and pre-processes the image data;

[0008] The grid division module, based on the Voronoi diagram division method, determines the number of seed points by the number of pixels, performs preliminary edge detection on the image to determine the grid side length, calculates the characteristic density of each grid for analysis and judgment, and subdivides the grid;

[0009] The region division module arranges the number of seed points based on the grid subdivision and forms a Voronoi local window. The intensity difference is calculated according to the gray value of the local window and a preliminary sorting is performed. The normalized value of the pixel point is calculated according to the sorting value for relative ranking. The ranking normalized values ​​of different frames are compared and the absolute difference and value between the two frames are calculated.

[0010] The marking area analysis module determines the target marking area based on the global consistency analysis of intensity difference and absolute difference and value;

[0011] The target classification module builds a deep learning model to classify and automatically annotate image data based on target marked areas;

[0012] The target tracking module communicates with the unmanned equipment based on the ROS software framework, continuously collects coordinates and performs path planning and target tracking;

[0013] The remote communication module sets the target loss strategy, performs target re-detection based on the target loss time threshold, connects to the remote control platform for data encryption transmission, and saves log files.

[0014] As a preferred solution of the unmanned equipment target autonomous identification and tracking system based on visual processing technology of the present invention, wherein: the acquisition of image data according to the acquisition device, frame synchronization acquisition based on the clock signal, and pre-processing of the image data include:

[0015] Select a high-performance camera to capture high-resolution images of the target area and obtain continuous frame data;

[0016] Select a high-performance FPGA, configure camera parameters, and perform image preprocessing on the acquired continuous frame data, including grayscale processing and noise removal;

[0017] While the camera is capturing images, it sends the data to the FPGA for frame synchronization through clock signal synchronization.

[0018] As a preferred solution of the unmanned equipment target autonomous identification and tracking system based on visual processing technology of the present invention, wherein: the Voronoi diagram-based partitioning method determines the number of seed points by the number of pixel points, performs preliminary edge detection on the image to determine the grid side length, calculates the feature density of each grid for analysis and judgment, and performs grid subdivision on the grid, including:

[0019] The total number of pixels is calculated by multiplying the image resolution and the acquisition frame rate based on the continuous frame data, and the maximum number of processed pixels per unit time is determined according to the processor clock speed of the FPGA;

[0020] Divide the total number of pixels by the maximum number of processed pixels to get the number of pixels that each seed point is responsible for processing. Further divide the total number of pixels by the number of pixels that each seed point is responsible for processing to get the number of seed points N.

[0021] Based on the total number of pixels of the image divided by the number of seed points N, and calculating its square root, the grid side length S of the image is obtained;

[0022] Based on the integrated Sobel operator edge detection method, the horizontal gradient and vertical gradient are calculated according to the gray value of the pixel position in the image and the horizontal convolution kernel of the Sobel operator, and the square root is further calculated based on the sum of the squares of the horizontal gradient and the vertical gradient to obtain the gradient amplitude;

[0023] Based on the gradient amplitude as the edge strength, the image is initially gridded according to the calculated grid side length S. The feature density D(i, j) of each grid is calculated by dividing the sum of the edge strengths of all pixel positions in the grid by the area of ​​the grid.

[0024] Based on the feature density D(i,j) of each grid, the feature densities of all grids are summed and divided by the total number of grids in the image to obtain the global average feature density, which is used as the feature density threshold;

[0025] If the grid feature density D(i,j) is greater than the feature density threshold, the selected grid feature density D(i,j) is divided by the feature density threshold and rounded up as the subdivision multiple n of the grid. The side length of the subdivided grid is obtained by dividing the grid side length S by the subdivision multiple n, and the original grid S is subdivided into smaller n×n sub-grids.

[0026] As a preferred solution of the unmanned equipment target autonomous identification and tracking system based on visual processing technology described in the present invention, wherein: the grid-based subdivision is used to arrange the seed number points and form a Voronoi local window, the intensity difference is calculated according to the gray value of the local window and a preliminary sorting is performed, the normalized value of the pixel point is calculated according to the sorting value for relative ranking, the ranking normalized values ​​of different frames are compared, and the absolute difference and value between the two frames are calculated, including,

[0027] In the subdivided grid, the center point of each subgrid is arranged as a new seed point;

[0028] Based on the OpenCV image processing open source library, the Voronoi diagram is generated according to the arrangement of seed points, and the image is divided into multiple regions. Each region is dominated by the nearest seed point, and a new local window V(si) is formed with the outline of the region dominated by the seed point as the boundary;

[0029] According to the Voronoi diagram, the local window is divided for processing and analysis. The absolute value of the difference between the grayscale value of all pixels in the local window V(si) and the grayscale average value of the local window V(si) is summed up to obtain the intensity difference ΔI in the Voronoi area V(si) V (s i );

[0030] For each Voronoi region V(si), the intensity differences of all pixels are Construct a matrix and arrange the pixels from small to large based on the pixel intensity difference value. Set the ranking value based on the pixel intensity difference value of the pixel point p(x,y) as the Rank(I(p(x,y))) value. Divide Rank(I(p(x,y))) by the total number of pixels in the corresponding Voronoi region V(si) to obtain the normalized ranking value of the pixel point p(x,y) in its region V(si);

[0031] A new relative ranking is generated by combining the ranking normalized values ​​of all pixels. By comparing the ranking normalized values ​​of the current frame with the reference frame, the minimum absolute difference between the two frames and the movement and position change of the detected target are calculated, which is expressed as:

[0032]

[0033] Among them, SAD(x,y) represents the sum of the absolute differences between the current frame and the previous frame at the pixel coordinates (x,y), k represents the radius of the calculation window, Rt(x+i,y+j) and Rt-1(x+i,y+j) represent the normalized ranking values ​​of the pixel at the coordinates (x+i,y+j), and i and j represent the offsets in the horizontal and vertical directions, respectively.

[0034] As a preferred solution of the unmanned equipment target autonomous identification and tracking system based on visual processing technology of the present invention, wherein: the global consistency of the intensity difference and the absolute difference and value analysis is used to determine the target marking area, including:

[0035] Based on the calculated intensity difference ΔI V (s i ) and SAD values, respectively calculate the intensity difference weight and global consistency within the Voronoi region, the intensity difference weight is obtained by dividing the intensity difference of the Voronoi region V(si) by the sum of the intensity differences of all Voronoi regions V(sj), and the global consistency is obtained by dividing the SAD value of the Voronoi region V(si) by the total SAD value of the image;

[0036] The sum of the intensity difference weight and the global consistency index of the SAD value divided by 2 is defined as the joint detection index of the Voronoi region V(si);

[0037] The average value and standard deviation of the joint detection index of all Voronoi regions are calculated, and the sum of the average value and standard deviation of the joint detection index is used as the global threshold. If the joint detection index is greater than the global threshold, the corresponding Voronoi region V(si) will be determined as the target marking area;

[0038] Based on the determined target marked area, the center position, area, bounding box and relative ranking information of the corresponding area are determined.

[0039] As a preferred solution of the unmanned equipment target autonomous identification and tracking system based on visual processing technology of the present invention, wherein: the deep learning model is constructed to classify and automatically annotate the image data based on the target marking area, including:

[0040] YOLOv5 was selected to build a deep learning model, and image blocks were extracted as input samples based on the target marked area, and image data under different environments, lighting conditions, and motion speeds were collected;

[0041] Use the LabelImg open source tool to label data, label the input data with classification labels, and label the bounding boxes;

[0042] The loss is calculated using the binary cross entropy loss function, and the Adam optimizer is used to update the model parameters through back propagation, gradually reducing the loss value until the loss no longer decreases significantly, then the iteration output model parameter update model is stopped;

[0043] The YOLOv5 model is used to analyze the images captured by the real-time camera, identify and classify the target, and output the bounding box of the target;

[0044] Use the automatic labeling tool combined with the analysis results of the YOLOv5 model to automatically label new data.

[0045] As a preferred solution of the unmanned equipment target autonomous identification and tracking system based on visual processing technology of the present invention, wherein: the information communication with the unmanned equipment based on the ROS software framework includes:

[0046] Real-time detection and extraction of target center coordinates based on deep learning models, and continuous collection and updating;

[0047] Install the ROS software framework on the main control system of the unmanned equipment, create a ROS workspace based on the ROS software framework, use the ROS package management tool to install MAVROS, and configure the MAVLink protocol to communicate with the unmanned equipment;

[0048] According to the continuous acquisition and update of the center position coordinates of the extracted target, based on the OpenCV computer vision library and the output of the deep learning model, the target is recognized and the image data is processed in real time.

[0049] As a preferred solution of the unmanned equipment target autonomous identification and tracking system based on visual processing technology of the present invention, the continuous acquisition of coordinates and path planning and target tracking include:

[0050] Based on the configuration of the lidar sensor for the unmanned equipment, the MoveBase and navigation function package are configured in ROS, and the center position coordinates of the extracted target are continuously collected and updated, and path planning and target tracking are performed. MoveBase automatically calculates the best path and controls the unmanned equipment to move along the path;

[0051] Pass the lidar data to MoveBase, and use MoveBase to adjust the motion trajectory of the unmanned equipment based on the lidar feedback data.

[0052] As a preferred solution of the unmanned equipment target autonomous identification and tracking system based on visual processing technology of the present invention, wherein: the setting of the lost target strategy and re-detection of the target through the lost target time threshold include:

[0053] Based on the system time timer, the target loss time threshold is set according to the historical records. If the unmanned equipment loses the tracking target and the time exceeds the target time threshold, re-detection is performed. The input image is re-detected through the deep learning model to identify the target position of the current image data.

[0054] As a preferred solution of the unmanned equipment target autonomous identification and tracking system based on visual processing technology described in the present invention, wherein: the connection to the remote control platform for data encryption transmission and saving log files refers to remote communication based on the remote control platform using the MAVLink protocol through wireless communication technology, encrypting all data transmissions using the encrypted communication protocol SSL, transmitting the camera data on the unmanned equipment to the remote control platform in real time, and synchronizing the target identification and tracking data of the unmanned equipment to the remote control platform, and recording and saving according to the task log.

[0055] The beneficial effects of the present invention are as follows: by using the edge detection method of the Sobel operator to further screen out the grid areas that need to be additionally subdivided, and in conjunction with the division method of the Voronoi diagram, the system can adaptively adjust the size and shape of the local window of the image, thereby achieving better adaptability to image features; by calculating the pixel intensity difference and the SAD value based on the Voronoi area, it is possible to effectively identify targets in the image that have both subtle changes and overall displacements, thereby enhancing the adaptability of the system in complex dynamic environments; by sorting the intensity difference values ​​and generating a relative ranking, a clear priority order is provided for the tracking and identification of multiple targets; by comprehensively analyzing the intensity difference and the SAD value, it is possible to comprehensively consider the subtle changes and global motions in the region, and in conjunction with the deep learning model for target identification, thereby improving the recognition efficiency in complex scenes with multiple targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0057] Figure 1 This is a schematic diagram of the structure of the unmanned equipment target autonomous recognition and tracking system based on visual processing technology.

[0058] Figure 2 This is a flow chart of the unmanned equipment target autonomous recognition and tracking system based on visual processing technology. DETAILED DESCRIPTION

[0059] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.

[0060] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0061] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments.

[0062] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, and which provides an unmanned equipment target autonomous identification and tracking system based on visual processing technology. The unmanned equipment target autonomous identification and tracking system based on visual processing technology includes:

[0063] S1, acquiring image data according to the acquisition device, performing frame synchronization acquisition based on the clock signal, and preprocessing the image data;

[0064] Preferably, the image data is acquired according to the acquisition device, frame synchronization acquisition is performed based on the clock signal, and the image data is preprocessed, including:

[0065] Choose a high-performance camera, such as the RealSense D435i depth camera, to capture high-resolution images of the target area and obtain continuous frame data;

[0066] Select a high-performance FPGA, such as the Xilinx Zynq-7000 series or the Intel Cyclone V series, configure camera parameters such as resolution and frame rate, and perform image preprocessing on the acquired continuous frame data, including grayscale processing and noise removal;

[0067] While collecting images, the camera sends data to the FPGA for frame synchronization through clock signal synchronization to achieve the transmission integrity of continuous frame data.

[0068] By using high-performance cameras such as depth cameras, high-resolution images can be captured to ensure that the acquired images are rich in details and the target area is accurately captured. The data is sent to the FPGA through clock signal synchronization to achieve synchronous processing of continuous frame data and ensure the integrity of image data transmission. By performing image preprocessing on the FPGA, the image processing speed can be greatly improved. Compared with traditional processor architectures, FPGA can complete the processing of large amounts of image data in a shorter time and reduce data processing delays.

[0069] S2, based on the Voronoi diagram partitioning method, the number of seed points is determined by the number of pixels, and the image is preliminarily edge detected to determine the grid side length, the feature density of each grid is calculated for analysis and judgment, and the grid is subdivided;

[0070] Preferably, based on the Voronoi diagram partitioning method, the number of seed points is determined by the number of pixels, and the image is preliminarily edge detected to determine the grid side length, the characteristic density of each grid is calculated for analysis and determination, and the grid is subdivided, including:

[0071] The total number of pixels is calculated by multiplying the image resolution and the acquisition frame rate based on the continuous frame data, and the maximum number of processed pixels per unit time is determined according to the processor clock speed of the FPGA;

[0072] The total number of pixels is divided by the maximum number of processed pixels to get the number of pixels that each seed point is responsible for processing. The total number of pixels is further divided by the number of pixels that each seed point is responsible for processing to get the number of seed points N, which is expressed as:

[0073]

[0074] Where Pt represents the number of pixels that each seed point is responsible for processing, Po represents the total number of pixels in the image, PPm represents the maximum number of pixels that the processor can process per unit time, PS represents the clock speed of the processor in GHz, CP represents the number of clock cycles required to process each pixel, which can be determined by dividing the actual processing time of the processor by the total number of pixels in the image and taking the average of multiple experimental measurements, and N represents the number of seed points;

[0075] Based on the total number of pixels of the image divided by the number of seed points N and calculating its square root, the grid side length S of the image is obtained, which is expressed as:

[0076]

[0077] Based on the integrated Sobel operator edge detection method, the horizontal gradient and vertical gradient are calculated according to the gray value of the pixel position in the image and the horizontal convolution kernel of the Sobel operator, and the square root of the square sum of the horizontal gradient and the vertical gradient is further calculated to obtain the gradient amplitude, which is expressed as:

[0078]

[0079] Where Gx(x,y) and Gy(x,y) represent the horizontal gradient and vertical gradient at the pixel position (x,y), respectively, G(x,y) represents the gradient amplitude of the image at the pixel position (x,y), I(x+i,y+j) represents the grayscale value of the image at the pixel position (x+i,y+j), Sx(i,j) represents the horizontal convolution kernel value of the Sobel operator, Sy(i,j) represents the vertical convolution kernel value of the Sobel operator, i and j represent the offset in the horizontal and vertical directions, respectively;

[0080] The horizontal convolution kernel value of the Sobel operator and the vertical convolution kernel value of the Sobel operator are expressed as:

[0081]

[0082] Based on the gradient amplitude as the edge strength, the image is initially gridded according to the calculated grid side length S, and the feature density D(i,j) of each grid is calculated;

[0083]

[0084] Where Grij represents the grid cell in the i-th row and j-th column in the image, It means to sum the edge strength G(x,y) of all pixel positions (x,y) belonging to the grid Grij, which means the sum of the edge strengths of all pixels in the grid;

[0085] Based on the feature density D(i,j) of each grid, the feature densities of all grids are summed and divided by the total number of grids in the image to obtain the global average feature density, which is used as the feature density threshold;

[0086] If the grid feature density D(i,j) is greater than the feature density threshold, the selected grid feature density D(i,j) is divided by the feature density threshold and rounded up as the subdivision multiple n of the grid. The side length of the subdivided grid is obtained by dividing the grid side length S by the subdivision multiple n, and the original grid S is subdivided into smaller n×n sub-grids.

[0087] By applying the Voronoi diagram partitioning method, the system can adaptively adjust the size and shape of the local window of the image, achieving better adaptability to image features. After completing the calculation of the number of seed points of the Voronoi diagram, the grid area that needs to be additionally subdivided is further screened out by using the edge detection method of the Sobel operator, so that a more subdivided grid division can be achieved for these features, and it plays an important auxiliary effect for the subsequent Voronoi diagram partitioning operation, so that the Voronoi diagram division can be effectively performed according to the feature density. In addition, the edges in the image are identified by calculating the gradient of the image. The calculation method of providing double derivatives in the horizontal and vertical directions is not common in traditional image processing methods. It can more effectively highlight the edge features in the image, especially better cope with noise interference in the image. Based on the convolution kernel of the Sobel operator, by considering the gradient changes of the pixel point and its neighborhood, the Sobel operator can not only detect the edge, but also capture the direction information of the edge, and further has a certain robustness to the noise in the image. Compared with the edge detection method that only considers the change of a single pixel, the Sobel operator can maintain a higher detection accuracy in a large noise environment.

[0088] S3, arrange the number of seed points based on the grid subdivision and form a Voronoi local window, calculate the intensity difference according to the gray value of the local window and perform preliminary sorting, calculate the normalized value of the pixel points according to the sorting value for relative ranking, compare the ranking normalized values ​​of different frames, and calculate the absolute difference and value between the two frames;

[0089] Preferably, the number of seed points are arranged based on the grid subdivision and a Voronoi local window is formed, the intensity difference is calculated according to the gray value of the local window and a preliminary sorting is performed, the normalized value of the pixel point is calculated according to the sorting value for relative ranking, the ranking normalized values ​​of different frames are compared, and the absolute difference and value between the two frames are calculated, including,

[0090] In the subdivided grid, the center point of each subgrid is arranged as a new seed point;

[0091] Based on the open source library of OpenCV image processing, the Voronoi diagram is generated according to the arrangement of seed points, and the image is divided into multiple regions. Each region is dominated by the nearest seed point, and a new local window V(si) is formed with the outline of the region dominated by the seed point as the boundary, which is expressed as:

[0092]

[0093] Where V(si) represents the window area of ​​the i-th seed point si(xi,yi) divided based on the Voronoi diagram, p(x,y) represents the pixel coordinates, d(p,si) represents the Euclidean distance from the pixel point p(x,y) to the i-th seed point si(xi,yi), R 2 represents a two-dimensional Euclidean space;

[0094] According to the Voronoi diagram, the local window is divided for processing and analysis. The absolute value of the difference between the grayscale value of all pixels in the local window V(si) and the grayscale average value of the local window V(si) is summed up to obtain the intensity difference within the Voronoi area V(si) It reflects the overall change of pixel intensity in the region and is expressed as:

[0095]

[0096] Where p(x,y) represents the pixel coordinates, I(p(x,y)) represents the grayscale value of the pixel. Represents the average gray value of pixels in the local window V(si);

[0097] For each Voronoi region V(si), the intensity differences of all pixels are Construct a matrix and arrange the pixels from small to large based on the pixel intensity difference value. Set the ranking value based on the pixel intensity difference value of the pixel point p(x,y) as the Rank(I(p(x,y))) value. Divide Rank(I(p(x,y))) by the total number of pixels in the corresponding Voronoi region V(si) to obtain the normalized ranking value of the pixel point p(x,y) in its region V(si);

[0098] A new relative ranking is generated by combining the ranking normalized values ​​of all pixels. By comparing the ranking normalized values ​​of the current frame with the reference frame, the minimum absolute difference between the two frames and the movement and position change of the detected target are calculated, which is expressed as:

[0099]

[0100] Among them, SAD(x,y) represents the sum of absolute differences between the current frame and the previous frame at the pixel coordinates (x,y), k represents the radius of the calculation window, Rt(x+i,y+j) and Rt-1(x+i,y+j) represent the normalized ranking values ​​of the pixel at the coordinates (x+i,y+j), and i and j represent the offsets in the horizontal and vertical directions respectively;

[0101] The k is determined based on the square root of the average area of ​​all Voronoi regions rounded up.

[0102] The geometric distance is calculated by the Voronoi diagram partitioning method, which is also applicable in the field of image processing, and can avoid the shortcomings of traditional fixed window partitioning, providing more refined and adaptive image processing. Compared with the existing technology, in traditional image processing, the window is usually fixed to a square or rectangle, and cannot be adjusted according to image features and processing requirements. By dynamically partitioning the window through the Voronoi diagram, the system can better process image features. Under the guidance of the edge intensity map generated by the calculation based on feature density and the Sobel operator, the window dynamically divided by the Voronoi diagram has an irregular polygonal shape, so that each area can adaptively cover the pixel group with similar features in the image, which is beneficial to subsequent image processing operations. At the same time, the dense distribution of seed points in feature-rich areas will make these feature-rich areas have more polygons, thereby refining the processing of feature areas. On the contrary, areas with fewer features will reduce the number of polygons due to sparse seed points, thereby reducing the computational burden of unmanned equipment.

[0103] By calculating the pixel intensity difference based on the area divided by the Voronoi diagram, it is possible to identify which areas in the image have undergone significant changes, which can help the system better identify subtle changes in the image and avoid missing important target information. By focusing on the intensity changes in each Voronoi area, the system can improve the detection sensitivity and accuracy of changes in a small range, especially in complex scenes or when there are multiple targets. This method can effectively avoid the interference of background noise. At the same time, by calculating the SAD value, the system can identify which pixels in the image have undergone significant displacement between two frames, further providing support for target tracking of mobile targets and unmanned equipment. Combined with the division of the Voronoi diagram, this positioning is not only more accurate, but can also be adjusted according to the feature density of the area, thereby improving the detection effect of moving targets. Traditional pixel The intensity difference calculation is performed in a fixed-size window, and after the Voronoi diagram is used for division, these calculations can be performed in an adaptively adjusted area, so that the size and shape of each area can better reflect the actual characteristics of the image. In addition, the intensity difference and SAD calculations can achieve complementarity. The pixel intensity difference calculation focuses on the pixel changes in the local area, while the SAD calculation can detect more significant global displacements. The combined calculation of the two can effectively identify targets with both subtle changes and overall displacements in the image, reflecting the overall changes in the Voronoi area, and enhancing the adaptability of the system in complex dynamic environments. Compared with traditional single change detection methods, it is difficult to achieve this complementary advantage, so it has better non-obviousness, which brings significant technical improvements to the target recognition and tracking system of unmanned equipment.

[0104] By sorting the intensity difference values ​​and generating relative rankings, the impact of external factors can be reduced. Traditional image processing methods usually rely on absolute intensity values ​​to judge changes in images. Converting the order of absolute intensity values ​​into relative rankings can help to more accurately detect and track multiple targets in complex scenes and avoid omissions or false alarms. Because larger means larger changes in consecutive frames, by sorting the intensity difference values ​​and generating relative rankings, the system can objectively evaluate the relative importance of each target area, providing a clear priority for multi-target tracking and identification. This serves as the basis for identification and tracking order, allowing unmanned equipment to dynamically allocate computing resources in multi-target tasks, give priority to target areas with significant changes, effectively improve the overall efficiency of multi-target tracking, and ensure that key targets are tracked more finely.

[0105] S4, global consistency analysis based on intensity difference and absolute difference sum value to determine the target marking area;

[0106] Preferably, the global consistency is analyzed based on the intensity difference and the absolute difference and value to determine the target marking area, including:

[0107] Based on calculated intensity differences The intensity difference weight and global consistency in the Voronoi region are calculated by dividing the intensity difference of the Voronoi region V(si) by the sum of the intensity differences of all Voronoi regions V(sj). The global consistency is obtained by dividing the SAD value of the Voronoi region V(si) by the total SAD value of the image, which is expressed as:

[0108]

[0109] Where W V (s i ) represents the intensity difference weight of the Voronoi region V(si), NV represents the total number of Voronoi regions, It represents the global consistency index of the SAD value of the Voronoi region V(si), Represents the proportion of the motion change of the Voronoi region in the entire image, which is also the global consistency index of the region;

[0110] The sum of the intensity difference weight and the global consistency index of the SAD value divided by 2 is defined as the joint detection index of the Voronoi region V(si);

[0111] The average value and standard deviation of the joint detection index of all Voronoi regions are calculated, and the sum of the average value and standard deviation of the joint detection index is used as the global threshold. If the joint detection index is greater than the global threshold, the corresponding Voronoi region V(si) will be determined as the target marking area;

[0112] Based on the determined target marked area, the center position, area, bounding box and relative ranking information of the corresponding area are determined.

[0113] The intensity difference calculation can capture subtle changes in the Voronoi area, while the SAD calculation focuses on significant movements in a larger range. By combining these two types of information, local and global changes can be considered simultaneously during detection, making target detection more accurate. This dual detection mechanism significantly improves the robustness of the system in complex scenes, and can also reduce false alarms caused by external factors such as lighting changes and noise. By calculating the joint detection index and the adaptive threshold of the joint detection index, the system can automatically adapt to different environmental conditions, whether it is a scene with intense movement or subtle changes in a static background, it can maintain a high detection accuracy, and perform intensity difference and SAD calculations simultaneously in the Voronoi area. Joint calculation avoids redundant pixel-level processing, thereby simplifying the calculation process, improving the execution efficiency of the algorithm, and enabling the system to complete real-time target detection tasks faster, especially in high-resolution images or multi-target scenes. Through the use of joint detection indicators, multiple target areas can be distinguished and marked more accurately. Especially when multiple targets are close to or overlapping, the system can avoid target confusion and improve the accuracy of multi-target detection by comprehensively considering subtle changes and global motion in the area. By directly calculating the center position, area and bounding box of the target based on the Voronoi region, the system can output more accurate target attribute information, providing a reliable foundation for subsequent target tracking and behavior analysis.

[0114] S5, building a deep learning model to classify and automatically annotate image data based on target marked areas;

[0115] Preferably, a deep learning model is constructed to classify and automatically annotate image data based on target marked areas, including:

[0116] Select YOLOv5 to build a deep learning model, including the input layer, backbone network, neck network, and detection head;

[0117] Extract image blocks as input samples based on the target marked area, and collect image data under different environments, lighting conditions, and motion speeds;

[0118] Use the LabelImg open source tool to label data, label the input data with classification labels, and label the bounding boxes;

[0119] The loss is calculated using the binary cross entropy loss function, and the Adam optimizer is used to update the model parameters through back propagation, gradually reducing the loss value until the loss no longer decreases significantly, then the iteration output model parameter update model is stopped;

[0120] The YOLOv5 model is used to analyze the images captured by the real-time camera, identify and classify the target, and output the bounding box of the target;

[0121] Use the automatic labeling tool combined with the analysis results of the YOLOv5 model to automatically label new data, thereby continuously improving model performance.

[0122] The YOLOv5 model can handle target detection tasks of different scales at the same time, adapt to changes in the size of the target, and enable the system to extract higher-level features (such as texture, shape, and semantic information) from the image, which greatly improves the recognition accuracy of complex targets, especially in the case of occlusion, lighting changes, or complex backgrounds. By combining the preliminary marked target areas for further identification, the system can continuously and stably track multiple targets and respond to the movement and changes of the targets in a timely manner, improving the reliability and robustness of target tracking. By regularly updating the data set and retraining, the YOLOv5 model can continuously adapt to changes in the environment and maintain high-precision target detection and classification capabilities.

[0123] S6, based on the ROS software framework, communicates with unmanned equipment, continuously collects coordinates, and performs path planning and target tracking;

[0124] Preferably, information communication with unmanned equipment is performed based on the ROS software framework, including:

[0125] Real-time detection and extraction of target center coordinates based on deep learning models, and continuous collection and updating;

[0126] Install the ROS software framework on the main control system of the unmanned equipment, create a ROS workspace based on the ROS software framework, use the ROS package management tool to install MAVROS, and configure the MAVLink protocol to communicate with the unmanned equipment;

[0127] According to the continuous acquisition and update of the center position coordinates of the extracted target, based on the OpenCV computer vision library and the output of the deep learning model, the target is recognized and the image data is processed in real time.

[0128] Through the information communication of the MAVLink protocol, the system can autonomously adjust the heading and speed to maintain accurate tracking of the target. By utilizing the real-time detection capability of the deep learning model, the system can continuously update the target's location information to ensure uninterrupted tracking of the target during its movement. No matter how complex the target's motion trajectory is, the system can autonomously adjust the motion path and maintain tracking. The system realizes fully automated operation from target detection to motion control, allowing unmanned equipment to automatically identify targets, adjust motion and continue tracking without human intervention. In addition, the modular design of ROS allows the system to flexibly expand functional modules, such as adding new sensor data processing nodes, integrating other control algorithms, etc. Through this scalability, the system can adapt to changing application needs and improve its versatility.

[0129] Furthermore, coordinates are continuously collected and path planning and target tracking are performed, including:

[0130] Based on the configuration of the lidar sensor for the unmanned equipment, the MoveBase and navigation function package are configured in ROS, and the center position coordinates of the extracted target are continuously collected and updated, and path planning and target tracking are performed. MoveBase automatically calculates the best path and controls the unmanned equipment to move along the path;

[0131] Pass the lidar data to MoveBase, and use MoveBase to adjust the motion trajectory of the unmanned equipment based on the lidar feedback data.

[0132] By setting up a lidar sensor to perceive the surrounding environment in real time and generate high-precision environmental point cloud data, MoveBase combines this data for path planning, which can ensure that unmanned equipment chooses the best path in a dynamic environment, avoids obstacles and reaches the target position safely. By continuously monitoring obstacles in the environment through lidar, MoveBase can adjust the movement trajectory in time, avoid obstacles, and prevent collisions. Using lidar data and visual data of the target center position, target tracking is more stable and accurate. Even when the target is temporarily blocked or the ambient light changes greatly, the position information provided by the lidar can still ensure the continuity of target tracking.

[0133] S7, set the target loss strategy, re-detect the target according to the target loss time threshold, connect to the remote control platform for data encryption transmission, and save the log file;

[0134] Preferably, a lost target strategy is set to perform target re-detection according to a lost target time threshold, including:

[0135] Based on the system time timer, the target loss time threshold is set according to the historical records. If the unmanned equipment loses the tracking target and the time exceeds the target time threshold, re-detection is performed. The input image is re-detected through the deep learning model to identify the target position of the current image data.

[0136] By setting the target loss time threshold based on historical records, the system can flexibly respond to different mission scenarios and environmental changes. When the unmanned equipment loses the target in its field of vision and exceeds the set loss time threshold, the system automatically triggers the re-detection mechanism and re-identifies the target position in time. Through precise control of the system time timer, the system can quickly determine whether re-detection is needed after the target is lost, avoiding unnecessary delays.

[0137] Furthermore, connecting to the remote control platform for encrypted data transmission and saving log files means using the MAVLink protocol for remote communication based on the remote control platform through wireless communication technology, encrypting all data transmissions with the encrypted communication protocol SSL, transmitting the camera data on the unmanned equipment to the remote control platform in real time, and synchronizing the target recognition and tracking data of the unmanned equipment to the remote control platform, and recording and saving them according to the task log.

[0138] By using the SSL encrypted communication protocol, all remote communication data is protected during transmission to prevent data from being stolen, tampered with or forged. Through the MAVLink protocol combined with SSL encryption, all control instructions and feedback data are encrypted for transmission to ensure data integrity and confidentiality. The target recognition and tracking data of the unmanned equipment can also be transmitted to the remote control platform in real time. The operator can clearly see the recognition status, position changes and tracking progress of the target, allowing the operator to make adjustments and maintenance during the use of the unmanned equipment. Through the automatic recording and preservation of the task log, all operating instructions, target recognition and tracking data, environmental conditions and other information can be completely saved.

[0139] Example 2

[0140] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0141] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0142] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0143] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0144] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. An unmanned equipment target autonomous recognition and tracking system based on visual processing technology, characterized by: include, The image acquisition module acquires image data according to the acquisition device, performs frame synchronization acquisition based on the clock signal, and pre-processes the image data; The grid division module, based on the Voronoi diagram division method, determines the number of seed points by the number of pixels, performs preliminary edge detection on the image to determine the grid side length, calculates the characteristic density of each grid for analysis and judgment, and subdivides the grid; The region division module arranges the number of seed points based on the grid subdivision and forms a Voronoi local window. The intensity difference is calculated according to the gray value of the local window and a preliminary sort is performed. The normalized value of the pixel is calculated according to the sort value for relative ranking. The sort value based on the pixel intensity difference value of the pixel point p(x,y) is set as the Rank(I(p(x,y))) value. The Rank(I(p(x,y))) is divided by the total number of pixels in the corresponding Voronoi region V(si) to obtain the normalized ranking value of the pixel point p(x,y) in the region V(si) to which it belongs. The normalized ranking values ​​of different frames are compared, and the absolute difference and value between the two frames are calculated. The marking area analysis module determines the target marking area based on the global consistency analysis of intensity difference and absolute difference and value; The target classification module builds a deep learning model to classify and automatically annotate image data based on target marked areas; The target tracking module communicates with the unmanned equipment based on the ROS software framework, continuously collects coordinates and performs path planning and target tracking; The remote communication module sets the target loss strategy, performs target re-detection based on the target loss time threshold, connects to the remote control platform for data encryption transmission, and saves log files.

2. The unmanned equipment target autonomous identification and tracking system based on visual processing technology as claimed in claim 1, characterized in that: The image data is acquired according to the acquisition device, frame synchronization acquisition is performed based on the clock signal, and the image data is preprocessed. include, Select a high-performance camera to capture high-resolution images of the target area and obtain continuous frame data; Select a high-performance FPGA, configure camera parameters, and perform image preprocessing on the acquired continuous frame data, including grayscale processing and noise removal; While the camera is capturing images, it sends the data to the FPGA for frame synchronization through clock signal synchronization.

3. The unmanned equipment target autonomous identification and tracking system based on visual processing technology as claimed in claim 2, characterized in that: The Voronoi diagram-based partitioning method determines the number of seed points by the number of pixels, performs preliminary edge detection on the image to determine the length of the grid side, calculates the feature density of each grid for analysis and judgment, and subdivides the grid, including: The total number of pixels is calculated by multiplying the image resolution and the acquisition frame rate based on the continuous frame data, and the maximum number of processed pixels per unit time is determined according to the processor clock speed of the FPGA; Divide the total number of pixels by the maximum number of processed pixels to get the number of pixels that each seed point is responsible for processing. Further divide the total number of pixels by the number of pixels that each seed point is responsible for processing to get the number of seed points N. Based on the total number of pixels of the image divided by the number of seed points N, and calculating its square root, the grid side length S of the image is obtained; Based on the integrated Sobel operator edge detection method, the horizontal gradient and vertical gradient are calculated according to the gray value of the pixel position in the image and the horizontal convolution kernel of the Sobel operator, and the square root is further calculated based on the sum of the squares of the horizontal gradient and the vertical gradient to obtain the gradient amplitude; Based on the gradient amplitude as the edge strength, the image is initially gridded according to the calculated grid side length S. The feature density D(i, j) of each grid is calculated by dividing the sum of the edge strengths of all pixel positions in the grid by the area of ​​the grid. Based on the feature density D(i,j) of each grid, the feature densities of all grids are summed and divided by the total number of grids in the image to obtain the global average feature density, which is used as the feature density threshold; If the grid feature density D(i,j) is greater than the feature density threshold, the selected grid feature density D(i,j) is divided by the feature density threshold and rounded up as the subdivision multiple n of the grid. The side length of the subdivided grid is obtained by dividing the grid side length S by the subdivision multiple n, and the original grid S is subdivided into smaller n×n sub-grids.

4. The unmanned equipment target autonomous identification and tracking system based on visual processing technology as claimed in claim 3, characterized in that: The method comprises: arranging the number of seed points based on the grid subdivision and forming a Voronoi local window, calculating the intensity difference according to the gray value of the local window and performing preliminary sorting, calculating the normalized value of the pixel points according to the sorting value for relative ranking, comparing the ranking normalized values ​​of different frames, and calculating the absolute difference and value between the two frames, including: In the subdivided grid, the center point of each subgrid is arranged as a new seed point; Based on the OpenCV image processing open source library, the Voronoi diagram is generated according to the arrangement of seed points, and the image is divided into multiple regions. Each region is dominated by the nearest seed point, and a new local window V(si) is formed with the outline of the region dominated by the seed point as the boundary; According to the Voronoi diagram, the local window is divided for processing and analysis. The absolute value of the difference between the grayscale value of all pixels in the local window V(si) and the grayscale average value of the local window V(si) is summed up to obtain the intensity difference within the Voronoi area V(si) For each Voronoi region V(si), the intensity differences of all pixels are A matrix is ​​constructed and arranged based on the pixel intensity difference value from small to large; A new relative ranking is generated by combining the ranking normalized values ​​of all pixels. By comparing the ranking normalized values ​​of the current frame with the reference frame, the minimum absolute difference between the two frames and the movement and position change of the detected target are calculated, which is expressed as: Among them, SAD(x,y) represents the sum of the absolute differences between the current frame and the previous frame at the pixel coordinates (x,y), k represents the radius of the calculation window, Rt(x+i,y+j) and Rt-1(x+i,y+j) represent the normalized ranking values ​​of the pixel at the coordinates (x+i,y+j), and i and j represent the offsets in the horizontal and vertical directions, respectively.

5. The unmanned equipment target autonomous identification and tracking system based on visual processing technology as claimed in claim 4, characterized in that: The global consistency analysis based on intensity difference and absolute difference and value is used to determine the target marking area, including: Based on calculated intensity differences The intensity difference weight and SAD value are calculated in the Voronoi region, respectively. The intensity difference weight is obtained by dividing the intensity difference of the Voronoi region V(si) by the sum of the intensity differences of all Voronoi regions V(sj). The global consistency is obtained by dividing the SAD value of the Voronoi region V(si) by the total SAD value of the image. The sum of the intensity difference weight and the global consistency index of the SAD value divided by 2 is defined as the joint detection index of the Voronoi region V(si); The average value and standard deviation of the joint detection index of all Voronoi regions are calculated, and the sum of the average value and standard deviation of the joint detection index is used as the global threshold. If the joint detection index is greater than the global threshold, the corresponding Voronoi region V(si) will be determined as the target marking area; Based on the determined target marked area, the center position, area, bounding box and relative ranking information of the corresponding area are determined.

6. The unmanned equipment target autonomous identification and tracking system based on visual processing technology as claimed in claim 5, characterized in that: The deep learning model is constructed to classify and automatically annotate the image data based on the target marked area, including: YOLOv5 was selected to build a deep learning model, and image blocks were extracted as input samples based on the target marked area, and image data under different environments, lighting conditions, and motion speeds were collected; Use the LabelImg open source tool to label data, label the input data with classification labels, and label the bounding boxes; The loss is calculated using the binary cross entropy loss function, and the Adam optimizer is used to update the model parameters through back propagation, gradually reducing the loss value until the loss no longer decreases significantly, then the iteration output model parameter update model is stopped; The YOLOv5 model is used to analyze the images captured by the real-time camera, identify and classify the target, and output the bounding box of the target; Use the automatic labeling tool combined with the analysis results of the YOLOv5 model to automatically label new data.

7. The unmanned equipment target autonomous identification and tracking system based on visual processing technology as claimed in claim 6, characterized in that: The information communication with the unmanned equipment based on the ROS software framework includes: Real-time detection and extraction of target center coordinates based on deep learning models, and continuous collection and updating; Install the ROS software framework on the main control system of the unmanned equipment, create a ROS workspace based on the ROS software framework, use the ROS package management tool to install MAVROS, and configure the MAVLink protocol to communicate with the unmanned equipment; According to the continuous acquisition and update of the center position coordinates of the extracted target, based on the OpenCV computer vision library and the output of the deep learning model, the target is recognized and the image data is processed in real time.

8. The unmanned equipment target autonomous identification and tracking system based on visual processing technology as claimed in claim 7, characterized in that: The continuous acquisition of coordinates and path planning and target tracking includes: Based on the configuration of the lidar sensor for the unmanned equipment, the MoveBase and navigation function package are configured in ROS, and the center position coordinates of the extracted target are continuously collected and updated, and path planning and target tracking are performed. MoveBase automatically calculates the best path and controls the unmanned equipment to move along the path; Pass the lidar data to MoveBase, and use MoveBase to adjust the motion trajectory of the unmanned equipment based on the lidar feedback data.

9. The unmanned equipment target autonomous identification and tracking system based on visual processing technology as claimed in claim 8, characterized in that: The setting of the lost target strategy and re-detection of the target by the lost target time threshold include: Based on the system time timer, the target loss time threshold is set according to the historical records. If the unmanned equipment loses the tracking target and the time exceeds the target time threshold, re-detection is performed. The input image is re-detected through the deep learning model to identify the target position of the current image data.

10. The unmanned equipment target autonomous identification and tracking system based on visual processing technology as claimed in claim 9, characterized in that: The connecting to the remote control platform for data encryption transmission and saving log files refers to using the MAVLink protocol for remote communication based on the remote control platform through wireless communication technology, encrypting all data transmission with the encrypted communication protocol SSL, transmitting the camera data on the unmanned equipment to the remote control platform in real time, and synchronizing the target recognition and tracking data of the unmanned equipment to the remote control platform, and recording and saving according to the task log.

Citation Information

Patent Citations

  • Target detection tracking method

    CN111383244A

  • Multi-robot autonomous exploration method and system based on mass center of unknown connected region

    CN116382307A