Target identification and positioning method based on non-training feature identification

By designing a non-training feature labeling method and utilizing image acquisition devices and Hough transform technology, real-time target localization without pre-training was achieved. This solves the problems of high labeling costs and decreased localization accuracy in dynamic scenes in traditional technologies, and improves the real-time response capability of the system.

CN121837360APending Publication Date: 2026-04-10HENAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional deep learning-based target recognition technologies suffer from problems such as high annotation costs, lengthy training cycles, poor adaptability to dynamic scenes, high recognition latency, and decreased positioning accuracy in dynamic scenarios. In particular, they cannot quickly reconstruct feature models in sudden task scenarios.

Method used

This paper adopts a non-training feature identification method. By designing distributed and clustered feature identification, images are acquired using an image acquisition device, preprocessed and edge detected, and features are identified based on Hough transform. The relative position of the target object is calculated through similar projection model or spatial constraint relationship, so as to achieve real-time target localization without pre-training.

Benefits of technology

It achieves real-time target localization without pre-training, improves the system's real-time response capability and positioning accuracy, adapts to rapid response in dynamic scenarios, and reduces dependence on pre-training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837360A_ABST
    Figure CN121837360A_ABST
Patent Text Reader

Abstract

The invention relates to a target identification and positioning method based on non-training type feature identification, which comprises the following steps of: pre-designing feature identification which comprises two types of distributed feature identification and clustered feature identification, and enabling one of the feature identification to be adsorbed or attached to a target object; acquiring an image containing the feature identifier and the target object by using an image acquisition device; performing preprocessing and edge detection on the image to obtain an edge image; based on Hough transform, circular features and linear features are identified from the edge image, and circular parameter information and rectangular parameter information are extracted; identifying the type of the feature identifier in the edge image through a preset discrimination condition; and executing a positioning calculation process based on the identified feature identification type. According to the invention, a training-free feature identification coding-decoding system is designed, the dependence degree of the system on pre-training data is reduced, and the real-time response capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target recognition, and more specifically to a target recognition and localization method based on non-trained feature identification. Background Technology

[0002] Traditional deep learning-based target recognition technologies, relying on pre-training methods with large-scale labeled datasets and rigid feature extraction frameworks, often face problems such as high annotation costs, lengthy training cycles, and poor adaptability to dynamic scenes. Furthermore, the spatial relationship between traditional pasted markers and targets depends on manual calibration, which can easily lead to millimeter-level positioning errors in dynamic scenarios such as vibration and deformation. In sudden task scenarios, existing technologies also suffer from high recognition latency and decreased positioning accuracy because they cannot quickly reconstruct the specific feature model to be recognized. Summary of the Invention

[0003] To address the aforementioned problems, this invention proposes a target recognition and localization method based on non-training feature identification, enabling real-time target localization without pre-training, thereby providing an innovative solution for rapid response in dynamic scenarios. The specific technical solution is as follows:

[0004] A target recognition and localization method based on non-trained feature identification includes the following steps:

[0005] Pre-design feature identifiers, which include two types: distributed feature identifiers and clustered feature identifiers, so that one type of feature identifier is adsorbed or attached to the target object;

[0006] Use an image acquisition device to acquire images containing feature identifiers and target objects;

[0007] The image is preprocessed and edge detected to obtain an edge image;

[0008] Based on Hough transform, circular and straight line features are identified from edge images, and circular and rectangular parameter information is extracted.

[0009] Based on the extracted feature parameter information, the type of feature identifier in the edge image is identified by a preset discrimination condition;

[0010] Based on the identified feature types, the localization solution process is executed: Specifically, for distributed feature identifiers, the relative position between the target object and the camera is calculated through a similar projection model; for clustered feature identifiers, spatial constraint relationships between feature points are established, and the spatial pose of the target object relative to the image acquisition device is solved by minimizing the reprojection error, thereby obtaining the spatial position coordinates of the target.

[0011] Optionally, the distributed feature marker consists of a cylinder and a cuboid located below the cylinder. Both the upper side of the cylinder and the upper side of the cuboid have black chamfered surfaces. The upper surface of the cylinder, the upper surface of the cuboid, and the side surface of the cylinder are all white, while the side surface of the cuboid is black, and the upper surface of the cuboid is a square.

[0012] Optionally, when the feature identifier is a clustered feature identifier, there are at least 6 feature identifiers; the clustered feature identifier consists of a cylinder and a hollow cylinder located on the outside of the cylinder. The cylinder is white and the hollow cylinder is black. The diameter of the cylinder is 0.8 times the outer diameter of the hollow cylinder, and the cylinder and the hollow cylinder have the same height.

[0013] Optionally, edge detection uses the Canny edge detection algorithm.

[0014] Optionally, the steps for extracting circular and rectangular parameter information based on Hough transform include: creating a two-dimensional accumulator array corresponding to the pixel resolution of the input image as the Hough space, the two-dimensional accumulator array including a circle accumulator, a line accumulator, and a combined accumulator; traversing each edge pixel in the edge image, traversing all possible circle center positions, radii, and line angle and distance parameters in the Hough space, and performing voting accumulation on the corresponding accumulators; finding local maxima in the Hough space where the voting results exceed a preset threshold, as candidate circle centers and candidate lines; performing non-maximum suppression on the candidate circles and candidate lines to filter out the final circular and rectangular parameter information.

[0015] Optionally, the discrimination criteria for feature identifiers are a weighted sum of the dimensional combination degree, rectangularity, and combination center deviation degree of the circular and rectangular contours they form.

[0016] The beneficial effects of this invention are as follows:

[0017] A training-free feature label encoding-decoding system was designed. By constructing a dynamically scalable feature label library and a lightweight geometric constraint solving module, the system effectively breaks the dependence of the data-driven model on prior knowledge, thereby significantly reducing the system's reliance on pre-training data and improving real-time response capabilities. Through the construction of a feature label system with strong generalization capabilities and a method for quickly determining the relative relationship between labels and targets, real-time target localization without pre-training is achieved, providing an innovative solution for rapid response in dynamic scenarios. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the adsorption of the dispersed feature markers described in the embodiment;

[0020] Figure 2 This is a schematic diagram of the clustered feature identifiers after adsorption as described in the embodiment;

[0021] Figure 3 This is a flowchart of the pose solving and inversion algorithm described in the embodiment. Detailed Implementation

[0022] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0023] The present invention provides the following specific implementation schemes:

[0024] like Figure 1 As shown, taking the fuel tank refueling task as an example, it is suitable to use distributed feature identification based on location and large area. This invention provides a target recognition and localization method based on non-trained feature identification, including the following steps:

[0025] S1. Pre-design a distributed feature marker, which consists of a cylinder and a cuboid located below the cylinder. Both the top of the cylinder and the top of the cuboid have black chamfered surfaces. The top surface of the cylinder, the top surface of the cuboid, and the side surface of the cylinder are white, while the side surface of the cuboid is black. The top surface of the cuboid is a square. The distributed feature marker is then attracted or attached to the target object. The number of distributed feature markers is one.

[0026] S2. Use an image acquisition device to acquire an image containing feature markers and target objects. In this embodiment, the image acquisition device is a camera.

[0027] S3. Perform preprocessing and edge detection on the image. Use the Canny edge detection algorithm to obtain the edge image. The Canny algorithm is as follows: edge_img = Canny(input_img);

[0028] S4. Based on the Hough transform, identify circular and straight line features from the edge image, and extract circular and rectangular parameter information, including the following steps:

[0029] S4.1 Parameter settings: Based on the characteristics of the fuel filler nozzle, set the diameter range of the fuel filler nozzle detection, set a smaller cumulative resolution, and set the minimum detection distance to the maximum value;

[0030] S4.2, Create Accumulators: Create three two-dimensional accumulator arrays, HoughSpace, whose size corresponds to the pixel resolution of the input image. Initialize to 0: HoughSpace = zeros(height, width), one for a circular accumulator, one for a linear accumulator, and one for a combined accumulator;

[0031] S4.3, Loop Detection: For each edge pixel in the input edge image, iterate through all possible center positions and line parameters in the Hough space in turn;

[0032] S4.4, Trigger accumulator increment: For each edge pixel (x,y) in edge_img, for each possible radius in radius_range, for each possible center position (a,b) in HoughSpace, if (xa) 2 +(yb) 2 =r 2 Then, accumulate on HoughSpace(a,b): HoughSpace(a,b)+=1; for each possible straight-line angle θin angle_range, for each possible straight-line distance ρin distance_range. If ρ=x·cos(θ)+y·sin(θ), accumulate on HoughSpace(θ,ρ): HoughSpace(θ,ρ)+=1;

[0033] S4.5 Circle detection: In Hough space, find the local maximum value where the voting result exceeds a preset threshold;

[0034] S4.6 For each point (a,b) in HoughSpace: if HoughSpace(a,b)>threshold: record the local maximum value (a,b): max_candidates.append(a,b); For each point (θ,ρ) in HoughSpace, if HoughSpace(θ,ρ)>threshold, record the local maximum value max_candidates.append(θ,ρ);

[0035] S4.7 For each candidate circle center (a,b) in max_candidates, extract the circle center coordinates. For each possible radius r in radius_range, if the accumulated value of (a,b) in HoughSpace equals r... 2 Then, extract the radius r of the circle; extract the parameters of the candidate lines using the parameters of the line. For each candidate line (θ, ρ) in max_candidates, extract the line angle angle: angle = θ, and calculate the line distance distance: distance = ρ;

[0036] S4.8 Non-maximum suppression: When multiple candidate circles may share the same center position and radius, the candidate circle with the highest voting value should be selected as the final result; when multiple line candidates may share the same line parameters, the candidate line with the highest voting value should be selected. Establish the intersection points of the lines, determine the right angle information of the intersection points, and filter the feature contours that meet the conditions. Finally, obtain the effective feature pixel position information, thereby obtaining all the parameter information of circles and rectangles: Circle parameter information: (x n ,y n ,r n ), Rectangle parameter information: (x m ,y m ,r m ).

[0037] S5. Based on the extracted feature parameter information, the type of feature identifier in the edge image is identified through preset discrimination conditions; specifically, the discrimination conditions are as follows:

[0038]

[0039] In the formula, D mn Q represents the degree of dimensional fit. m For rectangularity, E nm B represents the combined center deviation. n For the overall combination degree, f n For the overall assembly error, W1, W2, and W3 are the weighting coefficients for dimensional assembly degree, rectangularity, and assembly center deviation, respectively.m and h m Define the length and width of the rectangle feature.

[0040] S6. Using a robotic arm to control the camera, capture images of the dispersed feature markers and the target to be identified (fuel filler nozzle) within the visible range. Continue until the pixel error between the identified dispersed feature markers and the target to be identified (fuel filler nozzle) is within s units, then execute the localization calculation process.

[0041] S6.1 Based on the pixel position, the distributed feature identifier, and the pixel size of the target (fuel filler nozzle) to be identified, and combined with the size-pixel projection relationship, calculate the relative positional relationship between the distributed feature identifier and the target (fuel filler nozzle) to be identified.

[0042] S6.2 By combining similar projection relationships, the target point can be located. The principle is as follows:

[0043]

[0044] Where: f is the focal length of the camera (mm); (X,Y,D) Z (a,b) represents the position of the charging port relative to the camera (mm); (a,b) represents the number of pixels in the length and width directions of the camera, D w D represents the size (mm) of the target circular feature. i Let (m,n) be the size of the circular feature in pixels (px), (m,n) be the position of the feature in the image pixel coordinates, and p be the pixel size of the camera.

[0045] like Figure 2 As shown, taking the charging task of a charging port as an example, it is suitable to use clustered feature identification based on pose and small area. This invention provides a target recognition and localization method based on non-trained feature identification, including the following steps:

[0046] S1. A cluster-type feature identifier is pre-designed. The cluster-type feature identifier consists of a cylinder and a hollow cylinder located on the outside of the cylinder. The cylinder is white, and the hollow cylinder is black. The diameter of the cylinder is 0.8 times the outer diameter of the hollow cylinder, and the cylinder and the hollow cylinder have the same height. The cluster-type feature identifier is then attracted or attached to the target object. In this embodiment, the number of cluster-type feature identifiers is 10.

[0047] S2. Use an image acquisition device to acquire an image containing feature markers and target objects. In this embodiment, the image acquisition device is a camera.

[0048] S3. Perform preprocessing and edge detection on the image. Use the Canny edge detection algorithm to obtain the edge image. The Canny algorithm is as follows: edge_img = Canny(input_img);

[0049] S4. Based on the Hough transform, identify circular and straight line features from the edge image, and extract circular and rectangular parameter information, where the rectangular parameter information is 0. This includes the following steps:

[0050] S4.1 Parameter settings: Based on the characteristics of the fuel filler nozzle, set the diameter range of the fuel filler nozzle detection, set a smaller cumulative resolution, and set the minimum detection distance to the maximum value;

[0051] S4.2, Create Accumulators: Create three two-dimensional accumulator arrays, HoughSpace, whose size corresponds to the pixel resolution of the input image. Initialize to 0: HoughSpace = zeros(height, width), one for a circular accumulator, one for a linear accumulator, and one for a combined accumulator;

[0052] S4.3, Loop Detection: For each edge pixel in the input edge image, iterate through all possible center positions and line parameters in the Hough space in turn;

[0053] S4.4, Trigger accumulator increment: For each edge pixel (x,y) in edge_img, for each possible radius in radius_range, for each possible center position (a,b) in HoughSpace, if (xa) 2 +(yb) 2 =r 2 Then, accumulate on HoughSpace(a,b): HoughSpace(a,b)+=1; for each possible straight-line angle θin angle_range, for each possible straight-line distance ρin distance_range. If ρ=x·cos(θ)+y·sin(θ), accumulate on HoughSpace(θ,ρ): HoughSpace(θ,ρ)+=1;

[0054] S4.5 Circle detection: In Hough space, find the local maximum value where the voting result exceeds a preset threshold;

[0055] S4.6 For each point (a,b) in HoughSpace: if HoughSpace(a,b)>threshold: record the local maximum value (a,b): max_candidates.append(a,b); For each point (θ,ρ) in HoughSpace, if HoughSpace(θ,ρ)>threshold, record the local maximum value max_candidates.append(θ,ρ);

[0056] S4.7 For each candidate circle center (a,b) in max_candidates, extract the circle center coordinates. For each possible radius r in radius_range, if the accumulated value of (a,b) in HoughSpace equals r... 2 Then, extract the radius r of the circle; extract the parameters of the candidate lines using the parameters of the line. For each candidate line (θ, ρ) in max_candidates, extract the line angle angle: angle = θ, and calculate the line distance distance: distance = ρ;

[0057] S4.8 Non-maximum suppression: When multiple candidate circles may share the same center position and radius, the candidate circle with the highest voting value should be selected as the final result; when multiple line candidates may share the same line parameters, the candidate line with the highest voting value should be selected. Establish the intersection points of the lines, determine the right angle information of the intersection points, and filter the feature contours that meet the conditions. Finally, obtain the effective feature pixel position information, thereby obtaining all the parameter information of circles and rectangles: Circle parameter information: (x n ,y n ,r n ), Rectangle parameter information: (x m ,y m ,r m ).

[0058] S5. Based on the extracted feature parameter information, the type of feature identifier in the edge image is identified through preset discrimination conditions; specifically, the discrimination conditions are as follows:

[0059]

[0060] In the formula, D mn Q represents the degree of dimensional fit. m For rectangularity, E nm B represents the combined center deviation. n For the overall combination degree, f n For the overall assembly error, W1, W2, and W3 are the weighting coefficients for dimensional assembly degree, rectangularity, and assembly center deviation, respectively. m and h m Define the length and width of the rectangle feature.

[0061] S6. Using a robotic arm to control the camera, eight valid images containing clustered feature identifiers and target data are obtained. The localization calculation process is then executed. By representing points in 3D space using at least four control points linearly, the PnP problem is transformed into an optimization problem of control point coordinates, i.e., a pose calculation and inversion algorithm process. The positional relationship of these four control points in the world coordinate system and camera coordinate system is linked through a system of linear equations. The camera pose is solved by minimizing the reprojection error, thereby inverting the spatial position coordinates (x, y, x) of the feature identifiers and target image information. n ,y n ):

[0062] S6.1, such as Figure 3 As shown, the pose calculation and inversion algorithm flow is as follows: using pixel coordinates (x... apn ,y apn ) and its corresponding spatial coordinates (x on ,y on ,z on This allows us to obtain the position information (x, y) of the charging port coordinate origin relative to the camera center point. pos ,y pos ,z pos ) and attitude information (x ang ,y ang ,z ang For solving the PNP problem, different solution methods require different numbers of valid feature points.

[0063] Based on a vector set consisting of the three-dimensional coordinates of 10 feature points in space, the position of any coordinate point in space can be represented by setting weights. The principle is as follows:

[0064]

[0065] in It is a point with known three-dimensional coordinates in the world coordinate system. yes At which control point α in the world coordinate system? ij It is the weighting coefficient.

[0066] S6.2. Assume that in a task, n preferred features are selected, and the contour information of all possible feature points is obtained. Three special localization features are obtained from all contours, named Feature 1, Feature 2, and Feature 3 in order. The pixel positions of all contours that match the feature points and the bounding rectangle contour information are defined as: (x pn ,y pn ,w pn ,h pn ), (n=1,2…n), thus establishing the constraints for feature 1, which include the interaction relationships between features 1, 2, and 3:

[0067]

[0068] In the formula: D nm Let be the shortest distance (px) between the circumscribed surfaces of features n and m; The deviation coefficients for features n and m; c represents the deviation coefficients of features n and j. pn c pm c pj is the contour adjustment factor for features n, m and j; a is the adjustment matching trend coefficient, which is 10 here; R(n) is the matching degree under the nth combination.

[0069] S6.3. Based on the above constraints, we can obtain (x p1 ,y p1 ,w p1 ,h p1 ), (x p2 ,y p2 ,w p2 ,h p2 ) and (x p3 ,y p3 ,w p3 ,h p3 Based on the feature point information, matching pre-selected positions for all features can be constructed to obtain the spatial positions of different adsorption features under different states. The specific operation is as follows:

[0070]

[0071] In the formula: s represents the direction of the feature point; d o1 The distance between feature 1 and feature 2.

[0072] Therefore, based on different feature location information, it can be transformed into a PNP problem. By minimizing the reprojection error, the pose of the target relative to the camera can be solved, thereby inverting the spatial location coordinates (x, y, z, Rx, Ry, Rz) of the target's image information.

[0073] A training-free feature label encoding-decoding system was designed. By constructing a dynamically scalable feature label library and a lightweight geometric constraint solving module, the system effectively breaks the dependence of data-driven models on prior knowledge, thereby significantly reducing the system's reliance on pre-training data and improving real-time response capabilities. Through the construction of a feature label system with strong generalization capabilities and a method for quickly determining the relative relationship between labels and targets, real-time target localization without pre-training is achieved, providing an innovative solution for rapid response in dynamic scenarios. Furthermore, a local feature reconstruction algorithm based on topological continuity analyzes the spatial topological relationships of labels to complete the features of occluded parts, ensuring accurate localization even with high label occlusion rates. Ultimately, this provides an efficient and robust end-to-end solution for target recognition in unlabeled scenarios such as smart farms, industrial inspection, and logistics sorting.

[0074] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A target recognition and localization method based on non-trained feature identification, characterized in that... This includes the following steps: Pre-design feature identifiers, which include two types: distributed feature identifiers and clustered feature identifiers, so that one type of feature identifier is adsorbed or attached to the target object; Use an image acquisition device to acquire images containing feature identifiers and target objects; The image is preprocessed and edge detected to obtain an edge image; Based on Hough transform, circular and straight line features are identified from edge images, and circular and rectangular parameter information is extracted. Based on the extracted feature parameter information, the type of feature identifier in the edge image is identified by a preset discrimination condition; Based on the identified feature types, the localization solution process is executed: Specifically, for distributed feature identifiers, the relative position between the target object and the camera is calculated through a similar projection model; for clustered feature identifiers, spatial constraint relationships between feature points are established, and the spatial pose of the target object relative to the image acquisition device is solved by minimizing the reprojection error, thereby obtaining the spatial position coordinates of the target.

2. The target recognition and localization method based on non-trained feature identification according to claim 1, characterized in that: The distributed feature marker consists of a cylinder and a cuboid located below the cylinder. Both the top of the cylinder and the top of the cuboid have black chamfered surfaces. The top surface of the cylinder, the top surface of the cuboid, and the side surface of the cylinder are all white, while the side surface of the cuboid is black. The top surface of the cuboid is a square.

3. The target recognition and localization method based on non-trained feature identification according to claim 1, characterized in that: When the feature identifier is a clustered feature identifier, there are at least 6 feature identifiers. The clustered feature identifier consists of a cylinder and a hollow cylinder located on the outside of the cylinder. The cylinder is white and the hollow cylinder is black. The diameter of the cylinder is 0.8 times the outer diameter of the hollow cylinder, and the cylinder and the hollow cylinder have the same height.

4. The target recognition and localization method based on non-trained feature identification according to claim 1, characterized in that: Edge detection uses the Canny edge detection algorithm.

5. The target recognition and localization method based on non-trained feature identification according to claim 1, characterized in that, The steps for extracting parameter information for circles and rectangles based on Hough transform include: Create a two-dimensional accumulator array corresponding to the pixel resolution of the input image as the Hough space. The two-dimensional accumulator array includes a circular accumulator, a linear accumulator, and a combined accumulator. Iterate through each edge pixel in the edge image, iterate through all possible center positions, radii, and angle and distance parameters of the line in the Hough space, and perform voting accumulation on the corresponding accumulator; In the Hough space, find the local maximum value where the voting result exceeds a preset threshold, and use it as the candidate circle center and candidate line. Non-maximum suppression is applied to the candidate circles and candidate lines to filter out the final circle and rectangle parameter information.

6. The target recognition and localization method based on non-trained feature identification according to claim 1, characterized in that: The discrimination criterion for feature identifiers is a weighted sum of the dimensional combination degree, rectangularity, and combination center deviation degree of the circular and rectangular contours they form.