Target detection method based on sequence instance difference
By using the sequence instance difference method, the problems of missed detection and misidentification of moving targets under background deformation and imaging scenarios with different spatial positions are solved, and target detection is realized in various scenarios such as underwater and surveillance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-14
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional pixel difference algorithms cannot cope with the problems of missed detection and misidentification of moving targets in imaging scenarios with background deformation and different spatial positions.
The sequential instance difference method is adopted to achieve target matching and judgment by image preprocessing, segmentation, establishing a unified scene reference coordinate system, calculating the relative position difference of the target and the feature vector weight difference.
It effectively eliminates background distortion and false target interference, reduces missed detections and false detections, and is suitable for imaging scenarios in different spatial locations.
Smart Images

Figure CN121861284A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image recognition technology, specifically relating to a target detection method based on sequence instance difference. Background Technology
[0002] In target detection techniques such as digital subtraction and image background subtraction, traditional methods typically use pixel subtraction between one or more images and a fixed background image to remove the background and extract the target. However, this type of method has strict prerequisites: the subtracted images must be imaged in the same spatial location, and the background must remain stable.
[0003] However, in practical applications, image acquisition often faces the following problems: Spatial location differences: Images acquired from different spatial locations need to be differentially analyzed to identify moving targets (such as target matching from different perspectives in multi-camera surveillance); Background distortion or changes: The background may be distorted or interfered with due to environmental factors. For example, in underwater scenes, water flow and temperature changes cause background distortion, making it easy to misidentify plankton as targets; in aerial scenes, colorful clouds may be misidentified as flying objects; in industrial scenes, the wrinkles of flexible packaging, steam, or smoke can lead to increased background interference and false targets.
[0004] Traditional pixel difference algorithms cannot handle the above scenarios and are prone to target omission and false detection. Therefore, there is an urgent need for a target detection method that can adapt to background changes and imaging at different spatial locations. Summary of the Invention
[0005] The purpose of this invention is to provide a target detection method based on sequence instance difference, which solves the problems of missed detection and misidentification of moving targets in the prior art under background deformation and imaging scenarios with different spatial positions.
[0006] The technical solution adopted in this invention is a target detection method based on sequence instance difference, which specifically includes the following steps: Step 1: Acquire several images of the same scene and preprocess them; Step 2: Segment the preprocessed images and select the target set for each image based on preset feature vectors; Step 3: Establish a unified scene reference coordinate system for each image; Step 4: Calculate the coordinates of each target within each image relative to its own reference coordinate system independently; Step 5: Select two images, calculate the relative position difference between each target in one image and all targets in the other image, and establish the target matching relationship between the two images based on a preset first threshold, including matched targets and unmatched targets; Step 6: For matched targets, determine the background or moving target by using the feature vector weight difference result; for unmatched targets, apply the feature vector threshold to determine the background or moving target.
[0007] The invention is further characterized by: Preprocessing includes at least one of image cropping, scaling, rotation, grayscale correction, and color correction.
[0008] Step 2 specifically includes the following sub-steps: Step 2.1: Use any one of the following methods to segment the image, namely threshold segmentation, region segmentation, and edge segmentation, to obtain several candidate target regions; Step 2.2: For each candidate target region, extract multidimensional feature parameters and construct a feature vector. , represented as ,in, The three-dimensional coordinates of the target For the target area, For the target perimeter, For the purpose of achieving tightness, For the purpose of symmetry, For the target convexity, For the target roundness, For the target aspect ratio, The target pixel distribution entropy, The target pixel distribution is anisotropic; Choose coefficients for the parameters; Step 2.3: Preset the first feature vector Its components and Completely identical, and Each component serves as the filtering threshold for the corresponding feature; Step 2.4: For the feature vector of each candidate target region With the preset first feature vector Compare the components one by one: If the eigenvector All components are greater than The corresponding component determines whether the candidate target region is a valid target and includes it in the target set of the current image; otherwise, the candidate target region is invalid and is regarded as a background region.
[0009] In step 3, the origin of the scene reference coordinate system is fixed as the centroid of the preset rigid target or target set in the scene, and the origin and overall orientation of the scene reference coordinate system remain fixed.
[0010] The centroid of the target set is calculated using any of the following methods: The centroid coordinates of the target set are obtained by weighting the three-dimensional coordinates of all targets in the target set by using the target area or the total number of pixels of each target as the weight. The centroid coordinates of the target set are obtained by taking the arithmetic mean of the three-dimensional coordinates of all targets in the target set.
[0011] Step 5 specifically includes the following sub-steps: Step 5.1: Select two images to be matched, denoted as image A and image B; Step 5.2: Extract the target set from image A ,in m Let A be the number of targets in image A, and let the coordinates of each target relative to its own reference coordinate system be... ; Extract the target set of image B ,in k Let be the number of targets in image B, and let the coordinates of each target relative to its own reference coordinate system be . ; For each target in image A , in turn with each target in image B Perform pairing and calculate the first position in image A. i The target and image B are in the first j The relative positional difference of each target , means as follows:
[0012] Step 5.3: Set the first threshold ; like Determine the target and Establish a one-to-one correspondence for the matched target pairs; If the target in image A The relative position difference with all targets in image B is greater than ,determination Unmatched target; If the target in image B The relative position difference with all targets in image A is greater than ,determination Unmatched target; Uniqueness check: If the relative positional differences between multiple targets in image B corresponding to one target in image A are all no greater than 1, then the uniqueness check is as follows: If the relative position difference is the smallest, then the pair with the smallest relative position difference is selected as the final matching pair, and the rest are marked as unmatched; similarly, the case where one target in image B corresponds to multiple targets in image A is handled.
[0013] In step 6, for matched targets, the background or moving target is determined by the feature vector weight difference result, specifically as follows: First, for the feature vector Each component is assigned a weight ,in correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity; Secondly, for each matched target pair, calculate the weighted feature vector: Target in Image A i The weighted eigenvectors are represented as follows:
[0014] Target in image B j The weighted eigenvectors are represented as follows:
[0015] Calculation target i With the goal j The weighted eigenvector differences are calculated, and the L2 norm of the differences is used as the eigenvector weight difference norm. , means as follows:
[0016] In the formula, Representing the target i and target j The n One component; Finally, a preset difference norm threshold Q is set when At that time, determine the target i With the goal j For moving targets; when Determine the targeti With the goal j For the background.
[0017] Weight The value range is 0 to 1, and the total weight is 1.
[0018] In step 6, for unmatched targets, a feature vector threshold is applied to determine whether the target is a background or a moving target. Specifically: Preset second feature vector Its components and Completely identical, and Each component serves as the filtering threshold for the corresponding feature; The feature vector of the unmatched target is compared with the preset second feature vector. Compare the components one by one: If all components of the feature vector of the unmatched target are greater than The corresponding component determines whether the unmatched target is a moving target; otherwise, the unmatched target is the background.
[0019] Second eigenvector Each component value is greater than the first feature vector The corresponding components.
[0020] The beneficial effects of this invention are: This invention presents a target recognition calculation method based on sequence instance difference. It does not rely on a fixed background and effectively eliminates interference from background deformation and false targets (such as smoke and clouds) through target set segmentation and feature difference. This method supports imaging at different spatial locations and is applicable to various scenarios such as underwater and surveillance. The method combines coordinate system position matching and feature vector weight difference to achieve dual determination of position and features, reducing missed detections and false detections. Attached Figure Description
[0021] Figure 1 This is a flowchart of the target detection method using sequence instance difference according to the present invention; Figure 2 This is a flowchart of the target registration process for the target recognition calculation method of sequence instance difference in this invention; Figure 3 This is a diagram showing the detection results of an embodiment of the present invention. Detailed Implementation
[0022] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0023] Example 1 This embodiment provides a target detection method based on sequence instance difference, such as Figure 1 As shown, the specific steps include the following: Step 1: Acquire and preprocess several images of the same scene; image preprocessing includes at least one of cropping, scaling, rotation, grayscale correction, and color correction, laying the foundation for subsequent target recognition.
[0024] Step 2: Segment the preprocessed images and select the target set for each image based on preset feature vectors; Step 3: Establish a unified scene reference coordinate system for each image; Step 4: Calculate the coordinates of each target within each image relative to its own reference coordinate system independently; Step 5: Select two images, calculate the relative position difference between each target in one image and all targets in the other image, and establish the target matching relationship between the two images based on a preset first threshold, including matched targets and unmatched targets; Step 6: For matched targets, determine the background or moving target by using the feature vector weight difference result; for unmatched targets, apply the feature vector threshold to determine the background or moving target.
[0025] Example 2 Based on Example 1, step 2 specifically includes the following sub-steps: Step 2.1: Select any one of the following methods—threshold segmentation, region segmentation, and edge segmentation—to segment the preprocessed image based on the scene requirements, and obtain several candidate target regions; Step 2.2: For each candidate target region, extract multidimensional feature parameters and construct a feature vector. , represented as ,in, The three-dimensional coordinates of the target For the target area, For the target perimeter, For the purpose of achieving tightness, For the purpose of symmetry, For the target convexity, For the target roundness, For the target aspect ratio, The target pixel distribution entropy, The target pixel distribution is anisotropic; Choose coefficients for the parameters, i.e. At that time, the feature vector contains corresponding optional parameters. At this time, no corresponding optional parameters are included. Select the optional parameters according to the actual situation.
[0026] Step 2.3: Preset the first feature vector Its components and Completely identical (each component corresponds to the same feature parameters), and Each component serves as the filtering threshold for the corresponding feature; First eigenvector As an effective target selection threshold vector, the values of its components need to be set according to the specific application scenario. For example, in indoor monitoring scenarios, the area of targets such as pedestrians is usually greater than 100 pixels. In It can be set to 100; for industrial inspection scenarios, the target is such as the aspect ratio of a part. Usually greater than 0.5, for In It can be set to 0.5, etc.
[0027] Step 2.4: For the feature vector of each candidate target region With the preset first feature vector Compare the components one by one: If the eigenvector All components are greater than The corresponding components (i.e.) , , , … If the candidate target region is determined to be a valid target, it is included in the target set of the current image; otherwise, the candidate target region is invalid and is regarded as a background region.
[0028] Example 3 Step 3: Establish a unified scene reference coordinate system for each image; The scene reference coordinate system is a globally unified reference framework constructed to achieve consistent modeling of spatial relationships between objects across images. Essentially, it maps pixel coordinates or detection box coordinates, which originally depended on the local perspective of a single frame image, to a stable spatial reference system that does not drift with image acquisition posture, lens distortion, or background deformation. This coordinate system can be a three-dimensional rectangular coordinate system or a two-dimensional polar coordinate system, but it must satisfy the following requirements: strict consistency of coordinate axis directions, uniform scale units, and a stationary origin position among images. Its establishment does not rely on intermediate processes such as image registration or feature point matching, which are easily disturbed by dynamic backgrounds. Instead, it directly anchors to entities or statistical centers with spatial stability in the physical world, thereby avoiding the coordinate system instability problems caused by camera micro-movements, cloud drift, water surface ripples, and vapor obstruction in traditional methods.
[0029] Specifically, the origin of the scene reference coordinate system is fixed at the centroid of a pre-defined rigid target or target set in the scene, and the origin and overall orientation of the scene reference coordinate system remain fixed.
[0030] A rigid target is a fixed reference object pre-set in the image. It can be a target entity that exists in the imaging scene for a long time, whose geometric shape remains basically unchanged, or whose motion freedom is physically constrained, such as the catheter tip of interventional therapy.
[0031] The centroid of the target set is calculated using any of the following methods: The centroid coordinates of the target set are obtained by weighting the 3D coordinates of all targets in the target set using the target area or total number of pixels of each target as the weight; the specific calculation can be performed as follows: Obtain the 3D coordinates of each target and the corresponding target area , , The number of targets is represented by the area of each target. Using the weights, a weighted average is calculated for the three-dimensional coordinates of all targets, using the following formula: , ,
[0032] in, The coordinates of the centroid of the target set.
[0033] The barycenter coordinates of the target set can be obtained by taking the arithmetic mean of the three-dimensional coordinates of all targets in the target set, using the following calculation method: , ,
[0034] Depending on the actual needs, the centroid coordinates of the target set can also be calculated using methods such as target pixel weighting or target volume weighting.
[0035] Example 4 Based on the above embodiments, this embodiment provides a target registration process, such as... Figure 2 As shown.
[0036] Step 5 specifically includes the following sub-steps: Step 5.1: Select any two images to be matched, denoted as image A and image B. The selection of these two images does not depend on the strict continuity of the image acquisition time sequence; that is, it supports selecting adjacent frames in chronological order (e.g., the first...). t Frame and the t +1 frame), and also supports cross-frame skipping selection (such as the first frame). t (frames t+5) to adapt to the recognition needs of targets with different moving speeds.
[0037] Step 5.2: Extract the target set from image A ,in mLet A be the number of targets in image A, and let the coordinates of each target relative to its own reference coordinate system be... ; Extract the target set of image B ,in k Let be the number of targets in image B, and let the coordinates of each target relative to its own reference coordinate system be . ; For each target in image A , in turn with each target in image B Perform pairing and calculate the first position in image A. i The target and image B are in the first j The relative positional difference of each target , means as follows:
[0038] Similarly, the values in image A can be calculated in polar coordinates. i The target and image B are in the first j The relative positional difference of each target .
[0039] Step 5.3: Set the first threshold (This needs to be determined by factors such as image acquisition resolution, actual target size, and scene distance.) like Determine the target and Establish a one-to-one correspondence for the matched target pairs; If the target in image A The relative position difference with all targets in image B is greater than ,determination Unmatched target (image A side); If the target in image B The relative position difference with all targets in image A is greater than ,determination Unmatched target (image B side); The above two-way matching avoids the misjudgment problem caused by background drift.
[0040] Uniqueness check: If the relative positional differences between multiple targets in image B corresponding to one target in image A are all no greater than 1, then the uniqueness check is as follows: If the relative position difference is the smallest, then the pair with the smallest relative position difference is selected as the final matching pair, and the rest are marked as unmatched; similarly, the case where one target in image B corresponds to multiple targets in image A is handled.
[0041] Based on the above steps, spatial association of target instances across images is achieved under dynamic background conditions. This suppresses spurious matching caused by background deformation (such as cloud drift and water surface fluctuations), medium disturbances (such as vapor obstruction and smoke scattering), and changes in imaging pose (such as handheld shaking and gimbal rotation). This enables the system to stably distinguish between real moving targets and background disturbances, thus continuously outputting reliable target motion state determination results in high-dynamic scenarios such as underwater detection, industrial smoke environment monitoring, medical endoscopic image analysis, and low-altitude UAV inspection.
[0042] Example 5 Based on Example 4, moving targets or backgrounds are identified for matched target pairs and unmatched targets.
[0043] For a matched target, the background or moving target is determined by the feature vector weight difference result, specifically: First, for the feature vector Each component is assigned a weight This is used to characterize the relative sensitivity and contribution of each dimension of features in motion discrimination. correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Components; weights The value range is 0 to 1, and the total weight is 1.
[0044] In specific application scenarios, such as underwater scenarios, due to refraction... z Coordinate measurement has significant noise, which can be reduced. Suppressing the impact of depth error on the overall discrimination result; while compactness with roundness Highly correlated in rigid target identification, they can be jointly assigned higher weights to strengthen geometric consistency constraints.
[0045] Secondly, for each matched target pair, calculate the weighted feature vector: Target in Image A i The weighted eigenvectors are represented as follows:
[0046] Target in image B j The weighted eigenvectors are represented as follows:
[0047] Calculation target i With the goal j The weighted eigenvector differences are calculated, and the L2 norm of the differences is used as the eigenvector weight difference norm. , means as follows:
[0048] In the formula, Representing the target i and target j The n One component; The larger the value, the farther apart the two targets are in the weighted feature space, that is, the lower the attribute consistency.
[0049] Finally, a preset difference norm threshold Q is set. The value of Q is determined based on the dynamic intensity of the scene. For example, in an aerial remote sensing scene where clouds drift slowly, Q can be set to 0.35 to 0.45; while when detecting high-speed vibrating components on an industrial assembly line, Q should be set to 0.18 to 0.25 to improve response sensitivity.
[0050] when At that time, determine the target i With the goal j For moving targets; when Determine the target i With the goal j For the background.
[0051] For unmatched targets, a feature vector threshold is applied to determine whether the target is background or a moving target. Specifically: Preset second feature vector Its components and Completely identical, and Each component is the screening threshold for the corresponding feature; the second feature vector Each component value is greater than the first feature vector The corresponding components.
[0052] The feature vector of the unmatched target is compared with the preset second feature vector. Compare the components one by one: If all components of the feature vector of the unmatched target are greater than The corresponding component determines whether the unmatched target is a moving target; otherwise, the unmatched target is the background.
[0053] Example 6 The method of this invention detects impurities inside a liquid bag. First, several images of the liquid bag are acquired, including the bag body and the interface tube. The acquired images can be black-background or white-background images, etc. During the detection process, a coordinate system is constructed with the top of the interface tube as the origin and the vertical direction of the interface tube as the y-axis. Different detection thresholds are set according to the different types of impurities during the detection process.
[0054] The method of this invention is used to detect targets inside the liquid bag, and the Knapp test is followed. The human-machine comparison of the recognition results is shown below. Figure 3 Examples and Table 1 below are shown: Table 1 Test Results
[0055] Figure 3 Examples of dynamic target detection results for two time-division images are provided. In (a) and (b), the circled areas represent detected color block impurities; in (c) and (d), the circled areas represent detected hair impurities. As shown in Table 1 above, the method of this invention achieves a target (impurity) recognition rate of over 95% within the liquid bag, with a false detection rate of less than 3%. The Knapp energy efficiency ratio is greater than 1 (i.e., superior to manual detection).
Claims
1. A target detection method based on sequence instance difference, characterized in that, Specifically, the steps include the following: Step 1: Acquire several images of the same scene and preprocess them; Step 2: Segment the preprocessed images and select the target set for each image based on preset feature vectors; Step 3: Establish a unified scene reference coordinate system for each image; Step 4: Calculate the coordinates of each target within each image relative to its own reference coordinate system independently; Step 5: Select two images, calculate the relative position difference between each target in one image and all targets in the other image, and establish the target matching relationship between the two images based on a preset first threshold, including matched targets and unmatched targets; Step 6: For matched targets, determine the background or moving target by using the feature vector weight difference result; for unmatched targets, apply the feature vector threshold to determine the background or moving target.
2. The target detection method using sequence instance difference according to claim 1, characterized in that, The preprocessing described in step 1 includes at least one of image cropping, scaling, rotation, grayscale correction, and color correction.
3. The target detection method using sequence instance difference according to claim 1, characterized in that, Step 2 specifically includes the following sub-steps: Step 2.1: Use any one of the following methods to segment the image, namely threshold segmentation, region segmentation, and edge segmentation, to obtain several candidate target regions; Step 2.2: For each candidate target region, extract multidimensional feature parameters and construct a feature vector. , represented as ,in, The three-dimensional coordinates of the target For the target area, For the target perimeter, For the purpose of achieving tightness, For the purpose of symmetry, For the target convexity, For the target roundness, For the target aspect ratio, The target pixel distribution entropy, The target pixel distribution is anisotropic; Choose coefficients for the parameters; Step 2.3: Preset the first feature vector Its components and Completely identical, and Each component serves as the filtering threshold for the corresponding feature; Step 2.4: For the feature vector of each candidate target region With the preset first feature vector Compare the components one by one: If the eigenvector All components are greater than The corresponding component determines whether the candidate target region is a valid target and includes it in the target set of the current image; otherwise, the candidate target region is invalid and is regarded as a background region.
4. The target detection method using sequence instance difference according to claim 1, characterized in that, In step 3, the origin of the scene reference coordinate system is fixed as the centroid of a preset rigid target or target set in the scene, and the origin and overall orientation of the scene reference coordinate system remain fixed.
5. The target detection method using sequence instance difference according to claim 4, characterized in that, The centroid of the target set is calculated in any of the following ways: The centroid coordinates of the target set are obtained by weighting the three-dimensional coordinates of all targets in the target set by using the target area or the total number of pixels of each target as the weight. The centroid coordinates of the target set are obtained by taking the arithmetic mean of the three-dimensional coordinates of all targets in the target set.
6. The target detection method using sequence instance difference according to claim 3, characterized in that, Step 5 specifically includes the following sub-steps: Step 5.1: Select two images to be matched, denoted as image A and image B; Step 5.2: Extract the target set from image A ,in m Let A be the number of targets in image A, and let the coordinates of each target relative to its own reference coordinate system be... ; Extract the target set of image B ,in k Let be the number of targets in image B, and let the coordinates of each target relative to its own reference coordinate system be . ; For each target in image A , in turn with each target in image B Perform pairing and calculate the first position in image A. i The target and image B are in the first j The relative positional difference of each target , means as follows: Step 5.3: Set the first threshold ; like Determine the target and Establish a one-to-one correspondence for the matched target pairs; If the target in image A The relative position difference with all targets in image B is greater than ,determination Unmatched target; If the target in image B The relative position difference with all targets in image A is greater than ,determination Unmatched target; Uniqueness check: If the relative positional differences between multiple targets in image B corresponding to one target in image A are all no greater than 1, then the uniqueness check is as follows: If the relative position difference is the smallest, then the pair with the smallest relative position difference is selected as the final matching pair, and the rest are marked as unmatched; similarly, the case where one target in image B corresponds to multiple targets in image A is handled.
7. The target detection method using sequence instance difference according to claim 6, characterized in that, In step 6, for matched targets, the background or moving target is determined by the feature vector weight difference result, specifically as follows: First, for the feature vector Each component is assigned a weight ,in correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity, correspond Quantity; Secondly, for each matched target pair, calculate the weighted feature vector: Target in Image A i The weighted eigenvectors are represented as follows: Target in image B j The weighted eigenvectors are represented as follows: Calculation target i With the goal j The weighted eigenvector differences are calculated, and the L2 norm of the differences is used as the eigenvector weight difference norm. , means as follows: In the formula, Representing the target i and target j The n One component; Finally, a preset difference norm threshold Q is set when At that time, determine the target i With the goal j For moving targets; when Determine the target i With the goal j For the background.
8. The target detection method using sequence instance difference according to claim 7, characterized in that, The weight The value range is 0 to 1, and the total weight is 1.
9. The target detection method using sequence instance difference according to claim 6, characterized in that, In step 6, for unmatched targets, a feature vector threshold is applied to determine whether the target is a background or a moving target. Specifically: Preset second feature vector Its components and Completely identical, and Each component serves as the filtering threshold for the corresponding feature; The feature vector of the unmatched target is compared with the preset second feature vector. Compare the components one by one: If all components of the feature vector of the unmatched target are greater than The corresponding component determines whether the unmatched target is a moving target; otherwise, the unmatched target is the background.
10. The target detection method using sequence instance difference according to claim 9, characterized in that, The second feature vector Each component value is greater than the first feature vector The corresponding components.