Target detection method and system based on image processing technology
Through the historical annotation data-driven hybrid Gaussian model and lightweight regression network optimization bounding box, combined with teacher-student collaborative training, the problems of high labeling costs and high pseudo-label noise in the existing technology are solved, and efficient and accurate object detection is achieved.
Patent Information
- Application Number
- CN202510571534.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-08
AI Technical Summary
The existing object detection methods rely on large-scale and high-quality labeling data, with high labeling costs, high noise in pseudo labels, lagging in teacher model parameters, and lacking dynamic perception and adjustment.
By obtaining historical annotation data for preprocessing, a hybrid Gaussian model is used to dynamically generate candidate bounding boxes, combining lightweight regression networks and non-maximum suppression, a teacher-student collaborative training mechanism is built, and pseudo-labels are dynamically filtered and teacher models are updated.
Reduced dependence on large-scale rectangular box annotation, improved the recall and positioning accuracy of candidate bounding boxes, reduced the impact of noise samples, and improved the stability and generalization ability of semi-supervised learning.
Smart Images

Figure CN120451505A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and image processing technology, and in particular to a target detection method and system based on image processing technology. Background Art
[0002] With the rapid development of computer vision and deep learning technologies, object detection methods based on convolutional neural networks have made significant progress. For example, the R-CNN, SSD, and YOLO series have achieved a balance between accuracy and speed. However, the current mainstream detection frameworks still rely on large-scale, accurately labeled data, and usually use preset anchor boxes or sliding windows to generate candidate bounding boxes. This method has the following shortcomings: high labeling cost: obtaining large-scale, high-quality rectangular box annotations is not only time-consuming and labor-intensive, but also requires a high level of professionalism from the annotators; high pseudo-label noise: in weakly supervised or semi-supervised scenarios, traditional pseudo-label generation often lacks an effective noise filtering mechanism, which can easily introduce erroneous samples and affect the learning effect of the student model; model update lag: in teacher-student collaborative training, the update of teacher model parameters often relies on manually set time intervals or fixed strategies, and lacks dynamic perception and adjustment of the quality of generated pseudo-labels. Summary of the Invention
[0003] In order to overcome the shortcomings of high labeling cost and low efficiency, the present invention provides a target detection method and system based on image processing technology.
[0004] The technical solution of the present invention is: a target detection method based on image processing technology, comprising the following steps:
[0005] S1: Acquire historical annotation data of the target object and perform data preprocessing on the historical annotation data;
[0006] S2: Based on the pre-processed historical annotation data, candidate bounding boxes are dynamically generated through the pre-trained mixture Gaussian model;
[0007] S3: Optimize the candidate bounding box to obtain the final bounding box;
[0008] S4: Construct a teacher-student collaborative training mechanism to dynamically filter pseudo-labels and update the teacher model in an EMA manner.
[0009] Preferably, the acquiring of historical annotation data of the target object and the preprocessing of the historical annotation data include: acquiring historical annotation data according to a database of the region where the target object is located, the historical annotation data including the size data, position data or context features of the target object, and preprocessing the historical annotation data by removing outliers, deduplicating duplicate data and removing noise annotations.
[0010] Preferably, the method dynamically generates candidate bounding boxes based on the pre-processed historical annotation data through a pre-trained mixed Gaussian model, including: dynamically selecting at least one Gaussian component from the pre-trained mixed Gaussian model according to the pre-processed historical annotation data, generating a horizontal Gaussian kernel and a vertical Gaussian kernel based on the mean and variance of the selected Gaussian component, multiplying the response values of the horizontal Gaussian kernel and the vertical Gaussian kernel to generate a two-dimensional probability heat map, and extracting the candidate bounding box according to a preset threshold.
[0011] Preferably, the method dynamically selects at least one Gaussian component from the pre-trained mixed Gaussian model based on the pre-processed historical annotation data, including: performing cluster analysis on the target object size distribution in the historical annotation data to determine the number of Gaussian components, and synchronously updating the mean, variance and weight of the Gaussian components by jointly optimizing the loss function, wherein the optimized loss function is the intersection-over-union loss of the generated bounding box and the true annotation box and the KL divergence constraint loss of the Gaussian component parameters and the historical data distribution.
[0012] Preferably, the optimized loss function is the intersection-over-union loss between the generated bounding box and the true annotation box and the KL divergence constraint loss between the Gaussian component parameters and the historical data distribution, including: the optimized loss function is:
[0013] L=α1L IOU +α2L KL ;
[0014] Where, L IOU is the intersection-over-union loss between the generated bounding box and the true annotation box; L KL is the KL divergence constraint loss between the Gaussian component parameter and the historical data distribution; α1 and α2 are the weight adjustment coefficients of the optimization loss function;
[0015] L IOU =1-IOU(B gen , B true );
[0016] Where, L IOU B is the intersection-over-union loss between the generated bounding box and the true annotation box; gen is a candidate bounding box; B true is the real annotation box; IOU is the ratio of the overlapping area and the union area of the candidate bounding box and the real annotation box.
[0017] Preferably, the optimizing the candidate bounding box to obtain the final bounding box includes: adjusting the candidate bounding box based on the position and size offset of the predicted bounding box based on a lightweight regression network, dynamically adjusting the suppression threshold by combining non-maximum suppression with the heat map response value, retaining high confidence boxes, and enhancing the candidate bounding box in combination with the semantic features of the target object attribute prediction branch to obtain the final bounding box.
[0018] Preferably, the teacher-student collaborative training mechanism is constructed, pseudo labels are dynamically filtered and the teacher model is updated in an EMA manner, including: using precisely labeled data to train the initial teacher model, the teacher model generates candidate bounding boxes based on point labeled data and a mixed Gaussian model, the teacher model generates pseudo bounding boxes and category prediction results for the point labeled data, low-quality pseudo labels are dynamically filtered according to the weighted result of the confidence score of the pseudo bounding box and the heat map response value, the precisely labeled data and the filtered pseudo label data are input into the student model for joint training, and the teacher model parameters are periodically updated by exponential moving average.
[0019] An object detection system based on image processing technology, comprising:
[0020] A data acquisition module is used to acquire historical annotation data from a database of the area where the target object is located, and perform data preprocessing on the historical annotation data;
[0021] The candidate box generation module is used to dynamically generate candidate bounding boxes based on the mixed Gaussian model and output a two-dimensional probability heat map;
[0022] The bounding box optimization module is used to optimize the candidate bounding boxes through a lightweight regression network and a non-maximum suppression algorithm to obtain the final bounding box;
[0023] The collaborative training module is used to perform collaborative training of teacher and student models and dynamic filtering of pseudo labels, and update the teacher model parameters through EMA.
[0024] Preferably, the candidate box generation module includes:
[0025] a cluster analysis unit for performing cluster analysis on historical size data to determine the number of Gaussian components;
[0026] Dynamic selection unit for jointly optimizing Gaussian component parameters based on KL divergence constraints;
[0027] The heat map generation unit is used to multiply the horizontal and vertical Gaussian kernel response values to generate a two-dimensional probability heat map.
[0028] Preferably, the bounding box optimization module includes:
[0029] The semantic enhancement unit is used to enhance the candidate bounding box by combining the semantic features of the target attribute prediction branch to obtain the final bounding box.
[0030] Beneficial effects:
[0031] 1. Based on historical annotated data, this paper determines the number of Gaussian components through cluster analysis, and jointly optimizes model parameters with KL divergence constraints. It dynamically selects Gaussian components to generate horizontal and vertical Gaussian kernels, and then constructs a two-dimensional probability heat map to accurately extract candidate bounding boxes. This method does not require pre-setting the size and proportion of anchor boxes, and can better adapt to the target size distribution and scene diversity, thereby improving the recall rate and positioning accuracy of candidate bounding boxes.
[0032] 2. This paper uses a lightweight regression network to predict the displacement and size offset of candidate bounding boxes, and combines adaptive non-maximum suppression (dynamic adjustment of the suppression threshold) with heatmap response value weighting to retain high-confidence boxes. It also incorporates the semantic features of the target attribute prediction branch to enhance the candidate bounding boxes, further improving the accuracy and robustness of the final bounding box. This optimization process has low computational complexity and is easy to deploy, making it suitable for application on resource-constrained edge or mobile devices.
[0033] 3. In the collaborative training phase, the present invention generates pseudo bounding boxes and category predictions based on point annotations and a mixed Gaussian model by the teacher model. Low-quality pseudo labels are dynamically filtered using the weighted result of pseudo box confidence and heat map response value, thereby reducing the adverse effects of noise samples on the student model and improving the stability and generalization ability of semi-supervised learning.
[0034] 4. The present invention achieves high-performance target detection in weakly supervised or semi-supervised scenarios by making full use of historical annotation data and a small amount of precise annotations; the candidate bounding box generation and optimization process does not require human intervention, reducing the dependence on large-scale rectangular box annotations. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a flow chart of the target detection method based on image processing technology of the present invention;
[0036] Figure 2 This is a schematic diagram of the structure of the target detection system based on image processing technology of the present invention. DETAILED DESCRIPTION
[0037] Reference herein to an embodiment means that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The appearance of such a phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0038] Example 1: A target detection method based on image processing technology, such as Figure 1 As shown, the following steps are included:
[0039] S1: Acquire historical annotation data of the target object and perform data preprocessing on the historical annotation data;
[0040] According to the database of the area where the target object is located, historical annotation data is obtained, and the historical annotation data includes the size data, position data or context features of the target object, and the historical annotation data is preprocessed to eliminate outliers, duplicate data and noise annotations.
[0041] It should be explained that, according to the database of the area where the target object is located (such as a cloud-based geographic information management system or a local historical image annotation library), all historical annotation data related to the target object are retrieved through an API or database connection interface. The historical annotation data includes the size data of the target object at each moment (such as length, width, height in pixels or physical unit values), position data (such as center point coordinates, four-point coordinates of the bounding box) and context features (such as ambient light intensity, camera viewing angle and shooting distance). The retrieved historical annotation data is subjected to outlier elimination based on statistical methods: the mean μ of each dimension (size or position) and the sum of the two are calculated. The standard deviation σ is used, and samples exceeding μ±3σ are considered abnormal and automatically removed. For contextual features, if the value exceeds the preset engineering experience threshold range, it is also removed. For the remaining data, hash matching is performed according to the timestamp, camera ID and target ID triples. If exactly the same annotation record is found, one is retained and the remaining duplicates are deleted. Annotations with high overlap and close time are also considered duplicates and are merged or deduplicated. For manually or automatically annotated noise samples, a shallow classifier or a rule-based quality scoring mechanism is used to score each annotation. When its quality score is lower than the preset threshold, the noise annotation is removed.
[0042] S2: Based on the pre-processed historical annotation data, candidate bounding boxes are dynamically generated through the pre-trained mixture Gaussian model;
[0043] According to the preprocessed historical annotation data, at least one Gaussian component is dynamically selected from the pre-trained mixed Gaussian model. Based on the mean and variance of the selected Gaussian component, a horizontal Gaussian kernel and a vertical Gaussian kernel are generated respectively. The response values of the horizontal Gaussian kernel and the vertical Gaussian kernel are multiplied to generate a two-dimensional probability heat map, and the candidate bounding box is extracted according to the preset threshold.
[0044] It should be explained that, in the offline stage, the mixed Gaussian model is pre-trained using large-scale labeled samples to determine N Gaussian components, each of which contains a mean vector, a covariance matrix, and a weight. When used online, based on the pre-processed historical labeled data, at least one Gaussian component is dynamically selected by calculating the posterior probability of each sample for each Gaussian component: for component i, when its posterior probability is greater than the first preset probability threshold, the component is included in the subsequent candidate generation, and for each selected Gaussian component μi =(μ x , μ y , μ w , μ h ), in this embodiment, μ x is the average value of the horizontal center coordinates of the candidate bounding box in the image, indicating the location where the target object usually appears in the image in the historical annotation; μ y is the average value of the vertical center coordinates of the candidate bounding box in the image, indicating whether the target usually appears above, in the middle, or below the image; μ w is the mean width of the candidate bounding box, reflecting the typical horizontal size of the target object in historical data; μ h is the mean height of the candidate bounding box, reflecting the typical vertical size of the target object in the historical data, expressed in μ x , σ x Constructing a horizontal Gaussian kernel μ y , σ y Constructing a longitudinal Gaussian kernel The horizontal Gaussian kernel K x With the longitudinal Gaussian kernel K y The response values of are multiplied in the spatial domain to obtain a two-dimensional probability heat map H(x, y). The high-value area of the heat map is the high-probability area where the target appears. According to the preset threshold, the heat map H(x, y) is threshold segmented to extract all connected domains; the minimum enclosing rectangle of each connected domain is calculated to generate a preliminary candidate bounding box; for boxes with too high overlap (IOU), they are merged or the boxes with high heat map response are retained.
[0045] A cluster analysis is performed on the size distribution of target objects in the historical annotated data to determine the number of Gaussian components. The mean, variance, and weight of the Gaussian components are synchronously updated by jointly optimizing the loss function. The optimized loss function is the intersection-over-union loss between the generated bounding box and the true annotated box, and the KL divergence constraint loss between the Gaussian component parameters and the historical data distribution.
[0046] It should be explained that the size distribution (width, height) of the target objects in the pre-processed historical annotation data is clustered and analyzed, and a method such as K-means or spectral clustering is used. The optimal number of clusters is automatically determined according to the cluster silhouette coefficient or the elbow rule, that is, the number of Gaussian components of the mixed Gaussian model. The cluster center is used as the initial mean of each Gaussian component, the variance within each cluster is used as the initial variance, and the proportion of samples in each cluster is used as the initial weight. In each round of iteration, based on the candidate bounding box B gen and the true annotation box B tru Calculate L IOU , and calculate L based on the historical size data distribution KL, using gradient descent or EM algorithm, the initial mean of each Gaussian component, the variance within each cluster, and the proportion of samples in each cluster are simultaneously derived and updated to minimize the joint optimization loss function. After the update is completed, the posterior probability of a given sample x is calculated for each Gaussian component i. When the posterior probability is greater than or equal to the second preset probability threshold, it is considered that the component is highly representative of the current sample and the Gaussian component is included in the subsequent candidate bounding box generation; otherwise, it is discarded.
[0047] The optimized loss function is:
[0048] L=α1L IOU +α2L KL ;
[0049] Where, L IOU is the intersection-over-union loss between the generated bounding box and the true annotation box; L KL is the KL divergence constraint loss between the Gaussian component parameter and the historical data distribution; α1 and α2 are the weight adjustment coefficients of the optimization loss function;
[0050] L IOU =1-IOU(B gen , B true );
[0051] Where, L IOU B is the intersection-over-union loss between the generated bounding box and the true annotation box; gen is a candidate bounding box; B true is the real annotation box; IOU is the ratio of the overlapping area and the union area of the candidate bounding box and the real annotation box.
[0052] It needs to be explained that And define L IOU =1-IOU(B gen ,B true ), the larger the item is, the lower the overlap between the candidate bounding box and the real box is; for each Gaussian component i in the mixed Gaussian model, its parameter distribution is The actual size distribution P of the historical annotation data data (x), calculate its KL divergence: And sum them according to the weight of each component to form This item is used to constrain the consistency of Gaussian component parameters and historical size distribution. During the training process, μ i 、 ω i The gradients of the three parameters are calculated and updated at the learning rate until L converges or reaches the preset number of iterations.
[0053] S3: Optimize the candidate bounding box to obtain the final bounding box;
[0054] Based on the position and size offset of the bounding box predicted by the lightweight regression network, the candidate bounding box is adjusted. The suppression threshold is dynamically adjusted through non-maximum suppression combined with the heat map response value, high confidence boxes are retained, and the candidate bounding box is enhanced by combining the semantic features of the target object attribute prediction branch to obtain the final bounding box.
[0055] It should be explained that the output candidate bounding box is used as the input of the regression branch and sent to a lightweight regression network (for example, a regression head composed of four layers of convolution plus two layers of full connection). The network outputs a four-dimensional offset (Δx, αu, Δw, Δh), where Δx and αy represent the pixel offset of the center of the box in the horizontal and vertical directions respectively; αw and αh represent the scale change of width and height respectively. For each candidate bounding box B cand =(x, y, w, h) and make the following adjustments: x′=x+w×Δx; y′=y+h×Δy; w′=w×exp(Δw); h′=h×exp(Δh); generate the corrected frame B adj =(x′, y′, w′, h′), for all B adj Sort by confidence score from high to low, perform non-maximum suppression in turn, keep the box with the highest score, calculate its IOU with each remaining box, and if the IOU is greater than the threshold T nms , then it is removed; otherwise it is retained. At the same time, the corresponding position heat map response value H(x′, y′) is introduced as the confidence weighting factor, and dynamically adjusted according to the response value distribution Where T base is the basic suppression threshold, ρ is the adjustment coefficient, is the average response of the current heat map, which will be passed through T nms In the box set after dynamic threshold screening, only candidate bounding boxes with confidence scores greater than or equal to the preset candidate threshold are retained to form a preliminary output set; for each retained box, the semantic feature vector extracted by the target object attribute prediction branch (such as the category branch or the fine-grained attribute branch) is fused, and the vector is spliced with the intermediate features of the regression network, and further sent to a small fully connected layer or graph convolution layer for enhancement and correction, and the final box position and category score are output to form the final bounding box output set.
[0056] S4: Construct a teacher-student collaborative training mechanism to dynamically filter pseudo-labels and update the teacher model in an EMA manner.
[0057] An initial teacher model is trained using precisely labeled data. The teacher model generates candidate bounding boxes based on the point labeled data and a mixed Gaussian model. The teacher model generates pseudo bounding boxes and category prediction results for the point labeled data. Low-quality pseudo labels are dynamically filtered based on the weighted result of the pseudo bounding box confidence score and the heat map response value. The precisely labeled data and the filtered pseudo label data are input into the student model for joint training, and the teacher model parameters are periodically updated through exponential moving average.
[0058] It should be explained that a small amount of accurately labeled data (real rectangular boxes and corresponding category labels) is used to train the initial teacher model, which contains a mixture Gaussian candidate box generation module, a lightweight regression network, and a classification branch. The teacher model infers unlabeled or point-only labeled data and generates candidate bounding boxes based on the point annotation position and the mixture Gaussian model; the classification branch outputs the category prediction probability and position offset of each candidate bounding box; the offset box is recorded as a pseudo bounding box with a category confidence score, and for each pseudo bounding box, its confidence score S is calculated. cls The weighted score of the corresponding heat map response value H(x, y): S pseudo =λS cls +(1-λ)H(x, y), when S pseudo When the confidence level is less than a preset threshold, the pseudo-label is regarded as low-quality noise and discarded; only high-quality pseudo-labels are retained for subsequent training; the precisely labeled dataset is merged with the filtered pseudo-label dataset and input into the student model for joint training; the student model structure is the same as the teacher model, but the parameters are updated independently to learn the distribution characteristics of more data samples; finally, the teacher model parameters are periodically updated through exponential moving average.
[0059] Example 2: Based on Example 1, a target detection system based on image processing technology, such as Figure 2 Shown, including:
[0060] A data acquisition module is used to acquire historical annotation data from a database of the area where the target object is located, and perform data preprocessing on the historical annotation data;
[0061] The candidate box generation module is used to dynamically generate candidate bounding boxes based on the mixed Gaussian model and output a two-dimensional probability heat map;
[0062] The bounding box optimization module is used to optimize the candidate bounding boxes through a lightweight regression network and a non-maximum suppression algorithm to obtain the final bounding box;
[0063] The collaborative training module is used to perform collaborative training of teacher and student models and dynamic filtering of pseudo labels, and update the teacher model parameters through EMA.
[0064] The candidate frame generation module includes:
[0065] a cluster analysis unit for performing cluster analysis on historical size data to determine the number of Gaussian components;
[0066] Dynamic selection unit for jointly optimizing Gaussian component parameters based on KL divergence constraints;
[0067] The heat map generation unit is used to multiply the horizontal and vertical Gaussian kernel response values to generate a two-dimensional probability heat map.
[0068] The bounding box optimization module includes:
[0069] The semantic enhancement unit is used to enhance the candidate bounding box by combining the semantic features of the target attribute prediction branch to obtain the final bounding box.
[0070] While the present invention has been described with reference to exemplary embodiments, it is to be understood that the invention is not limited to the disclosed exemplary embodiments. The scope of the following claims is to be accorded the broadest interpretation so as to encompass all modifications and equivalent structures and functions.
Claims
1. A target detection method based on image processing technology, characterized in that: The following steps are involved: S1: Acquire historical annotation data of the target object and perform data preprocessing on the historical annotation data; S2: Based on the pre-processed historical annotation data, candidate bounding boxes are dynamically generated through the pre-trained mixture Gaussian model; S3: Optimize the candidate bounding box to obtain the final bounding box; S4: Construct a teacher-student collaborative training mechanism to dynamically filter pseudo-labels and update the teacher model in an EMA manner.
2. The target detection method based on image processing technology according to claim 1, characterized in that: The method of obtaining historical annotation data of the target object and performing data preprocessing on the historical annotation data includes: obtaining historical annotation data based on a database of the area where the target object is located, the historical annotation data including size data, position data or context features of the target object, and performing preprocessing on the historical annotation data to eliminate outliers, remove duplicate data and eliminate noise annotations.
3. The target detection method based on image processing technology according to claim 1, characterized in that: The method dynamically generates candidate bounding boxes based on the preprocessed historical annotation data through a pretrained mixed Gaussian model, including: dynamically selecting at least one Gaussian component from the pretrained mixed Gaussian model according to the preprocessed historical annotation data, generating a horizontal Gaussian kernel and a vertical Gaussian kernel based on the mean and variance of the selected Gaussian component, multiplying the response values of the horizontal Gaussian kernel and the vertical Gaussian kernel to generate a two-dimensional probability heat map, and extracting the candidate bounding box according to a preset threshold.
4. The target detection method based on image processing technology according to claim 3 is characterized in that: The method dynamically selects at least one Gaussian component from a pre-trained mixed Gaussian model based on the preprocessed historical annotated data, including: performing cluster analysis on the target object size distribution in the historical annotated data to determine the number of Gaussian components, and synchronously updating the mean, variance, and weight of the Gaussian components by jointly optimizing a loss function, wherein the optimized loss function is the intersection-over-union loss between the generated bounding box and the true annotated box and the KL divergence constraint loss between the Gaussian component parameters and the historical data distribution.
5. The target detection method based on image processing technology according to claim 4 is characterized in that: The optimization loss function is the intersection-over-union loss between the generated bounding box and the true annotation box and the KL divergence constraint loss between the Gaussian component parameters and the historical data distribution, including: the optimization loss function is: L=α1L IOU +α2L KL ; Where, L IOU is the intersection-over-union loss between the generated bounding box and the true annotation box; L KL is the KL divergence constraint loss between the Gaussian component parameter and the historical data distribution; α1 and α2 are the weight adjustment coefficients of the optimization loss function; L IOU =1-IOU(B gen ,B true ); Where, L IOU B is the intersection-over-union loss between the generated bounding box and the true annotation box; gen is a candidate bounding box; B true is the real annotation box; IOU is the ratio of the overlapping area and the union area of the candidate bounding box and the real annotation box.
6. The target detection method based on image processing technology according to claim 1 is characterized in that: The optimization of the candidate bounding boxes to obtain the final bounding box includes: adjusting the candidate bounding boxes based on the position and size offset of the predicted bounding boxes based on a lightweight regression network, dynamically adjusting the suppression threshold through non-maximum suppression combined with the heat map response value, retaining high-confidence boxes, and enhancing the candidate bounding boxes in combination with the semantic features of the target object attribute prediction branch to obtain the final bounding box.
7. The target detection method based on image processing technology according to claim 1 is characterized in that: The method constructs a teacher-student collaborative training mechanism, dynamically filters pseudo labels and updates the teacher model in an EMA manner, including: using precisely labeled data to train an initial teacher model, the teacher model generates candidate bounding boxes based on point labeled data and a mixed Gaussian model, the teacher model generates pseudo bounding boxes and category prediction results for the point labeled data, dynamically filters low-quality pseudo labels based on a weighted result of a confidence score of the pseudo bounding box and a heat map response value, inputs precisely labeled data and filtered pseudo label data into a student model for joint training, and periodically updates the teacher model parameters through exponential moving average.
8. A target detection system based on image processing technology, according to the target detection method based on image processing technology according to any one of claims 1 to 7, characterized in that: include: A data acquisition module is used to acquire historical annotation data from a database of the area where the target object is located, and perform data preprocessing on the historical annotation data; The candidate box generation module is used to dynamically generate candidate bounding boxes based on the mixed Gaussian model and output a two-dimensional probability heat map; The bounding box optimization module is used to optimize the candidate bounding boxes through a lightweight regression network and a non-maximum suppression algorithm to obtain the final bounding box; The collaborative training module is used to perform collaborative training of teacher and student models and dynamic filtering of pseudo labels, and update the teacher model parameters through EMA.
9. The target detection system based on image processing technology according to claim 8, characterized in that: The candidate frame generation module includes: a cluster analysis unit for performing cluster analysis on historical size data to determine the number of Gaussian components; Dynamic selection unit for jointly optimizing Gaussian component parameters based on KL divergence constraints; The heat map generation unit is used to multiply the horizontal and vertical Gaussian kernel response values to generate a two-dimensional probability heat map.
10. The target detection system based on image processing technology according to claim 8, characterized in that: The bounding box optimization module includes: The semantic enhancement unit is used to enhance the candidate bounding box by combining the semantic features of the target attribute prediction branch to obtain the final bounding box.