Shape rotation adaptive based low-quality image target counting method and system
Patent Information
- Application Number
- CN202610954681.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-11
AI Technical Summary
目前,因成像设备限制、环境干扰或传输损失等因素,导致图像质量显著下降,具体表现为图像模糊、低分辨率、高噪声干扰以及对比度下降等特征,从而使现有目标计数方法无法适用,在低质量图像中出现大量漏检和误检
1.利用深层语义特征预测的旋转与尺度参数构造旋转椭圆模板,对浅层特征进行对齐采样,使细节信息按目标主轴方向聚拢,提升了任意旋转目标的计数准确性;
Smart Images

Figure CN122737618A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and intelligent image analysis technology, and more specifically to a method and system for counting targets in low-quality images based on shape rotation adaptation. Background Technology
[0002] Target counting technology is widely used in traffic monitoring, public safety, urban management, industrial inspection, remote sensing interpretation, and biomedical analysis. Currently, limitations in imaging equipment, environmental interference, or transmission loss lead to a significant decline in image quality, manifesting as blurred images, low resolution, high noise interference, and reduced contrast. This renders existing target counting methods unsuitable, resulting in numerous missed and false detections in low-quality images.
[0003] For example, the MCNN method proposed by Zhang et al. uses multi-column convolutional networks to extract features at different scales, but still uses a fixed-scale Gaussian kernel to convert point annotations into density maps for supervision. When the target is rotated or has a large aspect ratio difference, the supervision signal suffers from a severe mismatch with the real target shape, leading to significant supervision bias. The CSRNet method proposed by Li et al. improves counting performance in dense scenes by increasing the receptive field through dilated convolution, but in low-quality images, shallow details are mixed with noise, and the lack of an effective noise suppression mechanism easily affects the stability and accuracy of counting. The Bayesian Loss method proposed by Ma et al. improves the loss function design, but its underlying supervision signal is still based on a fixed isotropic Gaussian kernel, which cannot accurately describe the true geometric shape of the target in scenes with arbitrarily rotated targets, and the supervision bias problem remains prominent.
[0004] Furthermore, although most existing methods employ rotation transformation data augmentation or multi-scale feature fusion, these rotation parameters are only used for data preprocessing or independent data augmentation stages, and do not form an optimization mechanism that shares parameters with the network's feature extraction process and supervision signals.
[0005] Therefore, existing technologies have not solved the core problems of shape detail alignment and shape parameter sharing between feature enhancement and supervision loss under low-quality image conditions, making it difficult to achieve stable and accurate target counting under complex conditions such as blurred noise and arbitrary rotation. Summary of the Invention
[0006] In view of the above problems, this invention proposes a low-quality image target counting method and system based on shape rotation adaptation, which aims to achieve robust counting of targets with arbitrary orientations through the collaborative training of a shape prior-driven deformable alignment enhancement mechanism and adaptive rotation region fitting loss.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] In a first aspect, embodiments of the present invention provide a method for counting targets in low-quality images based on shape rotation adaptation, the steps of which include: Feature extraction is performed on low-quality images to obtain multi-scale features that include at least deep features; Based on the deep features, shape parameters are obtained through shape priors; The sampling template is determined based on the shape parameters; The sampling template is used to perform aligned sampling on features other than deep features in the multi-scale features to obtain aligned detail features; Obtain the quality map corresponding to the low-quality image, and fuse the quality map with the alignment detail feature and the deep feature to obtain the enhanced feature; The enhanced features are decoded to obtain a target prediction density map and / or a target prediction shape response map.
[0009] In one embodiment, the multi-scale features include shallow features, mid-level features, and deep features.
[0010] In one embodiment, shape parameters are obtained based on the deep features through shape priors, including: The deep features are input into the shape parameter prediction head to obtain a dense parameter map; The corresponding shape parameters are determined based on the target parameters in the dense parameter graph. The shape parameters include rotation angle, major axis scale, minor axis scale, and offset parameter.
[0011] Rotation angle :
[0012] Major axis dimensions :
[0013] minor axis dimensions :
[0014] Offset parameter :
[0015] In the formula, The dense parameters output by the shape parameter prediction head for pixel z.
[0016] In one embodiment, determining the sampling template based on the shape parameters includes: According to the major axis dimension and minor axis dimensions Determine the major axis of the sampling template and short axis ;
[0017] According to the rotation angle Determine the rotation transformation matrix of the sampling template ;
[0018] And according to the rotation transformation matrix Determine the rotation coordinates of each grid point in the sampling template ;
[0019] Based on the rotation coordinates and offset parameters Determine the target sampling coordinates:
[0020]
[0021] in, The coordinates of the target center position are: The target sampling coordinates.
[0022] In one embodiment, obtaining the quality map corresponding to the low-quality image includes: The low-quality image is converted to grayscale space, and the Laplacian variance of pixel values is calculated within each local block of a preset size. The calculation results are then normalized. This yields a quality map characterizing the information richness or clarity of each region.
[0023] In one embodiment, the quality map is fused with the alignment detail feature and the deep feature to obtain an enhanced feature; including: The quality map is multiplied element-wise with the aligned detail feature to obtain the gated detail feature; The enhanced features are obtained by channel concatenation and convolution fusion of the gated detail features and the deep features.
[0024] In one embodiment, the method further includes constructing a joint loss function to train and optimize the counting process, wherein the joint loss function includes at least one of the following losses: shape fitting loss, instance consistency loss, and global consistency loss. The shape fitting loss is constructed based on the target predicted shape response map and the soft supervision map, wherein the soft supervision map is generated based on the shape parameters corresponding to the real labeled target. The instance consistency loss is used to constrain the density integral value within each soft-supervised region of the shape to correspond to a single target, and the total integral of the predicted density map is consistent with the total number of targets. The global consistency loss is constructed based on the predicted density map and point labels.
[0025] In one embodiment, the soft-supervised graph is generated based on the shape parameters corresponding to the real labeled target, including: A rotated elliptical region is constructed based on the shape parameters corresponding to each real labeled target, forming a soft-supervised shape region. The quality map is used as a weighting coefficient to perform pixel-by-pixel weighting on the shape soft supervision region; Spatial overlay and normalization are performed on the weighted shape soft supervision regions of all real labeled targets to generate a soft supervision map for supervising the entire image.
[0026] Secondly, embodiments of the present invention provide a low-quality image target counting system based on shape rotation adaptation. This system is used to implement the low-quality image target counting method based on shape rotation adaptation as described in any of the preceding claims, including: An encoder is used to extract features from low-quality images to obtain multi-scale features that include at least deep features; A shape parameter prediction head is used to obtain shape parameters based on the deep features and shape priors. The rotating ellipse alignment sampling module is used to determine a sampling template based on the shape parameters; and to use the sampling template to perform alignment sampling on features other than deep features in the multi-scale features to obtain aligned detail features. A quality estimator is used to obtain a quality map corresponding to the low-quality image; A quality-gated fusion module is used to fuse the quality map with the alignment detail features and the deep features to obtain enhanced features; A decoder is used to decode the enhanced features to obtain a target prediction density map and / or a target prediction shape response map.
[0027] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the low-quality image target counting method based on shape rotation adaptation as described in any of the preceding claims.
[0028] This invention provides a method and system for counting targets in low-quality images based on shape rotation adaptation. Compared with existing technologies, the beneficial effects include at least the following: 1. A rotating ellipse template is constructed using rotation and scale parameters predicted by deep semantic features. Shallow features are aligned and sampled to make detailed information cluster along the principal axis of the target, thus improving the counting accuracy of any rotating target. 2. The detail injection intensity is adjusted by the quality map, and the detail injection is automatically reduced in the blurred and low-contrast areas, so that the counting results are more stable under blurred noise and low resolution conditions; 3. The same set of shape parameters is used simultaneously for alignment sampling and rotation adaptive soft-supervised region generation, enabling feature enhancement and supervision signals to form a synergistic optimization training that converges faster and is more interpretable.
[0029] This application achieves superior counting accuracy compared to traditional density map methods in scenarios with fuzzy noise, low resolution, and arbitrary rotation. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0031] Figure 1 This is a flowchart of a low-quality image target counting method based on shape rotation adaptation provided in an embodiment of the present invention; Figure 2 This is a comparison image of the shape soft-supervised region generation effect provided in the embodiments of the present invention; Figure 3 This is a schematic diagram illustrating the joint loss function calculation process provided in an embodiment of the present invention; Figure 4 This is a framework diagram of a low-quality image target counting system based on shape rotation adaptation provided in an embodiment of the present invention. Detailed Implementation
[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] This invention discloses a method and system for counting targets in low-quality images based on shape rotation adaptation, aiming to solve the following technical problems in the prior art: (1) When the target in a low-quality image is in an arbitrary rotational posture and the boundary is blurred, the traditional fixed kernel density supervision does not match the real target shape, resulting in supervision bias. (2) In low-quality images, shallow details are highly mixed with noise, and conventional fusion methods are prone to noise leakage, which leads to unstable counting results; (3) Existing rotation enhancement methods do not share parameters with the feature extraction process, and the lack of isomorphic closed-loop mechanism leads to poor robustness of the model under low-quality conditions.
[0034] This application provides a method that differs from existing general methods that rely on conventional multi-scale fusion or random rotation transformation, and can form non-obvious structural improvements for low-quality imaging environments. Specifically, it adopts an encoder-decoder structure. The encoder extracts multi-scale features from the input low-quality image to obtain shallow, mid-level, and deep features. Using the deep features as input, it outputs a set of target orientation and shape parameters. Then, based on the shape parameter set, it constructs a rotated elliptical sampling template to align and sample the shallow and mid-level features to obtain aligned detail features. Furthermore, it fuses the aligned detail features with the deep features to output enhanced features. The decoder is used to output a predicted density map and a predicted shape response map based on the enhanced features.
[0035] The following is a description through specific embodiments.
[0036] Example 1 The low-quality image target counting method based on shape rotation adaptation provided in this embodiment, such as... Figure 1 As shown in the figure, this diagram illustrates the complete flow from low-quality image input to the final output predicted density map and predicted shape response map. In the figure, solid arrows represent the forward data flow path, and dashed arrows represent the parameter feedback path during the training phase. During training, the shape parameter set { After the shape prior prediction head outputs, on the one hand, the solid line path guides the rotating ellipse alignment sampling module to perform feature enhancement, and on the other hand, the dashed line path is used to generate the rotating adaptive soft supervision region and calculate the loss, forming a parameter sharing closed loop.
[0037] In one specific embodiment, the counting step includes: S1, for low-quality images Perform feature extraction to obtain multi-scale features that include at least deep features; In some implementations, the low-quality input images are first preprocessed to uniformly adjust the image resolution to 512. The resolution is 512. Then, normalization is performed to ensure that the values are distributed within the standard range.
[0038] Meanwhile, multi-scale features include shallow features Mid-layer characteristics and deep features For example, the encoder uses VGG16 as the backbone network, with shallow features... From the 4th convolutional block of VGG16, the downsampling factor is 4x, and the number of channels is 512; mid-layer features From the 5th convolutional block onwards, the downsampling factor is 8x, and the number of channels is 512; deep features After an additional 3 layers of dilated convolution processing, the downsampling factor is 16 times, and the number of channels is 512.
[0039] S2. Based on the deep features, shape parameters are obtained through shape priors; in this embodiment, a shape parameter prediction head is set to obtain shape parameters based on deep features. The shape parameter prediction head contains four layers of 3D models. Convolution and 1 layer 1 One convolution, with the number of channels ranging from 512 to 256 to 128 to 64, produces an output with dimension [missing information]. Dense parameter map, These represent the spatial height and width of the deep feature map, respectively.
[0040] Furthermore, based on the target parameters in the dense parameter map, corresponding shape parameters are determined. These shape parameters include rotation angle, major axis scale, minor axis scale, and offset parameters, used to characterize the target's orientation, extension scale, and optional center offset; wherein, Rotation angle :
[0041] Major axis dimensions :
[0042] minor axis dimensions :
[0043] Offset parameter :
[0044] In the formula, tanh is the hyperbolic tangent function, used to limit the rotation angle and center offset within a preset range; softplus is the smoothing positive activation function, used to ensure the major axis scale. and minor axis dimensions It is a positive value; ε is a very small positive number used to prevent scale degradation to 0 and ensure that the output is positive. The dense parameters output by the shape parameter prediction head for pixel z.
[0045] S3. Determine the sampling template based on the shape parameters; For each target center location and First, a 7×7 regular sampling grid is constructed in the local coordinate system, where grid points i,j∈{ 3, 2, 1,0,1,2,3}; Then, the ellipse scale is stretched according to the major axis and minor axis scales to obtain the major axis and minor axis of the sampling template;
[0046] The rotation transformation matrix of the sampling template is determined based on the rotation angle;
[0047] The rotation coordinates of each grid point in the sampling template are determined based on the rotation transformation matrix.
[0048] And determine the target sampling coordinates based on the target center location:
[0049]
[0050] in, The coordinates of the target center position are: For the target sampling coordinates, Offset parameter The components in the x and y directions are used for translation correction at the center of the sampling template. The coordinates for rotating the sampling template.
[0051] S4. Unlike the free offset learning of general deformable convolution, this embodiment uses the sampling template to perform alignment sampling on features other than deep features in the multi-scale features to obtain aligned detail features. In some implementations, differentiable bilinear interpolation is used to sample aligned detail features from shallow and mid-level features. .
[0052] In this way, shape parameters The gradient can be propagated back to the shape prior prediction head through the sampled coordinate expression, thereby achieving adaptive learning of the parameters. This means that the gradient backpropagation of the loss function will directly optimize the geometric parameters used for feature alignment, unlike the independent rotation parameters used at the data augmentation or annotation levels in existing technologies, thus significantly improving the convergence stability and robustness of target counting under low-quality image conditions.
[0053] S5. Obtain the quality map corresponding to the low-quality image, and fuse the quality map with the alignment detail feature and the deep feature to obtain the enhanced feature; this embodiment uses... The gating signal is weighted and modulated to align detail features, injecting significant detail only in high-quality areas and automatically suppressing it in low-quality areas, thereby reducing noise interference.
[0054] In some implementations, obtaining the quality map corresponding to the low-quality image includes: The low-quality image is converted to grayscale, and then processed at each preset size (e.g., 16). 16) Calculate the Laplacian variance of pixel values within a local block, and normalize the results to... This yields a quality map characterizing the information richness or clarity of each region. Among them, quality diagram The initial size is the input low-quality image. .
[0055] Furthermore, the quality map is fused with the alignment detail features and the deep features to obtain enhanced features; including: Align detail features The fused detail features are obtained by channel stitching; and the quality map is improved by bilinear interpolation. Spatial resolution upsampling to align detail features The same size is then multiplied element-wise with the fused detail feature to obtain the gated detail feature; The gated detail features and the deep features are channel-joined and 1 1. Convolutional fusion to obtain the enhanced features ,Right now .
[0056] S6. Decode the enhanced features to obtain a target prediction density map and / or a target prediction shape response map. In some implementations, bilinear interpolation is used to enhance the features. Perform stepwise upsampling, followed by 2 layers of 3 layers after each upsampling. 3. Convolution is used for feature refinement.
[0057] The decoded features are output in two ways: the density prediction head outputs a single-channel predicted density map D after two layers of 3×3 convolution. The density integral is achieved by summing all pixel values of the predicted density map D on the discrete image, and the sum is used as the final count result of the target in the image; the shape prediction outputs a single-channel predicted shape response map S after two layers of 3×3 convolution.
[0058] It should be noted that the step numbers above are for descriptive convenience only and do not constitute a strict limitation on the order of step execution. In some implementations, some steps can be executed in parallel or their order can be adjusted in an equivalent manner.
[0059] Furthermore, the innovation of this invention lies in the sharing and joint optimization of shape parameters between feature alignment and supervision loss to form a closed loop. That is, the same set of shape parameters is not only used for forward guidance of rotating ellipse alignment sampling, but also for backward generation of rotating adaptive soft supervision regions, so as to adopt a collaborative mechanism of rotating adaptive region fitting constraints and feature alignment enhancement based on parameter sharing during the training phase.
[0060] In the inference phase, each pixel position in the entire image is sampled; in the training phase, only the predicted values at the coordinate positions corresponding to the real point labels are selected for loss calculation and feature alignment guidance. The batch size is set to 8 during training, the Adam optimizer is used, the initial learning rate is set to 0.0001, and the cosine annealing strategy is used to decay the learning rate.
[0061] Specifically, during training, training images are first acquired and the locations of real target points are labeled; then, a supervision signal is constructed based on shape parameters, that is, the location of each target's real point is labeled. At this point, using the parameter set, an elliptic Gaussian transform is applied to generate a rotatable, stretchable, and offset soft-supervision region that perfectly matches the target's pose and scale. This solves the technical problem in existing technologies where a fixed Gaussian kernel cannot adapt to arbitrarily rotating targets, leading to misalignment between the monitoring signal and the target shape.
[0062] In this embodiment, the step of generating the shape soft supervision region includes: S1. Parameter Acquisition: Obtain the shape parameters of the actual target point. ,in Let be the rotation angle, and a and b be the radii of the major and minor axes of the ellipse, respectively. This is the center offset parameter.
[0063] S2. Elliptical Region Construction: Based on Parameters At the target marker point A rotating ellipse template is constructed around it. Specifically, the center of the ellipse is first corrected using offset parameters. , and then Establish a local coordinate system with the origin; the elliptical region satisfies
[0064] Where (x, y) represents the pixel relative to the corrected center. The coordinates.
[0065] S3. Gaussian Kernel Assignment and Noise Suppression: Assign Gaussian kernel weights within the rotated elliptical region. For pixel u within the region, rotate its relative coordinates by the specified angle. Transform to the coordinate system of the principal axes of the ellipse to obtain Let and represent the distances along the major axis and minor axis of the ellipse, respectively, and calculate the weights using the following formula:
[0066] in , where a is the major axis radius and b is the minor axis radius.
[0067] In this embodiment, the standard deviations of the elliptic Gaussian distribution along the principal and secondary axes are set to one-quarter of the major axis scale and one-quarter of the minor axis scale, respectively. This setting causes the Gaussian weights to be concentrated mainly near the center of the target and to decrease as the distance along the principal and secondary axes increases.
[0068] In this application, the closer a pixel is to the center of the target, the greater its weight; the farther it is from the center, the smaller its weight. Along the long axis... Control the decay rate, along the minor axis direction using By controlling the decay rate, an elliptical weight distribution is generated instead of a circular one. Weights outside the region are reset to 0, and weights within the region are normalized to obtain a soft-supervision region with a total weight of 1. This allows the monitoring signal to better match the target's orientation and scale characteristics, and reduces noise interference caused by low-quality imaging.
[0069] The generated effect is as follows Figure 2 As shown, the figure is divided into three parts for comparison. Part A shows the original point label position and the arbitrary rotation shape of the real target. Part B is the fixed circular Gaussian kernel supervision region used in traditional methods, which has a shape mismatch problem. Part C is the rotating elliptical soft supervision region proposed in this invention, which can adaptively adjust the rotation angle and major and minor axis scales according to the shape parameters, more accurately fit the real shape of the target, and effectively reduce supervision bias. The solid line represents the generation path from the point label to the soft supervision region, and the dashed line represents the input path of the shape parameter set P.
[0070] Furthermore, the soft supervision regions of all target instances are spatially overlaid, and the overlapping regions are normalized or truncated to prevent numerical overflow in the generated map and ensure the stability of the predicted shape response map S as a probability distribution, thus generating an overall soft supervision map. .
[0071] Then, a joint loss function is constructed to constrain and optimize the network parameters. In this embodiment, the joint loss function constraint mechanism includes at least one of the following losses: global consistency loss, shape fitting loss, and instance consistency loss. The global consistency loss is constructed based on the predicted density map and point labels.
[0072] Shape fitting constraints: Constructed based on the predicted shape response map and soft supervision map of the target; used to constrain the shape response map of the network output. With the aforementioned generated overall soft supervision graph To maintain consistency. This loss is designed to guide the network to autonomously learn and regress the principal axis direction and spatial coverage of the target under weak supervision, given only point annotations, thereby achieving shape perception.
[0073] Instance consistency loss constraint: used to constrain the density map of the network output. The integral characteristics. Specifically, each soft-supervised region... The density integral value within the image corresponds to a single target instance, while the sum of the density integrals of the entire image remains consistent with the total number of targets in the image, to ensure the accuracy of the count.
[0074] In some implementation schemes, refer to Figure 3 The figure illustrates the complete calculation process of the joint loss function, including the instance consistency loss branch, the shape fitting loss branch, and the global consistency loss branch, as shown below:
[0075] in: The global consistency loss term is used to constrain the integral result of the predicted density map D across the entire map to be consistent with the number of point labels. The calculation formula is: Where u represents the pixel position in the input image space. Let N represent the set of all pixel locations, where N is the total number of ground truth pixel labels, and D(u) represent the density value of the predicted density map D at pixel location u. The whole-image integral in a discrete image is the summation over all pixel locations, i.e. This represents the full-image integral result of the predicted density map; if using integral expression, where Ω represents the set of image pixels, This represents the integral summation of the density values at all pixel locations within Ω; The shape fitting loss term, used to constrain the predicted shape response map S to be consistent with the global soft supervision map R generated by the rotated elliptic Gaussian, is calculated as follows:
[0076] Where W(u) is the loss weighting coefficient, W(u)=α+(1-α)Q(u), Q(u) is the quality gate coefficient, and α is the minimum weighting coefficient.
[0077] This represents the instance consistency loss term, used to constrain the soft-supervised region corresponding to each target instance. The integral value of the internal prediction density map D is close to 1, specifically defined as each generated soft supervision region. The difference constraint between the integral value of the predicted density map D and the true label is used to ensure that shape parameter optimization is linked to the counting target; the calculation formula is:
[0078] This item links shape parameter optimization to a single instance count target.
[0079] In this embodiment, the formula represents the integral using the summation form in a discrete image. Since the density map is calculated on a pixel grid, continuous integration... In practical implementation, it is equivalent to discrete summation of the pixel set. Therefore, in this application, "integration" is implemented as pixel summation in discrete image scenarios.
[0080] further, Figure 3 The left side of the middle section represents the input for loss calculation. Density prediction D participates in both instance consistency loss and global consistency loss, while shape prediction S and shape supervision T jointly participate in shape fitting loss. Point annotations participate in global consistency loss. The three losses are weighted at the summation node Σ. , , The weighted summation yields the total loss used for end-to-end network optimization.
[0081] In this embodiment, The same set of shape prior parameters forms a parameter isomorphic shared closed loop between "rotation-adaptive soft-supervised region generation - region fitting loss constraint - feature alignment enhancement - gating fusion". Example 2 Based on the same inventive concept, this embodiment provides a low-quality image target counting system based on shape rotation adaptation. This system is an end-to-end trainable neural network structure, the framework of which is as follows: Figure 4 As shown in the figure, this diagram illustrates the overall network composition, including the encoder, shape parameter prediction head, rotation ellipse alignment sampling module, quality estimator, quality-gated fusion module, and decoder. Solid arrows represent the forward propagation and fusion paths of multi-scale features, while dashed arrows represent the propagation paths of the shape parameter set within the network. The shape parameter set drives the rotation ellipse alignment sampling module to align shallow and mid-level features via the solid path, and simultaneously participates in subsequent loss calculations via the dashed path, achieving parameter sharing between feature alignment and supervised loss.
[0082] In one embodiment, the system specifically includes: An encoder is used to extract features from low-quality images to obtain multi-scale features that include at least deep features; A shape parameter prediction head is used to obtain shape parameters based on the deep features and shape priors. The rotating ellipse alignment sampling module is used to determine a sampling template based on the shape parameters; and to use the sampling template to perform alignment sampling on features other than deep features in the multi-scale features to obtain aligned detail features. A quality estimator is used to obtain a quality map corresponding to the low-quality image; A quality-gated fusion module is used to fuse the quality map with the alignment detail features and the deep features to obtain enhanced features; A decoder is used to decode the enhanced features to obtain a target prediction density map and / or a target prediction shape response map.
[0083] The low-quality image target counting system based on shape rotation adaptation provided in this embodiment of the invention produces the same technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the aforementioned method embodiment, and will not be repeated here.
[0084] Example 3 This embodiment provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the low-quality image target counting method based on shape rotation adaptation as described in any of the preceding embodiments.
[0085] In this embodiment, the computer program can be called by a processor, which in some implementations may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units, microprocessors, digital processing chips, graphics processors and various control chips.
[0086] Any content or technical means not mentioned in the embodiments of this invention can be obtained by referring to the prior art. This disclosure does not limit the scope of the invention and therefore will not be elaborated further.
[0087] The embodiments of the present invention have been described in detail above, and the principles and implementation methods of the present invention have been explained. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of the present invention. Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, computer software program products, or electronic devices, etc. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0088] It should be noted that the word "comprising" does not exclude the presence of components or steps not listed in the claims. The words "a" or "an" preceding a component do not exclude the presence of a plurality of such components. This invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer.
[0089] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0090] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A low-quality image target counting method based on shape rotation self-adaptation, characterized in that, The counting steps include: Feature extraction is performed on low-quality images to obtain multi-scale features that include at least deep features; Based on the deep features, shape parameters are obtained through shape priors; The sampling template is determined based on the shape parameters; The sampling template is used to perform aligned sampling on features other than deep features in the multi-scale features to obtain aligned detail features; Obtain the quality map corresponding to the low-quality image, and fuse the quality map with the alignment detail feature and the deep feature to obtain the enhanced feature; The enhanced features are decoded to obtain a target prediction density map and / or a target prediction shape response map.
2. The shape rotation adaptive based low quality image target counting method of claim 1, wherein, The multi-scale features include shallow features, mid-level features, and deep features.
3. The low-quality image target counting method based on shape rotation adaptation as described in claim 1, characterized in that, Based on the deep features, shape parameters are obtained through shape priors, including: The deep features are input into the shape parameter prediction head to obtain a dense parameter map; The corresponding shape parameters are determined based on the dense parameters in the dense parameter diagram. The shape parameters include rotation angle, major axis scale, minor axis scale, and offset parameter.
4. The low-quality image target counting method based on shape rotation adaptation as described in claim 3, characterized in that, Determining the sampling template based on the shape parameters includes: According to the major axis dimension and minor axis dimensions Determine the major axis of the sampling template and short axis ; According to the rotation angle Determine the rotation transformation matrix of the sampling template ; And according to the rotation transformation matrix Determine the rotation coordinates of each grid point in the sampling template ; Based on the rotation coordinates and offset parameters Determine the target sampling coordinates: in, The coordinates of the target center position are: The target sampling coordinates.
5. The low-quality image target counting method based on shape rotation adaptation as described in claim 1, characterized in that, Obtaining the quality image corresponding to the low-quality image includes: The low-quality image is converted to grayscale space, and the Laplacian variance of pixel values is calculated within each local block of a preset size. The calculation results are then normalized. This yields a quality map characterizing the information richness or clarity of each region.
6. The low-quality image target counting method based on shape rotation adaptation as described in claim 1, characterized in that, The quality map is fused with the alignment detail features and the deep features to obtain enhanced features; including: The quality map is multiplied element-wise with the aligned detail feature to obtain the gated detail feature; The enhanced features are obtained by channel concatenation and convolution fusion of the gated detail features and the deep features.
7. The low-quality image target counting method based on shape rotation adaptation as described in claim 1, characterized in that, It also includes constructing a joint loss function to train and optimize the counting process. The joint loss function includes at least one of the following losses: shape fitting loss, instance consistency loss, and global consistency loss. The shape fitting loss is constructed based on the target predicted shape response map and the soft supervision map, wherein the soft supervision map is generated based on the shape parameters corresponding to the real labeled target. The instance consistency loss is used to constrain the density integral value within each soft-supervised region of the shape to correspond to a single target, and the whole-map integral of the predicted density map is consistent with the total number of targets. The global consistency loss is constructed based on the predicted density map and point labels.
8. The low-quality image target counting method based on shape rotation adaptation as described in claim 7, characterized in that, The soft-supervised graph is generated based on the shape parameters corresponding to the real labeled targets, including: A rotated elliptical region is constructed based on the shape parameters corresponding to each real labeled target, forming a soft-supervised shape region. The quality map is used as a weighting coefficient to perform pixel-by-pixel weighting on the shape soft supervision region; Spatial overlay and normalization are performed on the weighted shape soft supervision regions of all real labeled targets to generate a soft supervision map for supervising the entire image.
9. A low-quality image target counting system based on shape rotation adaptation, characterized in that, The method for counting low-quality images based on shape rotation adaptation as described in any one of claims 1 to 8 includes: An encoder is used to extract features from low-quality images to obtain multi-scale features that include at least deep features; A shape parameter prediction head is used to obtain shape parameters based on the deep features and shape priors. The rotating ellipse alignment sampling module is used to determine a sampling template based on the shape parameters; and to use the sampling template to perform alignment sampling on features other than deep features in the multi-scale features to obtain aligned detail features. A quality estimator is used to obtain a quality map corresponding to the low-quality image; A quality-gated fusion module is used to fuse the quality map with the alignment detail features and the deep features to obtain enhanced features; A decoder is used to decode the enhanced features to obtain a target prediction density map and / or a target prediction shape response map.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the low-quality image target counting method based on shape rotation adaptation as described in any one of claims 1 to 8.