Sugarcane tail real-time counting method based on multi-mode dynamic guidance and topology

Through multimodal dynamic guidance and topology methods, the problems of color similarity interference, dense occlusion and poor adaptability of dynamic scenes in the cane tail count are solved, and high-precision cane tail counting under complex lighting is achieved.

CN120339201AActive Publication Date: 2025-07-18GUANGXI YUEGUI GUANGYE HOLDINGS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510373584.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The prior art has problems such as color similarity interference, insufficient dense occlusion modeling, insufficient use of multimodal information and poor adaptability in dynamic scenes in the sucrose tail count, resulting in insufficient counting accuracy and adaptability.

Method used

Using multimodal dynamic guidance and topology methods, images are acquired through the camera, multimodal encoding and feature stitching are performed, and the total number of sugar tails is generated by combining spatial attention weight maps and maximum bulk density constraints.

Benefits of technology

It improves the ability to suppress color confusion under complex lighting, ensures the rationality of dense occlusion areas distribution, and enhances counting accuracy and dynamic adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339201A_ABST
    Figure CN120339201A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, in particular to a real-time sugarcane tail counting method based on multi-modal dynamic guidance and topology, which comprises the following steps: acquiring an original image of a sugarcane tail; performing multi-modal coding on the original image; extracting multi-scale features of the pre-processed image, and splicing the multi-scale features in a channel dimension to obtain a prediction density map; after the prediction density map is partitioned, continuous homology feature and topological feature fusion is calculated to generate a space attention weight map: an initial density map is generated through transposition convolution, material constraint is performed on the initial density map to obtain a constraint density map, and pixel-level accumulation is performed after the constraint density map is filtered to output the total number of sugarcane tails of the current frame; the method comprises the following steps: collecting a plurality of industrial field images, and constructing a multi-modal data set; and constructing an initial counting model and carrying out model training. According to the invention, the color confusion inhibition capability under complex illumination, the distribution rationality of dense shielding areas and the sugarcane tail counting precision under an industrial scene can be synchronously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a real-time sugarcane tail counting method based on multi-modal dynamic guidance and topology. Background Art

[0002] During the sugar-making process, sugarcane tail counting is mainly used to optimize raw material processing and quality control. The tails of sugarcane usually have a lower sugar content and more fibers, which may affect the sugar-making efficiency and the quality of the finished product. With the development of the automation of the sugar-making industry and the intelligentization of agricultural product quality inspection, the vision-based automatic sugarcane tail counting technology has become a key link to improve the efficiency of raw material statistics. The current mainstream dense object counting methods still have significant technical bottlenecks in the industrial-level sugarcane tail detection scenario:

[0003] Color similarity interference: Traditional convolutional neural networks (CNNs) rely on RGB color space feature extraction and are difficult to effectively distinguish the subtle color differences between sugarcane tails (yellow) and the reflections of metal notches (bright yellow), resulting in the misidentification of the reflective areas as sugarcane tail clusters. Existing improvement schemes use HSV color enhancement, but they cannot dynamically adapt to changes in lighting conditions.

[0004] Insufficient dense occlusion modeling: Methods based on density map regression, such as CSRNet, are prone to feature confusion in severely overlapping sugarcane tail regions. Existing spatial attention mechanisms only focus on local texture differences and lack explicit constraints on the topological connectivity of the targets, resulting in a fragmented distribution of the predicted density map.

[0005] Lack of physical rationality: Mainstream counting models, such as MCNN, directly regress density values without introducing physical prior constraints such as the maximum packing density. The prediction results may exceed the actual material stacking limit, such as predicting >5 sugarcane tails per single pixel, causing the cumulative counting error to be amplified.

[0006] Insufficient utilization of multi-modal information: Existing industrial detection schemes mostly use single visual input and do not effectively fuse scene description texts, such as "high density on the right side of the conveyor belt", with device coordinate information, resulting in the model being difficult to quickly adapt to changes in the camera position or sudden changes in the regional density distribution.

[0007] Poor adaptability to dynamic scenes: Traditional frequency domain enhancement methods, such as wavelet decomposition, use fixed filter banks and cannot adapt to changes in the sugarcane tail distribution pattern. The latest improvement schemes, such as frequency domain attention networks, although introduce learnable band selection, do not jointly optimize with topological persistence features and are difficult to balance between noise suppression and detail retention. Summary of the Invention

[0008] To solve the above problems, the present invention provides a real-time sugarcane tail counting method based on multi-modal dynamic guidance and topology, which can simultaneously improve the ability to suppress color confusion under complex lighting, the distribution rationality of dense occlusion regions, and the sugarcane tail counting accuracy in industrial scenarios.

[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0010] A real-time counting method for sugarcane tails based on multi-modal dynamic guidance and topology, comprising the following steps:

[0011] S1. Photograph the sugarcane tails on the conveyor belt through a camera to obtain the original images of the sugarcane tails;

[0012] S2. Perform multi-modal encoding on the original images to obtain preprocessed images with text encoding and coordinate encoding;

[0013] S3. Extract multi-scale features of the preprocessed images to obtain image features, and splice the image features with the text encoding and the coordinate encoding in the channel dimension to obtain a predicted density map;

[0014] S4. Calculate persistent homology features after dividing the predicted density map into blocks to screen topological features; fuse the divided topological features to generate a spatial attention weight map:

[0015] S5. Gradually upsample the attention weight map to the original resolution through transposed convolution to generate an initial density map, impose a maximum stacking density constraint on the density value of each pixel of the initial density map to obtain a constrained density map, and filter the constrained density map to obtain a filtered density map;

[0016] S6. Perform pixel-level accumulation on the filtered density map to output the total number of sugarcane tails in the current frame;

[0017] S7. Collect multiple industrial site images, mark the center points of the sugarcane tails in the site images, and generate density map labels through Gaussian kernels; associate scene description texts and associated camera coordinate parameters with the site images to construct a multi-modal data set;

[0018] S8. Construct an initial counting model through steps S2 - S6, and train the model through the multi-modal data set to obtain an optimized counting model.

[0019] Further, in step S1, the original images include RGB image size, text prompt words, and physical parameters. The RGB image size is H×W×3, where H is the height of the image; W is the width of the image; the text prompt words are scene descriptions and are used to provide semantic prior information; the physical parameters include notch size, conveyor belt speed, and maximum stacking density.

[0020] Further, in step S2, the original image is normalized before multi-modal encoding. The multi-modal encoding includes text encoding and coordinate encoding. The text encoding inputs the scene description according to the position of the original image, and converts the scene description into a 512-dimensional semantic vector through the CLIP text encoder; the coordinate encoding converts the pixel coordinates of the original image into normalized relative coordinates according to the installation position of the camera, and maps the relative coordinates into a 512-dimensional position encoding through a fully connected layer.

[0021] Further, in step S2, the original image is normalized according to the mean and standard deviation of the dataset:

[0022]

[0023] Wherein, is the original image, and its pixel value range is [0, 225]; is the mean value of all images in the dataset in the RGB three channels; is the standard deviation of all images in the dataset in the RGB three channels; I norn is the normalized original image;

[0024] The encoding method of the text encoding is:

[0025]

[0026] Wherein, Prompt is a natural language prompt; CLIP text is a pre-trained CLIP text encoder; t is a 512-dimensional semantic vector, encoding the semantic information of the text;

[0027] The encoding method of the coordinate encoding is:

[0028]

[0029] Wherein, p(x, y) is the normalized coordinate mapping; (x, y) is the input camera coordinate; H is the height of the RGB image of the original image; W is the width of the RGB image of the original image.

[0030] Further, in step S3, the preprocessed image is input into the ResNet-50 network to extract image features:

[0031]

[0032] Wherein, f ing is the image feature mapping; is the output resolution, and its number of channels is 2048;

[0033] The text encoding is passed through a text encoder to obtain text features:

[0034]

[0035] where f text is the text feature vector obtained after encoding the natural language prompt;

[0036] The coordinate encoding generates coordinate features through coordinate projection:

[0037]

[0038] where is a learnable weight matrix; is a learnable bias vector; f coord is the output 512-dimensional coordinate embedding vector, which represents the spatial information of a certain position in the image;

[0039] The image features, the text features, and the coordinate features are concatenated on the channel to obtain a predicted density map:

[0040]

[0041] where f fused is the feature map after multimodal fusion.

[0042] Furthermore, in step S4, the predicted density map is divided into 8×8 blocks:

[0043]

[0044] where is the predicted density map; K is the number of blocks;

[0045] Persistent homology features are calculated for each block, and connected components with long life cycles are screened and short life cycle noises are suppressed:

[0046] PD(D k ) = {(b i , d i ) | the connected component is born at threshold b i and dies at d i} Formula (9)

[0047] where PD(D k ) is the persistent homology diagram for sub-block D k ; (b i , d i ) is the birth and death of the connected component, and each point (b i , d i ) corresponds to a Betti-0 connected component at threshold bi is born at threshold d i and dies at time b, where Betti-0 is the number of connected components;

[0048] Through sub-block D k The life cycles (d i , b i ) of all connected components are weighted and summed to obtain the vector v k of the topological features of sub-block D k , then there is:

[0049]

[0050]

[0051] where v k is the topological feature vector; (d i - b i ) is the life cycle of the sub-block connected component; d PH is the dimension of the persistent homology vector; φ(b i , d i ) is the Gaussian kernel weighting function; μ is the life cycle mean; σ is the standard deviation, used to control the smoothness of the weighting function;

[0052] When (d i - b i ) ≈ μ and φ(b i , d i ) ≈ 1, it is a connected component with a long life cycle, representing a stable sugarcane tail mass; when (d i - b i ) has a large gap with μ and φ ≈ 0 or φ is very small, it is to suppress short life cycle noise, representing metal reflection.

[0053] Furthermore, in step S4, the screened topological features are calculated for attention weights, and the spatial attention weight map is obtained by element-wise multiplication:

[0054]

[0055] f attn = f fused ⊙ α Equation (13)

[0056] where is the final attention map; s(·) is the Sigmoid function, used to limit the attention value at each pixel position to the interval [0, 1]; is the learnable weight tensor; is the learnable bias term, and b a has the same dimension as α; is the spatial attention weight map output; is the feature map obtained by multi-modal fusion in the previous stage, and C is the number of channels.

[0057] Further, in step S5, the initial density map obtained by gradually upsampling the attention weight map to the original resolution:

[0058]

[0059] where, is the initial density map; ConvTranspose is the transposed convolution operation; is the spatial attention weight map; kernel_size = 4 is the convolution kernel size of 4×4; stride = 2 is the stride of 2; padding = 1 is the padding in the transposed convolution process;

[0060] Apply the maximum stacking density constraint to the density value of each pixel of the initial density map:

[0061] D pred = min(ReLU(D raw ), ρ max ) Equation (15)

[0062] where, D pred is the constrained density map; ReLU(x) is defined as max(0, x) to make the density value output greater than 0; ρ max is the maximum physical stacking density; min(ReLU(D raw ), ρ max ) performs a clip operation on all pixel values. If it exceeds the upper limit, it is pushed back to ρ max ;

[0063] Perform Gaussian filtering on the constrained density map along the horizontal movement direction of the conveyor belt to obtain the filtered density map.

[0064] Further, in step S8, the initial counting model is trained by basic training, topological fine-tuning, and joint optimization in sequence,

[0065] The basic training is performed by minimizing the pixel-level error of the density image and constraining the maximum density value;

[0066] The topological fine-tuning is performed by freezing the weights of the image encoder and introducing the persistent homology loss to optimize the topological structure of the density map;

[0067] The joint optimization is performed by minimizing the pixel-level error of the density image, constraining the maximum density value, introducing persistent homology loss to optimize the topological structure of the density map, and cross-modal alignment for training, and unfreezing all parameters. The joint optimization is trained until convergence to obtain an optimized counting model.

[0068] The beneficial effects of the present invention are as follows:

[0069] Through multi-modal encoding, the image integrates text semantics, image features, and spatial coordinate information. Through the cross-modal dynamic guidance mechanism, the robustness of the model to interference such as illumination changes and metal reflections is significantly improved, and the false detection problem caused by color confusion is effectively suppressed; based on the block-based persistent homology feature extraction technology, the topological connectivity of the cane tail distribution is explicitly constrained, avoiding the density map fragmentation problem generated by traditional methods in severely occluded areas, and ensuring the counting rationality in densely stacked scenarios; introducing the maximum stacking density threshold and directional post-processing filtering, forcing the density prediction value to conform to the actual stacking characteristics of the material, and solving the problem of over-prediction caused by the lack of physical rules in traditional models; the text prompt words and coordinate encoding cooperate to guide the attention distribution of the model, realizing the dynamic alignment of semantic prior and visual features, and enhancing the interpretability of the counting result and the efficiency of manual verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 is a flowchart of a real-time cane tail counting method based on multi-modal dynamic guidance and topology according to a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0073] Please refer to Figure 1 , a real-time cane tail counting method based on multi-modal dynamic guidance and topology according to a preferred embodiment of the present invention includes the following steps:

[0074] S1. The cane tails on the conveyor belt are photographed by a camera to obtain the original images of the cane tails.

[0075] In step S1, the original image includes an RGB image size, a text prompt, and physical parameters. The RGB image size is H×W×3, where H is the height of the image; W is the width of the image; the text prompt is a scene description, and the text prompt is used to provide semantic prior information, such as a high-density sugarcane tail at the right 1 / 3 of the conveyor belt; the physical parameters include the notch size, the conveyor belt speed, and the maximum stacking density, which are used to constrain the physical rationality of the model.

[0076] S2. Perform multi-modal encoding on the original image to obtain a preprocessed image with text encoding and coordinate encoding.

[0077] In step S2, the original image is normalized before multi-modal encoding. The multi-modal encoding includes text encoding and coordinate encoding. The text encoding inputs the scene description according to the position of the original image, and converts the scene description into a 512-dimensional semantic vector through the CLIP text encoder; the coordinate encoding converts the pixel coordinates of the original image into normalized relative coordinates according to the installation position of the camera, and maps the relative coordinates to 512-dimensional position encoding through a fully connected layer.

[0078] In step S2, the original image is normalized according to the mean and standard deviation of the dataset:

[0079]

[0080] where, is the original image, and its pixel value range is [0, 225]; is the mean of all images in the RGB three channels in the dataset; is the standard deviation of all images in the RGB three channels in the dataset; I norn is the normalized original image; by normalizing the original image, the illumination difference can be eliminated and the model convergence can be accelerated.

[0081] The encoding method of the text encoding is:

[0082]

[0083] where, Prompt is a natural language prompt, such as a high-density sugarcane tail at the right 1 / 3 of the conveyor belt; CLIP text is a pre-trained CLIP text encoder, which is a Transformer architecture; t is a 512-dimensional semantic vector, encoding the semantic information of the text;

[0084] The encoding method of the coordinate encoding is:

[0085]

[0086] Among them, p(x, y) is the normalized coordinate mapping; (x, y) is the input camera coordinates; H is the height of the RGB image of the original image; W is the width of the RGB image of the original image.

[0087] S3. Extract multi-scale features of the preprocessed image to obtain image features, and splice the image features with this encoding and coordinate encoding in the channel dimension to obtain a predicted density map.

[0088] In step S3, the preprocessed image is input into the ResNet-50 network to extract image features:

[0089]

[0090] Among them, f ing is the image feature mapping; is the output resolution, and its number of channels is 2048;

[0091] The residual structure of ResNet-50 in this embodiment can effectively alleviate the problem of gradient disappearance in deep networks and is suitable for processing high-resolution surveillance images. The shallow layers (stage1-2) are used to capture local details of the sugarcane tail, such as color and texture; the deep layers (stage4-5) are used to identify global distribution patterns, such as dense areas and occlusion relationships. In this embodiment, the output resolution is optimized, which can retain key spatial information through downsampling while reducing the computational amount; the large receptive field models the contrast relationship between the sugarcane tail and the background (metal notch, conveyor belt); the shallow layer details and deep layer semantics complement each other, enhancing the detection robustness of the occlusion area.

[0092] The text encoding passes through a text encoder to obtain text features:

[0093]

[0094] Among them, f text is the text feature vector obtained after encoding the natural language prompt words, and f text is used to represent the distribution position of the text in the semantic space; the contrastive pre-training of CLIP aligns the text and image features in the shared space. For example, for the text, the features of the high-density sugarcane tail are close to the visual features of the dense area in the image. In this embodiment, the sensitivity of the sugarcane tail color feature (yellow) is enhanced, reducing misjudgment of metal reflection, and at the same time visualizing the text-image similarity to verify the effectiveness of the semantic prompt.

[0095] The coordinate encoding generates coordinate features through coordinate projection:

[0096]

[0097] Among them, is the learnable weight matrix; is a learnable bias vector; f coord is a 512-dimensional coordinate embedding vector that represents the spatial information of a certain position in the image.

[0098] This embodiment distinguishes the distribution differences between the right 1 / 3 of the conveyor belt and the center of the slot, and maintains spatial perception consistency when the position of the camera changes.

[0099] The image features, text features, and coordinate features are concatenated on the channel to obtain a predicted density map:

[0100]

[0101] where f fused is the feature map after multimodal fusion. In this embodiment, the image features (2048-dimensional) are concatenated with the text and coordinate features (512-dimensional) on the channel, and f fused combines the features from the image encoder, text encoder, and coordinate projection layer together so that the subsequent network can further learn cross-modal interaction information.

[0102] The multimodal feature fusion of this embodiment has modal complementarity. Image features: sugarcane tail shape, texture; text features: semantic constraints, such as high density; coordinate features: spatial priors, such as the right region. Moreover, when the illumination changes, the multimodal features cooperate to correct the deviation.

[0103] S4. Calculate the persistent homology features after dividing the predicted density map into blocks to screen topological features; fuse the divided topological features to generate a spatial attention weight map.

[0104] In step S4, the predicted density map is divided into 8×8 blocks:

[0105]

[0106] where is the predicted density map; K is the number of blocks; by dividing the density map into 8×8 blocks, the computational complexity is reduced, and at the same time, the connected component features of different regions can be recognized, such as the dense packing area vs. the sparse area, and the complexity is reduced from O(N 3 ) to O((N / K) 3 ·K 2 ) to adapt to high-resolution inputs. is a single-channel density map obtained by network prediction or annotation, with a size of H×W. Each pixel value can be regarded as "the density of the sugarcane tail at this position" or "a certain intensity value"; after dividing it into smaller sub-blocks, the topological structure (number of connected components, duration, etc.) of these sub-blocks is calculated to reduce the amount of computation and allow the model to more finely focus on local distribution differences.

[0107] Perform continuous homology feature calculation on each block, and filter out connected components with long life cycles and suppress short-life cycle noises:

[0108] PD(D k )={(b i ,d i )∣Connected components are born at threshold b i and die at d i}. Formula (9)

[0109] Where PD(D k ) is the persistence homology diagram of sub-block D k ; (b i ,d i ) are the birth and death of the connected component. Each point (b i ,d i ) corresponds to a Betti-0 connected component that is born at threshold b i and dies at threshold d i , where Betti-0 is the number of connected components. Since the distribution of the sugarcane tail is mainly in connected regions (clump-like), annular or hollow structures rarely appear. In this embodiment, based on the Betti-0 feature selection, only the connected components are calculated, ignoring the rings or holes. Ignoring Betti-1 / Betti-2 can reduce computational redundancy and can directly reflect the physical distribution characteristics of the sugarcane tail.

[0110] PD(D k ) usually sets the pixel values of sub-block D k from low to high as thresholds in the calculation, observes how many independent connected components appear in which threshold interval, and when these connected components merge or disappear, the set of (b i ,d i ) can be obtained.

[0111] By performing a weighted sum of the life cycles (d k ,b i ) of all connected components of sub-block D i , a vector v k of the topological features of sub-block D k is obtained, and then there is:

[0112]

[0113]

[0114] Where v k is the topological feature vector; (d i -b i ) is the life cycle of the sub-block connected component; d PH is the dimension of the persistence homology vector; φ(bi ,d i ) is the Gaussian kernel weighted function; μ is the mean life cycle; σ is the standard deviation, which is used to control the smoothness of the weighted function;

[0115] When (d i -b i ) ≈ μ and φ(b i ,d i ) ≈ 1, it is a connected component with a long life cycle, representing a stable sugarcane tail mass; when (d i -b i ) has a large gap with μ and φ ≈ 0 or φ is very small, it is to suppress short life cycle noise, representing metal reflection.

[0116] By multiplying the life cycle (d i -b i ) of each connected component by a weight that varies according to its difference from μ, in order to highlight "components with a life cycle close to μ" or "components far from μ"; μ is the mean life cycle, which can be estimated from historical data or statistical priors, and μ is the average life cycle size of the main sugarcane tail masses in normal scenarios; σ is the standard deviation, which is used to control the smoothness of the weighted function. If σ is large, φ(b i ,d i ) has a more lenient tolerance for the life cycle; if σ is small, only components with a life cycle very close to μ will be given a high weight.

[0117] In this embodiment, through persistent homology feature extraction, noise can be suppressed, and isolated false high-density points (such as metal reflection) will be identified as short-lived connected components (d i -b i small), and its influence is suppressed by the low weight of the Gaussian kernel (φ ≈ 0). It can perform distribution rationality constraint, and long-lived connected components (d i -b i large), corresponding to stable sugarcane tail masses, ensuring that the density map conforms to the true stacking pattern, such as in clusters rather than scattered points.

[0118] In step S4, the screened topological features are calculated for attention weights, and the spatial attention weight map is obtained by element-wise multiplication:

[0119]

[0120] Among them, is the final attention map; s(·) is the Sigmoid function, which is used to limit the attention value at each pixel position to the interval [0,1]; is the learnable weight tensor; is the learnable bias term, and ba The same dimension as α; is the output spatial attention weight map; is the feature map obtained from multi-modal fusion in the previous stage, and C is the number of channels.

[0121] is to sequentially splice the topological feature vectors of all sub-blocks to form a vector of size K 2 ×d PH )-dim; is the persistent homology feature vector of the k-th sub-block (a total of K 2 sub-blocks); after splicing, a vector of length K 2 ·d PH is obtained, which contains the topological information of all sub-blocks; Map the topological features of all sub-blocks to the image coordinate system (H×W) to learn "which pixels are associated with which topological features (which block, which life cycle) and what weights are required; is to add a learnable "translation" to the attention distribution, allowing the model to more flexibly adjust the overall attention level, such as globally high or low, etc.; is used to play a role in pixel-level screening when fusing with features such as image-text in the follow-up; is the feature map obtained from multi-modal fusion in the previous stage, with the number of channels being C, which may be 2048 + 512 = 2560, etc.; ⊙ is element-wise multiplication, and α is broadcast to C channels, that is, the (H×W) features of each channel are weighted with the same attention coefficient; is the weighted output feature map, indicating that the model pays higher attention to the key regions in the density map, while suppressing the weights of the noise or unimportant regions, making their impact on subsequent predictions smaller.

[0122] Region Adaptive Focus: Long-life cycle components (high-density sugarcane tail clusters) obtain high weights through Gaussian kernel weighting (φ≈1), and the model gives priority to focusing on significant targets. The weights of short-life cycle components (noise or isolated points) are suppressed (φ≈0) to reduce false detections.

[0123] Cross-modal Collaboration: Text prompts such as "dense area" and PH features jointly guide the attention distribution: the text encoder outputs semantic features, such as dense, and jointly optimizes the weight allocation of W k with the PH feature vector v a . For example, when the text prompt is "high density on the right", the weights of the PH features of the right sub-block are automatically enhanced.

[0124] In this embodiment, the model is enabled with dynamic perception ability by calculating attention weights. When the illumination changes, the PH features provide anti-interference ability through topological stability (long-lived components) and complement the image features to correct the prediction deviation. Moreover, the interpretability is improved, and the visualization of attention weights can verify the synergistic effect between PH features and text prompts, such as the spatial consistency between high-weight regions and dense cane tip prompts.

[0125] S5. Gradually upsample the attention weight map to the original resolution through transposed convolution to generate an initial density map, impose a maximum packing density constraint on the density value of each pixel in the initial density map to obtain a constrained density map, and filter the constrained density map to obtain a filtered density map.

[0126] In step S5, the initial density map obtained after gradually upsampling the attention weight map to the original resolution:

[0127]

[0128] where is the initial density map; ConvTranspose is the transposed convolution operation; is the spatial attention weight map; kernel_size = 4 is the convolution kernel size of 4×4; stride = 2 is the stride of 2; padding = 1 is the padding in the transposed convolution process. Appropriate padding is performed during the transposed convolution to ensure that the output size is consistent with the target (i.e., restored to the resolution H×W of the original image).

[0129] Through transposed convolution, a feature map with a smaller resolution (such as ) can be restored to the original image size (H×W) to obtain pixel-level density prediction. Compared with bilinear interpolation or nearest-neighbor interpolation, transposed convolution contains a learnable convolution kernel, which can extract and combine high-level features while upsampling and better fit the target distribution; is an uncropped and unconstrained original predicted density map.

[0130] Impose a maximum packing density constraint on the density value of each pixel in the initial density map:

[0131] D pred = min(ReLU(D raw ), ρ max ) Formula (15)

[0132] where D pred is the constrained density map; ReLU(x) is defined as max(0, x) to make the density value output greater than 0; ρ max is the maximum physical packing density; min(ReLU(D raw ), ρmax ) To perform a clip operation on all pixel values, if it exceeds the upper limit, it is clamped back to ρ max .

[0133] ρ max For cane tail conveying, there is often a priori knowledge that there is an upper limit to the number of cane tails that can be carried by a single pixel. In actual scenarios, it is unlikely to stack indefinitely. Once the prediction exceeds ρ max , it is regarded as oversaturated and needs to be truncated. Ensure that it does not exceed the physical limit, jointly improve the adaptability to the real world, and reduce false predictions.

[0134] min(ReLU(D raw ), ρ max ) can suppress extreme predictions and avoid the network generating unreasonably large values locally.

[0135] Perform Gaussian filtering on the constrained density map along the horizontal movement direction of the conveyor belt to obtain the filtered density map.

[0136] S6. Accumulate the pixels of the filtered density map to output the total number of cane tails in the current frame.

[0137] In this embodiment, the value of each pixel (i, j) is the cane tail density corresponding to that position. The generated D pred can be directly used for visualization, such as a heat map, or further integrated to calculate the total number.

[0138] S7. Collect multiple industrial site images, mark the center points of the cane tails in the site images, and generate density map labels through a Gaussian kernel; associate the scene description text and the associated camera coordinate parameters with the site images to construct a multi-modal dataset;

[0139] S8. Construct an initial counting model through steps S2 - S6, and train the model through the multi-modal dataset to obtain an optimized counting model.

[0140] In step S8, the initial counting model is trained through basic training, topology fine-tuning, and joint optimization in sequence,

[0141] Basic training is performed by minimizing the pixel-level error of the density map and constraining the maximum density value;

[0142] Topology fine-tuning is performed by freezing the weights of the image encoder and introducing persistent homology loss to optimize the topology structure of the density map;

[0143] Joint optimization is performed by minimizing the pixel-level error of the density map, constraining the maximum density value, introducing persistent homology loss to optimize the topology structure of the density map, and cross-modal alignment, and unfreezing all parameters. The joint optimization is trained until convergence to obtain an optimized counting model.

[0144] In this embodiment, the loss function includes:

[0145] Total loss function L:

[0146]

[0147] MSE loss L MSE :

[0148]

[0149] where, : The true density map generated by blurring the point annotation through a Gaussian kernel;

[0150] Persistent homology loss L PH :

[0151]

[0152] where, W1: 1-Wasserstein distance, calculating the difference between two persistent homology diagrams; Block calculation: The density map is divided into K×K blocks, and each block is calculated independently;

[0153] Physical constraint loss L phys :

[0154]

[0155] where, the penalty term is a linear penalty for the predicted value exceeding the maximum density

[0156] Cross-modal contrast loss:

[0157] : Image feature vector (output of the CLIP image encoder);

[0158] : Positive / negative example text prompt features;

[0159] τ = 0.07: Temperature coefficient, adjusting the similarity distribution.

[0160] During basic training:

[0161] Enable loss: L MSE +L phys (λ1 = 1.0, λ3 = 0.5);

[0162] Optimizer: AdamW, learning rate 1×10 -4 and weight decay 1×10 -4 ; Objective: Initially fit the density distribution and satisfy the physical constraints.

[0163] During topological fine-tuning:

[0164] Enabled loss: (λ1 = 0.5, λ2 = 0.3, λ3 = 0.2);

[0165] Frozen parameters: The weights of the image encoder (ResNet-50) are fixed;

[0166] Learning rate: Reduced to 5×10^(-5);

[0167] Objective: Optimize the topology of the density map.

[0168] In joint optimization:

[0169] Enable all losses (λ1 = 0.7, λ2 = 0.1, λ3 = 0.1, λ4 = 0.1);

[0170] Unfreeze parameters: All parameters are trainable;

[0171] Learning rate: Further reduced to 1×10 -5 ;

[0172] Objective: Balance local accuracy and global consistency.

[0173] The inference process of this embodiment is as follows:

[0174] Input processing:

[0175] Image normalization: I norm =(I raw -μ) / σ;

[0176] Text encoding t = CLIP text (Prompt).

[0177] Obtain the prediction result:

[0178] D pred = Decoder(Attention(Encoder(I, t)))

[0179] This embodiment uses a high-definition industrial camera such as Hikvision, which is installed directly above the conveyor belt notch, and collects RGB images with a resolution of 1920×1080 at a rate of 30 frames per second. The computing unit uses the NVIDIA Jetson AGX Xavier embedded platform, which integrates image preprocessing, model inference, and post-processing modules.

Claims

1. A real-time counting method for sugarcane tails based on multi-modal dynamic guidance and topology, characterized in that, It includes the following steps: S1. Shoot the cane tails on the conveyor belt through a camera to obtain the original images of the cane tails; S2. Perform multimodal encoding on the original images to obtain preprocessed images with text encoding and coordinate encoding; S3. Extract multi-scale features of the preprocessed images to obtain image features, and splice the image features with the text encoding and the coordinate encoding in the channel dimension to obtain a predicted density map; S4. Calculate the persistent homology features after dividing the predicted density map into blocks to screen topological features; Fuse the divided topological features to generate a spatial attention weight map: S5. Gradually upsample the attention weight map to the original resolution through transposed convolution to generate an initial density map, apply a maximum stacking density constraint to the density value of each pixel of the initial density map to obtain a constrained density map, and filter the constrained density map to obtain a filtered density map; S6. Perform pixel-level accumulation on the filtered density map to output the total number of cane tails in the current frame; S7. Collect multiple industrial site images, mark the center points of the cane tails in the site images, and generate density map labels through a Gaussian kernel; associate scene description texts and associated camera coordinate parameters with the site images to construct a multimodal dataset; S8. Construct an initial counting model through steps S2 - S6, and train the model through the multimodal dataset to obtain an optimized counting model.

2. The real-time counting method of sugarcane tails based on multi-modal dynamic guidance and topology according to claim 1, characterized in that: In step S1, the original images include an RGB image size, text prompt words, and physical parameters. The RGB image size is H×W×3, where H is the height of the image; W is the width of the image; The text prompt words are scene descriptions and are used to provide semantic prior information; the physical parameters include notch size, conveyor belt speed, and maximum stacking density.

3. A real-time sugarcane tail counting method based on multi-modal dynamic guidance and topology according to claim 1, characterized in that: In step S2, the original images are normalized before multimodal encoding. The multimodal encoding includes text encoding and coordinate encoding. The text encoding inputs the scene description according to the position of the original image, and converts the scene description into a 512-dimensional semantic vector through a CLIP text encoder; the coordinate encoding converts the pixel coordinates of the original image into normalized relative coordinates according to the installation position of the camera, and maps the relative coordinates to a 512-dimensional position encoding through a fully connected layer.

4. A real-time sugarcane tail counting method based on multi-modal dynamic guidance and topology according to claim 3, characterized in that: In step S2, the original images are normalized according to the mean and standard deviation of the dataset: Among them, is the original image, and its pixel value range is [0, 225]; is the mean value of all images in the dataset in the three RGB channels; is the standard deviation of all images in the dataset in the three RGB channels; I norn is the original image after normalization; The encoding method of the text encoding is: Among them, Prompt is a natural language prompt; CLIP text is a pre-trained CLIP text encoder; t is a 512-dimensional semantic vector encoding the semantic information of the text; The encoding method of the coordinate encoding is: where p(x, y) is the normalized coordinate mapping; (x, y) is the input camera coordinate; H is the height of the RGB image of the original image; W is the width of the RGB image of the original image.

5. A real-time sugarcane tail counting method based on multi-modal dynamic guidance and topology according to claim 4, characterized in that: In step S3, input the preprocessed images into a ResNet-50 network to extract image features: Among them, f ing is the image feature mapping; is the output resolution, and its number of channels is 2048; The text encoding passes through a text encoder to obtain text features: where f text is the text feature vector obtained after encoding the natural language prompt; The coordinate encoding generates coordinate features through coordinate projection: Among them, is a learnable weight matrix; is a learnable bias vector; f coord is an output coordinate embedding vector of 512 dimensions, which represents the spatial information of a certain position in the image; Splice the image features with the text features and the coordinate features in the channel to obtain a predicted density map: Among them, f fused is the feature map after multi-modal fusion.

6. A real-time counting method for sugarcane tails based on multi-modal dynamic guidance and topology according to claim 1, characterized in that: In step S4, the predicted density map is divided into 8×8 blocks: Among them, is the predicted density map; K is the number of blocks; Persistent homology features are calculated for each block, and connected components with long life cycles are screened and short life cycle noises are suppressed: PD(D k ) = {(b i , d i ) | The connected component is born at threshold b i and dies at d i}. Equation (9) Among them, PD(D k ) is the persistence homology diagram for sub-block D k ; (b i , d i ) are the birth and death of connected components. Each point (b i , d i ) corresponds to a Betti-0 connected component that is born at threshold b i and dies at threshold d i , where Betti-0 is the number of connected components; Through sub-block D k The life cycles (d i , b i ) of all connected components of are weighted and summed to obtain the vector v k of the topological features of sub-block D k , then we have: where, v k is the topological feature vector; (d i - b i ) is the life cycle of the sub-block connected component; d PH is the dimension of the persistent homology vector; φ(b i , d i ) is the Gaussian kernel weighted function; μ is the life cycle mean; σ is the standard deviation, which is used to control the smoothness of the weighted function; When (d i -b i ) ≈ μ and φ(b i , d i ) ≈ 1, it is a connected component with a long life cycle, representing a stable sugarcane tail mass; when (d i -b i ) has a large gap from μ and φ ≈ 0 or φ is very small, it is to suppress short-life cycle noise, representing metal reflection.

7. A real-time counting method for sugarcane tails based on multi-modal dynamic guidance and topology according to claim 6, characterized in that: In step S4, the topological features obtained by screening are used to calculate attention weights, and a spatial attention weight map is obtained by element-wise multiplication for feature weighting: f attn = f fused ⊙α Formula (13) Among them, is the final attention map; s(·) is the Sigmoid function, which is used to limit the attention value at each pixel position to the interval [0, 1]; is a learnable weight tensor; is a learnable bias term, and b a has the same dimension as α; is the output spatial attention weight map; is the feature map obtained by multi-modal fusion in the previous stage, and C is the number of channels.

8. A real-time sugarcane tail counting method based on multimodal dynamic guidance and topology according to claim 7, characterized in that: In step S5, an initial density map is obtained by gradually upsampling the attention weight map to the original resolution: Among them, is the initial density map; ConvTranspose is the transposed convolution operation; is the spatial attention weight map; kernel_size = 4 is the convolution kernel size of 4×4; stride = 2 is the stride of 2; padding = 1 is the padding in the transposed convolution process; A maximum stacking density constraint is imposed on the density value of each pixel of the initial density map: D pred = min(ReLU(D raw ), ρ max ) Equation (15) Among them, D pred is the constraint density map; ReLU(x) is defined as max(0, x) to make the density value output greater than 0; ρ max is the maximum physical packing density; min(ReLU(D raw ), ρ max ) performs a clip operation on all pixel values, and if it exceeds the upper limit, it is pushed back to ρ max ; Gaussian filtering is performed on the constrained density map along the horizontal movement direction of the conveyor belt to obtain a filtered density map.

9. A real-time sugarcane tail counting method based on multi-modal dynamic guidance and topology according to claim 1, characterized in that: In step S8, the initial counting model is trained by basic training, topological fine-tuning, and joint optimization in sequence, The basic training is performed by minimizing the pixel-level error of the density map and constraining the maximum density value; The topological fine-tuning is performed by freezing the weights of the image encoder and introducing a persistent homology loss to optimize the topological structure of the density map; The joint optimization is performed by minimizing the pixel-level error of the density map, constraining the maximum density value, introducing a persistent homology loss to optimize the topological structure of the density map, and cross-modal alignment, and all parameters are unfrozen. The joint optimization is trained until convergence to obtain an optimized counting model.

Citation Information

Patent Citations

  • Image recognition model training method, image recognition method and device

    CN112183559A

  • Robust multi-modal image segmentation method and system based on instance perception query

    CN119672342A