Colloidal gold detection method and device based on ViT-S16 model
By pruning and compression optimization of the ViT-S16 model, combined with RGB images and colloidal gold reagents, the problems of limited detection accuracy, high equipment cost and large calculation overhead in the existing colloidal gold detection methods are solved, and efficient, economical and reliable pesticide residue detection is achieved.
Patent Information
- Application Number
- CN202510226235.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing colloidal gold detection methods have problems such as limited detection accuracy, high equipment cost and large calculation overhead, making it difficult to achieve efficient, economical and reliable detection in agricultural product pesticide residue detection.
The colloidal gold detection method based on the ViT-S16 model is adopted to optimize the model through pruning and compression, reduce model parameters and retain important features, and combine ordinary RGB images and colloidal gold reagents to achieve automated detection of pesticide residues.
It significantly improves detection accuracy and reliability, reduces equipment cost and calculation complexity, and makes the system highly economical and feasible in agricultural production and is suitable for a wide range of agricultural scenarios.
Smart Images

Figure CN120064639A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of colloidal gold detection, and specifically to a colloidal gold detection method and device based on the ViT-S16 model. Background Art
[0002] Currently, in the detection of pesticide residues in agricultural products, the colloidal gold immunochromatography method is often used for detection. Although this method has the advantages of rapidity and convenience, there are still the following problems:
[0003] (1) Limited detection accuracy: Traditional detection methods usually rely on manual visual judgment, which is greatly affected by human factors. Especially when the concentration of pesticide residues is low, it is difficult to accurately identify.
[0004] (2) High equipment cost: In order to improve the detection accuracy, it is usually necessary to equip high-cost hyperspectral imaging equipment or professional detection instruments, which limits the popularity and economy of large-scale agricultural product detection.
[0005] High computational overhead: Although some deep learning-based detection methods can improve the detection accuracy, due to the large number of model parameters and high computational cost, it is difficult to effectively deploy them in practical applications, especially for resource-constrained agricultural scenarios.
[0006] In view of the above technical defects, a solution is proposed now. Summary of the Invention
[0007] The object of the present invention is to solve the above-mentioned problems and propose a colloidal gold detection method and device based on the ViT-S16 model; improve the detection accuracy: by pruning and compressing and optimizing the ViT-S16 model, while reducing the number of model parameters, important features are retained, enabling the model to accurately distinguish the subtle differences in colloidal gold detection, and improving the accuracy and reliability of pesticide residue detection.
[0008] Reduce the equipment cost: By using ordinary RGB images combined with colloidal gold reagents and realizing the automatic detection of pesticide residues through a deep learning model, expensive professional detection equipment is avoided, making the system highly economical and feasible in agricultural production.
[0009] Model lightweight and deployment convenience: By pruning and compression techniques, the computational complexity of the model is significantly reduced, and the dependence on hardware is reduced, enabling the improved ViT-S16 model to be deployed on common low-power devices, such as embedded devices and mobile terminals, and applicable to a wide range of agricultural scenarios.
[0010] The object of the present invention can be achieved by the following technical solutions:
[0011] A colloidal gold detection method based on the ViT-S16 model. The specific steps of the colloidal gold detection method are as follows:
[0012] Image acquisition: Place the agricultural product to be tested under the image acquisition module and start the device to obtain the colloidal gold reaction image;
[0013] Image processing: The processing unit preprocesses the collected image, such as denoising and enhancing contrast, to ensure the accuracy of analysis;
[0014] Feature extraction and analysis: Use the ViT-S16 deep learning model to extract image features for classification and judgment;
[0015] Result display: The detection result is displayed through the interaction module, and the user can view the detected degree of pesticide residue.
[0016] As a preferred embodiment of the present invention, the data preprocessing process is as follows:
[0017] Collect the image dataset of colloidal gold detection and preprocess the images standardly; including adjusting the image size, normalizing, and data augmentation to improve the generalization ability of the model.
[0018] As a preferred embodiment of the present invention, the innovative pruning strategy process is as follows:
[0019] First, by calculating the importance of the attention weight matrix of each layer, combining the global attention information and local convolutional features, prune the redundant Transformer heads and weights;
[0020] In addition, during sparse regularization, an adaptive technique is adopted, and the pruning ratio r of each group of weights i Aggregates to the entire network structure, and the importance score calculated during the training process is used to dynamically adjust the pruning rate, and the nodes with smaller weights are grouped for pruning; the number of network parameters after pruning can be measured by the following formula:
[0021]
[0022] Where: P pruned Is the number of model parameters after pruning; P original Is the pruning ratio of the number of parameters of the original model; r is the pruning ratio, 0 ≤ r ≤ 1;
[0023] The importance score of each attention head is calculated using a function based on the absolute value of the weight:
[0024]
[0025] Where: A ij Is the activation value corresponding to the jth weight in the ith attention head; Si is the importance score for the i-th attention head; W ij is the j-th weight parameter in the i-th attention head; N is the total number of weight parameters in the attention head.
[0026] As a preferred embodiment of the present invention - model compression: joint quantization and low-rank decomposition. The compression process adopts a method of joint quantization and low-rank decomposition; first, the model parameters are reduced from 32-bit floating-point numbers to 8 bits or lower through quantization technology to reduce storage space and computational overhead; at the same time, the attention matrix is subjected to low-rank decomposition in combination with SVD; the specific formula is as follows:
[0027]
[0028] where q is the quantized integer value; f is the original floating-point value; fmin is the minimum value of the quantization range; in addition, Δ is the quantization step size, defined as:
[0029]
[0030] where f max is the maximum value of the quantization range; b is the number of quantization bits;
[0031] There is also low-rank decomposition: using singular value decomposition to perform low-rank approximation on the attention matrix, decomposing the high-dimensional matrix into the product of low-rank matrices to reduce computational complexity:
[0032]
[0033] where: A is the original attention matrix; U k , Σ k , V k T are the left singular matrix, singular value diagonal matrix, and the transpose of the right singular matrix obtained by SVD decomposition, retaining the first k singular values.
[0034] As a preferred embodiment of the present invention, the model training process is as follows:
[0035] Adopt the method of transfer learning to fine-tune the pruned and compressed ViT-S16 model; initialize the model with pre-trained weights and train it on the colloidal gold detection dataset;
[0036] Use a knowledge distillation strategy, taking the original unpruned model as the teacher model, and giving different weights to the outputs of different regions according to the task requirements (such as detecting pesticide residues), so that the student model can learn the feature information of this region;
[0037] The loss function based on region-weighted knowledge distillation is defined as:
[0038]
[0039] Among them, L KD is the knowledge distillation loss, which includes the weighted sum of the knowledge distillation losses of each region; a i is the weight of the i-th patch (token); N is the total number of Patches; generally, the Kullback-Leibler divergence is adopted; y student and y teacher are the outputs of the student model and the teacher model respectively, and y true is the true label; the loss weight coefficient α satisfies 0 ≤ α ≤ 1;
[0040] As a preferred embodiment of the present invention, the model evaluation and optimization process is as follows:
[0041] Evaluate the performance of the model on the validation set, including indicators such as accuracy, recall, precision, and F1 value; based on the evaluation results, further adjust the hyperparameters of the model; in order to enhance the robustness of the model, a strategy of randomly deactivating attention heads is adopted, and some attention heads are randomly deactivated during the training process to prevent the model from over-relying on specific features;
[0042] The calculation formulas of the evaluation indicators are as follows:
[0043] Accuracy, Precision, Recall, and F1 value.
[0044] As a preferred embodiment of the present invention, the real-time deployment and system integration process is as follows:
[0045] The present invention can combine with a smartphone to collect images on-site. The user can obtain the detection results in real time through a mobile terminal. The detected picture data is processed by the cloud, and a detailed report and relevant suggestions are provided using the optimized model. At the same time, a friendly user interface is designed, and the user can view the detection results in real time through the interface and perform corresponding operations. In addition, the optimized model is integrated into the colloidal gold detection device; in order to achieve real-time detection, an edge computing device is used for inference to accelerate the detection speed and reduce latency.
[0046] A colloidal gold detection device based on the ViT-S16 model, when the detection device runs, implements the above-mentioned colloidal gold detection method based on the ViT-S16 model.
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] In the present invention, the detection accuracy is high: By pruning and compressing and optimizing the ViT-S16 model, the model reduces redundant parameters while retaining important features, enabling more accurate detection of pesticide residues in agricultural product samples, and significantly improving the detection accuracy and reliability.
[0049] Real-time performance and high efficiency: By combining RGB images with colloidal gold color development and performing rapid analysis through a deep learning model, real-time detection of pesticide residues is achieved. Compared with traditional manual detection or analysis methods relying on complex equipment, the present invention can give results within a few seconds, greatly improving the detection efficiency.
[0050] Reducing the detection cost: Existing detection methods often rely on expensive hyperspectral imaging equipment, while the present invention uses a common RGB camera for image acquisition and combines a compressed deep learning model for processing, significantly reducing the detection cost and being more suitable for large-scale agricultural product detection applications.
[0051] Flexible device deployment: The ViT-S16 model optimized by pruning and compression has lightweight characteristics and can be deployed on embedded devices or mobile terminals. This means that the detection device can perform portable detection in the fields, facilitating tea farmers or other agricultural product producers to conduct on-site detection of pesticide residues without relying on laboratory conditions.
[0052] Adapting to the detection of various agricultural products: The method and device of the present invention are not only applicable to the detection of specific agricultural products such as tea, but also can optimize the model according to the characteristics of different agricultural products to adapt to the detection of pesticide residues in various types of agricultural products, having wide adaptability and practicability. Description of the Drawings
[0053] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the drawings.
[0054] Figure 1 It is the hardware architecture diagram of the present invention;
[0055] Figure 2 It is the working flow chart of the present invention;
[0056] Figure 3 It is the data flow schematic diagram of the present invention. Detailed Embodiments
[0057] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] As used herein, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0059] Please refer to Figure 1 - Figure 3 as shown below, the hardware part is as follows:
[0060] An image acquisition module S101, a calculation and processing unit S102, and a user interaction module S103 are provided inside the agricultural residue detection system center;
[0061] Image acquisition module S101: This module includes an RGB camera and a colloidal gold reagent detection device, and is used to acquire images of agricultural product samples to be detected. The acquired images contain both the reaction color development information of the colloidal gold and the surface features of the agricultural product samples;
[0062] Calculation and processing unit S102: This unit consists of an embedded device or a computer, and is built with a ViT-S16 deep learning model optimized by pruning and compression, and is used to process, extract features, and classify the acquired images to determine whether there is pesticide residue in the agricultural products and the degree of residue.
[0063] User interaction module S103: Includes a display screen or a mobile terminal, and is used to display the detection results and the pesticide residue levels to the user. It also has an input function for controlling the detection process and selecting the operation mode;
[0064] Structure and connection:
[0065] The image acquisition module is connected to the calculation and processing unit through USB or wireless communication to transmit the acquired image data in real time;
[0066] The calculation and processing unit displays the detection results in the form of graphics or text through the connection with the user interaction module, so that the user can intuitively understand the detection information;
[0067] A colloidal gold detection method based on the ViT-S16 model. The specific steps of the colloidal gold detection method are as follows:
[0068] S201: Image acquisition: Place the agricultural product to be tested under the image acquisition module and start the device to obtain the colloidal gold reaction image; This step corresponds to S301, and the acquisition of the agricultural product image serves as the starting point of the data stream.
[0069] S202: Image processing: The processing unit preprocesses the acquired image, such as denoising and enhancing contrast, to ensure the accuracy of the analysis; This step corresponds to S302, and the data stream is the image data transmission step.
[0070] S203: Feature extraction and analysis: Use the ViT-S16 deep learning model to extract image features and perform classification judgment; This step corresponds to S303 - S304, and the data stream is feature extraction, classification, and deep learning analysis.
[0071] S204: Result display: The detection result is displayed through the interaction module, and the user can view the detected pesticide residue level; This step corresponds to S305 - S306, and the data stream is the detection result acquisition and user interaction step.
[0072] Specifically:
[0073] Data preprocessing:
[0074] Collect the image dataset for colloidal gold detection and preprocess the images; This includes adjusting the image size (such as 224x224), normalizing, and data augmentation (such as rotation, flipping, cropping, etc.) to improve the generalization ability of the model.
[0075] Innovative pruning strategy:
[0076] During the pruning process, a pruning method of multi-scale importance evaluation is adopted.
[0077] First, by calculating the importance of the attention weight matrix of each layer, combining global attention information and local convolutional features, prune the redundant Transformer heads and weights.
[0078] In addition, a technique based on group sparse regularization is used to prune the nodes with smaller weights in groups; This method not only improves the pruning efficiency but also retains the key representation ability of the model.
[0079] The innovation of the pruning process lies in the adoption of an adaptive technique, and the pruning ratio r of each group of weights iAggregate to the entire network structure, calculate the importance scores during the training process, dynamically adjust the pruning rate, and group and prune the nodes with smaller weights; this not only retains the global attention information but also removes unnecessary redundant features, greatly reducing the model complexity; the number of network parameters after pruning can be measured by the following formula:
[0080]
[0081] Where: Ppruned is the number of model parameters after pruning; Poriginal is the number of parameters of the original model; the pruning ratio; r is the pruning ratio, 0 ≤ r ≤ 1.
[0082] The importance score of each attention head is calculated using a function based on the absolute value of the weights:
[0083]
[0084] Where: A ij is the activation value corresponding to the j-th weight in the i-th attention head; S i is the importance score of the i-th attention head; W ij is the j-th weight parameter in the i-th attention head; N is the total number of weight parameters in the attention head.
[0085] The pruning objective is to remove the attention heads whose importance scores Si are lower than the threshold θ to reduce the model complexity;
[0086] Model compression: Joint quantization and low-rank factorization
[0087] A method of joint quantization and low-rank factorization is adopted in the compression process;
[0088] First, the model parameters are reduced from 32-bit floating-point numbers to 8 bits or lower through quantization technology to reduce storage space and computational overhead;
[0089] At the same time, combined with SVD (Singular Value Decomposition), the attention matrix is decomposed into low-rank matrices, and the high-dimensional matrix is represented as the product of several low-rank matrices, thereby reducing the computational complexity of the model;
[0090] This method not only reduces the model parameters but also can greatly improve the inference speed while ensuring the accuracy;
[0091] The specific formula is as follows:
[0092]
[0093] Where q is the quantized integer value; f is the original floating-point value; fmin is the minimum value of the quantization range; in addition, Δ is the quantization step, defined as:
[0094]
[0095] f max The maximum value of the quantization range; b is the number of quantization bits;
[0096] Low-rank decomposition: Use singular value decomposition (SVD) to perform low-rank approximation on the attention matrix, decompose the high-dimensional matrix into the product of low-rank matrices, and reduce the computational complexity:
[0097]
[0098] Where: A is the original attention matrix; U k , Σ k , V k T Are the left singular matrix, singular value diagonal matrix, and transpose of the right singular matrix obtained by SVD decomposition, and the first k singular values are retained;
[0099] Model training:
[0100] Adopt the method of transfer learning to fine-tune the pruned and compressed ViT-S16 model. Initialize the model with pre-trained weights and train it on the colloidal gold detection dataset;
[0101] To overcome the problem of accuracy degradation caused by model pruning and compression, a knowledge distillation strategy is used. The original unpruned model is used as the teacher model, and the learning of the student model (the pruned and compressed model) is guided by soft labels;
[0102] The loss function of knowledge distillation is defined as:
[0103]
[0104] Where KL represents the Kullback-Leibler divergence, and L KD Is the knowledge distillation loss, which includes the weighted sum of the knowledge distillation losses of each region.
[0105] Model evaluation and optimization:
[0106] Evaluate the performance of the model on the validation set, including indicators such as accuracy, recall, precision, and F1 value; based on the evaluation results, further adjust the hyperparameters of the model; to enhance the robustness of the model, a strategy of randomly deactivating attention heads is adopted, and some attention heads are randomly deactivated during the training process to prevent the model from over-relying on specific features;
[0107] The calculation formulas of the evaluation indicators are as follows:
[0108] Accuracy, Precision, Recall, and F1 Score:
[0109] Real-time Deployment and System Integration
[0110] Integrate the optimized model into the colloidal gold detection device. To achieve real-time detection, an edge computing device (such as Jetson Nano) is used for inference to accelerate the detection speed and reduce latency. At the same time, a user-friendly interface is designed, allowing users to view the detection results in real-time and perform corresponding operations through the interface.
[0111] A colloidal gold detection device based on the ViT-S16 model. When the detection device is running, the above-mentioned colloidal gold detection method based on the ViT-S16 model is implemented.
[0112] The above formulas are obtained by collecting a large amount of data for software simulation and selecting a formula close to the true value. The coefficients in the formula are set by those skilled in the art according to the actual situation.
[0113] When the present invention is in use, first, the RGB image of the agricultural product sample is collected by the image acquisition module. The image contains the color change after the reaction of the colloidal gold reagent. Then, the image data is transmitted to the calculation and processing unit, and the ViT-S16 model improved by pruning and compression processes the image, extracts the relevant features of pesticide residues, and classifies them. Finally, the detection results are displayed through the user interaction module, including whether there are pesticide residues and their specific levels, as well as relevant prevention and control suggestions.
[0114] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only the specific embodiments. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principle and practical application of the present invention, so that those skilled in the art in the relevant technical field can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A colloidal gold detection method based on the ViT-S16 model, characterized in that: The specific steps of colloidal gold detection method are as follows: Image acquisition: Place the agricultural product to be tested under the image acquisition module and start the device to obtain the colloidal gold reaction image; Image processing: The processing unit pre-processes the collected images, such as denoising and contrast enhancement, to ensure the accuracy of the analysis; Feature extraction and analysis: Use the ViT-S16 deep learning model to extract image features and perform classification judgment; Result display: The test results are displayed through an interactive module, and users can view the detected pesticide residue levels.
2. A colloidal gold detection method based on the ViT-S16 model according to claim 1, characterized in that: The data preprocessing process is as follows: An image dataset for colloidal gold detection was collected, and the images were preprocessed for standardization, including adjusting the image size, normalization, and data enhancement to improve the generalization ability of the model. At the same time, background removal and local feature enhancement techniques were used to strengthen colloidal gold recognition and improve the accuracy of pesticide residue detection.
3. A colloidal gold detection method based on the ViT-S16 model according to claim 1, characterized in that: The innovative pruning strategy process is as follows: First, by calculating the importance of the attention weight matrix of each layer, the redundant Transformer heads and weights are pruned by combining the global attention information and local convolution features; In addition, during sparse regularization, adaptive technology is used to prune each set of weights at a ratio of r i Aggregated to the entire network structure, the importance score calculated during training, dynamically adjust the pruning rate, and prune nodes with smaller weights in groups; the number of network parameters after pruning can be measured by the following formula: Where: P pruned is the number of model parameters after pruning; P original is the parameter pruning ratio of the original model; r is the i-th pruning ratio; The importance score of each attention head is calculated using a function based on the absolute value of the weight: Among them: A ij is the activation value corresponding to the jth weight in the i-th attention head; S i Score the importance of the i-th attention head; W ij is the jth weight parameter in the i-th attention head; N is the total number of weight parameters in the attention head.
4. A colloidal gold detection method based on the ViT-S16 model according to claim 1, characterized in that: Model compression: Joint quantization and low-rank decomposition,The compression process adopts a joint quantization and low-rank decomposition method; First, the model parameters are reduced from 32-bit floating point numbers to 8 bits through quantization technology to reduce storage space and computational overhead; At the same time, SVD is combined to perform low-rank decomposition of the attention matrix; the specific formula is as follows: Where q is the quantized integer value; f is the original floating point value; f min is the minimum value of the quantization range; in addition, Δ is the quantization step size, which is defined as: f max The maximum value of the quantization range; b is the number of quantization bits; Low-rank decomposition: Use singular value decomposition to perform low-rank approximation on the attention matrix, decomposing the high-dimensional matrix into the product of low-rank matrices to reduce the computational complexity: Where: A is the original attention matrix; U k ,Σk,V k T It is the transpose of the left singular matrix, singular value diagonal matrix, and right singular matrix obtained by SVD decomposition, retaining the first k singular values.
5. A colloidal gold detection method based on the ViT-S16 model according to claim 1, characterized in that: The model training process is as follows: The pruned and compressed ViT-S16 model is fine-tuned using transfer learning. The model is initialized using pre-trained weights and trained on the colloidal gold detection dataset. A knowledge distillation strategy is used to take the original unpruned model as the teacher model, and different weights are given to different regions according to task requirements, so that the student model can learn the feature information of the region. The loss function based on region-weighted knowledge distillation is defined as: Among them, L KD is the knowledge distillation loss, which contains the weighted sum of the knowledge distillation loss in each region; a i is the weight of the ith patch (token); N is the total number of patches; Kullback-Leibler divergence is usually used; y student and teacher are the outputs of the student model and the teacher model, y true is the true label; the loss weight coefficient α satisfies 0≤α≤1.
6. A colloidal gold detection method based on the ViT-S16 model according to claim 1, characterized in that: The model evaluation and optimization process is as follows: The performance of the model was evaluated on the validation set, including indicators such as accuracy, recall, precision, and F1 value. Based on the evaluation results, the hyperparameters of the model were further adjusted. In order to enhance the robustness of the model, a random inactivation strategy of attention heads was adopted, which randomly inactivated some attention heads during the training process to prevent the model from over-relying on specific features. The calculation formula of the evaluation index is as follows: Accuracy, precision, recall and F1 value.
7. A colloidal gold detection method based on the ViT-S16 model according to claim 1, characterized in that: The real-time deployment and system integration process is as follows: The optimized model was integrated into the colloidal gold detection device. In order to achieve real-time detection, an edge computing device was used for inference to speed up detection and reduce latency. At the same time, a friendly user interface was designed, through which users can view the detection results in real time and perform corresponding operations.
8. A colloidal gold detection device based on the ViT-S16 model, characterized in that: A colloidal gold detection method based on the ViT-S16 model as described in any one of claims 1 to 7 above is used.