MXenes photo-thermal sterilization intelligent detection system combined with improved YOLOv11
By improving the YOLOv11 network model and RepViT architecture, optimizing computational efficiency and feature representation, and constructing a lightweight deep learning model, the real-time accuracy problem of MXenes photothermal sterilization detection was solved, and efficient detection of E. coli was achieved.
Patent Information
- Application Number
- CN202511253220.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies rely on traditional methods to detect MXenes photothermal sterilization processes, which cannot achieve real-time, accurate dynamic judgment and control, thus limiting their practical application.
By combining the improved YOLOv11 network model with the RepViT architecture, and through a multi-scale feature representation mechanism and a lightweight ViT design, we optimize computational efficiency and feature representation capabilities to build a lightweight deep learning model for E. coli detection.
It achieves high-precision detection of E. coli in complex scenarios, meets the deployment requirements of embedded devices, and improves detection efficiency and accuracy.
Smart Images

Figure BDA0005579823830000071 
Figure HDA0005579823840000011 
Figure HDA0005579823840000021
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of target detection. BACKGROUND
[0002] With the continuous improvement of modern medical and biological safety requirements, traditional sterilization technologies such as chemical disinfection and high-temperature treatment gradually expose many limitations in practical application. Chemical disinfection is often accompanied by harmful substance residues and secondary pollution risks, and high-temperature treatment has high energy consumption and low applicability to heat-sensitive materials, both of which face serious challenges in efficiency, safety and environmental impact. In recent years, MXenes, as a kind of two-dimensional material with excellent light-heat conversion performance, can rapidly heat up under near-infrared light excitation, showing high-efficiency and clean sterilization potential, providing an important direction for the development of new sterilization technology. However, the detection of its sterilization process still widely relies on traditional offline culture methods and manual microscope observation, which not only have slow response and cannot realize real-time feedback, but also are limited by the subjective experience of operators, making it difficult to dynamically and accurately judge and control the sterilization state, thereby greatly limiting the conversion of MXenes photothermal sterilization technology from the laboratory to practical application.
[0003] The use of deep learning technology to assist detection can significantly improve detection accuracy and efficiency. In automated target detection technology, the YOLO (You Only Look Once) series algorithm has become a research hotspot in the field of computer vision due to its architecture characteristics of simultaneously realizing multi-target positioning and classification through single forward propagation, excellent real-time performance and detection accuracy. The new generation of YOLOv11 version has made breakthrough progress in detection accuracy and model robustness, making it particularly suitable for Escherichia coli detection scenarios with strict requirements for efficiency and accuracy. However, in the face of special challenges such as the detection of small targets in complex backgrounds in Escherichia coli images, the existing model still needs to be improved to enhance its adaptability in complex scenarios.
[0004] Therefore, the present application develops a new generation of lightweight and high-precision target detection framework to significantly improve the accurate detection capability of targets in Escherichia coli images while ensuring the inference capability of embedded devices, overcoming the shortcomings of existing technologies. SUMMARY
[0005] The purpose of the present application is to optimize the computational efficiency and feature expression ability, so that the model can effectively detect targets in complex scenarios and meet the deployment requirements of embedded devices, providing a new generation of solution for lightweight Escherichia coli detection. An MXenes photothermal sterilization intelligent detection system combining improved YOLOv11 is proposed.
[0006] The specific process of the MXenes photothermal sterilization intelligent detection system combined with improved YOLOv11 is as follows:
[0007] Step one, obtain the data set; the specific process is as follows:
[0008] The data set is a Deepbacs data set; the Deepbacs data set is divided into a training set, a validation set and a test set in a ratio of 7:2:1. The training set contains 164 images, the validation set contains 47 images, and the test set contains 24 images. This division ratio ensures that the model has enough data for training in the training set, while also having an appropriate amount of data for parameter tuning and preventing overfitting, as well as independent data for final evaluation of model performance.
[0009] Step two, build a YOLOv11n-RepViT network model; the specific process is as follows:
[0010] The YOLOv11n network model includes a backbone network Backbone, a neck network Neck and a detection head Head;
[0011] The YOLOv11n architecture is improved by combining the backbone network Backbone with the RepViT network structure. The RepViT architecture integrates the lightweight ViT (Vision Transformer) design concept with the efficient feature extraction capability of CNN (Convolutional Neural Networks), and through a multi-scale feature representation mechanism, it captures global context information in shallow networks and accurately locates small target details in deep networks, optimizing computational efficiency and feature expression capability.
[0012] The working process of the YOLOv11n-RepViT network model is as follows:
[0013] The image is input into the backbone network Backbone, the neck network Neck and the detection head Head in turn, and the detection head outputs the target class and the two-dimensional label frame;
[0014] Step three, input the training data set into the YOLOv11n-RepViT network model to carry out model training, so as to build a lightweight deep learning model suitable for E. coli detection;
[0015] Step four, use the test data set to verify the performance of the trained YOLOv11n-RepViT network model, and comprehensively evaluate the detection effect of the model through multiple indicators;
[0016] Further, in the present application, in step one, the experimental data set Deepbacs used in the experiment contains a total of 235 images of E. coli. The Deepbacs data set is divided according to the ratio of 7:2:1, wherein 164 images constitute the training sample, 47 images constitute the test sample, and 24 images constitute the test sample.
[0017] Further, in the present application, in step two, the YOLOv11n-RepViT network model is constructed; the specific process is as follows:
[0018] The YOLOv11n network model includes a backbone network Backbone, a neck network Neck, and a detection head Head.
[0019] The YOLOv11n architecture is improved, combining the backbone network Backbone and the RepViT network structure. The RepViT architecture integrates the lightweight ViT design concept and the efficient feature extraction capability of CNN. Through the multi-scale feature representation mechanism, the global context information is captured in the shallow network, and the small target details are accurately positioned in the deep network. The core module adopts the cooperative design of depth separable convolution, GELU (Gaussian Error Linear Unit) activation function, SE (Squeeze-and-Excitation Networks) channel attention, and FFN (Feed-Forward Networks) layer, combined with the structure reparameterization strategy, maintaining the learning of complex patterns in the training stage. In the inference stage, it is compressed into a single-path efficient model. The calculation efficiency and feature expression ability are significantly optimized, so that the model can effectively detect small targets such as E. coli in detection, and also meet the strict deployment requirements of embedded devices.
[0020] In addition, RepViT enhances the network's expression ability through residual connection and SE mechanism in the training process, and ensures the integrity of the data and the stability of the network training. The optimized structure of the model makes it achieve excellent results in small target detection, and compared with traditional CNN models, RepViT can better cope with target detection problems in complex scenes through the combination of multiple convolution operations and attention mechanisms. The specific steps are as follows:
[0021] RepViT is a pure CNN architecture that integrates the efficient design principles of lightweight ViT. By improving the model's expression ability at different resolution levels, it effectively improves the extraction effect of small target features.
[0022] The RepViT stem extraction network is mainly realized by the following three key modules: Stem, Stage and down-sampling. The Stem module is mainly used for preliminary feature extraction. By performing two 3x3 convolution operations on the image, local features in the image can be effectively captured while maintaining computational efficiency. After the first convolution operation, a non-linear activation function GELU is followed to increase the non-linear expression ability of the network, thereby helping to learn more complex features.
[0023] The Stage module, as a key part of the RepViT stem feature extraction network, is composed of multiple RepViTSEBlock and RepViTBlock. First, the RepViTSEBlock module is processed by a 3x3 and 1x1 depth separable convolution and a residual connection, which realizes the multi-scale fusion of the input features. This structure not only captures rich spatial information, but also preserves the integrity of the features through residual connection. Next, the global average pooling operation of the SE channel attention mechanism is used to compress the feature map into a single value, which can capture the global spatial information of the entire feature map. Then, the full connection layer in the SE module is used to adaptively adjust the weights of each channel to enhance the attention to key features. The FFN layer then performs non-linear transformation on the features of the previous layer to enhance the expression effect of the network on deep features. The FFN layer here is composed of two 1x1 convolution operations and a GELU activation function. By introducing non-linear transformation, the model's ability to process complex features is improved, and the FFN layer uses residual connection to add the transformed features to the original features, ensuring data integrity and promoting the stability of deep network training.
[0024] In contrast, the RepViTBlock module omits the SE channel attention mechanism, making it more computationally efficient while still retaining the multi-scale feature fusion function achieved through depth separable convolution and residual connection. The down-sampling module is composed of a 3x3 depth separable convolution, a 1x1 pointwise convolution and an FFN module. The role of this module is to achieve precise spatial down-sampling while ensuring computational efficiency. First, the 3x3 depth separable convolution effectively preserves important spatial information while reducing computational complexity. Then, the 1x1 pointwise convolution adjusts the number of channels of the feature map, which not only enhances the network's ability to represent complex features, but also provides more rich feature information for subsequent layers. Finally, the FFN layer further optimizes the network's handling of complex relationships between features.
[0025] Further, in the present application, in step three, the training data set is input into the YOLOv11n-RepViT network model to carry out model training, so as to build a lightweight deep learning model suitable for E. coli detection; the method is as follows:
[0026] 1. Modify the configuration file: update the configuration file (cfg) of YOLOv11n, and adjust the number of classes in the corresponding yaml file to the actual number of classes in the data set.
[0027] 2. Set the hyperparameters: configure the hyperparameters of the network model, including the size of the input image, the number of training rounds (epochs), the number of samples in each batch (batchsize), the optimizer (optimizer), and the learning rate (learningrate).
[0028] Further, in the present application, in step four, the test data set is used to verify the performance of the trained YOLOv11n-RepViT network model, and the detection effect of the model is comprehensively evaluated through multiple indicators; specifically:
[0029] When the test data set is input into the YOLOv11n-RepViT model to carry out E. coli detection task, the model will analyze and process each test image, output the detection result, including class, position coordinate and confidence, etc. Then, the prediction result of the model is compared with the true annotation of the test set, and multiple performance indicators such as accuracy, recall rate and F1 score are calculated. This systematic test can comprehensively evaluate the overall performance of the model.
[0030] Through in-depth analysis of the test results, the actual application performance of the YOLOv11n-RepViT model can be fully mastered, which provides key technical support for subsequent optimization iteration and actual deployment. This systematic verification method can help objectively evaluate the performance of the model and ensure its reliability and effectiveness in E. coli detection. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 is the flow chart of the MXenes photothermal sterilization intelligent detection system of the present application method combined with improved YOLOv11;
[0032] Figure 2 is the network structure schematic diagram of the MXenes photothermal sterilization intelligent detection system of the present application method combined with improved YOLOv11;
[0033] Figure 3 is the principle schematic diagram of RepViT in the present application method;
[0034] Figure 4is a training indicator result graph on the training set, a is Metrics / Precision (B) of YOLOv11n-RepViT on the training set, Metrics / Precision (B) is the precision indicator; b is Metrics / recall (B) of YOLOv11n-RepViT on the training set, Metrics / recall (B) is the recall indicator; c is Metrics / mAP50 (B) of YOLOv11n-RepViT on the training set, Metrics / mAP50 (B) is the average precision under the condition of IoU (intersection over union) being 0.5; d is Metrics / mAP50-95 (B) of YOLOv11n-RepViT on the training set, Metrics / mAP50-95 (B) is the average value of mAP calculated on multiple IoU thresholds;
[0035] Figure 5 is a performance indicator curve graph on the validation set, a is the performance indicator curve graph of val / box_loss of YOLOv11n-RepViT on the validation set in the method of the application, val / box_loss represents the loss function of the validation set; b is the performance indicator curve graph of val / cls_loss of YOLOv11n-RepViT on the validation set in the method of the application, val / cls_loss represents the classification loss mean of the validation set; c is the performance indicator curve graph of val / dfl_loss of YOLOv11n-RepViT on the validation set in the method of the application, val / dfl_loss represents the loss mean of the validation set;
[0036] Figure 6 is a visualization detection result graph on the test set, a is a visualization detection result graph of YOLOv11n on the test set, b is a visualization detection result graph of YOLOv11n-RepViT on the test set in the method of the application; DETAILED DESCRIPTION
[0037] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application:
[0038] The improved YOLOv11n-RepViT network structure diagram is as shown in Figure 1 Figure 2 The method comprises the following steps:
[0039] Step a, the present application uses the disclosed Deepbacs dataset as the experimental data source, which contains a large number of E. coli images. The Deepbacs dataset is divided into training set, validation set and test set in the ratio of 7:2:1. The training set contains 164 images, the validation set contains 47 images, and the test set contains 24 images. This division ratio ensures that the model has enough data for training in the training set, while also having appropriate data for parameter tuning and preventing overfitting, as well as independent data for final evaluation of model performance.
[0040] Step b, the YOLOv11n network model includes a backbone network Backbone, a neck network Neck and a detection head Head;
[0041] The YOLOv11n architecture is improved by combining the backbone network Backbone with the RepViT network structure. The RepViT architecture integrates the lightweight ViT design concept with the efficient feature extraction capability of CNN. Through the multi-scale feature representation mechanism, it captures global context information in the shallow network and accurately locates small target details in the deep network. The core module adopts the cooperative design of depth separable convolution, GELU activation function, SE channel attention and FFN layer, combined with structure reparameterization strategy, to maintain the learning of complex patterns in the training stage of the multi-branch structure, and compresses it into a single-path efficient model during inference. The calculation efficiency and feature expression ability are significantly optimized, so that the model can effectively detect small targets such as E. coli and meet the strict deployment requirements of embedded devices.
[0042] RepViT enhances the network's expression ability and ensures the integrity of the data and the stability of the network during training through residual connection and SE mechanism. The optimized structure of this model makes it achieve excellent results in small target detection, and compared with traditional CNN models, RepViT can better cope with target detection problems in complex scenes through the combination of various convolution operations and attention mechanisms. The specific steps are as follows:
[0043] RepViT is a pure CNN architecture that integrates the efficient design principles of lightweight ViT, effectively improving the extraction of small target features by improving the model's expression ability at different resolution levels.
[0044] The RepViT backbone extraction network is mainly realized by the following Stem, Stage and down-sampling three key modules. The Stem module is mainly used for preliminary feature extraction, which can effectively capture local features in the image through two 3x3 convolution operations while maintaining computational efficiency. After the first convolution operation, a non-linear activation function GELU is followed to increase the non-linear expression ability of the network, thereby helping to learn more complex features.
[0045] The Stage module, as a key part of the RepViT backbone feature extraction network, is composed of multiple RepViTSEBlock and RepViTBlock modules. First, the RepViTSEBlock module is processed by a 3x3 and 1x1 depth separable convolution and a residual connection, which realizes the multi-scale fusion of the input features. This structure not only captures rich spatial information, but also preserves the integrity of the features through residual connection. Next, the global average pooling operation of the SE channel attention mechanism compresses the feature map into a single value, which can capture the global spatial information of the entire feature map. Then, the full connection layer in the SE module adjusts the weights of each channel adaptively to enhance the attention to key features. The FFN layer then performs nonlinear transformation on the features of the previous layer to enhance the expression effect of the network for deep features. The FFN layer here is composed of two 1x1 convolution operations and a GELU activation function. By introducing nonlinear transformation, the model's ability to process complex features is improved, and the FFN layer uses residual connection to add the transformed features to the original features, ensuring data integrity and promoting the stability of deep network training.
[0046] In contrast, the RepViTBlock module omits the SE channel attention mechanism, making it more computationally efficient while still retaining the multi-scale feature fusion function achieved through depth separable convolution and residual connection. The down-sampling module is composed of a 3x3 depth separable convolution, a 1x1 pointwise convolution, and an FFN module. The role of this module is to achieve precise spatial down-sampling while ensuring computational efficiency. First, the 3x3 depth separable convolution effectively preserves important spatial information while reducing computational complexity. Then, the 1x1 pointwise convolution adjusts the number of channels of the feature map, which not only enhances the network's ability to represent complex features, but also provides more rich feature information for subsequent layers. Finally, the FFN layer further optimizes the network's handling of complex relationships between features.
[0047] Step c, input the training data set into the YOLOv11n-RepViT network model to carry out model training, to build a lightweight deep learning model suitable for E. coli detection; the specific process is:
[0048] Step c1, the method for inputting the training data set into the YOLOv11n-RepViT network model to carry out model training to build a lightweight deep learning model suitable for E. coli detection is:
[0049] Modify the configuration file: update the configuration file (cfg) of YOLOv11n, adjust the number of classes in the corresponding yaml file to the actual number of classes in the data set.
[0050] Set hyperparameters: configure the hyperparameters of the network model, including the size of the input image, the number of training rounds (epochs), the number of samples per batch (batchsize), the optimizer (optimizer), and the learning rate (learningrate) during training.
[0051] Step c2, the evaluation system in the target detection method is used to evaluate the model, including the precision AP, recall, F1-score, accuracy (Precision) and average accuracy mean mAP@0.5 (mean average precision, IoU threshold is 0.5) of each class. The calculation formula is as follows:
[0052]
[0053] TP represents the number of correctly identified positive samples, TN represents the number of correctly identified negative samples, FP is the number of negative samples that are incorrectly identified as positive samples, and FN is the number of positive samples that are incorrectly identified as negative samples. The accuracy and recall can be used to further calculate the F1 score.f TP correctly identified positive samples, f TF correctly identified negative samples, f TN incorrectly identified negative samples. The area enclosed by the curve (P-R curve) composed of the accuracy (Precision) and recall (Recall) and the coordinate axes can be used to calculate the average precision (Average Precision, AP) of each target class. The detection speed of the model is represented by FPS, which means the number of images detected per second. Assuming that it takes x seconds to process one image, the value of FPS is The unit is frame / second.
[0054] The YOLOv11n-RepViT model is trained using the training set, and the training results are as follows Figure 4 and Figure 5The val_box represents the verification set bounding box, and the value becomes smaller and smaller after the training round number reaches 150 rounds and finally tends to be stable; the val_dfl represents the verification set loss mean, and the loss mean tends to be stable after the training round number reaches 200 rounds; the val_cls represents the verification set classification loss mean, and the loss mean basically converges after the training round number reaches 200 rounds; the Precision represents the percentage of the positive class found by the verification set, and the Precision value basically stabilizes when the training round number reaches 200 rounds; the Recall describes how many real positive examples are recalled by the binary classifier from the perspective of real results, and the recall rate gradually converges after the training round number reaches 150 rounds.
[0055] Overall, the model reaches a comprehensive stable state after 200 rounds of training, and the detection accuracy, target recall and loss control all reach the best balance point.
[0056] Step c3, the present application aims to improve the target detection ability of YOLOv11 algorithm, therefore we choose mAP@0.5 as the main evaluation index. In order to verify the improvement effect, we compare YOLOv11 with YOLOv11n-RepViT model of the present application on Deepbacs dataset.
[0057] The experimental results show that YOLOv11n-RepViT has a significant advantage in detection accuracy. Figure 4 The performance difference of the two models in average precision is intuitively presented, and the excellent performance of the present model is clearly shown. This comparison result strongly confirms that YOLOv11n-RepViT model can indeed effectively improve the target detection ability.
[0058] This performance comparison verifies the effectiveness of YOLOv11n-RepViT model in specific tasks, and provides some technical support for further optimization of target detection algorithm. The experimental results show that YOLOv11n-RepViT model brings a certain degree of performance improvement, which may have some inspiration for subsequent related research.
[0059] Step d, the performance of the trained YOLOv11n-RepViT model is verified by using the test dataset.
[0060] Step d1, the performance of the trained YOLOv11n-RepViT model is verified by using the test dataset, and the detection effect of the model is comprehensively evaluated by multiple indexes. The specific process is as follows:
[0061] After inputting the test set into the YOLOv11n-RepViT network model, the model detects E. coli in each image and outputs information such as its class, location, and confidence. Subsequently, we compare the model's prediction results with the true labels of the test set, and calculate core indicators such as accuracy, recall rate, and F1 score. This comprehensive evaluation not only measures the overall performance of the model, but also reveals its strengths and weaknesses in detecting different types of objects. Through detailed analysis of the test results, we can gain a deep understanding of the model's application performance in real-world scenarios, providing technical support for subsequent optimization and deployment. This systematic verification process objectively evaluates the model's actual performance, ensuring its stability and effectiveness.
[0062] Step d2, to verify the effectiveness of the improved model, the present application designs experiments for the improved model and the original model, and compares the performance of the model before and after lightweight. The evaluation indicators are selected as parameter quantity, computational quantity (GFLOPs), inference speed (FPS), and detection accuracy mAP50 (%) for quantitative analysis.
[0063] As shown in Table 1, on the Deepbacs dataset, YOLOv11n-RepViT shows significant efficiency optimization and precision improvement compared to the baseline model YOLOv11n. The parameter quantity is compressed from 2,590,620 to 2,112,672, a decrease of 18.45%, because of the RepViT's reparameterization convolution design. This technology enhances the representation ability through the multi-branch structure in the training stage, and combines into a single-path structure to realize parameter reduction during inference. The computational quantity (GFLOPs) decreases from 6.4 to 5.4, a decrease of 15.63%, and the inference speed (FPS) increases from 76.92 to 79.37, an increase of 3.19%. Moreover, under such significant lightweight improvement, mAP50 increases from 98.21% to 98.53%, an increase of 0.32%, indicating that the cross-stage feature reuse mechanism of RepViT effectively preserves the key discriminative information.
[0064] Table 1 Comparison of computational efficiency of improved YOLOv11n model and baseline model
[0065] Method YOLOv11n YOLOv11n-RepViT Delta (Δ) Parameter amount 2,590,620 2,112,672 -18.45%↓ Computational amount (GFLOPs) 6.4 5.4 -15.63%↓ Inference speed (FPS) 76.92 79.37 +3.19%↑ mAP50 (%) 98.21 98.53 +0.32%↑
[0066] The detection effect of the YOLOv11n-RepViT model proposed by the present application on the test set is as shown in Figure 6 It can be seen that the YOLOv11n-RepViT model can accurately detect E. coli and other targets. This proves the effectiveness and superiority of the present application in E. coli detection.
[0067] Although the present application is illustrated by way of specific embodiments, it is to be understood that these are presented by way of example only and should not be taken as limiting the scope of the application. Therefore, the embodiments can be modified or adapted in various ways without departing from the scope of the application as defined by the appended claims. In addition, the dependent claims and their technical features can be combined in different ways. Also, the specific features of one embodiment can be combined with features of other embodiments.
Claims
1. A smart detection system for photothermal sterilization using MXenes combined with an improved YOLOv11, characterized in that: The specific process of the method is as follows: Step 1: Obtain the dataset; the specific process is as follows: The dataset used is the Deepbacs dataset, which is divided into three parts in a 7:2:1 ratio: training, validation, and test sets. The training set contains 164 images, the validation set contains 47 images, and the test set contains 24 images. This division ensures that the model has sufficient data for training, an appropriate amount of data for parameter tuning and preventing overfitting, and independent data for final performance evaluation. Step 2: Construct the YOLOv11n-RepViT network model; the specific process is as follows: The YOLOv11n network model includes a backbone network, a neck network, and a head detection network. The YOLOv11n architecture is improved by combining the backbone network and the RepViT network structure. The RepViT architecture integrates the lightweight ViT (Vision Transformer) design concept with the efficient feature extraction capability of CNN (Convolutional Neural Networks). Through a multi-scale feature representation mechanism, global contextual information is captured in shallow networks, and small target details are accurately located in deep networks, thus optimizing computational efficiency and feature representation capability. The working process of the YOLOv11n-RepViT network model is as follows: The image is sequentially input into the backbone network, the neck network, and the head detection head. The head detection head outputs the target category and a two-dimensional bounding box. Step 3: Input the training dataset into the YOLOv11n-RepViT network model to train the model and build a lightweight deep learning model suitable for E. coli detection. Step 4: Validate the performance of the trained YOLOv11n-RepViT network model using the test dataset, and comprehensively evaluate the model's detection performance through multiple metrics.
2. A smart detection system for photothermal sterilization using MXenes combined with an improved YOLOv11, characterized in that, The specific steps for dividing the E. coli dataset in step one are as follows: The publicly available Deepbacs dataset was used as the experimental data source, containing a large number of E. coli images. The Deepbacs dataset was divided into three parts in a 7:2:1 ratio: training set, validation set, and test set. The training set contained 164 images, the validation set contained 47 images, and the test set contained 24 images. This division ensured that the model had sufficient data for training, an appropriate amount of data for parameter tuning and preventing overfitting, and independent data for final model performance evaluation.
3. A smart detection system for MXenes photothermal sterilization combined with an improved YOLOv11, characterized in that: The RepViT architecture integrates the lightweight ViT design philosophy with the efficient feature extraction capabilities of CNNs. It captures global contextual information in shallow layers and accurately locates small target details in deep layers through a multi-scale feature representation mechanism. The core modules employ a collaborative design of depthwise separable convolutions, GELU (Gaussian Error Linear Unit) activation functions, SE (Squeeze-and-Excitation Networks) channel attention, and FFN (Feed-Forward Networks) layers. Combined with a structural reparameterization strategy, it maintains a multi-branch structure to learn complex patterns during training and compresses it into a single-path, efficient model during inference. This significantly optimizes computational efficiency and feature representation capabilities, enabling the model to effectively detect small targets such as E. coli while meeting the stringent deployment requirements of embedded devices. Furthermore, RepViT enhances the network's expressive power and ensures data integrity and training stability during training through residual connections and SE mechanisms. Its optimized structure enables it to achieve excellent results in small object detection, and compared to traditional CNN models, RepViT, through the combination of various convolutional operations and attention mechanisms, is better able to handle object detection problems in complex scenes.
4. A smart detection system for MXenes photothermal sterilization combined with an improved YOLOv11, characterized in that: In step three, the training dataset is input into the YOLOv11n-RepViT network model for model training to construct a lightweight deep learning model suitable for E. coli detection; the specific process is as follows: The training set has an image size of 640×640, the input data size per batch is 32, the number of training rounds is 300, and the learning rate is 0.
01. This invention uses widely accepted evaluation metrics in the field of object detection to measure model performance, specifically including recall, precision, F1-score, and mean average precision mAP@0.5 (with an IoU threshold of 0.5).
5. A smart detection system for MXenes photothermal sterilization combined with an improved YOLOv11, characterized in that, Step four involves evaluating the trained model using a test set. The specific steps are as follows: To verify the effectiveness of the improved model, experiments were designed to compare the performance of the improved and original models before and after the weighting. Evaluation metrics included parameter count, computational cost (GFLOPs), inference speed (FPS), and detection accuracy mAP50 (%). Experimental results show that the improved detection model of this invention achieves significant improvements in both detection accuracy and speed compared to the baseline method.