Data processing method, data processing device, electronic equipment, storage medium and product
By enhancing farmland image data and training a target detection model, combined with insect-labeled data and a bidirectional routing attention mechanism, the problem of low insect recognition accuracy was solved, achieving high-precision insect detection and precision agricultural management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies have low accuracy in insect identification when image data quality is limited or insect features are not obvious.
By enhancing farmland image data, a target detection model is trained by combining a pre-trained model and insect-labeled data. A small target detection layer and a bidirectional routing attention mechanism are introduced to improve the model's insect detection capabilities.
It improves the accuracy of insect identification in complex environments, and enables high-precision detection of insect species, quantity and location, supporting precision agricultural management.
Smart Images

Figure CN121904448A_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology, and more particularly to a data processing method, a data processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the continuous development of smart agriculture, using image recognition and deep learning technologies for farmland pest and disease detection has become an important means to improve agricultural production efficiency. By collecting images of insects in farmland and combining them with environmental data, it is possible to intelligently identify and analyze the types, numbers, and distribution of insects, providing a scientific basis for agricultural pest control.
[0003] In related technologies, deep learning models are commonly used to process and identify images of pests and diseases. However, in practical applications, these methods still suffer from low accuracy when image data quality is limited or insect features are not obvious. Summary of the Invention
[0004] This application provides a data processing method, data processing device, electronic device, computer-readable storage medium, and computer program product that can improve recognition accuracy when image data quality is limited or insect features are not obvious.
[0005] The technical solution of this application embodiment is implemented as follows: This application provides a data processing method, the method comprising: Acquire image data of farmland and a pre-trained model; the image data is the enhanced version of the original image data. Obtain insect labeling data for the farmland; The image data and the insect annotation data are input into the pre-trained model for training to obtain a trained target detection model and the detection results of the image data output by the target detection model; wherein, the target detection model is used to detect insects in the farmland, and the detection results include one or more of the following: insect species, insect quantity, and insect location detected in the farmland.
[0006] This application provides a data processing apparatus, including: The acquisition unit is used to acquire image data of farmland and pre-trained models; The acquisition unit is used to acquire insect labeling data of the farmland; The processing unit is used to input the image data and the insect annotation data into the pre-trained model for training, and obtain the trained target detection model and the detection result of the image data output by the target detection model; wherein, the target detection model is used to detect insects in the farmland, and the detection result includes one or more of the following: insect species, insect quantity, and insect location detected in the farmland.
[0007] This application provides an electronic device, a communication interface, and a processor; wherein... The communication interface is used to acquire image data and pre-trained models of farmland; and to acquire insect annotation data of the farmland. The processor is used to input the image data and the insect annotation data into the pre-trained model for training, to obtain a trained target detection model and the detection results of the image data output by the target detection model; wherein, the target detection model is used to detect insects in the farmland, and the detection results include one or more of the following: insect species, insect quantity, and insect location detected in the farmland.
[0008] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the data processing method provided in this application when executed by a processor.
[0009] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the data processing method provided in this application.
[0010] The embodiments of this application have the following beneficial effects: First, by using image data enhancement and pre-trained models, the model's ability to identify insects in complex environments is improved. Second, by combining insect-labeled data for model training, the trained target detection model has higher insect detection accuracy, thereby enabling it to output one or more of the insect species, quantity, and location information more accurately. This solves the problem of low recognition accuracy in related technologies when image data quality is limited or insect features are not obvious. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating the data processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the pest and disease control process based on a cloud-based large model provided in an embodiment of this application; Figure 3 This is a schematic diagram of the network structure of the generative adversarial network provided in the embodiments of this application; Figure 4This is a schematic diagram of self-supervised learning provided in an embodiment of this application; Figure 5 This is a schematic diagram of the expansion of the large model training container provided in the embodiments of this application; Figure 6 This is a schematic diagram of the structure of the data processing apparatus provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0014] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0015] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0016] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0017] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0018] The data processing methods provided in the embodiments of this application can be executed by an electronic device, which may be a cloud server, an image processing terminal, or the like. That is, the data processing methods in the embodiments of this application can be executed by a cloud server, an image processing terminal, or a system that allows interaction between the cloud server and the image processing terminal.
[0019] Figure 1 This is a flowchart illustrating the data processing method provided in the embodiments of this application. The following will be combined with... Figure 1 The steps shown will be explained. It should be noted that... Figure 1 The method described uses a cloud server as the execution entity as an example for illustration. Figure 1 As shown, the method includes the following steps 101 to 103. Step 101: Obtain image data of farmland and a pre-trained model; the image data is the enhanced image data of the original image data.
[0020] Here, the farmland image data includes images of insects collected from farmland, which are then processed using a series of image enhancement techniques (such as random cropping, rotation, and brightness adjustment) to generate the image dataset. The purpose of image enhancement is to improve the model's adaptability to different lighting conditions, shooting angles, and other factors, thereby enhancing the robustness of object detection.
[0021] Here, the pre-trained model can be a network model (such as a deep neural network model) that has been pre-trained on a large-scale general image dataset. It has good feature extraction capabilities and can be adapted to specific tasks (such as farmland insect detection) through transfer learning, thereby reducing training time and improving model performance.
[0022] In practice, image enhancement operations typically include cropping, rotation, scaling, and color transformation. Image enhancement operations can increase the diversity of images without changing their content, enabling the model to maintain a high recognition accuracy when facing complex environments.
[0023] In practice, cloud servers can perform image enhancement on raw image data uploaded from image sensors or cameras and load pre-trained models for subsequent object detection training.
[0024] Step 102: Obtain insect labeling data for farmland.
[0025] Here, the insect annotation data for farmland refers to the manual or automatic annotation results of information such as the number, location, and species of insects in farmland images. This data typically includes bounding box coordinates, category labels, and quantity labels, used to guide the model in learning how to identify insects. The quality of the farmland insect annotation data directly affects the model's training performance; therefore, the annotations should be as accurate as possible, covering samples of various insect species and different postures and lighting conditions. In agriculture, there are numerous insect species, such as aphids, locusts, and moths, each with significant morphological differences. Therefore, high-quality farmland insect annotation data is crucial for the model's generalization ability.
[0026] In practice, the cloud server can obtain insect annotation data from farmland from expert annotation libraries or automated annotation systems, and match the insect annotation data with the enhanced image data to construct a training set. This matching process ensures that the model can accurately learn insect features during training, thereby improving detection accuracy.
[0027] Step 103: Input the image data and insect annotation data into the pre-trained model for training to obtain the trained target detection model and the detection results of the image data output by the target detection model; wherein, the target detection model is used to detect insects in the farmland, and the detection results include one or more of the following: insect species, insect quantity, and insect location detected in the farmland.
[0028] Here, the trained object detection model is a deep learning model capable of simultaneously performing one or more of the following: insect location detection, quantity detection, and species identification. The trained object detection model is obtained by inputting farmland image data, a pre-trained model, and labeled insect data from the farmland into the training process, followed by multiple iterations of optimization. The detection results include the insect species (e.g., aphids, locusts), insect quantity (e.g., 10 insects), and insect location (e.g., coordinate range in the image) detected by the trained object detection model in the farmland.
[0029] In practice, the cloud server uses enhanced image data, pre-trained models, and labeled insect data from farmland to train the model. Once training is complete, the trained target detection model can be deployed for real-time detection of insects in farmland. The detection results can be used to generate pest and disease reports and further combined with multi-dimensional environmental data for intelligent analysis, providing targeted prevention and control suggestions.
[0030] The data processing method provided in this application enhances farmland image data, combines a pre-trained model with labeled insect data from the farmland to train a pre-trained model, and obtains a trained target detection model and its output detection results on the image data. This achieves high-precision identification of insects in farmland and yields detection results containing information such as insect species, quantity, and location. The data processing method provided in this application further improves the intelligent level of agricultural pest and disease control, contributing to the realization of precision agricultural management.
[0031] In some embodiments, step 103 involves inputting image data and insect annotation data into a pre-trained model for training, obtaining a trained target detection model and the detection results of the image data output by the target detection model. This can be achieved through the following steps: Step 1031: Add a small target detection layer to the feature enhancement network layer of the pre-trained model to obtain the target model; the small target detection layer introduces a module designed based on a bidirectional routing attention mechanism.
[0032] Here, the small object detection layer refers to a neural network layer in the network structure specifically designed to identify small objects (such as insects). Because the images of insects on insect traps are small, these images are easily overlooked or misjudged in traditional object detection models. To improve the ability to identify such small targets, this embodiment adds a neural network layer for identifying small targets to the pre-trained model (feature enhancement network layer). This network layer strengthens the model's ability to perceive small targets such as insects.
[0033] In practical implementation, the small target detection layer employs a module designed based on a bidirectional routing attention mechanism. This module utilizes a query-aware sparse attention strategy, focusing only on the most valuable regions of the input image, thereby reducing computational complexity and improving processing efficiency. For example, the enhanced input image is divided into 160×160 small grid cells, allowing the model to capture the key features of each small target more precisely. Compared to traditional multi-head self-attention (MHSA) mechanisms, the module based on bidirectional routing attention avoids the high computational overhead of global attention while simultaneously improving the accuracy of small target localization.
[0034] Step 1032: Input the image data and insect annotation data into the target model for training to obtain the trained target detection model and detection results.
[0035] Here, the target model refers to the improved deep learning model, which adds a small target detection layer to the original pre-trained model and integrates a module designed based on a bidirectional routing attention mechanism. The target model not only has general image recognition capabilities, but is also specifically optimized for small targets (such as insects), improving recognition performance in low-resolution or blurry images.
[0036] Here, insect annotation data refers to a dataset of manually annotated insect images, where each image is labeled with the insect's location (bounding box), category, and possible other attributes (such as size, color, and quantity). This insect annotation data is used to supervise the training of the target model, enabling it to accurately identify different types of insects. During training, the target model continuously adjusts its parameters based on the difference between the predicted results and the actual annotations, ultimately obtaining a high-precision object detection model.
[0037] In practical implementation, by inputting high-quality image data and accurate insect annotation data into the target model for training, the system can automatically learn the visual characteristics of insects and efficiently and accurately detect new images during the testing phase. Because of these capabilities, agricultural workers can obtain real-time information on the distribution of insects in farmland, providing a scientific basis for subsequent pest and disease control.
[0038] In practical applications, the loss of the target detection model is as follows: after introducing a small target detection layer, the loss function becomes the detection loss, which includes classification and regression losses, and a regularization term is added.
[0039] In summary, by adding a small target detection layer to the pre-trained model and introducing a module designed based on a bidirectional routing attention mechanism, small targets such as insects can be identified more effectively. These improvements to the pre-trained model enhance the accuracy of the target detection model, enabling precise detection of the number and distribution of insects in farmland. This, in turn, provides users of the target detection system with timely and scientific pest and disease control recommendations.
[0040] In some embodiments, the insect labeling data includes insect species labeling data and insect quantity labeling data.
[0041] Here, insect species labeling data refers to the insect category information recorded after classifying insects captured in farmland using image recognition technology. For example, insects are classified into different categories such as aphids, planthoppers, and spider mites. Insect species labeling data helps the system determine which types of pests are present in farmland, thereby enabling the development of targeted control strategies. Insect species labeling data is usually stored in the form of tags, with each image sample corresponding to an identifier for one or more insect categories.
[0042] Here, insect count annotation data refers to the statistical recording of the number of each type of insect in an image. For example, if 5 aphids and 3 planthoppers are detected in an image, the insect count annotation data would be recorded as aphids: 5 and planthoppers: 3, respectively. Insect count annotation data can be used to assess pest density, determine whether the spraying threshold has been reached, and help predict future pest development trends. Insect count annotation data is usually stored in digital form and can be combined with timestamps to form dynamic monitoring data.
[0043] In practice, insect species labeling data is used to identify pest and disease types, while insect quantity labeling data is used to assess the severity of pest infestations. Combining insect species labeling data with insect quantity labeling data provides a more comprehensive picture of insect activity in farmland and improves the accuracy and efficiency of pest and disease control.
[0044] In this embodiment of the application, by subdividing insect labeling data into insect species labeling data and insect quantity labeling data, the system can more accurately identify different types of pests and the distribution density of insect species labeling data and insect quantity labeling data. Based on the insect species labeling data and insect quantity labeling data, agricultural managers can formulate prevention and control measures more scientifically, thereby improving the level of intelligence and sustainable development capability of agricultural production.
[0045] In some embodiments, acquiring image data of farmland in step 101 can be achieved through the following steps: Step 1011: Perform data preprocessing on the original image data of the farmland to obtain preprocessed image data.
[0046] In practice, data preprocessing is a crucial step before training image recognition and deep learning models. The purpose of data preprocessing is to remove noise from images, enhance image quality, and standardize image formats for subsequent model processing. Data preprocessing typically includes operations such as image cropping, rotation, translation, and brightness and contrast adjustment to increase image diversity and improve the model's generalization ability. For example, in outdoor farmland images, uneven lighting or weather conditions may cause problems such as shadows, blur, or color distortion, which can affect the model's recognition performance. Therefore, data preprocessing can improve image quality, providing more accurate foundational data for subsequent blur removal and super-resolution reconstruction.
[0047] In this embodiment, data preprocessing can reduce the impact of image interference. Data preprocessing can improve the accuracy of subsequent processing steps. Data preprocessing can improve the overall performance of pest and disease identification.
[0048] Step 1012: Deblur the preprocessed image data using a generative adversarial network to obtain the deblurred image data.
[0049] Here, a Generative Adversarial Network (GAN) is a deep learning architecture consisting of a generator and a discriminator. In this embodiment, a GAN is used to deblur pre-processed image data. Specifically, a generator is trained to produce sharp image data, while a discriminator evaluates the quality of the generated image data. The generator aims to produce results that are as close as possible to real, sharp image data, while the discriminator is responsible for determining whether the image data generated by the generator is truly sharp. Through an adversarial training mechanism between the generator and the discriminator, the generator progressively optimizes the sharpness of the image data, ultimately achieving effective recovery of blurry image data. For example, for blurry image data caused by camera shake or insufficient lighting, a GAN can automatically fill in details by learning a large number of mapping relationships between sharp and blurry image data, making the image data sharper and clearer.
[0050] In this embodiment, the problem of blurred image data can be effectively solved by using generative adversarial networks for deblurring; this deblurring method can improve the quality of image data; and this deblurring method can also enhance the accuracy of target detection and recognition.
[0051] Step 1013: Perform super-resolution reconstruction on the deblurred image data to obtain the image data.
[0052] Here, super-resolution reconstruction refers to the process of converting low-resolution image data into high-resolution image data to improve the detail and visual clarity of the image data. In this embodiment, super-resolution technology is used to further process the deblurred image data, enabling it to achieve a higher resolution and thus meeting the needs of deep learning models for high-quality image data. In practical applications, the super-resolution reconstruction process may be based on a Convolutional Neural Network (CNN) or Generative Adversarial Network (GAN) structure, which learns from and predicts low-resolution image data to generate higher-resolution image data. For example, the system can magnify the input blurry low-resolution image data by four times and output high-quality, clear image data, thereby significantly improving the recognizability of small targets such as insects.
[0053] In this embodiment, by employing super-resolution reconstruction technology, the system can improve the spatial detail representation of image data. The system can enhance the model's ability to identify pests and diseases. The system can improve the intelligence level of agricultural pest and disease identification and control systems.
[0054] In summary, in this embodiment of the application, the original farmland image data is preprocessed, deblurred, and super-resolution reconstructed using the methods described in this embodiment. This improves the quality of the image data, enhances the expression of pest and disease characteristics in the image data, and improves the accuracy and robustness of the recognition model, ultimately achieving more efficient identification and control of agricultural pests and diseases.
[0055] In practice, after using GANs to perform image denoising and super-resolution reconstruction, high-quality insect images are input into the target model for training. By adding a small target detection layer, the model can better understand the morphology, color, texture, and other features of insects, thereby significantly improving the recognition accuracy of the target detection model in field environments.
[0056] In some embodiments, obtaining the pre-trained model in step 101 can be achieved through the following steps: Step 1014: Obtain the reconstruction loss obtained by feature enhancement of the image data by the autoencoder.
[0057] Here, an autoencoder is an unsupervised learning model consisting of an encoder and a decoder, used to extract low-dimensional representations of input data and attempt to reconstruct the original data. In this application, the autoencoder extracts key features from pest and disease images by encoding and decoding them, and evaluates the image quality through the reconstruction process, thereby achieving feature enhancement. The reconstruction loss output by the autoencoder is a measure of the difference between the original image and the reconstructed image, which can be calculated using metrics such as mean squared error or cross-entropy.
[0058] In practice, using autoencoders for feature enhancement can effectively improve the consistency and quality of image data. When dealing with common problems such as blurred images or variations in lighting, this feature enhancement method helps neural network models better capture the key visual features of insects or pests. The feature enhancement step using autoencoders forms an important foundation for building high-quality image datasets and provides clearer and more representative input data for subsequent object detection and classification tasks.
[0059] In practice, by using an autoencoder to perform feature enhancement processing on image data, noise and interference factors in the image can be effectively removed, improving the image's recognizability and thus enhancing the generalization ability and robustness of the neural network model trained on the image.
[0060] Step 1015: Obtain the auxiliary loss obtained by training the image data using the auxiliary task.
[0061] Here, auxiliary task training refers to designing additional tasks to help the neural network model learn more meaningful features. These auxiliary tasks may include image rotation prediction, color transformation prediction, and contextual information recovery. By introducing these tasks, the neural network model can be trained without labeled data, thereby improving its ability to understand image content. The auxiliary loss generated during auxiliary task training is a quantitative indicator of how well the neural network model performs in completing the auxiliary task, and it usually shares some network parameters with the main task (such as pest and disease identification).
[0062] This approach not only improves the neural network model's ability to perceive image structure and semantic information, but also enhances its multi-task processing capabilities, enabling it to simultaneously handle multiple related but different task objectives. Furthermore, the inclusion of auxiliary tasks helps alleviate overfitting issues and improves the model's adaptability to real-world scenarios.
[0063] In this embodiment, by introducing auxiliary task training, the system can improve the neural network model's ability to extract image features, enhance the neural network model's multi-task processing capability, and reduce its dependence on a large amount of labeled data, thereby improving the overall performance and generalization ability of the neural network model.
[0064] Step 1016: Based on the reconstruction loss and auxiliary loss, adjust the model parameters of the network model to obtain the pre-trained model.
[0065] Here, the reconstruction loss generated by the autoencoder and the auxiliary loss generated by the auxiliary task training are weighted and combined as the overall training objective function to guide the parameter updates of the neural network model. The loss of the pre-trained model is a weighted sum of the reconstruction loss from feature enhancement and the auxiliary task loss, used for feature enhancement. Through gradient descent algorithms (such as the Adam optimizer), the neural network model continuously adjusts the weight parameters based on these two loss values to minimize the total loss, thereby gradually approaching the optimal model state. After multiple iterations, the neural network model can obtain good feature representation capabilities without relying on labeled data, forming a pre-trained model with strong versatility and high generalization performance.
[0066] The aforementioned joint optimization strategy enables the neural network model to simultaneously consider image reconstruction quality and task-relevant feature extraction, providing a solid foundation for subsequent supervised fine-tuning. The pre-trained model can be widely applied to various agricultural pest and disease identification scenarios, significantly improving the neural network model's adaptability and recognition accuracy under different environments.
[0067] In practice, combining reconstruction loss and auxiliary loss for training the neural network model can effectively improve the feature representation ability and task adaptability of the neural network model, and provide a high-performance pre-trained model foundation for subsequent pest and disease identification tasks.
[0068] In practical implementation, using autoencoders for image feature enhancement, combined with auxiliary task training, can effectively improve the quality of image data and the feature extraction capability of neural network models. By using autoencoders for image feature enhancement and combining them with auxiliary task training, the dependence on large amounts of labeled data can be reduced, thereby improving the generalization ability of neural network models and supporting various agricultural pest and disease identification tasks, ultimately achieving intelligent and high-precision pest and disease control.
[0069] In some embodiments, the above method further includes: Obtain multidimensional environmental feature vectors of farmland; By inputting the multidimensional environmental feature vector and detection results into the large model, one or more of the pest and disease analysis results and pest and disease control information output by the large model are obtained.
[0070] Here, a multidimensional environmental feature vector refers to a data structure in vector form that organizes various environmental data (such as temperature, humidity, wind speed, precipitation, carbon dioxide concentration, etc.) in farmland according to a time series. Multidimensional environmental feature vectors can comprehensively reflect the current environmental state of farmland, facilitating joint analysis with image recognition results.
[0071] For example, at a sampling point (SP), data such as temperature T=28℃, humidity H=75%, wind speed V=3m / s, precipitation P=0mm, and carbon dioxide concentration CO2=400ppm can be collected. These data can form a five-dimensional environmental feature vector [28,75,3,0,400]. By constructing a multi-dimensional environmental feature vector, the system can capture the real-time changing trends of farmland, thereby improving the accuracy of subsequent pest and disease assessments.
[0072] This application obtains multi-dimensional environmental feature vectors, which can comprehensively understand the environmental conditions of farmland, improve the accuracy of judging the conditions for the occurrence of pests and diseases, and thus provide more reliable data support for precision prevention and control.
[0073] Here, pest and disease analysis results refer to conclusions about pest and disease types, distribution, and severity derived from the joint analysis of image recognition results and multidimensional environmental feature vectors. For example, if the system detects a large number of aphids in a certain area, and combines this with the current high temperature and humidity multidimensional environmental feature vectors, it can infer that a certain area of farmland is prone to viral diseases, thus generating an analysis conclusion that a certain area of farmland may have a risk of virus transmission. Pest and disease control information consists of response suggestions based on the above analysis results, including specific measures such as which pesticides to use, the frequency of application, and whether human intervention is needed. For example, the system may recommend using highly effective and low-toxicity biological pesticides and spraying them every two days to control aphid numbers and prevent virus spread.
[0074] This application combines image detection results with multi-dimensional environmental feature vectors for analysis. The system can achieve intelligent diagnosis and scientific prevention and control of pests and diseases. The system can avoid misjudgment or omission caused by single-factor judgment. In this way, the system can improve the level of intelligence and prevention and control efficiency of agricultural production.
[0075] In this embodiment, the system acquires multi-dimensional environmental feature vectors of farmland and combines these vectors with image recognition results to generate pest and disease analysis and control information. By acquiring multi-dimensional environmental feature vectors of farmland and combining them with image recognition results, the system can more accurately determine the occurrence trend and impact range of pests and diseases. Based on the above analysis results, the system formulates more targeted control strategies, thereby significantly improving the management efficiency and crop yield of smart agriculture.
[0076] Based on the above, this application proposes an agricultural pest and disease control system based on a cloud-based large-scale model. This system significantly improves image quality by performing data augmentation on the original images (such as random cropping, rotation, brightness and contrast adjustment) and utilizing GANs for blur removal and super-resolution reconstruction, thereby enhancing the accuracy and robustness of the cloud-based large-scale model-based agricultural pest and disease control system. To improve the detection accuracy of small objects, this application improves the network architecture by adopting a two-layer routing attention mechanism combined with CNNs and Transformers, solving the computational resource consumption problem of traditional multi-head self-attention (such as MHSA) on large-scale datasets. The small object detection layer introduces a module designed based on a bidirectional routing attention mechanism (such as BiFormer). BiFormer uses a query-aware sparse attention strategy and divides the image into smaller grid units through a four-fold downsampling branch, improving the detection accuracy of small objects. In terms of natural language processing, this application uses a vertical large-scale model for data analysis. The vertical large-scale model is pre-trained with task-related data and then fine-tuned with instructions to enhance its understanding and execution capabilities for specific tasks.
[0077] Figure 2 This is a schematic diagram of a cloud-based large model-based pest and disease control process provided in an embodiment of this application, such as... Figure 2 As shown, the process includes steps 201 to 205. S201, Data Acquisition.
[0078] In practical applications, the system periodically receives photos of the insect traps uploaded by image sensors, as well as other data such as temperature and humidity in the farmland uploaded by various sensors. The system uses object storage technology to store the image data in the cloud and uses database technology to store basic farmland information in the cloud.
[0079] S202, Image Enhancement.
[0080] In practical applications, the system performs simple cropping, rotation, and translation operations on image data to enhance its richness. Since outdoor images may become blurry or unclear due to factors such as lighting, the system feeds the photos into a generative adversarial network (GAN) for blur removal, followed by super-resolution reconstruction, which quadruples the resolution of the original data. Blur removal and resolution reconstruction are implemented by the system through the GAN. In this embodiment, the GAN's network structure is as follows: Figure 3 As shown.
[0081] The working principle of generative adversarial networks is as follows: assuming there exists A sample with noise ,exist A normal sample Discriminator D receives input samples. ( for A sample with noise and (one of the normal samples), output a confidence value. This value represents the sample This represents the probability of a normal sample. The generator is used to generate samples. That is, generating a sample, which is a noise sample generated from a normal sample. It comes from the probability space The latent vector. Noisy samples are processed by a generator and a discriminator to obtain the output. Normal samples are processed by a discriminator to obtain the output. The loss function of the discriminator is calculated based on the system. Loss function of generator The adversarial loss function of this generative adversarial network is: ; in, It is binary cross-entropy loss. It is an adversarial loss mechanism that optimizes the loss function using the Adam gradient descent algorithm to obtain a generator model that can be denoising.
[0082] S203, a pre-trained model is obtained through self-supervised learning.
[0083] Here, self-supervised learning is a training method in machine learning. It encourages models to learn using unlabeled data, finding relationships between samples by mining the inherent features of the data. This application uses an autoencoder to enhance image features. The network is trained from a large amount of unlabeled data by designing an auxiliary task (pretext), resulting in a pre-trained model. Then, for new downstream tasks, the model parameters obtained through transfer learning are applied to the new task, and fine-tuning is performed to obtain the pre-trained model.
[0084] Figure 4 This is a schematic diagram of self-supervised learning provided in an embodiment of this application. Figure 4 The structure and data processing flow of the autoencoder are illustrated. The input image X first enters the encoder, which encodes the input image X, extracts its features, and compresses them into a representation called Z in the bottleneck layer. This achieves dimensionality reduction of the image data, removing redundant information and retaining key features, thus enabling the extraction of intrinsic features from large amounts of unlabeled data. Z then enters the decoder, which restores Z to (reconstructs the input) x. ’ The goal is to reconstruct an image as similar as possible to the original input X. In this embodiment, after obtaining Z, data dimensionality reduction and clustering are performed to mine sample relationships through auxiliary tasks.
[0085] S204, Object Detection Network Training.
[0086] Here, since harmful insects are relatively small targets, a small target strategy is used to improve the network when performing object detection. This application adds a small target detection layer to the feature enhancement network layer of the pre-trained model. This detection layer is based on a two-layer routing attention mechanism, combining the advantages of CNNs and Transformers to address the scalability challenges faced by MHSA. Traditional attention mechanisms require each query to interact with all key-value pairs, which can lead to significant computational and storage resource consumption when processing large-scale datasets. The small target detection layer added in this application addresses the problem of excessive computational and storage resource consumption caused by traditional attention mechanisms by employing a dynamic, query-aware sparse attention strategy. The small target detection layer uses a quadruple downsampling branch to divide the input image into 160x160 grid cells. These cells, compared to larger grid cells, can capture the features of small targets in the input image more finely, thereby improving the accuracy of small object detection. After introducing the small target detection layer, the loss function becomes a detection loss, including classification and regression losses, and a regularization term is added. The loss function of the object detection model (i.e., the object detection network) is: ; in, This is the predicted probability of the target being the anchor box. Here, the anchor box is a predefined bounding box used to locate and classify targets in an image. This represents the probability of the target's existence, i.e., the true probability. When adding a small target detection layer, the anchor box design needs to be optimized for the characteristics of small targets; for example, using smaller anchor box sizes to accommodate small targets. During training, AdamW, as the most up-to-date iterative optimizer, has fewer optimization iterations and higher computational efficiency in model training. The way weight decay regularization is incorporated is what distinguishes AdamW from other iterators. Unlike traditional L2 regularization, AdamW does not add the regularization term directly to the loss function. Instead, AdamW adds the gradient of the regularization term directly to the backpropagation formula, rather than the loss function itself, thus more directly affecting parameter updates, resulting in: ; in, It's the learning rate. It is the regularization coefficient. The weights from the previous step, This sets the current weights, which can more directly affect parameter updates, helping to reduce model complexity and improve the model's generalization ability.
[0087] In this embodiment, in order to build a fast and flexible training environment, this application uses mobile cloud hard disks as a big data storage method. Storage capacity can be quickly increased by creating and expanding virtual machine instances and container instances. The batch disk creation capability of block storage determines the elasticity of the training environment. If any problems occur during the training process, the system can quickly perform expansion and recovery operations, thereby improving the elasticity of the training environment.
[0088] Figure 5 A diagram illustrating the resizing of training container instances, as shown below. Figure 5 As shown, Customer Service refers to customer service-related business needs, the basic operating unit of the Elastic Container Instance (ECI Pod) is the container instance, and the dashed box shows the expansion of the container instance. In this application, the training environment increases storage capacity by expanding the container instance; Snapshot is used for data backup, etc., and Device can be a storage device. Together with the container instance, they constitute the storage system of the training environment, reflecting the construction of storage capacity through block storage and other methods, as well as the correlation and collaboration of each part in the training environment, so as to achieve elastic capabilities such as rapid expansion and recovery of the training environment.
[0089] By acquiring multi-feature data from the same sampling point at different times within different time periods, including but not limited to multi-dimensional environmental data such as temperature, humidity, wind speed, precipitation, and carbon dioxide concentration, the sampling points within the acquisition time period are compiled into a vector. This vector can be used... express, Indicates insect species, Indicates the number of insects, Represents a multi-dimensional feature vector. For time ( The system organizes the multidimensional feature data into training data, as shown in Table 1 below:
[0090] Table 1 Multi-feature data S205, Large Model Prediction.
[0091] Here, we can first train a Transformer model on a large amount of data. The Transformer is a deep learning model architecture for Natural Language Processing (NLP) tasks, primarily composed of multiple encoders and decoders, and incorporates a self-attention mechanism. This self-attention mechanism allows the model to consider all positions in the input sequence simultaneously, rather than processing them step-by-step like recurrent neural networks (RNNs) or convolutional neural networks (CNNs). The self-attention mechanism allows the model to assign different attention weights based on different parts of the input sequence, thereby better capturing semantic relationships.
[0092] Furthermore, a large vertical model is used for data analysis and recommendations, and task-related data is used for pre-training. The collected multidimensional feature data is treated as two-dimensional data to establish an agricultural pest and disease database. The model is then trained to learn statistical patterns and semantic information of language. Instruction tuning is then performed; it is a special form of supervised tuning designed to enable the model to understand and follow human instructions.
[0093] In practical applications, a series of Natural Language Processing (NLP) tasks are first prepared, and each task is transformed into instructions. Each instruction includes a human description of the task the model should perform and the expected output. Then, these instructions are used to supervise the learning of a pre-trained large language model. In this way, the model can learn and adapt to the instructions, thereby improving its performance on specific tasks.
[0094] This application also provides a data processing apparatus, such as... Figure 6 As shown, the data processing apparatus includes: The acquisition unit 601 is used to acquire image data of farmland and a pre-trained model; the image data is the enhanced image data of the original image data. Acquisition unit 601 is used to acquire insect annotation data of farmland; The processing unit 602 is used to input image data and insect annotation data into a pre-trained model for training, and obtain a trained target detection model and the detection results of the image data output by the target detection model; wherein, the target detection model is used to detect insects in farmland, and the detection results include one or more of the following: insect species, insect quantity, and insect location detected in the farmland.
[0095] In some embodiments, the processing unit 602 is used to add a small target detection layer to the feature enhancement network layer of the pre-trained model to obtain a target model; the small target detection layer introduces a module designed based on a bidirectional routing attention mechanism; image data and insect annotation data are input into the target model for training to obtain a trained target detection model and detection results.
[0096] In some embodiments, the insect labeling data includes insect species labeling data and insect quantity labeling data.
[0097] In some embodiments, the processing unit 602 is used to preprocess the original image data of the farmland to obtain preprocessed image data; to deblur the preprocessed image data through a generative adversarial network to obtain deblurred image data; and to perform super-resolution reconstruction on the deblurred image data to obtain image data.
[0098] In some embodiments, the processing unit 602 is used to obtain the reconstruction loss obtained by the autoencoder performing feature enhancement on the image data; obtain the auxiliary loss obtained by training the image data for an auxiliary task; and adjust the model parameters of the network model based on the reconstruction loss and the auxiliary loss to obtain a pre-trained model.
[0099] In some embodiments, the processing unit 602 is used to obtain a multidimensional environmental feature vector of the farmland; input the multidimensional environmental feature vector and the detection results into a large model to obtain one or more of the pest and disease analysis results and pest and disease control information output by the large model.
[0100] This application also provides an electronic device, such as... Figure 7 As shown, the electronic device 700 includes: a communication interface 701 and a processor 702; wherein, The communication interface 701 enables information exchange with other devices; The processor 702 is connected to the communication interface 701 to enable information interaction with other devices and to execute the methods provided by one or more technical solutions on the electronic device side when running a computer program. Memory 703 stores computer programs that can run on processor 702.
[0101] Communication interface 701 is used to acquire image data of farmland and pre-trained models; the image data is the augmented image data of the original image data; and to acquire insect annotation data of farmland. The processor 702 is used to input image data and insect annotation data into a pre-trained model for training, and obtain a trained target detection model and the detection results of the image data output by the target detection model; wherein, the target detection model is used to detect insects in farmland, and the detection results include one or more of the following: insect species, insect quantity, and insect location detected in the farmland.
[0102] In some embodiments, the processor 702 is used to add a small target detection layer to the feature enhancement network layer of the pre-trained model to obtain a target model; the small target detection layer introduces a module designed based on a bidirectional routing attention mechanism; image data and insect annotation data are input into the target model for training to obtain a trained target detection model and detection results.
[0103] In some embodiments, the insect labeling data includes insect species labeling data and insect quantity labeling data.
[0104] In some embodiments, the processor 702 is configured to preprocess the original image data of the farmland to obtain preprocessed image data; deblur the preprocessed image data using a generative adversarial network to obtain deblurred image data; and perform super-resolution reconstruction on the deblurred image data to obtain image data.
[0105] In some embodiments, the processor 702 is configured to obtain the reconstruction loss obtained by the autoencoder performing feature enhancement on the image data; obtain the auxiliary loss obtained by training the image data for an auxiliary task; and adjust the model parameters of the network model based on the reconstruction loss and the auxiliary loss to obtain a pre-trained model.
[0106] In some embodiments, the processor 702 is used to acquire a multidimensional environmental feature vector of farmland; input the multidimensional environmental feature vector and detection results into a large model to obtain one or more of the pest and disease analysis results and pest and disease control information output by the large model.
[0107] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the data processing method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0108] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the data processing method described in this application.
[0109] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the data processing method provided in this application. For example, ... Figure 1 The data processing method is shown.
[0110] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0111] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0112] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0113] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0114] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A data processing method, characterized in that, include: Acquire image data of farmland and a pre-trained model; the image data is the enhanced version of the original image data. Obtain insect labeling data for the farmland; The image data and the insect annotation data are input into the pre-trained model for training to obtain a trained target detection model and the detection results of the image data output by the target detection model; wherein, the target detection model is used to detect insects in the farmland, and the detection results include one or more of the following: insect species, insect quantity, and insect location detected in the farmland.
2. The method according to claim 1, characterized in that, The step of inputting the image data and the insect annotation data into the pre-trained model for training, to obtain the trained target detection model and the detection results of the image data output by the target detection model, includes: A small target detection layer is added to the feature enhancement network layer of the pre-trained model to obtain the target model; the small target detection layer introduces a module designed based on a bidirectional routing attention mechanism. The image data and the insect annotation data are input into the target model for training to obtain the trained target detection model and the detection results.
3. The method according to claim 1, characterized in that, The insect labeling data includes insect species labeling data and insect quantity labeling data.
4. The method according to claim 1, characterized in that, The acquisition of farmland image data includes: The original image data of the farmland is preprocessed to obtain preprocessed image data; The preprocessed image data is deblurred using a generative adversarial network to obtain the deblurred image data. The deblurred image data is then reconstructed using super-resolution to obtain the image data.
5. The method according to any one of claims 1 to 4, characterized in that, Obtaining a pre-trained model includes: Obtain the reconstruction loss obtained by feature enhancement of the image data using an autoencoder; Obtain the auxiliary loss obtained by training the image data using an auxiliary task. Based on the reconstruction loss and the auxiliary loss, the model parameters of the network model are adjusted to obtain the pre-trained model.
6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtain the multidimensional environmental feature vector of the farmland; The multidimensional environmental feature vector and the detection results are input into the large model to obtain one or more of the pest and disease analysis results and pest and disease control information output by the large model.
7. A data processing apparatus, characterized in that, include: The acquisition unit is used to acquire image data of farmland and pre-trained models; The acquisition unit is used to acquire insect labeling data of the farmland; The processing unit is used to input the image data and the insect annotation data into the pre-trained model for training, and obtain the trained target detection model and the detection result of the image data output by the target detection model; wherein, the target detection model is used to detect insects in the farmland, and the detection result includes one or more of the following: insect species, insect quantity, and insect location detected in the farmland.
8. An electronic device, characterized in that, Includes communication interfaces and processors; among which, The communication interface is used to acquire image data and pre-trained models of farmland; and to acquire insect annotation data of the farmland. The processor is used to input the image data and the insect annotation data into the pre-trained model for training, to obtain a trained target detection model and the detection results of the image data output by the target detection model; wherein, the target detection model is used to detect insects in the farmland, and the detection results include one or more of the following: insect species, insect quantity, and insect location detected in the farmland.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.