Forestry pest identification and science popularization method and device based on image-text understanding
Through the method based on graphic and text understanding, a model of graphic and text alignment is constructed to realize the identification and popularization of forestry pests, and the problem of traditional pest recognition requires a lot of human resources, reducing costs and improving identification efficiency.
Patent Information
- Application Number
- CN202510295117.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, forestry pest identification requires a lot of manpower and material resources, resulting in the high cost of pest and disease protection and control.
Using a method based on graphic and text understanding, the identification and popular science of forestry pests are realized by constructing training data sets, model architecture, model training and model packaging for graphic and text feature alignment. This method uses artificial intelligence technology, combining image and text information to perform pest recognition and answer generation.
By reducing dependence on human resources, the cost of pest identification and control is reduced, the accuracy and efficiency of identification are improved, and the smart construction of urban areas and the development of ecological civilization are promoted.
Smart Images

Figure CN120125914A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of pest identification, and in particular to a method and device for forest pest identification and popular science based on graphic and text understanding. Background Art
[0002] The identification of forest pests is a highly professional task. In traditional garden maintenance work, a large number of professionals with professional knowledge and experience are usually required to conduct on-site investigations and collect pest samples. Due to the large area of park scenic spots and the high coverage rate of tree vegetation, a large amount of manpower, material resources and financial resources are required to complete this work, resulting in the high cost of pest control and management work. Summary of the Invention
[0003] Aiming at the deficiencies of the prior art, the present application provides a method and device for forest pest identification and popular science based on graphic and text understanding to solve the above technical problems that a large amount of manpower, material resources and financial resources are required, resulting in the high cost of pest control and management work.
[0004] To achieve the above object, the present application provides the following technical solutions: A method for forest pest identification and popular science based on graphic and text understanding, comprising: S1: Construct a training data set First, collect live pests and collect pest images. Use a 12-megapixel resolution camera to take pictures of all live pests from different angles, distances, lighting conditions and living states, and integrate and process the photos. Each photo is used as a sample, and multiple photos finally form a pest sample library. Scale and rotate the pest images, and fuse the pest photos with various types and multiple background template images to ensure the random distribution of pests in the fused images and the authenticity of the pest sizes at a specific distance. At least 5000 sample images are generated, and each image includes at least two pests, with at least two of each type of pest. Then, manually add text information to each sample. The text information added in step S1: Construct a training data set includes: questions and corresponding answers, and finally generate training samples. A training sample is composed of an image, question text and corresponding answer text.
[0005] S2: Model architecture for graphic and text feature alignment The model of image-text feature alignment consists of two parts: image-text feature extraction and alignment and feature retrieval to generate answers. The image-text feature extraction and alignment consists of two network structures: UniversalSentenceEncoder and FiLM EfficientNet. The feature retrieval to generate answers consists of a Transformer network structure, and there are at least twenty groups of Transformer networks. The UniversalSentenceEncoder network is used to encode the input text-based questions and generate an embedded representation in the form of a vector, that is, to extract the features in the text and convert them into vector representations, which are recorded as USE(que). The FiLM EfficientNet network structure uses the pre-trained EfficientNet as the basis, encodes the input image data, and adds a FiLM layer to fuse the encoded image and the encoded text feature vector USE(que) to obtain a Tokens representation that fuses image and text features. Multiple Transformer networks use Tokens that fuse image-text features as input, retrieve image-text features through the attention mechanism, and output text-based answers. Using the training data set described above, training this model and adjusting the parameters of each network structure can enable the model to have the ability to align images and texts and retrieve and generate text answers from images.
[0006] S3: Model training According to the training data set in the construction, any training sample is taken, including: First, the text question is marked as Q, the pest picture is marked as P, and the text answer is marked as A. Q and P are used as the input of the model, and A is used as the sample label; During training, Q and P are input into the model at the same time. After the forward propagation of the model, an output containing the model parameters is obtained. The distance between the wrong output answer and the correct output A is the loss function. By minimizing the loss function and backpropagating, the model parameters are updated, and the UniversalSentenceEncoder network is adjusted to tend to extract the number of pests as the main feature. The FiLM EfficientNet network is adjusted to tend to associate the number of pests in the text with the same number of pests in the picture, and this association is generated into Tokens to achieve alignment of image and text features. Finally, multiple Transformer network structures are used to further realize the retrieval of the number of pests in the image. During training, the minibatch method is used to select samples in batches to complete the above forward calculation and backpropagation, and the model parameters are continuously updated until convergence.
[0007] S4: Model packaging The trained and optimized model is packaged, the packaged file is read into the memory and integrated into the dynamic link library together with the code, and finally the packaged model is deployed inside the forestry pest identification and popular science device based on graphic understanding, and real-time prediction or batch processing is performed during use, and the performance of the model is monitored.
[0008] Preferably, the living pests include common pests such as American pest, two-tailed boat moth, willow toxic moth, pine tip borer, citrus fruit fly, Sophora japonica looper, glabripennis beetle, scale insect, yellow thorn moth and tabulaeformis. The above pests can comprehensively cover most of the pests in the urban area, thereby greatly improving the accuracy of the model.
[0009] Preferably, in step S3: in model training, cross entropy loss is used for model training. Cross entropy loss is very sensitive to changes in probability values and can better guide the update of model parameters. In classification problems, it reflects the distance between the probability distribution predicted by the model and the probability distribution of the true label. Therefore, the smaller the cross entropy loss, the closer the two probability distributions are, and the higher the prediction accuracy of the model. When using mean square error loss (MSE) as the loss function of the classification problem, since MSE squares the predicted probability, when the predicted probability is close to 0 or 1, its gradient will tend to 0, resulting in the gradient vanishing problem, affecting the training efficiency of the model. However, cross entropy loss does not have this problem. It can always provide effective gradient information for the model, thereby avoiding gradient vanishing and making model training more stable. Cross entropy loss calculates the loss by taking the logarithm, avoiding the numerical instability problem that may occur in some extreme cases, making the model more stable and reliable during training. The calculation method of the cross entropy loss function is relatively simple and can be directly implemented through a standard mathematical library. At the same time, its intuitiveness allows us to easily understand the difference between model predictions and actual situations, so as to better tune the model. The cross entropy loss function can handle multi-category classification problems well. It calculates the loss of each category separately and sums them up to get the total loss, which gives it a significant advantage in dealing with multi-classification problems.
[0010] Preferably, in step S1: constructing a training data set, pest images are collected by taking photos from the front angle, side angle, oblique side angle, back angle, and overhead and upward angles, and at the same time, close-up, near shot, mid shot and long shot are taken at different distances, and the photos are taken under natural light, hard light and soft light, artificial light and front light, side light and back light respectively. By shooting at multiple angles, distances and light and shadow environments, the model's accuracy in identifying pests is improved.
[0011] Preferably, in step S4: in model encapsulation, the model operating environment is: Intel i9-10900X processor, NVIDIA GeForce RTX 3090 GPU, 128G memory, and Python 3.6 and Pytorch 1.4.0 are deployed. With the above hardware, the operating speed of the device can be greatly improved, thereby enhancing the user experience when using the model.
[0012] Preferably, in step S1: in constructing the training dataset, the model collects at least 35,000 training samples. Each training sample consists of an image, question text, and corresponding answer text. The image covers at least 16 common forestry pests in urban areas, and the size of each picture is 324x324. Through multiple groups of training samples, the recognition speed and accuracy of the model for pests are greatly improved.
[0013] A device for forestry pest identification and popular science based on graphic and text understanding, comprising: a base, a telescopic arm, a host device, and a protective top cover. The telescopic arm is assembled on the top of the base, the host device is assembled on the front of the telescopic arm, and the protective top cover is assembled on the top of the host device. A groove is provided on the front of the host device. Ground nails are evenly distributed on the top of the base. The base can support the telescopic arm, enabling the telescopic arm to be stably assembled in an outdoor open environment. The ground nails can fix the base. The telescopic arm can be telescoped, thereby changing the overall height of the device, facilitating the storage of the device when not in use and changing the height when in use. At the same time, the telescopic arm is a common telescopic rod device in the prior art, and its length can be adjusted and fixed. The host device is a common server device in the prior art, and also includes a communication module, a motherboard, a central processing unit, a memory, a hard disk, a RAID card, a network card, a power supply, and a cooling fan inside. The protective top cover can protect the host device, preventing rain and the sun from affecting the service life of the internal components of the host device. The ground nails can fixedly assemble the base on the ground, and a control mechanism is assembled on the top of the protective top cover.
[0014] Preferably, the control mechanism includes a sensor group and a top plate. The sensor group is assembled on the top of the protective top cover, and the top plate is assembled inside the protective top cover. Hoists are connected to both sides of the top of the top plate, and the hoists are connected to the sensor group through wires. The output end of the hoist is connected to a weight plate, and a protective belt is connected between the weight plate and the top plate. The sensor group is a common sensor device in the prior art and is used to work through a rainwater control signal amplifier and a hoist. The hoist is a common hoisting device in the prior art and is used to drive the weight plate to move. The protective belt is a common elastic waterproof fabric in the prior art to prevent rainwater from damaging the main equipment. The weight of the weight plate can pull the weight plate to move. When the weight plate returns to its position, it can be engaged with the protective top cover, thereby greatly increasing the overall length of the protective top cover. At the same time, the weight plate and the top plate can squeeze the protective belt to drain the external water source.
[0015] Preferably, a signal amplifier is assembled at the bottom of the sensor group, and the signal amplifier is connected to the sensor group through a wire. The signal amplifier is a common signal amplification device in the prior art. The switch of the signal amplifier is controlled by the sensor group through a wire. During rainy weather, the signal amplifier can increase the signal strength of the main equipment. At the same time, the signal amplifier can be turned off on sunny days, thereby avoiding affecting the mobile phone signals of outdoor visitors.
[0016] Preferably, drain holes are evenly arranged on the top of the weight plate. The inlet of the drain hole is designed with a wide mouth, and the outlet of the drain hole is designed to slope downward. And at least six groups of drain holes are provided. The drain holes can drain the water source discharged by the protective belt. The wide-mouth design of the inlet facilitates the collection of water sources, and the sloping design of the outlet facilitates the discharge of the collected water sources. At the same time, multiple groups of drain holes can improve the drainage speed.
[0017] In summary, compared with the prior art, the present application provides a method and device for forest pest identification and popular science based on graphic understanding, which has the following beneficial effects: 1. The method and device for forest pest identification and popular science based on graphic understanding utilize the latest multimodal understanding technology in the field of artificial intelligence. It can not only provide popular science knowledge to tourists and improve the fun of the tour, but also obtain first-hand information on pest distribution based on the popular science interaction of tourists, giving full play to the characteristics of wide distribution and large number of tourists and promptly discovering pest information, saving a large amount of human and material resources in traditional maintenance work; 2. The method and device for forestry pest identification and popular science based on graphic and text understanding can provide popular science knowledge about forestry pests to tourists. At the same time, park managers can obtain early warning information about the occurrence of pests in the park area, reduce the cost of pest control work, utilize the characteristics of wide distribution and large number of tourists in the park scenic area, provide pest-related popular science knowledge to users through an interactive Q&A method, and obtain pest information in the park area at the same time, improve the efficiency of pest control work in the park scenic area, save labor costs, improve the intelligent level of garden management, and promote the development of urban intelligent construction and ecological civilization; 3. The method and device for forestry pest identification and popular science based on graphic and text understanding, when it is rainy, through the added control mechanism, automatically shields and protects the outside of the device to avoid damage to the device caused by rain. At the same time, through the added signal amplifier, when it is rainy, the signal of the device is automatically turned on and expanded to avoid the influence of rain on the signal transmission of the device and reduce the user experience. At the same time, when the control mechanism returns to its position, it can also discharge the external water source by extrusion, so as to avoid the normal use of the device being affected by the continuous dripping of water from the control mechanism outside the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the flowchart of the method of the present invention.
[0019] Figure 2 is the working flowchart of the present invention.
[0020] Figure 3 is the front view schematic diagram of the present invention.
[0021] Figure 4 is the partial schematic diagram of the control mechanism of the present invention.
[0022] Description of the Reference Numerals in the Drawings: 1. Base; 11. Ground nail; 2. Telescopic arm; 3. Host device; 31. Groove; 4. Protective top cover; 5. Control mechanism; 51. Sensor group; 52. Signal amplifier; 53. Top plate; 54. Winch; 55. Weight plate; 56. Protection belt; 57. Drainage port. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] Embodiment 1 Please refer to Figure 1, a method for forest pest identification and popular science based on graphic and text understanding, including: S1: Construct a training data set First, collect live pests and collect pest images. Use a 12-megapixel resolution camera to take pictures of all live pests from different angles, distances, lighting conditions, and living states. Then integrate and process the photos to form a pest sample library. Scale and rotate the pest images, and fuse the pest photos with various types and multiple background template images to ensure the random distribution of pests in the fused images and the authenticity of pest sizes at a specific distance. At least 5000 sample images are generated, with at least two pests on each image and at least two of each type of pest. Manually add text information to each sample. In step S1: Construct a training data set, the added text information includes: questions and corresponding answers. Finally, training samples are generated. One training sample consists of an image, question text, and corresponding answer text.
[0025] S2: Model architecture for graphic and text feature alignment The model for graphic and text feature alignment consists of two parts: graphic and text feature extraction and alignment, and feature retrieval to generate answers. Graphic and text feature extraction and alignment are composed of two network structures, UniversalSentenceEncoder and FiLMEfficientNet. Feature retrieval to generate answers is composed of a Transformer network structure, and there are at least twenty groups of Transformer networks. The UniversalSentenceEncoder network is used to encode the input text-form questions to generate an embedded representation in vector form, that is, extract the features in the text and convert them into vector representations, and this vector is denoted as USE(que); the FiLMEfficientNet network structure is based on the pre-trained EfficientNet, encodes the input image data, and adds a FiLM layer to fuse the encoded image and the encoded text feature vector USE(que) to obtain a Tokens representation that fuses image and text features; multiple Transformer networks take the Tokens that fuse graphic and text features as input, realize the retrieval of graphic and text features through the attention mechanism, and output text-form answers. Using the training data set described above, train this model and adjust the parameters of each network structure, so that the model can have the ability of graphic and text alignment, and retrieve and generate text answers from images.
[0026] S3: Model training Randomly select any training sample from the training data set in the constructed training data set, which includes: First, mark the question in text form as Q, the picture of the pest as P, and the answer in text form as A. Q and P are used as the input of the model, and A is used as the sample label; During training, input Q and P into the model simultaneously. Through the forward propagation of the model, an output containing model parameters is obtained. The distance between the wrongly output answer and the correct output A is the loss function. By minimizing the loss function through backpropagation, the model parameters are updated, adjusting the Universal Sentence Encoder network to tend to extract the number of pests as the main feature, and adjusting the FiLM EfficientNet network to tend to establish an association between the number of pests in the text and the same number of pests in the picture, and generating Tokens representations of this association to achieve the alignment of text and image features. Finally, through multiple Transformer network structures, the retrieval of the number of pests in the image is further realized. During training, samples are selected in batches using the minibatch method to complete the above forward calculation and backpropagation, and continuously update the model parameters until convergence.
[0027] S4: Model encapsulation Package the trained and optimized model, read the packaged file into memory and integrate it with the code into a dynamic link library. Finally, deploy the packaged model inside the forestry pest identification and popular science device based on text and image understanding, and perform real-time prediction or batch processing during use, and monitor the performance of the model.
[0028] The garden management department equips the machine with an APP as the user interface of the device. Tourists use their mobile phones to take pictures of pests, input the pictures and questions into the APP. The APP sends the input pictures and questions to the device through the Internet, and receives the answers returned by the device and presents them to the users.
[0029] The pest in vivo includes common pests such as American pests, Cerura menciana Moore, Stilpnotia salicis Linnaeus, Dioryctria rubella Hampson, Bactrocera dorsalis Hendel, Semiothisa cinerearia Bremer et Grey, Anoplophora glabripennis Motschulsky, Drosicha corpulenta Kuwana, Cnidocampa flavescens Walker, and Dendrolimus tabulaeformis Tsai et Liu. Through the above pests, most pests in the urban area can be comprehensively covered, thus greatly improving the accuracy of the model.
[0030] Step S3: During model training, cross-entropy loss is used for calculation. Cross-entropy loss is very sensitive to changes in probability values and can better guide the update of model parameters. In classification problems, it reflects the distance between the probability distribution predicted by the model and the probability distribution of the true labels. Therefore, the smaller the cross-entropy loss, the closer the two probability distributions are, and the higher the prediction accuracy of the model. When using mean squared error loss (MSE) as the loss function for classification problems, since MSE squares the predicted probability, when the predicted probability approaches 0 or 1, its gradient will tend to 0, resulting in the problem of gradient disappearance and affecting the training efficiency of the model. However, cross-entropy loss does not have this problem. It can always provide effective gradient information for the model, thus avoiding gradient disappearance and making the model training more stable. Cross-entropy loss calculates the loss by taking the logarithm, avoiding the numerical instability problems that may occur in some extreme cases, making the model more stable and reliable during the training process. The calculation method of the cross-entropy loss function is relatively simple and can be directly implemented through standard mathematical libraries. At the same time, its intuitiveness enables us to easily understand the difference between the model prediction and the actual situation, thus better optimizing the model. The cross-entropy loss function can handle multi-class classification problems well. By calculating the loss of each class separately and summing them to obtain the total loss, this gives it significant advantages in dealing with multi-classification problems.
[0031] Step S1: In constructing the training dataset, pest images are collected by taking photos at front, side, oblique side, back, top-down and bottom-up angles. At the same time, photos are taken at different distances through close-up, medium shot, long shot and extreme long shot, and are taken respectively under the light and shadow environments of natural light, hard light and soft light, artificial light, front light, side light and back light. By taking photos under various angles, distances and light and shadow environments, the recognition accuracy of the model for pests is improved.
[0032] Step S4: In model encapsulation, the model running environment is: Intel Core i9-10900X processor, NVIDIA GeForce RTX 3090 GPU, 128G of memory, and Python 3.6 and Pytorch 1.4.0 are deployed. Through the above hardware, the running speed of the device can be greatly improved, and then the user experience when using the model can be increased.
[0033] Step S1: In constructing the training dataset, the model collects at least 35,000 training samples. Each training sample consists of an image, question text and corresponding answer text. The image covers at least 16 common forestry pests in urban areas, and the size of each picture is 324x324. Through multiple groups of training samples, the recognition speed and accuracy of the model for pests are greatly improved.
[0034] Please refer to Figure 3, a device for forestry pest identification and popular science based on graphic and text understanding, comprising: a base 1, a telescopic arm 2, a host device 3, and a protective top cover 4. The telescopic arm 2 is assembled on the top of the base 1, the host device 3 is assembled on the front of the telescopic arm 2, and the protective top cover 4 is assembled on the top of the host device 3. A groove 31 is provided on the front of the host device 3. Ground nails 11 are evenly distributed on the top of the base 1. The base 1 can support the telescopic arm 2, enabling the telescopic arm 2 to be stably assembled in an outdoor open environment. The ground nails 11 can fix the base 1. The telescopic arm 2 can be telescoped, thereby changing the overall height of the device, facilitating the storage of the device when not in use and adjusting the height when in use. At the same time, the telescopic arm 2 is a common telescopic rod device in the prior art, with adjustable length and the ability to fix the adjusted length. The host device 3 is a common server device in the prior art, and also includes a communication module, a main board, a central processing unit, a memory, a hard disk, a RAID card, a network card, a power supply, and a cooling fan inside. The protective top cover 4 can protect the host device 3, preventing rainwater and the sun from affecting the service life of the internal components of the host device 3. The ground nails 11 can fixedly assemble the base 1 on the ground.
[0035] Embodiment 2 The difference between this Embodiment 2 and Embodiment 1 is: Please refer to Figure 4 , a control mechanism 5 is assembled on the top of the protective top cover 4. The control mechanism 5 includes a sensor group 51 and a top plate 53. The sensor group 51 is assembled on the top of the protective top cover 4, and the top plate 53 is assembled inside the protective top cover 4. Hoists 54 are connected to both sides of the top of the top plate 53, and the hoists 54 are connected to the sensor group 51 through wires. The output end of the hoist 54 is connected to a weight plate 55, and a protective belt 56 is connected between the weight plate 55 and the top plate 53. The sensor group 51 is a common sensor device in the prior art, used to control the signal amplifier 52 and the hoist 54 to work through rainwater. The hoist 54 is a common hoisting device in the prior art, used to drive the weight plate 55 to move. The protective belt 56 is a common elastic waterproof fabric in the prior art, preventing rainwater from damaging the host device 3. The weight of the weight plate 55 can pull the weight plate 55 to move. When the weight plate 55 returns to its position, it can be engaged with the protective top cover 4, thereby greatly increasing the overall length of the protective top cover 4. At the same time, the weight plate 55 cooperates with the top plate 53 to squeeze the protective belt 56, excluding external water sources.
[0036] The bottom of the sensor group 51 is equipped with a signal amplifier 52. The signal amplifier 52 is connected to the sensor group 51 through a circuit. The signal amplifier 52 is a common signal amplification device in the prior art. The switch of the signal amplifier 52 is controlled by the sensor group 51 through a circuit. In rainy weather, the signal amplifier 52 can increase the signal strength of the host device 3. At the same time, the signal amplifier 52 can be turned off on sunny days, thus avoiding affecting the mobile phone signals of outdoor visitors.
[0037] Drainage ports 57 are evenly distributed on the top of the weight plate 55. The inlet of the drainage port 57 is designed with a wide mouth, and the outlet of the drainage port 57 is designed to be inclined downward. And at least six groups of drainage ports 57 are provided. The drainage ports 57 can drain the water source discharged by the protective belt 56. The wide-mouth designed inlet is convenient for collecting the water source, and the inclined-designed outlet is convenient for discharging the collected water source. At the same time, multiple groups of drainage ports 57 can improve the drainage speed.
[0038] Please refer to Figure 2 , send a signal to the Internet connection model through the mobile phone APP, and input text and pictures. The model extracts and aligns the text and picture features, generates an answer through feature retrieval, and then outputs the answer to the Internet. Finally, it is transmitted to the user's mobile phone through the Internet. APP download module: When visitors enter the park, they can download and install the dedicated APP for free through the official park channels or QR code scanning. Photo upload module: During the tour, once visitors find pests, they can take pictures of the pests through the photo function of the APP and upload the pictures along with questions or descriptions to the system server. Instant feedback module: After receiving the photos and questions uploaded by visitors, the system server quickly gives corresponding answers or treatment suggestions through the built-in pest recognition algorithm or manual review. The corresponding process of one question and answer takes about 5 seconds to ensure that visitors can obtain feedback immediately. Data analysis module: The system server also has a data analysis function, which can statistically analyze the pest information uploaded by visitors and provide decision-making support for park management personnel.
[0039] In this solution, first, the base 1 is assembled outdoors in an open space so that there are no large objects around the device to block it. Fix and limit the position of the base 1 through the ground nails 11, adjust the height of the telescopic arm 2, and lock the height. Connect the host device 3 to power, detect and calibrate the electronic components inside the protective top cover 4, and paste the QR code connected to the inside of the host device 3 outside the groove 31. Users can scan the groove 31 to download the APP installed inside the host device 3. When entering the park for a tour, they can take pictures of the pests found at any time, upload the pictures and questions in the APP, and get corresponding answers. The corresponding process of one question and answer takes about 5 seconds.
[0040] During rainy weather, rain falls on the top of the sensor group 51, which then drives the sensor group 51 to move, controls the winch 54 to start, and at this time, the pulling on the weighted disk 55 is released. At this time, the weighted disk 55 descends by its weight, driving the protective belt 56 to cover the outside of the base 1, preventing rain from damaging the device. At the same time, the sensor group 51 controls the signal amplifier 52 to start, increasing the signal strength of the device. When the rain stops, the winch 54 is driven to bring the weighted disk 55 back to its original position. At this time, the protective belt 56 is squeezed, squeezing the rainwater outside it onto the top of the weighted disk 55, and finally discharging it through the drain port 57. At the same time, the signal amplifier 52 is turned off to prevent the signal amplifier 52 from continuously working and affecting the mobile phone signals of surrounding tourists.
[0041] Using the latest multi-modal understanding technology in the field of artificial intelligence, not only can popular science knowledge be provided to tourists to enhance the fun of the tour, but also first-hand information on the distribution of pests can be obtained based on the popular science interactions of tourists. By giving full play to the characteristics of wide distribution and large number of tourists, pest information can be discovered in a timely manner, saving a large amount of human and material resources in traditional maintenance work. While tourists are receiving popular science on forestry pests, park managers can obtain early warning information about the appearance of pests in the park, reducing the cost of pest prevention and control work. By taking advantage of the characteristics of wide distribution and large number of tourists in the park scenic area, popular science knowledge about pests is provided to users through an interactive Q&A method, while obtaining pest information in the park area, improving the efficiency of pest control work in the park scenic area, saving labor costs, enhancing the intelligent level of garden management, and promoting the development of urban intelligent construction and ecological civilization. During rainy weather, through the added control mechanism 5, the outside of the device is automatically protected from being damaged by rain. At the same time, through the added signal amplifier 52, the signal of the device is automatically amplified during rainy weather to prevent rain from affecting the signal transmission of the device and reducing the user experience. At the same time, when the control mechanism 5 returns to its original position, it can also discharge the external water source by squeezing, thus preventing the outside of the device from being continuously dripped with water by the control mechanism 5 and affecting the normal use of the device.
[0042] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0043] Although embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for identifying and popularizing forest pests based on graphic and text understanding, characterized by: include: S1: Construct training dataset First, live pests are collected and pest images are collected. The live pests are photographed with a camera according to their angle, distance, lighting and living status. The photos are then integrated and processed. Each photo is used as a sample. Multiple photos eventually form a pest sample library. The pest images are scaled and rotated. The pest photos are fused with species and background template images to ensure the random distribution of pests and diseases in the fused image and the authenticity of the pest and disease size at a specific distance. Text information is then manually added to each sample to eventually generate multiple training samples. S2: Model architecture for image-text feature alignment The model for image-text feature alignment consists of two parts: image-text feature extraction and alignment and feature retrieval and answer generation. Image-text feature extraction and alignment consists of two network structures: UniversalSentenceEncoder and FiLM EfficientNet. Feature retrieval and answer generation consists of a Transformer network structure, and there are at least twenty groups of Transformer networks. The UniversalSentenceEncoder network is used to encode the input textual questions and generate an embedding representation in the form of a vector, that is, to extract the features in the text and convert them into a vector representation, which is denoted as USE. The FiLM EfficientNet network structure is based on the pre-trained EfficientNet, encodes the input image data, and adds a FiLM layer to fuse the encoded image and the encoded text feature vector USE to obtain a Tokens representation that combines image and text features. Multiple Transformer networks use tokens that combine image and text features as input, retrieve image and text features through the attention mechanism, and output answers in text form. Using the training dataset, train this model and adjust the parameters of each network structure to enable the model to align images and texts, and retrieve and generate text answers from images. S3: Model training According to the training data set in the construction, any training sample is taken, including: First, the text question is marked as Q, the pest picture is marked as P, and the text answer is marked as A. Q and P are used as the input of the model, and A is used as the sample label; During training, Q and P are input into the model at the same time. After the forward propagation of the model, an output containing the model parameters is obtained. The distance between the wrong output answer and the correct output A is the loss function. By minimizing the loss function and backpropagating, the model parameters are updated, the UniversalSentenceEncoder network is adjusted to tend to extract the number of pests as the main feature, and the FiLM EfficientNet network is adjusted to tend to associate the number of pests in the text with the same number of pests in the picture, and this association is generated into Tokens to achieve alignment of image and text features. Finally, multiple Transformer network structures are used to further realize the retrieval of the number of pests in the image. During training, the minibatch method is used to select samples in batches to complete the above forward calculation and backpropagation, and the model parameters are continuously updated until convergence. S4: Model packaging The trained and optimized model is packaged, read into the memory and integrated with the code into the dynamic link library. Finally, the model is deployed inside the device for forestry pest identification and popular science based on graphic understanding, and real-time prediction or batch processing is performed during use, and the performance of the model is monitored.
2. The method for identifying and popularizing forest pests based on graphic and text understanding according to claim 1, characterized in that: The step S1: constructs a training data set, where a training sample is composed of three parts: an image, a question text, and a corresponding answer text.
3. The method for identifying and popularizing forest pests based on graphic and text understanding according to claim 1, characterized in that: The step S3: in model training, cross entropy loss calculation is adopted in model training.
4. The method for identifying and popularizing forest pests based on graphic and text understanding according to claim 1, characterized in that: The step S1: constructs a training data set, wherein pest images are collected by taking photos from the front angle, side angle, oblique side angle, back angle, and overhead and upward angles, and at the same time, photographing at different distances through close-up, near shot, mid shot and distant shot, and photographing under natural light, hard light and soft light, artificial light and front light, side light and back light respectively.
5. The method for identifying and popularizing forest pests based on graphic and text understanding according to claim 1, characterized in that: In the step S4: model packaging, the model operating environment is: Intel i9-10900X processor, NVIDIA GeForce RTX 3090 GPU, 128G memory and deployment of Python 3.6 and Pytorch 1.4.
0.
6. The method for identifying and popularizing forest pests based on graphic and text understanding according to claim 1, characterized in that: The step S1: constructs a training data set, in which the model collects at least 35,000 training samples, and the images cover at least 16 common forestry pests in urban areas.
7. A device for identifying and popularizing forest pests based on image and text understanding, comprising a method for identifying and popularizing forest pests based on image and text understanding as described in any one of claims 1 to 6, the device comprising: A base (1), a telescopic arm (2), a host device (3) and a protective top cover (4), characterized in that: the telescopic arm (2) is mounted on the top of the base (1), the host device (3) is mounted on the front of the telescopic arm (2), and the protective top cover (4) is mounted on the top of the host device (3), a groove (31) is provided on the front of the host device (3), the top of the base (1) is evenly distributed with ground nails (11), and the top of the protective top cover (4) is equipped with a control mechanism (5).
8. The device for identifying and popularizing forest pests based on graphic and text understanding according to claim 7, characterized in that: The control mechanism (5) comprises a sensor group (51) and a top plate (53), wherein the sensor group (51) is mounted on the top of the protective top cover (4), and the top plate (53) is mounted inside the protective top cover (4). Both sides of the top of the top plate (53) are connected to a winch (54), and the winch (54) and the sensor group (51) are connected via a line. The output end of the winch (54) is connected to a weighted plate (55), and a protective belt (56) is connected between the weighted plate (55) and the top plate (53).
9. The device for identifying and popularizing forest pests based on graphic and text understanding according to claim 8, characterized in that: A signal amplifier (52) is installed at the bottom of the sensor group (51), and the signal amplifier (52) is connected to the sensor group (51) via a line.
10. The device for identifying and popularizing forest pests based on graphic and text understanding according to claim 8, characterized in that: Drainage ports (57) are evenly distributed on the top of the weighted plate (55); the inlet of the drainage ports (57) is designed to be wide-mouthed; the outlet of the drainage ports (57) is designed to be inclined downward; and at least six groups of drainage ports (57) are provided.
Citation Information
Patent Citations
Automatic identification method and device for forestry diseases and insect pests
CN116188872A
Disease and pest knowledge generation type question answering method based on big and small model collaboration
CN118551845A
Crop disease visual question-answering method and system, computer equipment and medium
CN118779834A
Agricultural pest control method fusing computer vision and language model
CN119443286A
Rainproof structure of switch cabinet body
CN220754042U