Weed identification method, device, equipment and storage medium
Patent Information
- Application Number
- CN202211348333.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-10-31
AI Technical Summary
但是,上述方法识别的杂草种类十分有限,使得杂草的识别效果较差
[0020] The process involves acquiring weed images and descriptions of all weeds within those images. Based on a pre-built weed category database and the descriptions, the weed categories in the images are identified to obtain a first predicted weed category. Then, based on the weed images and descriptions, a second predicted weed category is obtained. This process predicts the category of all weeds in the images from multiple dimensions. Finally, the first and second predicted weed categories from different dimensions are fused to more accurately identify the weed categories.
Smart Images

Figure CN115761481B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular to a method, apparatus, device and storage medium for identifying weeds. Background Technology
[0002] With the advent of precision agriculture, weed identification and control technologies in fields are gradually developing towards mechanization and intelligence. Accurate weed identification is a core technology for improving weed control accuracy and efficiency. Currently, digital image processing and spectral characteristic analysis have become common methods for weed identification. However, these methods can only identify a limited number of weed species, resulting in poor identification effectiveness. Summary of the Invention
[0003] To address the aforementioned problems, this application proposes a method, apparatus, device, and storage medium for weed identification, which can effectively improve the weed identification results.
[0004] According to a first aspect of the embodiments of this application, a method for identifying weeds is provided, comprising:
[0005] Acquire images of weeds and descriptions of all weeds in the images;
[0006] Based on a pre-built weed category database and the weed description information, the weed category in the weed image is predicted to obtain a first predicted weed category;
[0007] Based on the weed image and the weed description information, weed identification processing is performed to obtain a second predicted weed category;
[0008] The weed identification result is determined based on the first predicted weed category and the second predicted weed category.
[0009] According to a second aspect of the embodiments of this application, a weed identification device is provided, comprising:
[0010] The acquisition module is used to acquire weed images and weed description information of all weeds in the weed images;
[0011] The first processing module is used to predict the weed category in the weed image based on a pre-built weed category database and the weed description information, and obtain a first predicted weed category.
[0012] The second processing module is used to perform weed identification processing based on the weed image and the weed description information to obtain a second predicted weed category;
[0013] The identification module is used to determine the weed identification result based on the first predicted weed category and the second predicted weed category.
[0014] A third aspect of this application provides an electronic device, comprising:
[0015] Memory and processor;
[0016] The memory is connected to the processor and is used to store programs;
[0017] The processor implements the aforementioned weed identification method by running the program in the memory.
[0018] The fourth aspect of this application provides a storage medium storing a computer program, which, when run by a processor, implements the aforementioned weed identification method.
[0019] One embodiment of the above application has the following advantages or beneficial effects:
[0020] The process involves acquiring weed images and descriptions of all weeds within those images. Based on a pre-built weed category database and the descriptions, the weed categories in the images are identified to obtain a first predicted weed category. Then, based on the weed images and descriptions, a second predicted weed category is obtained. This process predicts the category of all weeds in the images from multiple dimensions. Finally, the first and second predicted weed categories from different dimensions are fused to more accurately identify the weed categories. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating a method for identifying weeds according to an embodiment of this application;
[0023] Figure 2 This is a schematic diagram combining relevant problem information with weed images and weed description information according to another embodiment of this application;
[0024] Figure 3 This is a flowchart illustrating a weed identification method according to another embodiment of this application;
[0025] Figure 4 This is a flowchart illustrating a weed identification method according to another embodiment of this application;
[0026] Figure 5This is a schematic diagram illustrating the fusion of image features and semantic features at multiple scales according to an embodiment of this application;
[0027] Figure 6 This is a block diagram illustrating the generation of a second predicted weed category according to an embodiment of this application;
[0028] Figure 7 This is a schematic diagram illustrating the generation of image features at multiple scales according to an embodiment of this application;
[0029] Figure 8 This is a block diagram of a weed identification device according to another embodiment of this application;
[0030] Figure 9 This is a block diagram of an electronic device used to implement the weed identification method of the embodiments of this application. Detailed Implementation
[0031] The technical solutions of this application are applicable to various image recognition scenarios, such as human-computer interaction. Using the technical solutions of this application, weed categories can be identified more accurately.
[0032] The technical solutions of this application can be exemplarily applied to hardware devices such as processors, electronic devices, and servers (including cloud servers), or packaged as software programs and run. When the hardware device executes the processing procedure of the technical solutions of this application, or when the aforementioned software program is run, it is possible to predict the category of all weeds in a weed image from multiple dimensions, thereby obtaining the purpose of weed identification results. This application only provides exemplary descriptions of the specific processing procedure of the technical solutions of this application and does not limit the specific implementation form of the technical solutions of this application. Any technical implementation form that can execute the processing procedure of the technical solutions of this application can be adopted by this application.
[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0034] Exemplary methods
[0035] Figure 1 This is a flowchart of a weed identification method according to an embodiment of this application. In an exemplary embodiment, a weed identification method is provided, including:
[0036] S110. Obtain weed images and weed description information for all weeds in the weed images;
[0037] S120. Based on the pre-constructed weed category database and the weed description information, predict the weed category in the weed image to obtain the first predicted weed category;
[0038] S130. Based on the weed image and the weed description information, perform weed identification processing to obtain a second predicted weed category;
[0039] S140. Determine the weed identification result based on the first predicted weed category and the second predicted weed category.
[0040] In step S110, exemplarily, the weed image is used to represent an image containing weeds, which can be captured by a camera on any device. Weed description information is used to represent relevant information about the weeds in the weed image. For example, spatial information about the weeds, temporal information about the weeds, weed distribution information, and visual characteristic information about the weeds. The spatial information of the weeds can be represented by geographical location, such as latitude and longitude. The temporal information of the weeds can be represented by the season information (e.g., spring, summer, autumn, winter) or a specific date (e.g., year, month, day). The weed distribution information can be represented by information about the plants growing around the weeds in the image. The visual characteristic information of the weeds is used to represent the content visible to the eye in the weed image, such as the appearance of the weeds, leaf shape, stem, growing location, and height.
[0041] Optionally, the weed description information can be obtained by staff manually annotating the weed images. Optionally, the weed description information can be obtained by staff describing all the weeds in the weed image based on text or voice. Optionally, the weed description information can also be obtained by combining the information annotated when the weed images were taken (such as location information) with the staff's description of all the weeds in the weed image.
[0042] In this embodiment, after acquiring the weed image, if the image contains only one type of weed, a single-label is applied without marking the target bounding box or outline points. If the image contains multiple types of weeds, multi-labeling is applied, and all weed categories in the image are labeled. The weeds seen in the image are described in text or voice, mainly describing their appearance, leaf shape, stem, growth location, height, and other relevant information, thus outputting visual characteristic information of the weeds. When the weeds are photographed in a field, the crop information needs to be labeled on the weed image; when the weeds are photographed outside a field, the crops planted in the nearest field or some common plants in the surrounding area need to be labeled on the weed image, thus outputting weed distribution information. This allows for more accurate identification of weed categories by observing the plants growing alongside the weeds.
[0043] When acquiring weed images, the shooting location of the weed images is simultaneously extracted and its specific location is marked. Alternatively, the location of the weed images can be marked using the map location of the shooting terminal device, thus outputting spatial information. The season in which the weed images were taken is marked, which can be in the form of spring, summer, autumn, or winter, or in the format of year, month, and day, thus outputting time information. Then, weed descriptive information is determined based on the visual characteristics, time information, spatial information, and weed distribution information. Optionally, keywords can be extracted from the visual characteristics, time information, spatial information, and weed distribution information of the weeds, and the extracted keywords can be used as weed descriptive information. The keyword extraction method can be through a pre-trained keyword model, such as a TF-IDF model (Term Frequency–Inverse Document Frequency) combined with a text keyword extraction model (Text Rank).
[0044] In step S120, for example, in order to better identify uncommon weeds, a weed category database is pre-constructed. This database stores weed types, weed descriptions, and the corresponding relationships between them. Optionally, the data in the weed category database can be crawled from various web pages, drawn from staff knowledge reserves, or obtained from a combination of multiple sources. The data source for the weed category database is not limited here, allowing for a more comprehensive database. Optionally, the weed category database can contain multiple weed category data, stored in the form of weed name-alias-geographical location (spatial information)-seasonal information (temporal information)-accompanying crops (weed distribution information)-main characteristics (leaf shape, stem, growth location, height, etc.).
[0045] Specifically, the similarity between the weed description information and the data of each weed category in the weed category database is calculated. Based on the similarity, weed categories that are impossible to appear and weed categories that are possible to appear are determined. The weed categories that are possible to appear are taken as the first predicted weed categories.
[0046] In this embodiment, a similarity threshold can be preset. The similarity of each weed category data in the weed category database with the weed description information is compared to obtain the weed similarity score. The weed similarity score is then compared to the similarity threshold. If it is not less than the threshold, the weed category data is considered a possible weed category; if it is less than the threshold, the weed category data is considered an impossible weed category. The possible weed categories are then sorted according to their similarity, and the sorted result is used as the first predicted weed category.
[0047] In step S130, for example, the weed image and weed description information are fused together, and the corresponding second predicted weed category is determined based on the fusion result. In this way, the weed category can be predicted more accurately by combining image and text information.
[0048] Specifically, the neural network can be pre-trained using training images of weeds, along with corresponding weed descriptions and categories, to obtain a trained neural network. Then, inputting the weed images and descriptions into the trained neural network will yield a second predicted weed category. This second predicted weed category can include multiple weed categories.
[0049] In step S140, optionally, the set of the first predicted weed category and the second predicted weed category can be taken, and all the predicted weed categories can be used as the final weed identification result.
[0050] Preferably, step S140 includes: using the first predicted weed category to filter the second predicted weed category, and determining the filtered second predicted weed category as the weed identification result.
[0051] Specifically, the intersection of the first predicted weed category and the second predicted weed category is determined, i.e., the same predicted weed category in both categories, and this same predicted weed category is taken as the weed identification result. As the above analysis shows, since the first predicted weed category is determined based on a weed category database, it can be guaranteed that weed categories not present in the first predicted weed category are definitely not present in the weed image. Thus, filtering the second weed category using the first predicted weed category further ensures the accuracy of the identified weed categories.
[0052] Furthermore, if the first predicted weed category and the second predicted weed category are both different, the first predicted weed category is determined as the weed identification result.
[0053] Specifically, the first predicted weed category is used to filter the second predicted weed category. If the first predicted weed category is completely different from the second predicted weed category, it indicates that all weed categories in the second predicted weed category are incorrect weed categories. In this case, the first predicted weed category is output as the weed identification result, thus ensuring the accuracy of the output weed identification result. Furthermore, the first predicted weed category can be further filtered. For example, a threshold can be set to filter the first predicted weed category, selecting weed species in the first predicted weed category that exceed the threshold as the final weed identification result, thereby further improving the accuracy of the weed identification result.
[0054] In the technical solution of this application, weed images and weed descriptions of all weeds in the images are acquired. Based on a pre-built weed category database and the weed descriptions, the weed categories in the images are identified to obtain a first predicted weed category. Then, based on the weed images and weed descriptions, the weed categories in the images are identified again to obtain a second predicted weed category. In this way, the categories of all weeds in the images are predicted from multiple dimensions, and the first and second predicted weed categories from different dimensions are then fused to more accurately identify the weed categories.
[0055] In some embodiments, the step of performing weed identification processing based on the weed image and the weed description information to obtain a second predicted weed category includes:
[0056] The relevant problem information of the weeds, the weed images and the weed description information are subjected to feature interaction, and a second predicted weed category is determined based on the feature interaction results;
[0057] The relevant information about weeds includes questions about the categories of weeds in the weed images.
[0058] For example, the relevant questions about weeds are pre-set questions. Each time weed identification is performed, the input questions can be the same or different. The relevant questions can include one question or multiple questions. For example, the relevant questions could be "What is the weed category?" or "How many weed categories are there?", or "What is the weed category and how many weed categories are there?". In this embodiment, the relevant questions about weeds are the same, specifically including "What is the weed category?" and "How many weed categories are there?". Therefore, because the relevant questions about weeds include weed categories and / or the number of categories, information about the species can be extracted more accurately from the weed description information and weed images. Furthermore, because the relevant questions interact with the weed images and weed description information, weed identification can also be achieved through voice or text interaction.
[0059] Optionally, the neural network model can be trained based on relevant question information about weeds (i.e., "What are the weed categories?" and "How many categories of weeds are there?"), weed categories, training images of weeds, and their corresponding descriptive information. During model training, the weed images, descriptive information, and input question information about weeds interact, allowing the model to learn the correlation between these two elements and acquire more important features for weed identification. This results in more accurate weed predictions using the trained neural network model.
[0060] Thus, after acquiring weed images, weed descriptions, and related question information, the system searches a weed category database for possible weed categories based on the weed descriptions to generate a first predicted weed category. The related question information, weed images, and weed descriptions are then input into a trained neural network model to obtain a second predicted weed category. Finally, the first and second predicted weed categories are fused to obtain the weed identification result.
[0061] Preferably, such as Figure 3 As shown, feature interaction is performed on the acquired weed-related problem information, the weed image, and the weed description information, and a second predicted weed category is determined based on the feature interaction results, including:
[0062] S310. The weed image and the weed description information are fused to obtain weed image-text fusion features;
[0063] S320. Based on the correlation between the weed image-text fusion features and the relevant problem information of the weeds, determine the second predicted weed category.
[0064] Specifically, features can be extracted by fusing the weed image and the weed description information to obtain weed image-text fusion features. Alternatively, features can be extracted from the weed image and the weed description information separately, and then fused to obtain weed image-text fusion features. Optionally, features can be extracted from the weed image and the weed description information separately using different models, or features can be extracted from the weed image and the weed description information separately using the same model. For example, features of the weed image can be extracted using two superimposed Swin transformer models, and features of the weed description information can be extracted using a transformer model. Another example is extracting features from the weed image and the weed description information separately using a transformer model.
[0065] After determining the weed image-text fusion features, the correlation between these features and relevant information about weeds is determined. Weed image-text fusion features with a correlation higher than a correlation threshold are used as the second predicted weed category, thereby enabling the identification of all weed categories in the weed image. The correlation threshold is set according to actual needs and is not limited here.
[0066] In some implementations, such as Figure 4 As shown, the step of performing weed identification processing based on the weed image and the weed description information to obtain a second predicted weed category includes:
[0067] S410. Perform feature fusion on the weed image and the weed description information to obtain weed image-text fusion features;
[0068] S420. Determine the second predicted weed category based on the weed image-text fusion features.
[0069] For example, features can be extracted after fusing the weed image and the weed description information to obtain the weed image-text fusion feature. Alternatively, features can be extracted from the weed image and the weed description information separately, and then the extracted features can be fused to obtain the weed image-text fusion feature.
[0070] The neural network model can be pre-trained based on the weed image-text fusion features corresponding to the weed training images and the weed category, resulting in a trained neural network model. Then, by inputting the weed image-text fusion features into the trained neural network model, a second predicted weed category can be obtained, thereby achieving weed category identification.
[0071] In some implementations, the weed image and the weed description information are fused to obtain weed image-text fusion features, including:
[0072] Multiple scale image features were extracted from the weed image;
[0073] Extract semantic features from the weed description information;
[0074] The image features at each of the multiple scales are fused with the semantic features to obtain the weed image-text fusion feature.
[0075] For example, image features at multiple scales are used to represent the image features corresponding to weed images of different sizes. Specifically, feature extraction of the weed image can be performed using Feature Pyramid Networks (FPN) to obtain image features at multiple scales. Furthermore, to make the extracted image features at multiple scales more accurate, preliminary feature extraction can be performed on the weed image before extracting the image features at multiple scales, and then the image features at multiple scales can be extracted based on the preliminary features.
[0076] For example, since the weed description information may include information related to multiple weeds, a unified feature extraction can be performed on the weed description information first based on a semantic segmentation model. Alternatively, the relevant information in the weed description information can be extracted separately, and then the multiple features can be fused to obtain a fused feature. Finally, the semantic features can be extracted from the fused feature.
[0077] In this embodiment, as Figure 5As shown, image features at each scale of the multi-scale image features are fused with semantic features. The multi-scale image features are four different scales: C2, C3, C4, and C5. The extracted semantic features are fused with these four scales of image features respectively. Semantic features show different levels of attention to image features at different scales, meaning that image features at different scales have different importance to semantic features. Therefore, semantic features and image features at different scales are input into different cross-attention modules. Specifically, the input to the first attention mechanism (Cross Attention 1) is C5 and semantic features; the input to the second attention mechanism (Cross Attention 2) is C4 and semantic features; the input to the third attention mechanism (Cross Attention 3) is C3 and semantic features; and the input to the fourth attention mechanism (Cross Attention 4) is C2 and text features. This yields four different sizes of text-image fusion features, which are then mapped to a unified scale through a 1*1 convolution and finally merged according to the channel dimension to obtain a new fusion feature NF. To enhance the interaction between different channels, i.e., between different scales, global pooling is used to obtain 1*1*C, which is then passed through a multilayer fully connected neural network model (Multilayer Perceptron, MLP). Activation functions (such as sigmoid) are then connected to obtain the weight values of different channels. Finally, the fusion feature NF is multiplied to obtain the weed image-text fusion feature TIF.
[0078] Furthermore, such as Figure 6 As shown, to better determine the correlation between the weed image-text fusion features and the related question information about weeds, when acquiring the related question information about weeds, an encoder is used to encode the relevant question information to obtain a question encoding vector. The encoder can use a word2vec model. Since the length of the related question information about weeds is fixed, the sliding window size in the word2vec model can be directly set to better extract the semantic information from the related question information. Then, feature extraction is performed on the question encoding vector to obtain the related question features TC. Specifically, feature extraction can be performed using a transformer model, or other models can be used; no limitation is made here.
[0079] The relevant problem feature TC and the weed image-text fusion feature TIF are interacted. The TIF is compressed into a single channel, and then attention is calculated for each row of TC and TIF. The attention calculation includes dot product and concatenation of each row of TC and TIF. Then, an activation function (such as softmax) is used to normalize the calculated results to obtain the relevance scores of different rows of features in TIF to the number and types of weed categories. Row features with relevance scores greater than a relevance threshold are then extracted. The above row features are merged and then classified through a fully convolutional and softmax layer to obtain a multi-classified weed category result, which is denoted as the second predicted weed category.
[0080] Preferably, multiple scale image features are extracted from the weed image, including:
[0081] Extract the corresponding local features and corresponding global features from the weed image;
[0082] The local features and the global features are fused to obtain the image features at multiple scales.
[0083] For example, local features are used to represent the details of weeds in a weed image, while global features are used to represent the weeds as a whole in the image. Local and global features can be extracted using different methods; for example, a Convolutional Neural Network (CNN) can be used to extract local features, while a Swin transformer model can be used to extract global features. Therefore, extracting both the corresponding local and global features simultaneously from a weed image results in a more comprehensive feature extraction process.
[0084] In this embodiment, as Figure 7 The weed image is first processed by two superimposed Swin Transformer models for preliminary feature extraction. These preliminary features are then input into two branches for further feature extraction. One branch uses multiple CNNs combined with a First-Party Network (FPN) to extract features at different scales; that is, the preliminary features are passed through multiple CNNs before being input into the FPN. The CNNs include convolutional layers, batch normalization (BN) layers, and activation layers. The other branch uses the backbone network of the Swin Transformer model combined with the FPN to extract features at different scales; that is, the preliminary features are passed through the backbone network of the Swin Transformer model before being input into the FPN. Finally, the local and global features obtained from these two branches are concatenated to output multi-scale image features.
[0085] In some embodiments, the weed description information includes: time information, spatial information, weed distribution information, and visual characteristic information;
[0086] Accordingly, the extraction of semantic features from the weed description information includes:
[0087] The time information, the spatial information, and the weed distribution information are encoded to obtain a first encoding result;
[0088] The visual feature information is encoded to obtain a second encoding result;
[0089] The first encoding result and the second encoding result are merged to obtain the third encoding result;
[0090] The semantic features are extracted based on the third encoding result.
[0091] Specifically, the temporal information, spatial information, and weed distribution information are encoded by an encoder to obtain a first encoding result. The encoder can employ a global vector model (GloVe). In this embodiment, the temporal information, spatial information, and weed distribution information can be treated as a set. The local and global information vectors of this set are extracted using a global vector model, and then concatenated to obtain the first encoding result.
[0092] The visual feature information is encoded by an encoder to obtain text feature encoding (i.e., the second encoding result). The text feature encoding is a semantic information vector with low-dimensional dense features. The encoder can employ a word2vec model. The first and second encoding results are merged and then processed by a feature extractor to extract features. This yields multi-dimensional fused semantic features, which is helpful for subsequent weed category prediction. Optionally, the feature extractor can employ semantic segmentation models, attention models, etc., such as the Transformer model. The self-attention mechanism in the Transformer model extracts key features from long visual feature information while effectively utilizing the contextual sequence information in the text.
[0093] Exemplary device
[0094] Correspondingly, Figure 8 This is a schematic diagram of a weed identification device according to an embodiment of this application. In an exemplary embodiment, a weed identification device is provided, comprising:
[0095] The acquisition module 810 is used to acquire weed images and weed description information of all weeds in the weed images;
[0096] The first processing module 820 is used to predict the weed category in the weed image based on a pre-built weed category database and the weed description information, and obtain a first predicted weed category.
[0097] The second processing module 830 is used to perform weed identification processing based on the weed image and the weed description information to obtain a second predicted weed category;
[0098] The identification module 840 is used to determine the weed identification result based on the first predicted weed category and the second predicted weed category.
[0099] In some embodiments, the second processing module 830 includes:
[0100] The feature interaction module is used to perform feature interaction on the acquired weed-related problem information, the weed image and the weed description information, and determine the second predicted weed category based on the feature interaction results;
[0101] The information related to weeds includes questions about the types of weeds in the weed images.
[0102] In some embodiments, the second processing module 830 includes:
[0103] The first feature fusion module is used to fuse the weed image and the weed description information to obtain weed image-text fusion features;
[0104] The first prediction module is used to determine the second predicted weed category based on the weed image-text fusion features.
[0105] In some implementations, the feature interaction module includes:
[0106] The second feature fusion module is used to fuse the weed image and the weed description information to obtain weed image-text fusion features.
[0107] The second prediction module is used to determine the second predicted weed category based on the correlation between the weed image-text fusion features and the relevant problem information of the weeds.
[0108] In some implementations, the weed image and the weed description information are fused to obtain weed image-text fusion features, including:
[0109] Multiple scale image features were extracted from the weed image;
[0110] Extract semantic features from the weed description information;
[0111] The image features at each of the multiple scales are fused with the semantic features to obtain the weed image-text fusion feature.
[0112] In some embodiments, the extraction of image features at multiple scales from the weed image includes:
[0113] Extract the corresponding local features and corresponding global features from the weed image;
[0114] The local features and the global features are fused to obtain the image features at multiple scales.
[0115] In some embodiments, the weed description information includes: time information, spatial information, weed distribution information, and visual characteristic information;
[0116] Accordingly, the extraction of semantic features from the weed description information includes:
[0117] The time information, the spatial information, and the weed distribution information are encoded to obtain a first encoding result;
[0118] The visual feature information is encoded to obtain a second encoding result;
[0119] The first encoding result and the second encoding result are merged to obtain the third encoding result;
[0120] The semantic features are extracted based on the third encoding result.
[0121] In some embodiments, the identification module 840 is further configured to:
[0122] The first predicted weed category is used to filter the second predicted weed category, and the filtered second predicted weed category is determined as the weed identification result.
[0123] In some embodiments, the apparatus further includes:
[0124] The judgment module is used to determine the first predicted weed category as the weed identification result when the first predicted weed category and the second predicted weed category are both different.
[0125] The weed identification device provided in this embodiment belongs to the same concept as the weed identification method provided in the above embodiments of this application. It can execute the weed identification method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the weed identification method. Technical details not described in detail in this embodiment can be found in the specific processing content of the weed identification method provided in the above embodiments of this application, and will not be repeated here.
[0126] Exemplary electronic devices
[0127] Another embodiment of this application also provides an electronic device, see [link to relevant documentation] Figure 9 As shown, the device includes:
[0128] Memory 900 and processor 910;
[0129] The memory 900 is connected to the processor 910 and is used to store programs;
[0130] The processor 910 is used to implement the weed identification method disclosed in any of the above embodiments by running the program stored in the memory 900.
[0131] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 920, an input device 930, and an output device 940.
[0132] The processor 910, memory 900, communication interface 920, input device 930, and output device 940 are interconnected via a bus. Among them:
[0133] A bus can include a pathway for transmitting information between various components of a computer system.
[0134] The processor 910 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0135] The processor 910 may include a main processor, as well as a baseband chip, modem, etc.
[0136] The memory 900 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 900 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.
[0137] Input device 930 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.
[0138] Output device 940 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0139] The communication interface 920 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0140] The processor 910 executes the program stored in the memory 900 and calls other devices, which can be used to implement the various steps of any of the weed identification methods provided in the above embodiments of this application.
[0141] Exemplary computer program products and storage media
[0142] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the weed identification methods according to various embodiments of this application described in the "Exemplary Methods" section of this specification.
[0143] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0144] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor in the steps of the weed identification methods according to various embodiments of this application described in the "Exemplary Methods" section above.
[0145] The specific working content of the aforementioned electronic device, as well as the specific working content of the aforementioned computer program product and the computer program on the storage medium being run by the processor, can all be found in the content of the aforementioned method embodiments, and will not be repeated here.
[0146] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0147] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0148] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.
[0149] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.
[0150] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0151] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.
[0152] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.
[0153] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0154] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0155] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0156] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying weeds, characterized in that, include: Acquire weed images and weed description information for all weeds in the weed images. The weed description information is obtained by describing all weeds in the weed images and includes: time information, spatial information, weed distribution information, and visual characteristic information. The time information, spatial information, and weed distribution information are encoded to obtain a first encoding result; the visual feature information is encoded to obtain a second encoding result; the first encoding result and the second encoding result are merged to obtain a third encoding result; semantic features are extracted based on the third encoding result. Based on a pre-built weed category database and the weed description information, the weed category in the weed image is predicted to obtain a first predicted weed category; Based on the semantic features corresponding to the weed image and the weed description information, weed identification processing is performed to obtain a second predicted weed category; The weed identification result is determined based on the first predicted weed category and the second predicted weed category.
2. The method according to claim 1, characterized in that, The step of performing weed identification processing based on the semantic features corresponding to the weed image and the weed description information to obtain a second predicted weed category includes: The semantic features corresponding to the weed image and the weed description information are fused to obtain the weed image-text fusion feature; The second predicted weed category is determined based on the weed image-text fusion features.
3. The method according to claim 2, characterized in that, Based on the semantic features corresponding to the weed image and the weed description information, weed identification processing is performed to obtain a second predicted weed category, including: The semantic features corresponding to the weed image and the weed description information are fused to obtain the weed image-text fusion feature; Based on the correlation between the weed image-text fusion features and the relevant question information of the weeds, the second predicted weed category is determined; wherein, the relevant question information of the weeds includes question information that asks about the weed category in the weed image.
4. The method according to claim 2, characterized in that, The semantic features corresponding to the weed image and the weed description information are fused to obtain weed image-text fusion features, including: Multiple scale image features were extracted from the weed image; The image features at each of the multiple scales are fused with the semantic features to obtain the weed image-text fusion feature.
5. The method according to claim 4, characterized in that, The extraction of image features at multiple scales from the weed image includes: Extract the corresponding local features and corresponding global features from the weed image; The local features and the global features are fused to obtain the image features at multiple scales.
6. The method according to claim 1, characterized in that, The step of determining the weed identification result based on the first predicted weed category and the second predicted weed category includes: The first predicted weed category is used to filter the second predicted weed category, and the filtered second predicted weed category is determined as the weed identification result.
7. The method according to claim 6, characterized in that, The method further includes: If the first predicted weed category and the second predicted weed category are both different, the first predicted weed category is determined as the weed identification result.
8. A weed identification device, characterized in that, include: The acquisition module is used to acquire weed images and weed description information of all weeds in the weed images. The weed description information is obtained by describing all weeds in the weed images and includes: time information, spatial information, weed distribution information and visual characteristic information. A first processing module is used to encode the time information, the spatial information, and the weed distribution information to obtain a first encoding result; encode the visual feature information to obtain a second encoding result; merge the first encoding result and the second encoding result to obtain a third encoding result; extract semantic features based on the third encoding result; and predict the weed category in the weed image according to a pre-constructed weed category database and the weed description information to obtain a first predicted weed category. The second processing module is used to perform weed identification processing based on the weed image and the weed description information to obtain a second predicted weed category; The identification module is used to determine the weed identification result based on the first predicted weed category and the second predicted weed category.
9. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor, by running the program in the memory, implements the weed identification method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a computer program, which, when run by a processor, implements the weed identification method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Image recognition method, device and equipment and storage medium
CN111178349A
Target detection method and device
CN113837257A
Visual question and answer method based on deep reasoning attention mechanism
CN114398471A