Cordyceps sinensis identification method and system based on LLaVA multi-mode large model
Through the Cordyceps recognition method based on the LLaVA multimodal large model, combined with image and text features, the problem of semantic information loss in traditional Cordyceps recognition technology is solved, and high-precision fine-grained Cordyceps recognition and automated report generation is achieved.
Patent Information
- Application Number
- CN202510448736.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
The existing Cordyceps recognition technology methods rely on visual features and cannot integrate semantic background information. The recognition accuracy is low, and the model has poor adaptability to Cordyceps appearance and the difference in fine-grained appearance.
Using the LLaVA multimodal large model, multi-angle images of Cordyceps are collected through image acquisition equipment and equipped with text descriptions. After data preprocessing, the LLaVA model is used to extract image features and process text descriptions, and feature fusion is performed in combination with the multimodal reasoning mechanism, and finally fine-grained classification and classification report are generated.
It significantly improves the accuracy and stability of Cordyceps recognition, enhances the robustness and generalization capabilities of the model, realizes automatic classification report generation, reduces manual judgment dependence, and improves the visualization and user-friendliness of the system.
Smart Images

Figure CN120298799A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision, natural language processing, and artificial intelligence, and specifically provides a cordyceps identification method and system based on the LLaVA multimodal large model. Background Art
[0002] Cordyceps (such as Ophiocordyceps sinensis, etc.) is a high-value traditional Chinese medicine, and its quality directly determines its economic value and medicinal effects. However, due to the large variety of cordyceps, the appearance differences between different species are very subtle (such as characteristics like color, texture, annulations, etc.). Traditional manual classification and identification methods face the following problems: Manual identification relies on experienced experts, but the judgments of experts are often affected by subjective factors, and the stability of the identification results is insufficient; Manual classification requires checking samples one by one, with low efficiency and difficulty in meeting the requirements of large-scale commercial production; Most existing automated identification systems are based on single image feature extraction methods, making it difficult to distinguish cordyceps species with similar morphologies, especially in fine-grained recognition tasks, the effect is not ideal.
[0003] In recent years, the rise of deep learning technology has provided new ideas for cordyceps identification. For example, visual features in cordyceps images are extracted through convolutional neural networks (CNNs), and support vector machines or fully connected neural networks are used for classification. However, these traditional visual models still have the following limitations: Existing methods only use image features as input, ignoring a large amount of potential semantic auxiliary information (such as the origin, growth environment, color description, etc. of cordyceps), resulting in insufficient judgment ability of the model when facing small differences between species; When there are changes in image acquisition conditions such as lighting, angle, occlusion, etc., the generalization ability and robustness of the model decrease significantly; When there are only extremely small appearance differences between cordyceps species (such as color gradation, texture density, annulation distribution), traditional image models often cannot effectively capture and utilize these differences for high-precision classification. Existing cordyceps identification technologies have obvious deficiencies in terms of identification accuracy, stability, and practicality, and there is an urgent need for a cordyceps identification method that integrates multi-source information, has strong generalization ability, and high robustness. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the technical problems solved by the present invention are: Existing cordyceps identification technology methods have problems of relying on visual features and being unable to integrate semantic background information, with low identification accuracy, poor adaptability of the model to fine-grained differences in cordyceps appearance, and the problem of how to deeply integrate image and text features to achieve multi-modal high-precision identification.
[0006] To solve the above technical problems, the present invention provides the following technical solution: A Cordyceps identification method based on the LLaVA multimodal large model, including collecting Cordyceps data and preprocessing the Cordyceps data; establishing a feature extraction algorithm to extract and fuse features from the preprocessed Cordyceps data; classifying Cordyceps according to the fusion result and generating a Cordyceps classification report.
[0007] As a preferred embodiment of the Cordyceps identification method based on the LLaVA multimodal large model of the present invention, wherein: the collection of Cordyceps data includes collecting images of different perspectives of Cordyceps through an image acquisition device, and each image is equipped with a text description.
[0008] The text description includes the variety name, collection location, ecological environment, and color characteristics of Cordyceps.
[0009] As a preferred embodiment of the Cordyceps identification method based on the LLaVA multimodal large model of the present invention, wherein: the preprocessing of the Cordyceps data includes denoising, enhancing, normalizing, and image segmentation of the Cordyceps images.
[0010] As a preferred embodiment of the Cordyceps identification method based on the LLaVA multimodal large model of the present invention, wherein: the feature extraction algorithm includes constructing a feature extraction algorithm based on the joint reasoning of Cordyceps images and text descriptions, extracting image features through the LLaVA model, processing text description information, the LLaVA model extracts visual features from Cordyceps images through a convolutional neural network, and uses a pre-trained language model to parse the text description and extract the variety name, color characteristics, and ecological environment of Cordyceps from the text description.
[0011] As a preferred embodiment of the Cordyceps identification method based on the LLaVA multimodal large model of the present invention, wherein: the extraction and fusion of features from the preprocessed Cordyceps data includes using the multimodal reasoning mechanism of the LLaVA model to fuse the features of Cordyceps images and text descriptions, and output a joint feature vector, which includes the variety, external shape features, and fine-grained features of Cordyceps.
[0012] As a preferred embodiment of the Cordyceps identification method based on the LLaVA multimodal large model of the present invention, wherein: the classification of Cordyceps according to the fusion result includes performing fine-grained classification on Cordyceps according to the fused feature vector, generating a variety classification result of Cordyceps, and the classification result includes the variety name, external shape features, and color information of Cordyceps.
[0013] As a preferred embodiment of the Cordyceps identification method based on the LLaVA multi-modal large model of the present invention, wherein: the generation of the Cordyceps classification report includes generating a Cordyceps variety report according to the classification result, and the Cordyceps variety report includes the classification label and text description of the Cordyceps, and feeding back the classification result to the user.
[0014] Another object of the present invention is to provide a Cordyceps identification system based on the LLaVA multi-modal large model, which can jointly reason and fuse image and text features, and solves the problems of missing semantic information and weak fine-grained recognition ability in the current Cordyceps image recognition technology.
[0015] As a preferred embodiment of the Cordyceps identification system based on the LLaVA multi-modal large model of the present invention, wherein: it includes a Cordyceps data collection and preprocessing module, a feature extraction and fusion module, and a Cordyceps classification and result output module; the Cordyceps data collection and preprocessing module includes a multi-modal data collection unit and an image preprocessing unit. The multi-modal data collection unit is used to collect multi-angle images of Cordyceps and generate corresponding text descriptions. The image preprocessing unit is used to denoise, enhance, standardize, and segment the images; the feature extraction and fusion module includes a feature extraction unit and a multi-modal feature fusion unit. The feature extraction unit is used to extract image features and text semantic information through the LLaVA model, and the multi-modal feature fusion unit is used to fuse the image and text features into a unified feature vector by using the multi-modal reasoning mechanism of the LLaVA model; the Cordyceps classification and result output module includes a fine-grained classification unit and a report generation unit. The fine-grained classification unit is used to perform fine-grained classification on the fused feature vector, and the report generation unit is used to generate a Cordyceps identification report according to the classification result.
[0016] A computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the Cordyceps identification method based on the LLaVA multi-modal large model.
[0017] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the Cordyceps identification method based on the LLaVA multi-modal large model.
[0018] Advantages of the present invention: The cordyceps identification method based on the LLaVA multimodal large model provided by the present invention collects multi-angle images of cordyceps through an image acquisition device, and generates a text description containing information such as cordyceps variety, origin, color, ecological environment, etc. The collected data is preprocessed, significantly improving the clarity and consistency of the image data, realizing a more complete multimodal sample input than traditional pure image acquisition, enhancing the quality of the recognition basic data. By establishing a feature extraction algorithm, multi-modal feature extraction and fusion are performed on the cordyceps data, solving the single-modal limitation of traditional models that only rely on vision, realizing the joint reasoning of cordyceps images and semantic information; significantly enhancing the model's perception ability of fine-grained features; enhancing the robustness and generalization ability of the model in complex acquisition environments. By performing fine-grained classification on the fused joint feature vectors and organizing the recognition information into a structured classification report and feeding it back to the user, the recognition accuracy is significantly improved, realizing the generation of an automated classification report, reducing the dependence on manual judgment, and enhancing the system visualization and user-friendliness. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 It is the overall flowchart of a cordyceps identification method based on the LLaVA multimodal large model provided by the first embodiment of the present invention.
[0021] Figure 2 It is the overall flowchart of a cordyceps identification system based on the LLaVA multimodal large model provided by the third embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0023] Example 1, referring to Figure 1 This is an embodiment of the present invention, providing a cordyceps identification method based on the LLaVA multimodal large model, including:
[0024] S1: Collect cordyceps data and preprocess the cordyceps data.
[0025] Furthermore, collecting cordyceps data includes collecting images of cordyceps from different perspectives through an image acquisition device, and each image is equipped with a text description.
[0026] The text description includes the variety name, collection location, ecological environment, and color characteristics of the cordyceps.
[0027] It should also be noted that a high-resolution camera, industrial camera, or microscopic imaging device is used to capture cordyceps samples from multiple angles and scales. The cordyceps data collection scenario includes various cordyceps species such as Ophiocordyceps sinensis and cordyceps fungi; the shooting angles include the top view, side view, and local close-up of the cordyceps; the shooting light conditions include shooting under natural light, fluorescent lamps, and incandescent lamps to enhance the model's adaptability to light changes. An image set covering different angle views is formed, and each cordyceps image is equipped with structured text description information. The text can include attribute fields such as the variety name, collection location, growth environment, color, and shape of the cordyceps, constituting a data pair corresponding to the text and image.
[0028] Furthermore, preprocessing the cordyceps data includes denoising, enhancing, normalizing, and image segmentation of the cordyceps images.
[0029] It should also be noted that noise filtering is performed on the noise present in the original image. For example, algorithms such as median filtering and Gaussian filtering are used to remove image noise; at the same time, enhancement means such as image sharpening and histogram equalization can be used to improve the expressiveness of image detail features. All images are adjusted to a unified resolution to ensure data consistency. Data augmentation techniques (such as rotation, cropping, and brightness adjustment) are used to increase data diversity and improve the generalization ability of the model. Edge detection is combined to locate the cordyceps region in the image, and it is cropped to remove background interference.
[0030] S2: Establish a feature extraction algorithm to extract and fuse features from the preprocessed cordyceps data.
[0031] Furthermore, the feature extraction algorithm includes constructing a feature extraction algorithm based on the joint inference of cordyceps images and text descriptions. The LLaVA model is used to extract image features and process text description information. The LLaVA model extracts visual features from cordyceps images through a convolutional neural network and uses a pre-trained language model to parse the text description to extract the variety name, color characteristics, and ecological environment of the cordyceps from the text description.
[0032] It should be noted that a preferred solution for parsing text descriptions using a pre-trained language model is as follows: Generate semantic descriptions for each cordyceps image through GPT-4. Take the cordyceps image and the labeled classification tags as inputs, and the output is the semantic description. For example, the image description is that the image is a certain type of cordyceps, with a yellowish color, delicate texture, and shallow annulations. And perform semantic enhancement by combining the medicinal value of cordyceps (such as anti-inflammatory effects) to provide more context information for the model.
[0033] The extracted features include color features, texture features, and morphological features. Color features include the overall color distribution and local color changes of cordyceps; texture features include calculating the texture gradient of cordyceps through the Sobel operator, and morphological features include the length, width, number of annulations, etc. of cordyceps.
[0034] Calculating the cordyceps texture gradient is expressed as:
[0035]
[0036] Among them, G is the overall gradient intensity of a certain pixel point in the cordyceps image, reflecting the significant degree of change of this point in texture, which is the basic quantity for texture feature modeling. I is the input cordyceps image. is the gray-scale change rate of the cordyceps image in the horizontal direction, used to detect the horizontal texture changes on the surface of cordyceps. is the gray-scale change rate of the cordyceps image in the vertical direction, used to detect the longitudinal structural features of cordyceps, such as annulations, stripes, etc. x is the horizontal direction coordinate in the image, and y is the vertical direction coordinate in the image.
[0037] Furthermore, feature extraction and fusion of the preprocessed cordyceps data include using the multi-modal reasoning mechanism of the LLaVA model to fuse the features of the cordyceps image and the text description, and output a joint feature vector. The feature vector includes the variety, external shape features, and fine-grained features of cordyceps.
[0038] It should also be noted that the present invention uses the multi-modal attention mechanism of the LLaVA model to fuse and model the visual features extracted from the cordyceps image and the language features extracted from the semantic description, and output a joint feature vector containing image and semantic information. This feature vector not only contains the variety information of cordyceps, but also fuses its external shape features (such as annulations, texture) and other fine-grained difference features, which is expressed as:
[0039] F = Attention(V, T)
[0040] Among them, F is the fused multi-modal feature vector, representing the unified representation after the combination of image features and text semantic features through the attention mechanism. Attention(.) is the attention mechanism function, V is the high-dimensional vector extracted from the Cordyceps image through the deep vision network, encoding color, texture, and morphological information, and T is the representation after encoding the Cordyceps description statement.
[0041] Specifically, the attention mechanism function is expressed as:
[0042]
[0043] Among them, softmax(.) is used to assign the association weights with language semantics to each visual position, and W Q 、W K 、W V are trainable linear mapping matrices, which respectively convert the input features into query (Q), key (K), and value (V) vectors, and d is the feature dimension.
[0044] Through the multi-modal feature fusion scheme based on the attention mechanism, the present invention realizes the deep coupling of image and semantic information in the fine-grained recognition of Cordyceps, overcomes the limitation that the traditional single-modal recognition model only relies on images and ignores background semantics; significantly enhances the model's perception ability of fine-grained differences (such as subtle texture changes between similar Cordyceps varieties), improves the accuracy and stability of Cordyceps recognition; improves the generalization ability of the model in complex environments, such as the influence of acquisition angle changes, light interference, etc. on the recognition performance is reduced; provides a higher expressive feature input for subsequent classification tasks.
[0045] S3: Classify Cordyceps according to the fusion result and generate a Cordyceps classification report.
[0046] Furthermore, classifying Cordyceps according to the fusion result includes performing fine-grained classification on Cordyceps according to the fused feature vector to generate the variety classification result of Cordyceps, and the classification result includes the variety name, appearance characteristics, and color information of Cordyceps.
[0047] It should be noted that after the multi-modal fusion of image and semantic information, the system takes the fused feature vector as the input and executes the fine-grained classification task of Cordyceps, not only identifying the basic types of Cordyceps, but also comprehensively analyzing the details such as its appearance and color. Based on the fused joint feature vector F, a multi-layer classifier or support vector machine, fully connected network model is constructed to perform fine-grained classification on Cordyceps, and the fine-grained classification identifies the variety name of Cordyceps, and extracts and outputs the corresponding appearance characteristics (such as length, number of annulations, texture density, etc.) and color information (such as brightness, hue, color gamut distribution, etc.).
[0048] The results output by the classification process constitute the structured classification information of cordyceps, including: the variety name of cordyceps (such as Ophiocordyceps sinensis, Cordyceps fungi, etc.); the external feature information (such as having continuous annular markings, slender morphology); the color feature information (such as being yellowish-brown, uniform color).
[0049] After completing the multi-modal fusion of the image and semantic information, the system takes the fused feature vector F as the input of the fine-grained classification module and performs the fine-grained recognition task of cordyceps. The feature vector F integrates the image features of cordyceps (including color distribution, texture structure, morphological contour, etc.) and the text semantic features (including variety description, origin information, color vocabulary, etc.), and has strong representation ability. The system can construct a multi-layer classification network or a support vector machine model (SVM) based on the fused feature vector F to perform step-by-step refined recognition on cordyceps samples.
[0050] The fine-grained classification process includes, in the input stage, inputting the fused feature vector F ∈ R d into the classifier module, using a fully connected neural network structure with more than three layers, or a kernel function mapping structure based on SVM to perform non-linear feature partitioning on F. The network output includes not only the basic types of cordyceps (such as Ophiocordyceps sinensis, Cordyceps fungi, etc.), but also its external feature labels (such as length, width, number of annular markings, texture density) and color information labels (such as brightness, hue, color uniformity, etc.). Through joint training of the cross-entropy loss function and the multi-label loss function, the recognition accuracy of subtle appearance differences is improved. During the training process, methods such as Grad-CAM and Attention Map can be used to backtrack the discriminant area in the reverse direction to further enhance the model's attention to the key visual difference areas of cordyceps. The fine-grained classification outputs a structured recognition result, including the variety name of cordyceps, key external parameters (such as cordyceps length, number of annular markings, texture density), and color features (such as hue, saturation, lightness), and feeds it back to the end user as part of the recognition report.
[0051] Through the above classification mechanism, the system can effectively distinguish highly similar cordyceps varieties, especially between samples with similar colors, textures, and slightly different morphologies, and still has stable discriminant ability, significantly improving the practicality and accuracy of the cordyceps recognition system.
[0052] The cordyceps classification mechanism of the present invention significantly improves the system's ability to distinguish cordyceps with similar morphologies, overcoming the shortcoming of insufficient accuracy in traditional methods when dealing with fine-grained image differences.
[0053] Furthermore, generating a cordyceps classification report includes, according to the classification results, generating a cordyceps variety report, and the cordyceps variety report includes the classification label and text description of cordyceps, and feeding back the classification results to the user.
[0054] It should also be noted that based on the structured results such as the name of the cordyceps variety, external characteristics, and color information obtained through the classification steps, a classification report of cordyceps is generated. The report is in the form of a combination of text and images, including but not limited to classification labels: standardized naming of cordyceps species, such as Ophiocordyceps sinensis, Cordyceps fungi, etc.; text description: semantic description generated based on the fused features, such as the sample has a yellowish color, continuous annular patterns, a dense texture, and is suspected to be Ophiocordyceps sinensis produced in the Qinghai-Tibet Plateau; the generated report is displayed to the user through a terminal device, such as a PC interface or a mobile APP interface, for convenient manual review, archiving, or direct use in commercial transactions.
[0055] It should also be noted that the cordyceps classification report is not only used for information feedback but also serves as an information basis for quality inspection reports, product labels, traceability vouchers, etc., which helps to promote the implementation of standardized and digital identification management in the Chinese herbal medicine industry.
[0056] Example 2, which is an embodiment of the present invention, provides a cordyceps identification method based on the LLaVA multimodal large model. To verify the beneficial effects of the present invention, scientific demonstrations are carried out through economic benefit calculations and simulation experiments.
[0057] The experiment first selects 8 representative cordyceps categories from cordyceps physical samples (covering different origins, morphologies, colors, annular pattern characteristics, etc.) and uses a high-resolution industrial camera to collect multi-angle images under a standard shooting environment, including top views, side views, close-up views, etc. 3 images are taken for each cordyceps sample, and senior Chinese herbal medicine experts write structured semantic descriptions based on information such as the origin, growth environment, appearance characteristics, and medicinal uses of the cordyceps. After the collected data is preprocessed by image denoising, standardization, enhancement, and semantic description, a multimodal dataset of image-text pairing is constructed.
[0058] Subsequently, a cordyceps feature extraction and fusion algorithm based on the LLaVA model is constructed. The image input extracts high-dimensional visual features (such as color gradients, texture distributions, edge information, etc.) through a convolutional neural network, and the text information is encoded into semantic vectors by the BERT model. Through the LLaVA attention mechanism, deep fusion is carried out, and a unified feature vector is output. The fused feature vector is fed into a support vector machine (SVM) classifier for variety identification, and the cordyceps species, morphological analysis results, and confidence scores are output, and a complete cordyceps classification report is generated. To verify the effect, a control group is simultaneously constructed in the experiment, which only uses image features for identification, and the differences between the two groups of systems in key indicators such as recognition accuracy, confidence, and time consumption are evaluated. The experimental data is shown in Table 1.
[0059] Table 1 Data table of the cordyceps identification multimodal fusion experiment
[0060]
[0061]
[0062] As can be seen from the tabular data, the Cordyceps identification method that combines image and text information for fusion is significantly superior to the traditional identification method that only relies on images in terms of identification accuracy and confidence. The average identification accuracy reaches 95.4%. Among them, the classification accuracy of samples with more complete image quality and semantic descriptions (such as samples 1, 4, and 8) reaches or exceeds 97%, and the prediction confidence is as high as 0.99, indicating that the LLaVA model can make full use of semantic information to optimize the classification boundary. In contrast, for samples with slightly worse image clarity or semantic information quality (such as samples 3 and 7), although the accuracy decreases, it still remains above 88%, and the confidence is still better than the traditional method. This shows that the system of the present invention still has strong fault tolerance and robustness under incomplete information. In addition, the effective rate of the fused feature dimension reaches more than 90% on average, verifying the effectiveness of the attention mechanism in modeling the semantic association between images and texts. The identification time consumption in the experiment is stably controlled between 120 ms and 145 ms, meeting the real-time requirements of actual deployment, and reflecting the superiority of the present invention in engineering implementation.
[0063] Example 3, referring to Figure 2 , which is an embodiment of the present invention, provides a Cordyceps identification system based on the LLaVA multi-modal large model, including a Cordyceps data acquisition and preprocessing module 100, a feature extraction and fusion module 200, and a Cordyceps classification and result output module 300.
[0064] Among them, S4: The Cordyceps data acquisition and preprocessing module 100 includes a multi-modal data acquisition unit 101 and an image preprocessing unit 102. The multi-modal data acquisition unit 101 is used to collect multi-angle images of Cordyceps and generate corresponding text descriptions. The image preprocessing unit 102 is used to denoise, enhance, standardize, and segment the images.
[0065] It should also be noted that the Cordyceps data acquisition unit 101 obtains multi-angle Cordyceps images and corresponding text information and sends them to the image preprocessing unit 102. The image preprocessing unit 102 preprocesses the collected images, and the processed image-text data is used as a structured sample and transmitted to the feature extraction unit 201.
[0066] S5: The feature extraction and fusion module 200 includes a feature extraction unit 201 and a multi-modal feature fusion unit 202. The feature extraction unit 201 is used to extract image features and text semantic information through the LLaVA model. The multi-modal feature fusion unit 202 is used to use the multi-modal inference mechanism of the LLaVA model to fuse the image and text features into a unified feature vector.
[0067] It should also be noted that after receiving the image and text data, the feature extraction unit 201 extracts the image feature vector and the semantic feature vector respectively, and transmits them to the multi-modal feature fusion unit 202. The multi-modal feature fusion unit 202 aligns and fuses the image and text features based on the attention mechanism to generate the feature vector F. This feature vector F is used as the classification input and transmitted to the fine-grained classification unit 301.
[0068] S6: The cordyceps classification and result output module 300 includes a fine-grained classification unit 301 and a report generation unit 302. The fine-grained classification unit 301 is used to perform fine-grained classification on the fused feature vector, and the report generation unit 302 is used to generate a cordyceps identification report according to the classification result.
[0069] It should also be noted that the fine-grained classification unit 301 receives the fused feature vector F, completes the classification of cordyceps varieties and features, and generates classification labels and relevant attribute information. The classification result is transmitted to the report generation unit 302, and the generation unit 302 collates and forms a complete cordyceps identification report, which is output to the user interface or the interface system.
[0070] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0071] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0072] More specific examples (a non-exhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer diskettes (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.
[0073] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A Cordyceps identification method based on the LLaVA multi-modal large model, characterized in that It includes: Collect Cordyceps data and preprocess the Cordyceps data; Establish a feature extraction algorithm to extract and fuse features from the preprocessed Cordyceps data; Classify Cordyceps according to the fusion result and generate a Cordyceps classification report.
2. The cordyceps identification method based on the LLaVA multi-modal large model according to claim 1, wherein: The collection of Cordyceps data includes collecting images of different perspectives of Cordyceps through an image acquisition device, and each image is equipped with a text description; The text description includes the variety name, collection location, ecological environment, and color characteristics of Cordyceps.
3. The cordyceps identification method based on the LLaVA multi-modal large model according to claim 3, wherein: The preprocessing of Cordyceps data includes denoising, enhancing, normalizing, and image segmentation of Cordyceps images.
4. The cordyceps identification method based on the LLaVA multimodal large model according to any one of claims 1, 2 or 4, characterized in that: The feature extraction algorithm includes constructing a feature extraction algorithm based on the joint reasoning of Cordyceps images and text descriptions, extracting image features through the LLaVA model, processing text description information, the LLaVA model extracts visual features from Cordyceps images through a convolutional neural network, and uses a pre-trained language model to parse the text description to extract the variety name, color characteristics, and ecological environment of Cordyceps from the text description.
5. The cordyceps identification method based on the LLaVA multi-modal large model according to claim 5, wherein: The extraction and fusion of features from the preprocessed Cordyceps data includes using the multimodal reasoning mechanism of the LLaVA model to fuse the features of Cordyceps images and text descriptions, and output a joint feature vector, which includes the variety, shape features, and fine-grained features of Cordyceps.
6. The cordyceps identification method based on the LLaVA multimodal large model according to any one of claims 1, 2, 4 or 6, characterized in that: The classification of Cordyceps according to the fusion result includes performing fine-grained classification on Cordyceps according to the fused feature vector to generate the variety classification result of Cordyceps, and the classification result includes the variety name, shape features, and color information of Cordyceps.
7. The cordyceps identification method based on the LLaVA multi-modal large model according to any one of claims 1, 2, 4 or 6, characterized in that: The generation of the Cordyceps classification report includes generating a Cordyceps variety report according to the classification result, the Cordyceps variety report includes the classification label and text description of Cordyceps, and feedback the classification result to the user.
8. A cordyceps identification system using the LLaVA multimodal large model as described in any one of claims 1 to 8, characterized in that: It includes a Cordyceps data collection and preprocessing module (100), a feature extraction and fusion module (200), and a Cordyceps classification and result output module (300); The Cordyceps data collection and preprocessing module (100) includes a multimodal data collection unit (101) and an image preprocessing unit (102). The multimodal data collection unit (101) is used to collect multi-angle images of Cordyceps and generate a text description accordingly. The image preprocessing unit (102) is used to denoise, enhance, normalize, and segment the image; The feature extraction and fusion module (200) includes a feature extraction unit (201) and a multimodal feature fusion unit (202). The feature extraction unit (201) is used to extract image features and text semantic information through the LLaVA model. The multimodal feature fusion unit (202) is used to use the multimodal reasoning mechanism of the LLaVA model to fuse the image and text features into a unified feature vector; The Cordyceps classification and result output module (300) includes a fine-grained classification unit (301) and a report generation unit (302). The fine-grained classification unit (301) is used to perform fine-grained classification on the fused feature vector, and the report generation unit (302) is used to generate a Cordyceps recognition report according to the classification result.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the cordyceps identification method based on the LLaVA multi-modal large model described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the cordyceps identification method based on the LLaVA multi-modal large model described in any one of claims 1 to 8 are implemented.