A Diatom Detection Method and System Based on Multimodal Large Models
Through a multimodal large model based on Transformer architecture, the problems of low detection efficiency and high labeling cost in forensic diatom testing are solved, and the intelligence and automation of diatom detection are realized, and the diatom size is automatically calculated, which reduces the impurity error detection rate.
Patent Information
- Application Number
- CN202510558928.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In forensic science, diatom testing has problems such as low detection efficiency and insufficient automation, the existing deep learning solutions are expensive to label and the detection results are out of touch with physical measurements.
Using a multimodal large model based on Transformer architecture, the training data of image-text pairs is constructed, and the diatom morphology dictionary is injected into the cross-modal alignment of visual features and text features is optimized, and the target rectangular box parameters are fine-tuned through knowledge distillation and quantization model are output to calculate the diatom size.
The intelligent and automated diatom detection is realized, the detection efficiency is improved, the labeling cost is reduced, the impurity error detection rate is reduced, and the actual size of the diatom is automatically calculated.
Smart Images

Figure CN120088782B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of forensic medicine inspection and computer vision. Specifically, it relates to a diatom detection method and system based on a multimodal large model. Background Art
[0002] In forensic diagnosis of drowning, diatom examination, as the acquisition of key evidence, highly depends on manual identification and measurement under a microscope. Traditional methods require examiners to visually screen diatoms one by one and compare them with a microscopic micrometer, suffering from technical bottlenecks such as low detection efficiency and large subjective judgment deviations. Existing automated solutions based on image processing mostly adopt traditional edge detection algorithms or morphological operations, but face two major limitations in practical applications: First, diatom samples are often mixed with impurity particles, and traditional algorithms are difficult to stably distinguish non - target objects with similar shapes. Second, diatoms are tiny in size and diverse in shape, and conventional detection methods are prone to feature omission or misdetection. In recent years, although object detection technologies based on deep learning have made progress in fields such as industrial inspection and medical image analysis, their application in forensic diatom examination is still restricted by two core problems: First, the data annotation cost is too high. Second, there is inconsistency in annotation standards. Different annotators have significant differences in determining the boundaries of diatoms, resulting in label noise during model training and affecting the detection robustness. In addition, existing technologies have not effectively solved the problem of connecting the detection results with physical size measurement, and still require manual pixel - micrometer conversion, restricting the degree of full - process automation. Summary of the Invention
[0003] Aiming at the problems of low efficiency and insufficient automation of traditional diatom inspection methods, as well as high annotation costs and disconnection between detection results and physical measurements in existing deep - learning solutions, the purpose of the present invention is to provide a diatom detection technology based on a multimodal large model, aiming to provide an intelligent and automated solution for forensic diatom inspection.
[0004] To achieve the above - mentioned technical purpose, the present application provides a diatom detection method based on a multimodal large model, including the following steps:
[0005] Based on a multimodal large model of the Transformer architecture, according to the constructed image - text pair training data, by retaining the general feature extraction ability of the pre - trained visual encoder, injecting a dictionary of the diatom morphology field as a learnable embedding vector into the text encoder, and optimizing the cross - modal alignment of visual features and text features using a contrastive learning loss function, for fine - tuning;
[0006] Based on the fine-tuned multi-modal large model, after knowledge distillation and quantization of the model, the collected diatom microscopic images with known magnification are fused with text embedding vectors through spatial attention, and the target rectangle box parameters are output through the self-attention weight screening mechanism of the cross-modal decoder. Then, based on the known magnification, the size of the diatom is detected and calculated.
[0007] Preferably, in the process of constructing the image-text pair training data, the lableme annotation tool is used to annotate the diatom rectangle box on the scanning electron microscope image, generate a JSON file, convert the annotation data into the COCO format, associate the morphological description text of the diatom, and construct an image-text pair dataset.
[0008] Preferably, when fine-tuning the multi-modal large model based on the Transformer architecture, the morphological text of the diatom is injected into the text encoder, and the cross-modal alignment is optimized through the contrastive learning loss. Prior knowledge constraints on the aspect ratio and size range of the diatom are established to suppress the response of non-target impurities.
[0009] Preferably, when performing knowledge distillation, the cross-modal attention matrix of the teacher model is extracted and transferred to the student model through the KL divergence loss.
[0010] Preferably, when quantifying the model, TensorRT INT8 symmetric quantization is performed on the detection head network layer.
[0011] Preferably, when quantifying the model, the FP16 precision of the text encoder is retained to maintain semantic consistency.
[0012] Preferably, when detecting and calculating the size of the diatom, based on the target rectangle box parameters, the width and height parameters of the rectangle box are multiplied by the magnification respectively, and the larger value of the two is taken as the actual major axis, and the smaller value is taken as the actual minor axis.
[0013] The present invention discloses a diatom detection system based on a multi-modal large model, including:
[0014] A dataset construction module, which is used to use the lableme annotation tool to annotate the diatom rectangle box on the scanning electron microscope image, generate a JSON file, convert the annotation data into the COCO format, associate the morphological description text of the diatom, and construct an image-text pair dataset;
[0015] The model fine-tuning module is used to fine-tune the multimodal large model based on the Transformer architecture. Among them, for the multimodal large model based on the Transformer architecture, according to the image-text pair training data, by retaining the general feature extraction ability of the pre-trained visual encoder, injecting a dictionary in the field of diatom morphology as a learnable embedding vector in the text encoder, and using a contrastive learning loss function to optimize the cross-modal alignment of visual features and text features for fine-tuning;
[0016] The detection module is used to, based on the fine-tuned multimodal large model, after knowledge distillation and quantization of the model, perform spatial attention fusion on the collected diatom microscopic images with known magnification and the text embedding vector, output the target rectangle box parameters through the self-attention weight screening mechanism of the cross-modal decoder, and calculate the size of the diatom according to the known magnification.
[0017] Preferably, the detection module is also used to extract the cross-modal attention matrix of the teacher model and transfer it to the student model through the KL divergence loss.
[0018] Preferably, the detection module is also used to perform TensorRT INT8 symmetric quantization on the detection head network layer and retain the FP16 precision of the text encoder to maintain semantic consistency.
[0019] The present invention discloses the following technical effects:
[0020] The detection efficiency of the present invention is greatly improved compared with manual measurement; the model can be fine-tuned based on hundreds of labeled samples, and the labeling cost is greatly reduced compared with the traditional scheme; by dynamically adjusting the visual feature weight through text prompts, the impurity misdetection rate is significantly reduced compared with the traditional scheme; the lightweight model accelerated by TensorRT greatly reduces the required computing resources; automatically calculates the actual size of diatoms; provides a new intelligent and automated solution for forensic diatom inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 is the flow schematic diagram of the present invention;
[0023] Figure 2 is the schematic diagram of the Transformer architecture of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. The components of the embodiments of this application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.
[0025] As Figure 1-2 shown, the present invention provides a diatom detection method based on a multimodal large model, including the following steps:
[0026] Based on the multimodal large model of the Transformer architecture, according to the constructed image-text pair training data, by retaining the general feature extraction ability of the pre-trained visual encoder, injecting a diatom morphology domain dictionary as a learnable embedding vector into the text encoder, and using a contrastive learning loss function to optimize the cross-modal alignment of visual features and text features, for fine-tuning;
[0027] Based on the fine-tuned multimodal large model, after knowledge distillation and quantization of the model, the collected diatom microscopic images with known magnification are fused with the text embedding vector in terms of spatial attention, and the target rectangle box parameters are output through the self-attention weight screening mechanism of the cross-modal decoder, and the size of the diatom is detected and calculated according to the known magnification.
[0028] Further preferably, for the diatom detection method based on a multimodal large model provided by the present invention, during the process of constructing the image-text pair training data, the lableme annotation tool is used to annotate the diatom rectangle box on the scanning electron microscope image, generate a JSON file, convert the annotation data into the COCO format, associate the morphological description text of the diatom, and construct an image-text pair dataset.
[0029] Further preferably, for the diatom detection method based on a multimodal large model provided by the present invention, when fine-tuning the multimodal large model based on the Transformer architecture, inject the diatom morphology text into the text encoder, optimize the cross-modal alignment through contrastive learning loss, establish prior knowledge constraints on the aspect ratio and size range of the diatom, and suppress the response of non-target impurities.
[0030] Further preferably, for a diatom detection method based on a multimodal large model provided by the present invention, when performing knowledge distillation, the cross-modal attention matrix of the teacher model is extracted and transferred to the student model through the KL divergence loss.
[0031] Further preferably, for a diatom detection method based on a multimodal large model provided by the present invention, when quantifying the model, TensorRT INT8 symmetric quantization is performed on the detection head network layer.
[0032] Further preferably, for a diatom detection method based on a multimodal large model provided by the present invention, when quantifying the model, the FP16 precision of the text encoder is retained to maintain semantic consistency.
[0033] Further preferably, for a diatom detection method based on a multimodal large model provided by the present invention, when detecting and calculating the size of diatoms, based on the target rectangle frame parameters, the width and height parameters of the rectangle frame are respectively multiplied by the magnification factor, and the larger value of the two is taken as the actual major axis, and the smaller value is taken as the actual minor axis.
[0034] The present invention discloses a diatom detection system based on a multimodal large model, including:
[0035] A dataset construction module, which is used to label the diatom rectangle frame on the scanning electron microscope image using the lableme annotation tool, generate a JSON file, convert the annotation data into the COCO format, associate the morphological description text of the diatoms, and construct an image-text pair dataset;
[0036] A model fine-tuning module, which is used to fine-tune the multimodal large model based on the Transformer architecture. Among them, for the multimodal large model based on the Transformer architecture, according to the image-text pair training data, by retaining the general feature extraction ability of the pre-trained visual encoder, a diatom morphology domain dictionary is injected into the text encoder as a learnable embedding vector, and a contrast learning loss function is used to optimize the cross-modal alignment of visual features and text features for fine-tuning;
[0037] A detection module, which is used to, based on the fine-tuned multimodal large model, after knowledge distillation and model quantization, fuse the collected diatom microscopic images with known magnification factors with the text embedding vector through spatial attention, output the target rectangle frame parameters through the self-attention weight screening mechanism of the cross-modal decoder, and detect and calculate the size of the diatoms based on the known magnification factor.
[0038] Further preferably, the detection module of a diatom detection system based on a multimodal large model disclosed by the present invention is further used to extract the cross-modal attention matrix of the teacher model and transfer it to the student model through the KL divergence loss.
[0039] Further preferably, the detection module of a diatom detection system based on a multimodal large model disclosed by the present invention is further configured to perform TensorRT INT8 symmetric quantization on the detection head network layer, and retain the FP16 precision of the text encoder to maintain semantic consistency.
[0040] Embodiment: The present invention provides a diatom cheek detection technology based on a multimodal large model, which specifically includes the following processes:
[0041] S1: Convert the diatom image data set generated by the annotation tool and the diatom morphological feature text into a format to generate multimodal training data including image features and morphological text descriptions.
[0042] The format conversion in step S1 specifically includes: parsing the JSON file annotated by lableme to extract the rectangular box coordinates and converting them into the COCO annotation format; associating the morphological description texts of each diatom to construct image-text pair training data.
[0043] S2: Based on the multimodal data set constructed in step S1, adopt a domain adaptation fine-tuning strategy to optimize the parameters of the multimodal large model based on the Transformer architecture, and construct a multimodal detection model sensitive to diatom features.
[0044] The domain adaptation fine-tuning strategy in step S2 includes: freezing the visual encoder network parameters of the multimodal large model based on the Transformer architecture; injecting a diatom morphology domain dictionary into the text encoder as a learnable embedding vector; adopting a contrastive learning loss function to optimize the cross-modal alignment of visual features and text features.
[0045] S3: Transfer the cross-modal feature alignment ability of the fine-tuned model to the lightweight student model through a knowledge distillation framework to form a detection architecture with real-time inference ability.
[0046] The knowledge distillation framework in step S3 includes: extracting the cross-modal attention matrix of the teacher model as a supervision signal; constructing a visual-text feature interaction constraint loss function for the student model; realizing transfer learning of attention weights through minimizing KL divergence.
[0047] S4: Perform low-bit quantization compression on the detection model obtained in step S3 to generate a lightweight deployment version.
[0048] The specific implementation of the low-bit quantization in step S4 is: performing TensorRT INT8 symmetric quantization on the detection head network layer; retaining the FP16 floating-point calculation precision of the text encoder to maintain the stability of semantic alignment.
[0049] S5: Obtain diatom microscopic images with a known magnification through a scanning electron microscope.
[0050] S6: Input the microscopic image and the text of diatom morphological features into the detection model in step S4, and output the parameters of the horizontal rectangular box with pixel-level positioning through the cross-modal decoder.
[0051] The working process of the cross-modal decoder in step S6 includes: performing spatial attention fusion on the visual feature map and the text embedding vector; outputting the parameters of the target rectangular box through the self-attention weight screening mechanism of the cross-modal decoder.
[0052] S7: Automatically calculate the actual major and minor axis dimensions of the diatom based on the geometric parameters of the rectangular box output in step S6 and the image magnification.
[0053] The dimension calculation in step S7 includes: multiplying the width and height parameters of the rectangular box by the magnification factor respectively, and taking the larger value of the two as the actual major axis and the smaller value as the actual minor axis.
[0054] The present invention analyzes the JSON file labeled by lableme to extract the rectangular box coordinates and convert them into the COCO annotation format, associates the morphological description text of each diatom, constructs image-text pair training data; fine-tunes the multi-modal large model based on the Transformer architecture, retains the general feature extraction ability of the pre-trained visual encoder, injects the diatom morphological text into the text encoder, optimizes the cross-modal alignment through the contrastive learning loss, establishes the prior knowledge constraints on the aspect ratio and size range of the diatom, and suppresses the response of non-target impurities; extracts the cross-modal attention matrix of the teacher model and transfers it to the student model through the KL divergence loss; performs INT8 symmetric quantization on the detection head, and the text encoder retains the FP16 precision to maintain semantic consistency; performs spatial attention fusion on the visual feature map and the text embedding vector to output the target rectangular box; combines the image magnification to output the actual major and minor axis dimensions of the diatom. The detection efficiency of the present invention is greatly improved compared with manual measurement; the model can be fine-tuned based on hundreds of labeled samples, and the labeling cost is greatly reduced compared with the traditional scheme; the visual feature weight is dynamically adjusted through text prompts, and the false detection rate of impurities is significantly reduced compared with the traditional scheme; the lightweight model accelerated by TensorRT greatly reduces the required computing resources; automatically calculates the actual size of the diatom.
[0055] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0056] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0057] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A diatom detection method based on a multimodal large model, characterized in that, Including the following steps: Based on the multi-modal large model of the Transformer architecture, according to the constructed image-text pair training data, by retaining the general feature extraction ability of the pre-trained visual encoder, injecting the diatom morphology domain dictionary as a learnable embedding vector into the text encoder, and using the contrastive learning loss function to optimize the cross-modal alignment of visual features and text features, for fine-tuning; Based on the fine-tuned multi-modal large model, after knowledge distillation and quantization of the model, the collected diatom microscopic images with known magnification are subjected to spatial attention fusion with the text embedding vector, and the target rectangle box parameters are output through the self-attention weight screening mechanism of the cross-modal decoder, and based on the known magnification, the size of the diatom is detected and calculated.
2. The diatom detection method based on a multi-modal large model according to claim 1, wherein: In the process of constructing the image-text pair training data, use the lableme annotation tool to annotate the diatom rectangle box on the scanning electron microscope image, generate a JSON file, convert the annotation data into the COCO format, associate the morphological description text of the diatom, and construct the image-text pair data set.
3. The diatom detection method based on a multi-modal large model according to claim 2, wherein: When fine-tuning the multi-modal large model based on the Transformer architecture, inject the diatom morphology text into the text encoder, optimize the cross-modal alignment through the contrastive learning loss, establish the prior knowledge constraints of the diatom aspect ratio and size range, and suppress the response of non-target impurities.
4. The diatom detection method based on a multi-modal large model according to claim 3, wherein: When performing knowledge distillation, extract the cross-modal attention matrix of the teacher model and transfer it to the student model through the KL divergence loss.
5. The diatom detection method based on a multi-modal large model according to claim 4, wherein: When performing model quantization, perform TensorRT INT8 symmetric quantization on the detection head network layer.
6. The diatom detection method based on a multi-modal large model according to claim 5, wherein: When performing model quantization, retain the FP16 precision of the text encoder to maintain semantic consistency.
7. The diatom detection method based on a multi-modal large model according to claim 6, wherein: When detecting and calculating the size of the diatom, based on the target rectangle box parameters, multiply the width and height parameters of the rectangle box by the magnification respectively, and take the larger value of the two as the actual major axis and the smaller value as the actual minor axis.
8. A diatom detection system based on a multimodal large model, characterized in that, Including: A data set construction module, which is used to use the lableme annotation tool to annotate the diatom rectangle box on the scanning electron microscope image, generate a JSON file, convert the annotation data into the COCO format, associate the morphological description text of the diatom, and construct an image-text pair data set; A model fine-tuning module for fine-tuning a multi-modal large model based on the Transformer architecture. Among them, for the multi-modal large model based on the Transformer architecture, according to the image-text pair training data, by retaining the general feature extraction ability of the pre-trained visual encoder, injecting a diatom morphology domain dictionary as a learnable embedding vector into the text encoder, and using a contrastive learning loss function to optimize the cross-modal alignment of visual features and text features for fine-tuning; A detection module for, based on the fine-tuned multi-modal large model, after knowledge distillation and quantization of the model, performing spatial attention fusion on the collected diatom microscopic images with known magnification factors and text embedding vectors, outputting target rectangle box parameters through the self-attention weight screening mechanism of the cross-modal decoder, and calculating the size of diatoms according to the known magnification factors.
9. The diatom detection system based on a multi-modal large model according to claim 8, wherein: The detection module is further configured to extract the cross-modal attention matrix of the teacher model and transfer it to the student model through the KL divergence loss.
10. The diatom detection system based on a multi-modal large model according to claim 9, wherein: The detection module is further configured to perform TensorRT INT8 symmetric quantization on the detection head network layer and retain the FP16 precision of the text encoder to maintain semantic consistency.
Citation Information
Patent Citations
Diatom size automatic measurement method based on directed target detection
CN116735463A
Underwater target detection method based on vision-language model knowledge distillation
CN119445353A