Analysis system based on AI visual interpretation
Through AI visual interpretation technology, the automatic identification and interpretation of blood type cards is realized, the accuracy and misjudgment of traditional blood type detection is solved, the detection efficiency and accuracy are improved, and the multiple blood type cards are adapted to.
Patent Information
- Application Number
- CN202510463097.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional blood type detection relies on manual operation, which has problems such as insufficient physical characteristic recognition, poor fault tolerance of position offset, manual verification of information accuracy is prone to errors and low interpretation accuracy, resulting in high risk of detection failure and misjudgment.
An analysis system based on AI visual interpretation is adopted, including an image acquisition module, a multi-task model of deep convolutional neural network CNN, and a control execution module, to realize the automatic identification and interpretation of blood type cards. The system uses the physical characteristic recognition submodel, the position offset detection submodel, the information accuracy verification submodel and the erythrocyte agglutination analysis submodel, combined with optical character recognition and barcode recognition technology to collect and process blood type card images in real time, and drive the robotic arm to operate.
It improves the accuracy of automated detection and identification of blood type analysis, reduces the risk of misjudgment, improves detection efficiency and accuracy, reduces manual intervention, adapts to blood type cards from different manufacturers and batches, and has a high hardware resource utilization rate.
Smart Images

Figure CN120388371A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of artificial intelligence (AI) and medical detection technology, and particularly to an analysis system based on AI visual interpretation. Background Art
[0002] Traditional blood type detection relies on manual operation. It is necessary to observe the red blood cell agglutination reaction on the blood type card with the naked eye and manually judge the blood type result. Although some automated devices can assist in detection, there are the following problems: insufficient physical property recognition, unable to automatically distinguish different batches or shape differences of blood type cards, and it is easy to cause detection failure due to incorrect card slot adaptation; poor tolerance to position offset. When the robotic arm grasps the blood type card and there is a position offset, the existing system is difficult to correct it in real time, affecting subsequent operations; the information accuracy depends on manual work. The label information on the blood type card (such as blood type identification, patient ID) needs to be manually checked, and it is easy to cause data errors due to operational negligence; the interpretation accuracy is low. The judgment of the red blood cell agglutination state by manual or semi - automated systems is easily interfered by subjective factors, and the risk of misjudgment is relatively high. Summary of the Invention
[0003] In view of this, the present invention provides an analysis system based on AI visual interpretation to improve the accuracy of automated detection and blood type analysis recognition and reduce the risk of misjudgment.
[0004] In a first aspect, the present invention provides an analysis system based on AI visual interpretation, and the system includes:
[0005] An image acquisition module, an AI visual interpretation algorithm module, namely a multi - task model based on a deep convolutional neural network CNN, and a control execution module; the image acquisition module uses a high - resolution industrial camera to collect images of the blood type card and the reaction area in real time, and performs denoising and normalization processing on the collected images; the multi - task model is used to obtain an interpretation result according to the images of the blood type card and the reaction area, and the multi - task model includes a physical property recognition sub - model, a position offset detection sub - model, an information accuracy verification sub - model, and a red blood cell agglutination analysis sub - model; the control execution module is used to drive the robotic arm to complete the grasping, placement, sorting of the blood type card or the isolation of incorrect blood type cards according to the interpretation result.
[0006] Optionally, the physical property recognition sub - model is used to judge the batch number and shape parameters of the blood type card through image segmentation and feature extraction; the position offset detection sub - model uses a key - point matching algorithm, namely scale - invariant feature transform SIFT, to calculate the offset between the actual position and the preset position of the blood type card in real time, and feeds it back to the control execution module to control the robotic arm for dynamic adjustment.
[0007] Optionally, the information accuracy verification sub-model combines optical character recognition (OCR) and barcode recognition technology to extract the text and barcode information on the blood type card, and automatically compares it with the system database or the test results. If the comparison is inconsistent, an alarm is triggered; the red blood cell agglutination analysis sub-model uses a U-Net network to segment the agglutination area, combines morphological algorithms to quantify the agglutination intensity, and determines the blood type.
[0008] Optionally, the physical property recognition sub-model first performs image preprocessing, including grayscale conversion, filtering and noise reduction, and contrast enhancement; Grayscale conversion: Turn on the front light source and the back light source, collect the image, and convert the color blood type card image into a grayscale image to reduce the computational amount and eliminate the interference of color factors on subsequent processing at the same time. The conversion is carried out through the expression Gray = 0.299R + 0.587G + 0.114B, where R, G, and B are the red, green, and blue channel values of the image respectively; Filtering and noise reduction: Use median filtering or Gaussian filtering to remove the noise in the image and smooth the image; Contrast enhancement: Use histogram equalization or contrast limited adaptive histogram equalization (CLAHE) to enhance the contrast of the image;
[0009] Secondly, the batch number is recognized, including region localization, character segmentation, feature extraction and recognition; Region localization includes template matching and edge detection and contour analysis. Template matching: If the position of the batch number on the blood type card is relatively fixed, a template of the number region is pre-made, and the number region is located in the preprocessed image through the template matching algorithm; Edge detection and contour analysis: Use the Canny edge detection algorithm to extract the edge information of the image, and then through contour search and screening, locate the region containing the batch number according to the characteristics such as the area and aspect ratio of the contour; Character segmentation includes binary processing and connected component analysis. Binary processing: Perform binary processing on the located number region to convert the image into a black-and-white image for subsequent character segmentation. The Otsu automatic threshold method is used to determine the binary threshold; Connected component analysis: The characters in the binary image are segmented into individual character regions through connected component analysis, and the noise and small connected regions are removed, and only the character regions are retained; Feature extraction and recognition include feature extraction and machine learning or deep learning recognition. Feature extraction: Extract the geometric features, projection features, and moment features of the characters. Geometric features include area, perimeter, and aspect ratio; Projection features include horizontal projection and vertical projection; Moment features include Hu moments; Machine learning or deep learning recognition: Use learning algorithms such as support vector machine (SVM) and K-nearest neighbor (KNN) or deep learning algorithms such as convolutional neural network (CNN) to train and recognize the extracted features, and classify the characters into corresponding numbers or letters to obtain the batch number;
[0010] Finally, the shape parameters are extracted, including card slot area segmentation and parameter calculation; the card slot area segmentation includes color segmentation and edge detection and contour fitting. Color segmentation: If there is a color difference between the card slot and the background, the image is converted to the HSV color space, and the card slot area is segmented by setting color thresholds. Edge detection and contour fitting: The Canny edge detection algorithm is used to detect the edges of the card slot, and then through contour search and fitting, geometric shapes such as the minimum bounding rectangle or ellipse are used to approximate the shape of the card slot. Parameter calculation includes dimension measurement and position and angle calculation. Dimension measurement: According to the fitted geometric shape, the dimension parameters of the length, width, and diameter of the card slot are calculated. Position and angle calculation: The central position coordinates and rotation angle of the card slot are calculated to determine the position and orientation of the card slot on the blood type card.
[0011] Optionally, the position offset detection sub-model first generates SIFT key point detection and descriptors, including scale space extreme value detection, key point localization, direction assignment, and descriptor generation. Scale space extreme value detection: The scale space is constructed through a Gaussian pyramid to detect extreme points at different scales and obtain potential key points. Key point localization: The Hessian matrix is used to exclude low-contrast points and edge response points and retain stable key points. Direction assignment: The gradient histogram of the key point neighborhood is calculated to assign the main direction to make the feature rotation invariant. Descriptor generation: With the key point as the center, a 16×16 pixel neighborhood is constructed, and the gradient amplitude and direction in 8 directions are calculated to generate a 128-dimensional SIFT descriptor.
[0012] Secondly, feature matching and outliers are filtered, including brute-force matching BF or fast approximate nearest neighbor matching FLANN and the RANSAC algorithm to remove outliers. Brute-force matching BF: The Euclidean distance between descriptors is calculated, and the matching pairs with a distance ratio less than ().7 are retained, and the incorrect matches are filtered out. FLANN matching: Suitable for large-scale data, the search is accelerated through a KD tree to improve the matching speed. RANSAC algorithm to remove outliers: Using the random sample consensus algorithm, the homography matrix H is fitted to filter out the incorrect matching points that do not conform to the transformation model and retain the inliers.
[0013] Finally, calculate the coordinate transformation and offset, and preset the position calibration: Collect the blood type card image under ideal conditions, extract the coordinates of the reference key points, and establish a preset coordinate system; Obtain the actual position coordinates: Extract the key point coordinates from the real-time image; Decompose the homography matrix: Decompose the translation vector, rotation angle, and scaling factor through the homography matrix: Through offset calculation, confirm the allowable range of the offset: ±0.5mm for the X / Y axis; Generate the robotic arm control instruction, coordinate system transformation: Convert the image pixel coordinate system to the robotic arm base coordinate system, and pre-calibrate the conversion parameters through the calibration board; Instruction mapping: According to the robotic arm kinematic model, convert the offset into joint angle or end effector displacement instruction; Dynamic adjustment: Through real-time closed-loop feedback, control the robotic arm to complete translation and rotation adjustments to ensure that the position of the blood type card is consistent with the preset position.
[0014] Optionally, in the optical character recognition OCR of the information accuracy verification sub-model, first perform text area localization. Locate the text area with a fixed format through template matching or edge detection. For text at non-fixed positions, use the EAST text detection algorithm to dynamically locate the text box; Secondly, perform OCR recognition. Basic solution: Use the Tesseract OCR engine and configure a custom dictionary; Deep learning solution: Adopt the convolutional recurrent neural network CRNN or TrOCR model to perform end-to-end recognition of complex fonts, and normalize the input image to a fixed size of 32×128.
[0015] Optionally, the U-Net network structure in the red blood cell agglutination analysis sub-model includes an encoder, a decoder, and an output layer; Encoder: 4 convolutional blocks, namely 3×3 convolution, ReLU, batch normalization, 2×2 max pooling, used to extract multi-scale features; Decoder: 3 transposed convolutional blocks, namely 2×2 upsampling, 3×3 convolution, skip connection, used to restore the spatial resolution; Output layer: 1×1 convolution, activation function Sigmoid, used to generate a pixel-level agglutination probability map. When the value ∈[0,1] and >0.5, it is determined as the agglutination area;
[0016] Training strategy of the U-Net network:
[0017] Data augmentation: Rotate the annotated image by ±15°, scale it by 0.8 - 1.2 times, and add Gaussian noise to improve the generalization ability; Loss function: Joint optimization of Dice loss and cross-entropy loss to solve the class imbalance problem; Inference process: Normalize the input image, and its pixel value is 0 - 1; The network outputs a binary mask, with the agglutination area being white and the background being black;
[0018] Agglutination intensity index:
[0019] The number of particles N: The total number of aggregated particles per unit area, which is used to reflect the degree of aggregation dispersion; The average area S_avg: The average pixel area of a single aggregated particle, which is used to measure the particle size; The aggregation ratio R: The ratio of the area of the aggregation region to the total area of the reaction region; The roundness C: The degree to which the particle approaches a circle, C = 4π×area / (perimeter²), and the closer the value is to 1, the more regular the particle is.
[0020] In the technical solution provided by the present invention, the system includes an image acquisition module, an AI vision interpretation algorithm module, i.e., a multi-task model based on the deep convolutional neural network CNN, and a control execution module; The image acquisition module uses a high-resolution industrial camera to collect images of the blood type card and the reaction area in real time, and performs denoising and normalization processing on the collected images; The multi-task model is used to obtain an interpretation result based on the images of the blood type card and the reaction area. The multi-task model includes a physical property recognition sub-model, a position offset detection sub-model, an information accuracy verification sub-model, and a red blood cell aggregation analysis sub-model; The control execution module is used to drive the robotic arm to complete the grasping, placement, sorting of the blood type card or the isolation of the incorrect blood type card according to the interpretation result. This system improves the accuracy of automatic detection and blood type analysis recognition, and reduces the risk of misjudgment. Brief Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a schematic diagram of the analysis system based on AI vision interpretation provided by the embodiment of the present invention;
[0023] Figure 2 It is a schematic diagram of another analysis system based on AI vision interpretation provided by the embodiment of the present invention. Detailed Embodiments
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0025] It should be clear that the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present invention.
[0026] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0027] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0028] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0029] The present invention provides an analysis system based on AI visual interpretation, such as Figure 1 and Figure 2 shown, the system includes: an image acquisition module, an AI visual interpretation algorithm module, namely a multi-task model based on the deep convolutional neural network CNN, and a control execution module; the image acquisition module uses a high-resolution industrial camera to collect images of the blood type card and the reaction area in real time, and performs denoising and normalization processing on the collected images; the multi-task model is used to obtain an interpretation result according to the images of the blood type card and the reaction area, and the multi-task model includes a physical property recognition sub-model, a position offset detection sub-model, an information accuracy verification sub-model, and a red blood cell agglutination analysis sub-model; the control execution module is used to drive the robotic arm to complete the grasping, placing, sorting of the blood type card or the isolation of the wrong blood type card according to the interpretation result.
[0030] In the embodiments of the present invention, the physical property recognition sub-model is used to judge the batch number and shape parameters of the blood type card through image segmentation and feature extraction; the position offset detection sub-model uses the key point matching algorithm, namely the scale-invariant feature transform SIFT, to calculate the offset between the actual position and the preset position of the blood type card in real time, and feeds it back to the control execution module to control the robotic arm for dynamic adjustment.
[0031] In the embodiments of the present invention, the information accuracy verification sub-model combines optical character recognition (OCR) and barcode recognition technology to extract the text and barcode information on the blood type card, such as blood type identification and patient ID, and automatically compares it with the system database or test results. If the comparison is inconsistent, an alarm is triggered; the red blood cell agglutination analysis sub-model uses a U-Net network to segment the agglutination area, combines morphological algorithms to quantify the agglutination intensity, and determines the blood type.
[0032] In the embodiments of the present invention, the physical property recognition sub-model first performs image preprocessing, including grayscale conversion, filtering and noise reduction, and contrast enhancement; Grayscale conversion: Turn on the front light source and back light source, collect the image, and convert the color blood type card image into a grayscale image to reduce the computational amount and eliminate the interference of color factors on subsequent processing. The conversion is performed through the expression Gray = 0.299R + 0.587G + 0.114B, where R, G, and B are the red, green, and blue channel values of the image respectively; Filtering and noise reduction: Use median filtering or Gaussian filtering to remove the noise in the image and smooth the image; for example, median filtering replaces the grayscale value of each pixel point with the median of the pixel grayscale values in its neighborhood, effectively removing salt-and-pepper noise; Contrast enhancement: Use histogram equalization or contrast-limited adaptive histogram equalization (CLAHE) to enhance the contrast of the image, making features such as batch numbers and card slots clearer;
[0033] Secondly, the batch number is recognized, including region localization, character segmentation, feature extraction, and recognition; Region localization includes template matching and edge detection and contour analysis. Template matching: If the position of the batch number on the blood type card is relatively fixed, a template of the number region is pre-made, and the number region is located in the preprocessed image through the template matching algorithm; Edge detection and contour analysis: Use the Canny edge detection algorithm to extract the edge information of the image, and then through contour search and screening, locate the region containing the batch number according to the characteristics of the contour, such as area, aspect ratio, etc.; Character segmentation includes binary processing and connected component analysis. Binary processing: Perform binary processing on the located number region, convert the image into a black-and-white image for subsequent character segmentation, and use the automatic threshold method Otsu to determine the binary threshold; Connected component analysis: Segment the characters in the binary image into individual character regions through connected component analysis, remove noise and small connected regions, and only retain the regions of the characters; Feature extraction and recognition include feature extraction and machine learning or deep learning recognition. Feature extraction: Extract the geometric features, projection features, and moment features of the characters. Geometric features include area, perimeter, and aspect ratio; Projection features include horizontal projection and vertical projection; Moment features include Hu moments; Machine learning or deep learning recognition: Use learning algorithms such as support vector machine (SVM) and K-nearest neighbor (KNN) or deep learning algorithms such as convolutional neural network (CNN) to train and recognize the extracted features, and classify the characters into corresponding numbers or letters to obtain the batch number;
[0034] Finally, extract the shape parameters, including the slot area segmentation and parameter calculation; the slot area segmentation includes color segmentation and edge detection and contour fitting. Color segmentation: If there is a color difference between the slot and the background, convert the image to the HSV color space and segment the slot area by setting color thresholds. Edge detection and contour fitting: Use the Canny edge detection algorithm to detect the edges of the slot, and then through contour searching and fitting, use geometric shapes such as the minimum bounding rectangle or ellipse to approximate the shape of the slot. Parameter calculation includes dimension measurement and position and angle calculation. Dimension measurement: According to the fitted geometric shape, calculate the dimension parameters of the length, width, and diameter of the slot; for example, for the minimum bounding rectangle, its length and width are the length and width of the slot. Position and angle calculation: Calculate the center position coordinates and rotation angle of the slot to determine the position and orientation of the slot on the blood type card.
[0035] In the embodiment of the present invention, the position offset detection sub-model first generates SIFT key point detection and descriptors, including scale space extreme value detection, key point localization, direction assignment, and descriptor generation. Scale space extreme value detection: Construct a scale space through a Gaussian pyramid, detect extreme points at different scales, and obtain potential key points. Key point localization: Use the Hessian matrix to exclude low-contrast points and edge response points and retain stable key points. Direction assignment: Calculate the gradient histogram of the key point neighborhood and assign the main direction to make the feature rotation-invariant. Descriptor generation: With the key point as the center, construct a 16×16 pixel neighborhood, calculate the gradient magnitude and direction in 8 directions, and generate a 128-dimensional SIFT descriptor.
[0036] Secondly, filter the feature matching and outliers, including brute-force matching BF or fast approximate nearest neighbor matching FLANN and the RANSAC algorithm to remove outliers. Brute-force matching BF: Calculate the Euclidean distance between descriptors, retain the matching pairs with a distance ratio less than 0.7, and filter out incorrect matches. FLANN matching: Suitable for large-scale data, accelerate the search through the KD tree, and improve the matching speed. RANSAC algorithm to remove outliers: Use the random sample consensus algorithm to fit the homography matrix H, filter out the incorrect matching points that do not conform to the transformation model, and retain the inliers.
[0037] Finally, calculate the coordinate transformation and offset, and preset the position calibration: collect the blood type card image in the ideal state, extract the coordinates of the reference key points, and establish the preset coordinate system; obtain the actual position coordinates: extract the key point coordinates from the real-time image; decompose the homography matrix: decompose the translation vector, rotation angle, and scaling factor through the homography matrix; confirm the allowable range of the offset through offset calculation: ±0.5mm for the X / Y axis; generate the robotic arm control instruction, and perform coordinate transformation: convert the image pixel coordinate system to the robotic arm base coordinate system, and pre-calibrate the transformation parameters through the calibration plate; instruction mapping: convert the offset into joint angle or end effector displacement instruction according to the robotic arm kinematic model; dynamic adjustment: control the robotic arm to complete translation and rotation adjustments through real-time closed-loop feedback to ensure that the position of the blood type card is consistent with the preset position.
[0038] In the embodiment of the present invention, for the optical character recognition OCR in the information accuracy verification sub-model, first perform text region localization. Locate the text region with a fixed format through template matching or edge detection. For text at non-fixed positions, use the EAST text detection algorithm to dynamically locate the text box; secondly, perform OCR recognition. Basic solution: Use the Tesseract OCR engine and configure a custom dictionary to improve the accuracy; Deep learning solution: Adopt the convolutional recurrent neural network CRNN or TrOCR model to perform end-to-end recognition on complex fonts (such as handwritten fonts and low-contrast texts), and normalize the input image to a fixed size of 32×128.
[0039] In the embodiment of the present invention, the U-Net network structure in the red blood cell agglutination analysis sub-model includes an encoder, a decoder, and an output layer; Encoder: 4 convolutional blocks, namely 3×3 convolution, ReLU, batch normalization, and 2×2 max pooling, used to extract multi-scale features; Decoder: 3 transposed convolutional blocks, namely 2×2 upsampling, 3×3 convolution, and skip connection, used to restore the spatial resolution; Output layer: 1×1 convolution and activation function Sigmoid, used to generate a pixel-level agglutination probability map. When the value ∈[0,1] and >0.5, it is determined as the agglutination region;
[0040] Training strategy of the U-Net network:
[0041] Data augmentation: Rotate the annotated image by ±15°, scale it by 0.8 - 1.2 times, and add Gaussian noise to improve the generalization ability; Loss function: Joint optimization of Dice loss and cross-entropy loss to solve the class imbalance problem; Inference process: Normalize the input image, and its pixel value is 0 - 1; The network outputs a binary mask, with the agglutination region being white and the background being black;
[0042] Agglutination intensity index:
[0043] The number of particles N: The total number of aggregated particles per unit area, used to reflect the degree of aggregation dispersion; The average area S_avg: The average pixel area of a single aggregated particle, used to measure the particle size; The aggregation ratio R: The ratio of the area of the aggregation region to the total area of the reaction region; The roundness C: The degree to which the particle approaches a circle, C = 4π × area / (perimeter^2), and the closer the value is to 1, the more regular the particle is.
[0044] In the embodiments of the present invention, the blood type determination logic is shown in Table 1, and the ABO blood type determination is as follows:
[0045] Table 1
[0046] Agglutination intensity of anti-A serum Agglutination intensity of anti-B serum Blood type determination Strong (R > 50% and S_avg > 200) Weak (R = 10% - 30%) Type A Weak Strong Type B Strong Strong Type AB None (R < 5%) None Type O ;
[0047] Rh factor determination: Detect the reaction of anti-D serum. If the aggregation intensity R > 20% and N > 10, it is determined as Rh positive (+), otherwise it is negative (-). Rule priority: First determine the ABO blood type, and then combine the Rh factor to output the complete blood type (such as A+, B-).
[0048] In the embodiments of the present invention, the system further includes:
[0049] 1. Establish a dataset:
[0050] Core objective: Build a training data environment with strong robustness.
[0051] Technical implementation: The dataset is 100,000 labeled blood type card images, including different batches, shapes, offset scenarios, and label information; Multi-scenario coverage, including real scenario data such as different printing batches (ink depth differences), card shapes (square cards / circular cards / irregularly shaped cards), text offsets (rotation ±15°, translation ±10%), and lighting changes (overexposure / shadow); Intelligent annotation enhancement, using a semi-automatic annotation process: first pre-label with YOLOv5, manually review, and add error samples to the training set for iterative optimization; The label information includes the character coordinates of the ABO blood type (such as "A" / "B" / "O"), the Rh symbol ("+" / "-"), and special marked areas; Data augmentation strategy: Expand the training data by adding operations such as random transformation, color jitter, Gaussian blur, and perspective transformation through software, enabling the model to learn more text variations and enhancing the generalization ability.
[0052] 2. OCR training:
[0053] Core objective: Achieve accurate recognition of medical-specific characters.
[0054] Technical implementation: Font library customization, collecting 20 special fonts for blood type cards. Generating adversarial samples: Using CycleGAN to convert general fonts into device font styles; Tesseract tuning: Improving its recognition accuracy and performance in specific scenarios; Post-processing verification: Establishing a blood type rule engine: Automatically correcting the recognition result of "AB-" to "AB-"; Introducing an N-gram language model and combining blood type distribution probabilities for error correction (such as correcting "0" to "O").
[0055] 3. Training framework:
[0056] Training the PyTorch model with a model framework: a dual-channel feature fusion network, using Adam as the optimizer and a learning rate of 0.001;
[0057] Technical implementation: Importing necessary libraries, first importing PyTorch and its related libraries for subsequent operations; Preparing the dataset, creating a custom dataset class with input data types of images and text features; Defining a dual-channel feature fusion network model, establishing channel models for images and text respectively; Defining the loss function and optimizer, loss function: Using nn.MSELoss() to implement, determining whether the content described by the blood type card image and text belongs to the manufacturer or batch; Optimizer: Using torch.optim.SGD() to define, updating model parameters by calculating the gradients of each sample or mini-batch of samples. Training the model: In the training loop, sequentially complete the operations of forward propagation, calculating the loss, backpropagation, and parameter update; Saving the model: After training, save the model for subsequent use.
[0058] The process of the system in the present invention is: image acquisition, preprocessing (denoising, normalization), multi-task model inference, information verification by comparing with the database, control instruction generation, and robotic arm execution.
[0059] The present invention conducts information verification and database association: Through the patient ID or batch number in the barcode, obtaining preset information such as patient name and expected blood type from the hospital information system HIS or laboratory information system LIS. Triple comparison logic: Barcode and OCR cross-verification: Checking whether the patient ID extracted by OCR contains the core fields in the barcode (such as the last 6 digits); Detection result comparison: Conducting consistency verification on the blood type identification extracted by OCR and the result output by the detection device; Format compliance check: Verifying whether the text information conforms to the preset format.
[0060] Parameter examples in the present invention:
[0061] Rh factor determination: If the agglutination intensity R > 20% and N > 10, it is determined as Rh positive; Allowable range of offset: ±0.5 mm for X / Y axis, and correction is triggered when exceeding the limit; Information verification tolerance: When the label does not match the test result, an audible and visual alarm is immediately triggered and the process is paused.
[0062] The technical solution provided by the present invention is realized through AI vision interpretation technology: automatic identification of the physical characteristics (batch, shape) of the blood type card; real-time detection and correction of the offset of the grasping position; automatic verification of the blood type card information (label, barcode) to ensure consistency with the test result; high-precision interpretation of the red blood cell agglutination state to ensure the accuracy of the blood type result and the full automation of the detection process.
[0063] The technical solution provided by the present invention, the system has the following beneficial effects:
[0064] Multi-task fusion model: A single algorithm simultaneously processes physical characteristic identification, position correction, information verification, and blood type determination, reducing the occupancy of hardware resources; Dynamic offset compensation: Based on visual feedback, millimeter-level position correction is achieved, and the grasping success rate is increased to over 99.5%; Information consistency guarantee: Through the linkage of OCR and the database, automatic verification of the label information and the test result is realized, and the error rate < 0.01%; Anti-interference interpretation: The model trained by data augmentation (simulating light changes, stain interference) still maintains an interpretation accuracy of > 98% in complex environments; Full-process automation: Reducing manual intervention, the detection efficiency is increased by 60%; High precision and accuracy: The accuracy of physical characteristic identification ≥ 99%; The error of position offset correction < 0.1 mm; The accuracy of blood type interpretation > 99.8%, superior to manual interpretation (about 95%); The accuracy of information verification ≥ 99.9%, eliminating the risk of label-result mismatch. Strong compatibility: Supports blood type cards from different manufacturers and batches, and is compatible with various hardware of blood type analyzers.
[0065] In the technical solution provided by the present invention, the system includes an image acquisition module, an AI vision interpretation algorithm module, i.e., a multi-task model based on the deep convolutional neural network CNN, and a control execution module; The image acquisition module uses a high-resolution industrial camera to collect images of the blood type card and the reaction area in real time, and performs denoising and normalization processing on the collected images; The multi-task model is used to obtain an interpretation result based on the images of the blood type card and the reaction area. The multi-task model includes a physical characteristic identification sub-model, a position offset detection sub-model, an information accuracy verification sub-model, and a red blood cell agglutination analysis sub-model; The control execution module is used to drive the robotic arm to complete the grasping, placement, sorting of the blood type card, or the isolation of the wrong blood type card according to the interpretation result. The system improves the accuracy of automatic detection and blood type analysis identification and reduces the risk of misjudgment.
[0066] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0067] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An analysis system based on AI vision interpretation, characterized in that, The system includes: an image acquisition module, an AI vision interpretation algorithm module, i.e., a multi-task model based on the deep convolutional neural network CNN, and a control execution module; the image acquisition module uses a high-resolution industrial camera to collect images of blood type cards and reaction areas in real time, and denoise and normalize the collected images; the multi-task model is used to obtain an interpretation result based on the images of blood type cards and reaction areas, and the multi-task model includes a physical property recognition sub-model, a position offset detection sub-model, an information accuracy verification sub-model, and a red blood cell agglutination analysis sub-model; the control execution module is used to drive the robotic arm to complete the grasping, placement, sorting of blood type cards or the isolation of incorrect blood type cards according to the interpretation result.
2. The system according to claim 1, wherein The physical property recognition sub-model is used to judge the batch number and shape parameters of the blood type card through image segmentation and feature extraction; the position offset detection sub-model uses a key point matching algorithm, i.e., Scale-Invariant Feature Transform (SIFT), to calculate the offset between the actual position and the preset position of the blood type card in real time, and feedback it to the control execution module to control the robotic arm for dynamic adjustment.
3. The system according to claim 1, wherein The information accuracy verification sub-model combines optical character recognition (OCR) and barcode recognition technology to extract the text and barcode information on the blood type card, and automatically compare it with the system database or the detection result. If the comparison is inconsistent, an alarm is triggered; the red blood cell agglutination analysis sub-model uses a U-Net network to segment the agglutination area, combines morphological algorithms to quantify the agglutination intensity, and determines the blood type.
4. The system according to claim 2, wherein The physical property recognition sub-model first performs image preprocessing, including grayscale conversion, filtering and denoising, and contrast enhancement; Grayscale conversion: Turn on the front light source and back light source, collect the image, and convert the color blood type card image into a grayscale image to reduce the amount of calculation and eliminate the interference of color factors on subsequent processing at the same time. The conversion is carried out through the expression Gray = 0.299R + 0.587G + 0.114B, where R, G, and B are the red, green, and blue channel values of the image respectively; Filtering and denoising: Use median filtering or Gaussian filtering to remove the noise in the image and smooth the image; Contrast enhancement: Use histogram equalization or Contrast Limited Adaptive Histogram Equalization (CLAHE) to enhance the contrast of the image. Secondly, identify the batch number, including region localization, character segmentation, feature extraction and recognition; Region localization includes template matching and edge detection and contour analysis. Template matching: If the position of the batch number on the blood type card is relatively fixed, a template of the number region is pre-made, and the number region is located in the pre-processed image through the template matching algorithm. Edge detection and contour analysis: Use the Canny edge detection algorithm to extract the edge information of the image, and then locate the region containing the batch number according to the characteristics of the contour, such as area, aspect ratio, etc. through contour search and screening. Character segmentation includes binarization processing and connected component analysis. Binarization processing: Perform binarization processing on the located number region, convert the image into a black and white image for subsequent character segmentation, and use the automatic threshold method Otsu to determine the binarization threshold. Connected component analysis: Segment the characters in the binary image into individual character regions through connected component analysis, remove noise and small connected regions, and only retain the regions of the characters. Feature extraction and recognition include feature extraction and machine learning or deep learning recognition. Feature extraction: Extract the geometric features, projection features and moment features of the characters. Geometric features include area, perimeter, aspect ratio. Projection features include horizontal projection and vertical projection. Moment features include Hu moments. Machine learning or deep learning recognition: Use learning algorithms such as support vector machine (SVM), K-nearest neighbor (KNN) or deep learning algorithms such as convolutional neural network (CNN) to train and recognize the extracted features, classify the characters into corresponding numbers or letters, and thus obtain the batch number; Finally, extract the shape parameters, including card slot region segmentation and parameter calculation; Card slot region segmentation includes color segmentation and edge detection and contour fitting. Color segmentation: If there is a color difference between the card slot and the background, convert the image to the HSV color space and segment the card slot region by setting color thresholds. Edge detection and contour fitting: Use the Canny edge detection algorithm to detect the edge of the card slot, and then locate and fit the contour, and use geometric shapes such as the minimum bounding rectangle or ellipse to approximate the shape of the card slot. Parameter calculation includes dimension measurement and position and angle calculation. Dimension measurement: Calculate the dimension parameters of the length, width and diameter of the card slot according to the fitted geometric shape. Position and angle calculation: Calculate the central position coordinates and rotation angle of the card slot to determine the position and orientation of the card slot on the blood type card.
5. The system according to claim 2, wherein The position offset detection sub-model first generates SIFT key point detection and descriptors, including scale space extreme value detection, key point localization, direction assignment and descriptor generation; Scale space extreme value detection: Construct a scale space through a Gaussian pyramid, detect extreme points at different scales, and obtain potential key points; Key point localization: Use the Hessian matrix to exclude low-contrast points and edge response points and retain stable key points; Direction assignment: Calculate the gradient histogram of the key point neighborhood and assign the main direction to make the feature rotation invariant; Descriptor generation: With the key point as the center, construct a 16×16 pixel neighborhood, calculate the gradient amplitude and direction in 8 directions, and generate a 128-dimensional SIFT descriptor; Secondly, perform feature matching and outlier filtering, including brute-force matching BF or fast approximate nearest neighbor matching FLANN and using the RANSAC algorithm to remove outliers; Brute-Force Matching BF: Calculate the Euclidean distance between descriptors, retain the matching pairs with a distance ratio less than 0.7, and filter out incorrect matches; FLANN Matching: Suitable for large-scale data, accelerate the search through KD trees, and improve the matching speed; RANSAC algorithm to remove outliers: Use the random sample consensus algorithm to fit the homography matrix H, filter out the incorrect matching points that do not conform to the transformation model, and retain the inliers; Finally, calculate the coordinate transformation and offset. Preset position calibration: Collect the blood type card image in an ideal state, extract the coordinates of the reference key points, and establish a preset coordinate system; Obtain the actual position coordinates: Extract the key point coordinates from the real-time image; Homography matrix decomposition: Decompose the homography matrix to obtain the translation vector, rotation angle, and scaling factor: Through offset calculation, confirm the allowable range of the offset: ±0.5mm on the X / Y axis; Generate the robotic arm control instruction, coordinate system transformation: Convert the image pixel coordinate system to the robotic arm base coordinate system, and pre-calibrate the transformation parameters through the calibration board; Instruction mapping: According to the robotic arm kinematic model, convert the offset into joint angle or end effector displacement instructions; Dynamic adjustment: Through real-time closed-loop feedback, control the robotic arm to complete translation and rotation adjustments to ensure that the position of the blood type card is consistent with the preset position.
6. The system according to claim 3, characterized in that, In the information accuracy verification sub-model, for optical character recognition OCR, first perform text area localization. Locate the text area with a fixed format through template matching or edge detection. For text at non-fixed positions, use the EAST text detection algorithm to dynamically locate the text box; Secondly, perform OCR recognition. Basic solution: Use the Tesseract OCR engine and configure a custom dictionary; Deep learning solution: Adopt a convolutional recurrent neural network CRNN or TrOCR model to perform end-to-end recognition of complex fonts, and normalize the input image to a fixed size of 32×128.
7. The system according to claim 3, characterized in that, The U-Net network structure in the red blood cell agglutination analysis sub-model includes an encoder, a decoder, and an output layer; Encoder: 4 convolutional blocks, namely 3×3 convolution, ReLU, batch normalization, 2×2 max pooling, used to extract multi-scale features; Decoder: 3 transposed convolutional blocks, namely 2×2 upsampling, 3×3 convolution, skip connection, used to restore the spatial resolution; Output layer: 1×1 convolution, activation function Sigmoid, used to generate a pixel-level agglutination probability map. When the value ∈[0,1] and >0.5, it is determined as the agglutination area; Training strategy of the U-Net network: Data augmentation: Rotate the annotated image by ±15°, scale it by 0.8 - 1.2 times, and add Gaussian noise to improve the generalization ability; Loss function: Joint optimization of Dice loss and cross-entropy loss to solve the class imbalance problem; Inference process: Normalize the input image, and its pixel value is 0 - 1; The network outputs a binary mask, with the agglutination area being white and the background being black; Agglutination intensity index: The number of particles N: The total number of aggregated particles per unit area, used to reflect the degree of aggregation dispersion; The average area S_avg: The average pixel area of a single aggregated particle, used to measure the particle size; The aggregation ratio R: The ratio of the area of the aggregation region to the total area of the reaction region; The roundness C: The degree to which the particle approaches a circle, C = 4π × area / (perimeter²), and the closer the value is to 1, the more regular the particle is.