Medical Form Data Recognition Method and System Based on OCR and MLLM

Through the combination of OCR and MLLM, the multi-scale visual feature extraction, low text recognition efficiency and professional knowledge constraints in medical form data recognition are solved, efficient and accurate medical form data recognition and structured processing are achieved, and the digital efficiency and reliability of medical data are improved.

CN119672743BActive Publication Date: 2025-07-04NANJING ARTIFICIAL INTELLIGENCE CHIPS RES INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411807834.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-07-04
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient multi-scale visual feature extraction, low text recognition efficiency, lack of professional knowledge constraints, inaccurate cross-modal feature alignment, and inadequate data verification in data recognition, resulting in insufficient recognition efficiency and reliability.

Method used

Using an OCR and MLLM-based method, through adaptive contrast enhancement and regional importance weight matrix, combined with multi-resolution image processing, text and spatial position features are extracted, text-spatial attention maps are constructed, text recognition and feature alignment are performed, relationship reasoning and data verification are performed in combination with professional knowledge bases, and normalized data is generated.

Benefits of technology

It realizes efficient identification and accurate structure of medical forms, improves the digital efficiency and accuracy of medical data, and ensures the professionalism and logic of data associations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672743B_ABST
    Figure CN119672743B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for medical form data recognition based on OCR and MLLM. The method includes: receiving original image data, generating an optimized image sequence through adaptive contrast enhancement, regional importance weighting, and multi-resolution optimization; performing feature extraction based on the optimized image sequence, constructing a text-spatial attention map to achieve text recognition; performing visual coding and text coding on the optimized image sequence and the text recognition results, and obtaining a unified feature representation through feature alignment; constructing an information hierarchy graph based on the unified feature representation, performing relationship reasoning, and combining professional knowledge verification to obtain structured feature information; performing data verification on the structured feature information to obtain normalized data. The present invention realizes the efficient recognition and accurate structuring of medical forms, and improves the digitalization efficiency of medical data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data recognition, and in particular, relates to a method and system for recognizing medical form data based on OCR and MLLM. Background Art

[0002] The intelligent recognition and structuring of medical form data are important foundations for the construction of medical informatization. A large number of medical forms such as inspection reports, prescription forms, and medical records contain key health information of patients. Converting these unstructured form data into standardized electronic information is of great significance for improving medical efficiency, supporting clinical decision-making, and promoting medical big data analysis. Automated form data recognition technology can reduce the workload of manual entry, reduce data error rates, accelerate the digitalization process of medical data, and provide data support for the development of intelligent healthcare.

[0003] Currently, the recognition of medical form data mainly adopts a method that combines traditional OCR technology with rule templates. The typical processing flow includes steps such as image preprocessing, text detection, character recognition, and template matching. Some studies have improved the accuracy of text recognition by introducing deep learning technologies, such as using CNN for feature extraction, LSTM for sequence modeling, or Transformer to enhance the processing ability for long texts. In terms of form structure analysis, it mainly relies on predefined layout analysis rules and fixed table parsing templates, and realizes the recognition of table lines and the division of cells through computer vision technologies such as edge detection and connected component analysis.

[0004] However, the existing technologies still face many challenges in processing medical forms: First, the extraction of multi-scale visual features in medical forms is insufficient. For areas that simultaneously contain visual elements of different sizes such as table lines, text, and numbers, it is difficult to balance the feature expression of the global structure and local details. Second, the spatial position information is ignored during the text recognition process, resulting in low serial recognition efficiency of the text within table cells and prone to misalignment errors. Third, there is a lack of professional medical knowledge constraints, and it is impossible to accurately understand the semantic associations between test items and corresponding values, units, and reference values. Fourth, the alignment of cross-modal features is not precise enough, and the mapping relationship between visual features and text features is fuzzy, affecting the accuracy of structured information. Fifth, the data verification and correction lack overall consistency consideration, and local data correction may lead to the destruction of the overall logical relationship. These technical problems restrict the efficiency and reliability of medical form data recognition. Summary of the Invention

[0005] The object of the invention is to provide a method and system for recognizing medical form data based on OCR and MLLM to solve the above problems existing in the prior art.

[0006] Technical solution, a method for identifying medical form data based on OCR and MLLM, includes the following steps:

[0007] S1. Receive the original image data of the professional field form, perform contrast enhancement to obtain an enhanced image; based on the enhanced image, calculate the region weights to obtain a region importance weight matrix; based on the region importance weight matrix, adjust the resolution to obtain a multi-resolution image sequence; perform edge optimization processing on the multi-resolution image sequence to obtain a multi-resolution optimized image sequence;

[0008] S2. Based on the multi-resolution optimized image sequence, extract feature vectors to obtain text feature vectors and spatial position feature vectors; based on the text feature vectors and spatial position feature vectors, construct an attention association to obtain a text-spatial attention map; based on the text-spatial attention map, perform text recognition to obtain an initial text recognition result and a confidence matrix; based on the initial text recognition result and the confidence matrix, extract low-confidence regions and optimize them to obtain an updated text recognition result;

[0009] S3. Perform visual encoding on the multi-resolution optimized image sequence to obtain a visual feature tensor; perform text encoding on the updated text recognition result to obtain a text feature tensor; based on the visual feature tensor and the text feature tensor, calculate feature alignment to obtain a cross-modal alignment matrix; based on the cross-modal alignment matrix, fuse the visual feature tensor and the text feature tensor to obtain a unified feature representation;

[0010] S4. Based on the unified feature representation, construct a hierarchical graph structure to obtain an information hierarchical graph; perform relationship reasoning on the information hierarchical graph to obtain a structure relationship matrix; based on the structure relationship matrix and a pre-stored professional knowledge base, verify to obtain a relationship confidence matrix; integrate the structure relationship matrix and the relationship confidence matrix to obtain structured domain feature information;

[0011] S5. Perform data verification on the structured domain feature information to obtain a verification result matrix; based on the verification result matrix, generate normalized data.

[0012] A medical form data recognition system based on OCR and MLLM includes:

[0013] At least one processor; and,

[0014] A memory communicatively connected to at least one of the processors; wherein,

[0015] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method for identifying medical form data based on OCR and MLLM.

[0016] Beneficial effects: The present invention solves the imaging quality problem of diversified medical forms, providing a high-quality image basis for subsequent recognition; accurately captures the correlation features between the text content and position in the form, breaking through the limitations of traditional OCR in processing complex tables; realizes the deep integration of visual layout information and text semantic information, providing strong support for accurately understanding the form content; conducts in-depth reasoning by combining professional knowledge, ensuring the professionalism and logic of data correlation and realizing the automatic guarantee of data quality; improves the efficiency and accuracy of medical data digitization, providing reliable technical support for medical informatization construction. Brief Description of the Drawings

[0017] Figure 1 It is a flowchart of the method of the present invention.

[0018] Figure 2 It is a flowchart of step S1 of the present invention.

[0019] Figure 3 It is a flowchart of step S2 of the present invention.

[0020] Figure 4 It is a flowchart of step S3 of the present invention.

[0021] Figure 5 It is a flowchart of step S4 of the present invention.

[0022] Figure 6 It is a flowchart of step S5 of the present invention. Detailed Embodiment

[0023] As Figure 1 shown, the present application proposes a medical form data recognition method based on OCR and MLLM, including the following steps:

[0024] S1. Receive the original image data of the form in the professional field, perform contrast enhancement to obtain an enhanced image; calculate the regional weight based on the enhanced image to obtain a regional importance weight matrix; adjust the resolution based on the regional importance weight matrix to obtain a multi-resolution image sequence; perform edge optimization processing on the multi-resolution image sequence to obtain a multi-resolution optimized image sequence;

[0025] S2. Extract feature vectors based on the multi-resolution optimized image sequence to obtain text feature vectors and spatial position feature vectors; construct an attention association based on the text feature vectors and spatial position feature vectors to obtain a text-space attention map; perform text recognition based on the text-space attention map to obtain an initial text recognition result and a confidence matrix; extract low-confidence regions based on the initial text recognition result and the confidence matrix and perform optimization to obtain an updated text recognition result;

[0026] S3. Visually encode the multi - resolution optimized image sequence to obtain a visual feature tensor; perform text encoding on the updated text recognition result to obtain a text feature tensor; calculate feature alignment based on the visual feature tensor and the text feature tensor to obtain a cross - modal alignment matrix; fuse the visual feature tensor and the text feature tensor based on the cross - modal alignment matrix to obtain a unified feature representation;

[0027] S4. Based on the unified feature representation, construct a hierarchical graph structure to obtain an information hierarchy graph; perform relationship reasoning on the information hierarchy graph to obtain a structure relationship matrix; verify based on the structure relationship matrix and a pre - stored professional knowledge base to obtain a relationship confidence matrix; integrate the structure relationship matrix and the relationship confidence matrix to obtain structured domain feature information;

[0028] S5. Perform data verification on the structured domain feature information to obtain a verification result matrix; generate normalized data based on the verification result matrix.

[0029] As Figure 2 shown, according to one aspect of the present application, step S1 is further:

[0030] S11. Receive the original image data of the professional field form, calculate the gray - level histogram distribution characteristics of each region of the image; dynamically adjust the local contrast parameter based on the gray - level histogram distribution characteristics, perform adaptive histogram equalization processing, and generate an enhanced image;

[0031] S12. Based on the enhanced image, extract an image feature map through a convolutional neural network; calculate the attention scores of each region in the image feature map; construct a regional importance distribution based on the attention scores, and generate a regional importance weight matrix representing the importance degree of each region;

[0032] S13. Based on the weight values in the regional importance weight matrix, resample different regions in the enhanced image at different resolution sampling rates to generate a multi - resolution image sequence containing a predetermined number of different resolution versions;

[0033] S14. Based on the multi - resolution image sequence, use an anisotropic diffusion filtering algorithm with an adaptive threshold to suppress image noise, and at the same time use an edge enhancement algorithm based on local gradients to strengthen the edge features of the image, and generate a multi - resolution optimized image sequence.

[0034] In an embodiment of the present application, the process of performing adaptive histogram equalization is specifically as follows: Enhance the image Ie(x, y) = φ(x, y)·[H(x, y) - μ(x, y)] / σ(x, y) + μ(x, y); where H(x, y) = ∑[w(i, j)·I(i, j)] / (∑w(i, j)), H(x, y) is the local weighted histogram mean, w(i, j) = exp(-||(i, j) - (x, y)|| 2 / λ 2 ) is the spatial weight function, and I(i, j) is the original image; φ(x, y) = 1 + η·log(1 + var(x, y)) is the adaptive enhancement coefficient, var(x, y) is the local variance; μ(x, y) is the local mean, σ(x, y) is the local standard deviation; λ is the spatial weight decay parameter; η is the enhancement intensity parameter; x, y are pixel coordinates; the local statistics are calculated within an r×r window, and r is the local window radius.

[0035] The process of different resolution sampling rates is specifically as follows: The sampling rate R(x, y) = Rmin + (Rmax - Rmin)·[1 - exp(-θ·W(x, y))]; where W(x, y) is the regional importance weight; Rmin is the minimum sampling rate, Rmax is the maximum sampling rate; θ is the sampling rate adjustment parameter; the coordinates of the sampled image (x', y') = (x·R(x, y), y·R(x, y)); x, y are the coordinates of the original image; the sampling rate R(x, y) ∈ [Rmin, Rmax] and increases monotonically with the regional importance; θ controls the steepness of the change in the sampling rate.

[0036] The process of generating the regional importance weight matrix is specifically as follows: The regional importance weight W(x, y) = α·A(x, y) + β·S(x, y) + γ·C(x, y); where A(x, y) = softmax(V(x, y)·K(x, y) T / sqrt(d))·V(x, y), A(x, y) is the spatial attention score, V(x, y) is the regional visual feature, K(x, y) is the regional query feature, d is the feature dimension; S(x, y) = exp(-||P(x, y) - P(x', y')|| 2 / σ 2), where S(x, y) is the spatial smoothness, P(x, y) is the position coordinate, and σ is the smoothing factor; C(x, y) = ∑(w_i·f_i(x, y)) / (∑w_i), where C(x, y) is the channel importance, f_i(x, y) is the feature response of the i-th channel, and w_i is the channel weight; α, β, and γ are weight coefficients, and α + β + γ = 1; x, y are the image coordinate positions; the weight coefficients are obtained by optimizing through gradient descent.

[0037] Anisotropic diffusion algorithm: The image optimization function I'(x, y, t) = div[c(x, y, t)·▽I(x, y, t)]; where c(x, y, t) = exp(-||▽I(x, y, t)|| 2 / k 2 ) is the diffusion coefficient, and k is the gradient threshold parameter; ▽I is the image gradient, and div is the divergence operator; the edge enhancement term E(x, y) = I(x, y) + λ·sign(▽ 2 I)·|▽I|, where ▽ 2 I is the Laplacian operator; t is the iteration time step; x, y are the pixel coordinates; λ is the edge enhancement intensity parameter.

[0038] In this embodiment, by constructing an adaptive image enhancement and optimization process, robust image preprocessing is achieved for the diverse imaging quality problems commonly found in medical forms (such as uneven illumination, low contrast, noise interference, etc.). First, adaptive histogram equalization is performed by dynamically adjusting the local contrast parameter, effectively improving the detail visibility of different regions; second, the region importance weights calculated by the attention mechanism are used to accurately locate and enhance the key information regions (such as test values, reference values, etc.), reducing the interference of irrelevant backgrounds; third, through the importance-based multi-resolution sampling strategy, while maintaining the high resolution of the key regions, downsampling is performed on the secondary regions, reducing the consumption of computing resources; finally, the combination of anisotropic diffusion filtering with an adaptive threshold and edge enhancement based on local gradients is used to maintain the clarity of the text edges while suppressing noise, providing a high-quality image basis for subsequent optical character (OCR) recognition. This embodiment is particularly suitable for the characteristics of medical forms and can effectively handle the enhancement requirements of various visual elements such as numbers, texts, and table lines in the form.

[0039] According to one aspect of the present application, step S12 is further as follows:

[0040] S121. Based on the enhanced image, extract hierarchical visual features through a multi-scale convolutional network; based on the hierarchical visual features, construct a feature pyramid to generate a multi-scale feature tensor containing different resolution information;

[0041] S122. Based on the multi-scale feature tensors, use the adaptive feature aggregation algorithm to calculate the attention weights for each scale layer; perform weighted fusion on the multi-scale feature tensors and the corresponding attention weights to generate the feature aggregation tensor;

[0042] S123. Based on the feature aggregation tensor, construct a spatial attention network, and generate an attention feature map through double modeling of channel attention and position attention;

[0043] S124. Based on the attention feature map, use the pre-configured probability graph model to model the spatial dependence relationship to obtain the spatial dependence relationship model; based on the spatial dependence relationship model, use the variational inference method to calculate the conditional probability distribution between regions to generate the spatial dependence matrix;

[0044] S125. Based on the spatial dependence matrix and the attention feature map, construct a regional importance evaluation network, and generate a regional score matrix through comprehensive calculation of multi-dimensional indicators;

[0045] S126. Based on the regional score matrix, use the adaptive threshold segmentation algorithm to stratify the scores, and perform smoothing processing through density estimation to generate the final regional importance weight matrix.

[0046] In this embodiment, by constructing a multi-level regional importance evaluation mechanism, in view of the information value differences in different regions of the medical form, accurate attention allocation is achieved. First, through the multi-scale convolutional network and the feature pyramid structure, the multi-resolution expression of image features is realized, and the hierarchical visual features from the overall layout to the detailed text are accurately captured; second, the adaptive feature aggregation algorithm is used to dynamically adjust the weights of different scale layers, effectively balancing the importance of the global structure and local details; third, through the dual attention modeling mechanism, the information distribution in both the channel dimension and the spatial dimension is considered simultaneously, and the key regions in the form are accurately located; then, the probability graph model is used to deeply model the spatial dependence relationship, and the association strength between different regions is accurately grasped; finally, through the smoothing processing of adaptive threshold segmentation and density estimation, a stable and continuous importance weight distribution is generated. This embodiment is particularly suitable for the differential processing requirements of different types of information such as test items, numerical values, and reference values, and provides important region-oriented information for subsequent recognition and parsing.

[0047] As Figure 3 shown, according to one aspect of the present application, step S2 is further as follows:

[0048] S21. Based on the multi-resolution optimized image sequence, extract the visual features of the text region through the first convolutional neural network to obtain the text feature vector; at the same time, extract the spatial layout information of the text through the second convolutional neural network to obtain the spatial position feature vector;

[0049] S22. Calculate the similarity matrix between the text feature vector and the spatial position feature vector; based on the similarity matrix, perform normalization processing through the softmax function to generate a text-spatial attention map representing the association strength between the text region and the spatial region;

[0050] S23. Weightedly fuse the text-spatial attention map with the text feature vector and the spatial position feature vector to obtain a fused feature vector; input the fused feature vector into a pre-configured text recognition model for decoding to generate an initial text recognition result; based on the initial text recognition result, calculate the confidence score of each recognition result to construct a confidence matrix;

[0051] S24. Based on the initial text recognition result, extract the image patches corresponding to the regions in the confidence matrix that are lower than the preset threshold to obtain low-confidence region image patches; based on the low-confidence region image patches, re-recognize them using a pre-configured high-precision text recognition model to obtain a re-recognition result; replace the low-confidence region image patches with the re-recognition result to obtain an updated text recognition result.

[0052] In an embodiment of the present application, the process of generating the text-spatial attention map is specifically: the attention score At(i, j) = ρ·exp(Qt(i)·Ks(j) T / τ)·M(i, j); where Qt(i) is the text feature query vector, Ks(j) is the spatial feature key vector; M(i, j) = sigmoid(Wm·[Qt(i); Ks(j)] + bm) is the multi-modal fusion gate, Wm is the fusion weight matrix, bm is the bias term; τ is the temperature parameter that controls the smoothness of the attention distribution; ρ is the attention normalization factor to ensure that ∑jAt(i, j) = 1; i is the text sequence position index, j is the spatial feature position index; [;] represents the feature concatenation operation.

[0053] The pre-configured high-precision text recognition model is specifically: the confidence score C(x)=χ*·Pc(x) + (1-χ*)·Ps(x); where Pc(x) = max(softmax(Wc·fx + bc)) is the character-level confidence, fx is the feature vector; Ps(x)= exp(-D(x, T) / ν) is the structural similarity, D(x, T) is the minimum distance from the template library T; χ* is the balance factor; Wc is the classification weight matrix, bc is the bias vector; ν is the similarity scale parameter; x is the text segment to be recognized.

[0054] In this embodiment, by constructing a text recognition framework that integrates spatial location information, high-precision text localization and recognition are achieved for the complex layout and strict data association requirements in medical forms. First, a two-stream convolutional neural network is used to extract text visual features and spatial layout features respectively, and an attention mechanism is used to establish an association mapping between text content and spatial location, effectively solving the serial recognition problem of traditional OCR when processing table structures. Secondly, based on the weighted feature fusion mechanism of the text-spatial attention map, the recognition accuracy of specific regions in the form (such as test item names and numerical regions) is improved. Thirdly, through hierarchical confidence evaluation and adaptive re-recognition strategies, refined recognition is performed on low-confidence regions, improving the recognition accuracy of difficult-to-recognize texts (such as medical terms and special symbols). This embodiment is particularly suitable for complex layout characteristics, can accurately understand the spatial relationship between data items, and provides a reliable text basis for subsequent semantic parsing.

[0055] According to one aspect of the present application, step S23 is further as follows:

[0056] S231. Use an adaptive grid division algorithm to divide the text-spatial attention map into blocks, determine the optimal segmentation threshold according to the distribution characteristics of attention values, and generate an attention region grid;

[0057] S232. Based on the attention distribution of the attention region grid, calculate the weight coefficients of the text feature vector and the spatial location feature vector; based on the weight coefficients, use the weighted average method for feature fusion to generate a fusion feature matrix;

[0058] S233. Based on the fusion feature matrix, calculate the local statistical features of the fusion features in each grid; based on the local statistical features and the pre-stored character feature template library, perform template matching to generate a character similarity matrix;

[0059] S234. Based on the character similarity matrix, construct a spatial constraint relationship graph between characters; based on the spatial constraint relationship graph, calculate the optimal character combination scheme through a graph matching algorithm to generate a text combination scheme matrix;

[0060] S235. Based on the text combination scheme matrix, use a pre-configured conditional random field model to perform probability modeling on the character sequence, calculate the conditional probability distribution of different text combinations, and generate a sequence probability matrix;

[0061] S236. Based on the sequence probability matrix, use a dynamic programming algorithm to find the optimal text sequence path to obtain the initial text recognition result; based on the initial text recognition result, calculate the confidence score of each recognition result to generate a confidence matrix.

[0062] In an embodiment of the present application, the process of partitioning the text-spatial attention map is specifically as follows: Determine the grid weight Gr(i, j) = Ψ(i, j)·exp(-Dt(i, j) / ε)·Ns(i, j); where Ψ(i, j) = ∑[At(x, y)·Δ(x, y∈Bij)] / (|Bij|) is the average attention value within the grid cell Bij, At(x, y) is the attention map value, and Δ(·) is the indicator function; Dt(i, j) = ||Ct(i) - Ct(j)|| 2 is the distance between grid centers, Ct(i) is the centroid coordinate of grid i; Ns(i, j) = exp(-|Wi - Wj| / κ) is the grid similarity, Wi is the feature vector of grid i; ε is the distance decay parameter; κ is the similarity adjustment parameter; i, j are grid indices; |Bij| represents the number of pixels in the grid cell.

[0063] The process of performing probability modeling is specifically as follows: The character sequence probability P(Y|X) = ∏[ψt(yt)·φt(yt, yt-1)·ωt(yt, X)]; where ψt(yt) = softmax(Wp·ht + bp) is the character output probability, ht is the hidden state vector; φt(yt, yt-1) = exp(Wt·[e(yt); e(yt-1)] / υ) is the transition probability, e(·) is the character embedding vector; ωt(yt, X) = sigmoid(Wa·[ht; ct] + ba) is the attention weight, ct is the context vector; Wp, Wt, Wa are weight matrices, bp, ba are bias vectors; υ is the transition temperature parameter; X is the input feature sequence, Y is the output character sequence; t is the sequence position index, and ∏ is the product operator.

[0064] In this embodiment, by constructing a topology-aware text sequence decoding mechanism, high-precision text recognition and organization are achieved for the strict text layout and association requirements in medical forms. First, an adaptive grid division algorithm is used to accurately divide the attention map into blocks. By calculating the centroid with attention weighting, the topological structure of the text area is maintained, effectively avoiding the text breakage problem caused by traditional fixed grid division. Second, based on the feature fusion strategy of attention distribution, the adaptive integration of text features and spatial features is realized, accurately capturing the correspondence between the text and its position in the form. Third, through the deep matching of local statistical features and pre-stored character templates, combined with the global optimization of the spatial constraint relationship graph, the recognition accuracy of professional terms and special symbols is improved. Finally, a conditional random field model is used to probabilistically model the character sequence, and the optimal text path is found through a dynamic programming algorithm, realizing the accurate recognition and reliability evaluation of continuous text in medical forms. This embodiment is particularly suitable for the requirements of standardized layout, can accurately restore the spatial organizational structure of the text, and provides a reliable text basis for subsequent semantic understanding.

[0065] According to one aspect of the present application, step S231 is further as follows:

[0066] S2311. Receive the text-spatial attention map At, calculate the global statistical features of the attention map, including the attention mean μ and standard deviation σ, and determine a threshold based on the adaptive method of μ + 0.5σ to generate an adaptive threshold matrix Th.

[0067] S2312. Receive the adaptive threshold matrix Th and the text-spatial attention map At, divide the attention map into 8×8 grid cells, calculate the attention-weighted average value for each grid cell, and generate a grid attention matrix Ga.

[0068] S2313. Receive the grid attention matrix Ga and the text-spatial attention map At, calculate the attention-weighted centroid for each grid cell, obtain the centroid coordinates by taking the attention-weighted average of the pixel coordinates within the grid, and generate a grid centroid matrix Gc.

[0069] S2314. Receive the grid centroid matrix Gc and the grid attention matrix Ga, and calculate the topological relationship between grid cells. For any two grid cells, calculate the topological association strength based on their centroid distance and attention value similarity, where the distance uses the Euclidean distance and is attenuated by an exponential function, and the attention similarity is calculated by the normalized difference, to generate a topological relationship matrix Tp.

[0070] S2315. Receive the topological relationship matrix Tp and the text-spatial attention map At, and extract features for each grid cell. First, calculate the local statistical features, including the average attention, the standard deviation of attention, the maximum attention value, and the total attention value. Then, extract the spatial distribution features, including the first-order moments and the second-order central moments of the attention-weighted horizontal and vertical directions. Finally, extract the topological relationship features, including the mean, standard deviation, and significant connection ratio of the topological weights, to generate the grid feature matrix Gf.

[0071] S2316. Receive the grid feature matrix Gf and the topological relationship matrix Tp, and through the topology-aware feature aggregation operation, weight and combine the features according to the topological relationship to preserve the local structure information and generate the attention region grid Gr.

[0072] As Figure 4 shown, according to one aspect of the present application, step S3 is further as follows:

[0073] S31. Based on the multi-resolution optimized image sequence, extract multi-level visual features through the vision transformer network and perform feature aggregation to generate a visual feature tensor representing the image content.

[0074] S32. Based on the updated text recognition result, encode the text sequence through the text transformer network to generate a text feature tensor representing the text semantics.

[0075] S33. Calculate the cosine similarity between the visual feature tensor and the text feature tensor, and through normalization processing, generate a cross-modal alignment matrix representing the corresponding relationship between the visual feature and the text feature.

[0076] S34. Based on the corresponding relationship in the cross-modal alignment matrix, weight and fuse the visual feature tensor and the text feature tensor to generate a unified feature representation.

[0077] In an embodiment of the present application, the process of generating the cross-modal alignment matrix is specifically: alignment matrix M(i, j) = π·exp(-D(i, j) / γ)·Z(i, j); where D(i, j) = ||Sv(i) - St(j)|| 2 / ξ + α·[1 - cos(Sv(i), St(j))] is the feature distance metric, where Sv(i) is the visual feature vector and St(j) is the text feature vector; Z(i, j) = sigmoid(Wz·[Sv(i); St(j)] + bz) is the modality fusion gate; cos(·, ·) is the cosine similarity; π is the doubly stochastic matrix after iterative optimization by the Sinkhorn algorithm; γ is the temperature parameter; ξ is the distance scale factor; α is the similarity weight; i is the visual feature index and j is the text feature index; Wz is the fusion weight matrix, bz is the bias vector, and sigmoid( ) is the activation function.

[0078] In this embodiment, by designing a cross-modal feature alignment and fusion mechanism, for the deep correlation features between the text content and visual layout in medical forms, efficient multi-modal information integration is achieved. First, a visual transformer network is used to hierarchically encode the multi-resolution image sequence, and the visual feature dependence relationships at different scales are captured through the multi-head self-attention mechanism, effectively extracting the hierarchical visual representation of the form. Second, a text transformer network is used to semantically encode the recognition results, and the context correlation of the text sequence is established through the position-aware attention mechanism, accurately grasping the semantic relationships of medical terms and numerical values. Third, by designing a cross-modal alignment mechanism based on cosine similarity and combining with the graph attention network of structural similarity, an accurate mapping relationship between visual features and text features is established, solving the alignment deviation problem in traditional methods when dealing with heterogeneous information. Finally, through an adaptive weighted feature fusion strategy, while maintaining modality specificity, complementary enhancement of information is achieved, providing a unified feature representation for subsequent structured parsing. This embodiment is particularly suitable for multi-modal characteristics and can make full use of the complementary information of visual layout and text semantics.

[0079] According to one aspect of the present application, step S31 is further as follows:

[0080] S311. Receive the multi-resolution optimized image sequence Io, divide the image into a sequence of blocks of a fixed size through an image block encoder, and embed position encoding information to generate an image block feature sequence Ps.

[0081] S312. Receive the image block feature sequence Ps, calculate the attention weights between blocks through a multi-head self-attention module, and process the attention output in combination with a feed-forward neural network to generate a self-attention feature matrix Sa.

[0082] S313. Receive the self-attention feature matrix Sa, use a cross-scale transformer encoder to construct the association between feature blocks of different resolutions, and generate a cross-scale feature matrix Cs through multi-level information transmission.

[0083] S314. Receive the cross-scale feature matrix Cs, extract global semantic information through a hierarchical Transformer decoder, and generate a hierarchical decoding feature matrix Hd by combining a multi-layer attention mechanism.

[0084] S315. Receive the hierarchical decoding feature matrix Hd and the cross-scale feature matrix Cs, use a bidirectional Transformer network to fuse the encoding and decoding features, and generate a fused feature matrix Mf through a residual connection.

[0085] S316. Receive the fused feature matrix Mf, perform feature alignment and aggregation through a feature normalization Transformer network, and generate a final visual feature tensor Fv by using adaptive pooling.

[0086] In an embodiment of the present application, the process of generating the cross-scale feature matrix is specifically as follows: Feature update F'l(p)= Fl(p) + ∑[G(s)·T(Fl+s(N(p)))]; where Fl(p) is the feature at position p in layer l; G(s) is the scale offset weight function, and s∈{-1, 0, 1} represents adjacent layers; T(·) = MLP(LayerNorm(·)) is the feature transformation function, and MLP is a multi-layer perceptron; N(p) is the corresponding region of position p in the adjacent layer; T(·) includes a parameter matrix WT and a bias vector bT; l is the feature layer index, and p is the feature position coordinate; the feature dimension remains consistent across all layers.

[0087] The process of fusing the encoding and decoding features is specifically as follows: Fused feature M(p) = ∑[αi·Fi(p) + βi·Bi(p)]; where Fi(p) is the forward feature and Bi(p) is the backward feature; αi and βi are adaptive weight coefficients, calculated by αi =softmax(Wa·Fi(p)) and βi = softmax(Wb·Bi(p)); Wa and Wb are weight matrices; p is the feature position; i is the feature layer index; the feature dimension is unified through a linear transformation.

[0088] In this embodiment, by constructing a hierarchical visual feature extraction framework, efficient multi-scale information encoding is achieved for the complex visual structural features in medical forms. First, an image block encoder is used to convert the image into a sequence of fixed-size blocks, and the spatial position information is maintained through position encoding, effectively solving the memory bottleneck problem of traditional methods when dealing with large-size images. Second, the multi-head self-attention mechanism is used to deeply model the features of the image blocks, accurately capturing the long-range dependencies between different regions and enhancing the understanding ability of the overall structure of the form. Third, the cross-scale transformer encoder is used to establish the association between features of different resolutions, and combined with the multi-level information transmission mechanism, a unified expression from the global layout to local details is achieved. Fourth, the bidirectional transformer network is used to fuse the encoded and decoded features, and the original visual information is effectively maintained through residual connection. Finally, feature normalization and adaptive pooling processing are adopted to generate a unified visual feature representation. This embodiment is particularly suitable for multi-level visual elements and can effectively process different types of visual information such as table lines, text, and graphics.

[0089] According to one aspect of the present application, step S32 is further as follows:

[0090] S321. Receive the text recognition result Rt, perform tokenization processing on the text through a tokenization encoder, and generate a text encoding sequence Te by combining position encoding and semantic type encoding.

[0091] S322. Receive the text encoding sequence Te, calculate the attention distribution between tokens using a multi-head self-attention transformer layer, and generate an attention feature matrix At by processing the attention output through a feed-forward network.

[0092] S323. Receive the attention feature matrix At, construct the association between different semantic levels through a hierarchical transformer encoder, and generate a hierarchical semantic matrix Ht using a multi-layer cross-attention mechanism.

[0093] S324. Receive the hierarchical semantic matrix Ht, integrate the long-range semantic dependencies using a deep transformer decoder, and generate a semantic dependency matrix Dt through a bidirectional feature propagation network.

[0094] S325. Receive the semantic dependency matrix Dt and the hierarchical semantic matrix Ht, fuse the local and global semantic features through a cascaded transformer network, and generate a fused semantic matrix Mt using a residual connection.

[0095] S326. Receive the fused semantic matrix Mt, perform dimension alignment and feature aggregation through a feature aggregation transformer network, and generate a final text feature tensor Ft by combining an adaptive pooling operation.

[0096] In one embodiment of the present application, the process of generating the semantic dependency matrix is specifically as follows: The dependency strength D(i, j) = ρ·tanh(Wd·[si; sj; si - sj; si⊕sj] + bd); where si and sj are node semantic vectors; tanh( ) is the hyperbolic tangent function; ⊕ is the vector outer product operation; Wd is the dependency weight matrix, bd is the bias vector; ρ is the normalization factor; [; ] represents vector concatenation; i and j are sequence position indices; the dependency tree is constructed by the maximum spanning tree algorithm.

[0097] In this embodiment, by designing a deep text semantic encoding framework, in view of the particularity of professional terms and numerical expressions in medical forms, a high-quality text feature representation is achieved. First, through the word segmentation encoder combined with position encoding and semantic type encoding, accurate marking of different types of texts such as medical terms, numerical values, and units is realized, effectively solving the limitations of traditional word segmentation methods in dealing with professional vocabulary; second, the multi-head self-attention transformer layer is used to deeply model the token sequence, and the long-distance semantic dependencies are accurately captured through the attention mechanism, improving the understanding ability of complex medical expressions; third, the hierarchical transformer encoder is used to establish the association between different semantic levels, combined with the multi-layer cross-attention mechanism, to achieve hierarchical representation from the word level to the semantic level; then, through the deep transformer decoder and the bidirectional feature propagation network, the perception ability of the context is enhanced, and the relationship between the test items and their corresponding numerical values is accurately understood; finally, the cascaded transformer network and the feature aggregation operation are used to generate a unified text feature representation. This embodiment is particularly suitable for strict professional expression requirements and can accurately understand and express complex medical information.

[0098] According to one aspect of the present application, step S33 is further as follows:

[0099] S331. Receive the visual feature tensor Fv and the text feature tensor Ft, and use projection transformation to map the two feature tensors into a common semantic space to generate the visual semantic vector Sv and the text semantic vector St.

[0100] S332. Receive the visual semantic vector Sv and the text semantic vector St, calculate the cosine similarity between the two vectors, and adjust the similarity distribution through the temperature scaling factor to generate the initial similarity matrix Si.

[0101] S333. Receive the initial similarity matrix Si and the visual semantic vector Sv, construct a local neighborhood graph, and calculate the structural similarity between nodes through the graph attention network to generate the structural similarity matrix Ss.

[0102] S334. Receive the initial similarity matrix Si and the structural similarity matrix Ss, calculate the comprehensive similarity through an adaptive weighted fusion algorithm, and perform non-maximum suppression to generate a fused similarity matrix Sf.

[0103] S335. Receive the fused similarity matrix Sf, calculate the optimal matching relationship between cross-modal features using the optimal transport algorithm, and generate a matching relationship matrix Rm.

[0104] S336. Receive the matching relationship matrix Rm and the fused similarity matrix Sf, screen reliable matching pairs through bidirectional consistency checking, and perform normalization processing to generate a final cross-modal alignment matrix M.

[0105] In step S334, when calculating the comprehensive similarity, not only focus on strongly correlated (high similarity) information, but also retain some weakly correlated (low similarity) information; use a soft threshold mechanism to preserve the weakly correlated information.

[0106] In this embodiment, by constructing an accurate cross-modal alignment mechanism, for the deep fusion requirement of visual information and text information in medical forms, reliable feature mapping is achieved. First, use projection transformation to map visual features and text features to a unified semantic space, and adjust the similarity distribution through a temperature scaling factor, effectively solving the problem of inconsistent scales of different modal features; second, use a graph attention network to construct a local neighborhood graph, accurately capture the structural similarity between nodes, and enhance the recognition ability of related elements in the form; third, calculate the comprehensive similarity through an adaptive weighted fusion algorithm, combined with non-maximum suppression technology, effectively filter out noisy matches, and maintain the clarity of key information; fourth, use the optimal transport algorithm to calculate the optimal matching relationship between cross-modal features, ensuring the accurate correspondence between visual elements and text content; finally, screen reliable matching pairs through bidirectional consistency checking to generate a stable cross-modal alignment result. This embodiment is particularly suitable for strict information correspondence requirements and can accurately establish the mapping relationship between the visual layout and the text content.

[0107] As Figure 5 shown, according to one aspect of the present application, step S4 is further as follows:

[0108] S41. Based on the unified feature representation, use a hierarchical clustering algorithm for grouping to obtain a grouping result; based on the grouping result, construct a hierarchical graph structure including inspection items, numerical values, units, and reference value nodes, and generate an information hierarchy graph representing the information hierarchy relationship.

[0109] S42. Based on the information hierarchy graph, perform message passing and information aggregation on the nodes in the graph through a graph neural network to obtain updated node representations; based on the updated node representations, infer the association relationships and dependence degrees between nodes, and generate a structure relationship matrix representing the structural relationships between nodes.

[0110] S43. Call the pre-stored professional knowledge base to verify whether the structural relationships in the structure relationship matrix conform to professional knowledge, and obtain a verification result; based on the verification result, calculate the credibility score of each relationship and generate a relationship confidence matrix.

[0111] S44. Based on the confidence levels in the relationship confidence matrix, screen and optimize the structural relationships in the structure relationship matrix to obtain structural relationships that meet the requirements; integrate the structural relationships that meet the requirements into a standardized data structure to generate structured domain feature information.

[0112] In an embodiment of the present application, the process of constructing a hierarchical graph structure is specifically as follows: The node association strength E(i, j) = β·H(i, j) + (1 - β)·L(i, j); where H(i, j) = tanh(Wh·[hi; hj] + bh) is the high-level semantic association, and hi is the semantic feature of node i; L(i, j) = σ(di, j)·exp(-||pi - pj|| 2 / μ) is the low-level structural association, di, j is the node type compatibility, pi is the node position vector; σ(·) is the structure compatibility function; β is the hierarchical balance factor; Wh is the semantic transformation matrix, bh is the bias vector; μ is the spatial scale parameter; i, j are graph node indices.

[0113] In this embodiment, by establishing a hierarchical information structure parsing framework, high-reliability semantic understanding and structured representation are achieved for the complex data association relationships and professional domain knowledge constraints in medical forms. First, an initial hierarchical structure is constructed using an adaptive feature decomposition and spectral clustering method, and the hierarchical dependence relationships between elements such as test items, values, and units are accurately captured through semantic association analysis; second, the feature clusters are mapped to predefined medical concept nodes using a concept ontology mapping network, and a structured representation that conforms to the professional knowledge system is established through dynamic graph construction and edge weight learning; third, based on a multi-hop inference network and hierarchical attention pooling, deep information transmission between nodes is realized, and the implicit associations between test results are accurately inferred; finally, the necessity and sufficiency of the associations are verified through a causal inference network to ensure the logical consistency of the structured information. This embodiment is particularly suitable for strict logical requirements and can accurately understand and maintain complex medical data associations.

[0114] According to one aspect of the present application, step S41 is further as follows:

[0115] S411. Receive the unified feature representation Fu, divide the feature space into semantic subspaces through an adaptive feature decomposition network, calculate the hierarchical structure of feature clusters using a spectral clustering algorithm, and generate an initial hierarchical matrix Hi.

[0116] S412. Receive the initial hierarchical matrix Hi, identify the hierarchical dependence relationship between feature clusters through a semantic association analysis module, and generate a hierarchical dependence matrix Hd using a bottom-up iterative aggregation strategy.

[0117] S413. Receive the hierarchical dependence matrix Hd, construct a concept ontology mapping network, map the feature clusters to predefined professional domain concept nodes (such as test items, numerical values, units, reference values), and generate a concept mapping matrix Cm.

[0118] S414. Receive the concept mapping matrix Cm and the hierarchical dependence matrix Hd, create an initial connection relationship between nodes through a dynamic graph builder, and generate a graph connection matrix Gc using an edge weight adaptive learning algorithm.

[0119] S415. Receive the graph connection matrix Gc, prune and enhance the connection relationship using a graph structure optimization network, calculate the importance score of edges through a graph attention mechanism, and generate an optimized graph structure matrix Go.

[0120] S416. Receive the optimized graph structure matrix Go and the concept mapping matrix Cm, integrate the node semantic information and topological structure information through a hierarchical graph fusion network, and generate a final information hierarchical graph Gh using a graph embedding technique.

[0121] In an embodiment of the present application, the graph structure optimization network is specifically: the edge weight update W'(i, j) = W(i, j) + Δ(i, j)·R(i, j); where Δ(i, j) = sigmoid(WΔ·[hi; hj] + bΔ) is the edge update intensity, hi, hj are node features; R(i, j) = max(0, θ - min(k≠i, j)W(i, k)) is the sparsification regularization term; WΔ is the update weight matrix, bΔ is the bias vector; θ is the sparsification threshold; i, j, k are graph node indices.

[0122] In this embodiment, by constructing a hierarchical information organization framework driven by professional knowledge, a highly reliable semantic hierarchical representation is achieved for the complex data association structure in medical forms. First, an adaptive feature decomposition network is used to divide the unified feature space into different semantic subspaces, and the hierarchical structure of feature clusters is accurately calculated through spectral clustering algorithm, effectively solving the confusion problem of traditional methods in dealing with heterogeneous medical information; second, a semantic association analysis module is used to identify the hierarchical dependence relationships between feature clusters from bottom to top, accurately grasping the logical associations among elements such as test items, numerical values, and units; third, an abstract feature is mapped to specific medical concept nodes through a concept ontology mapping network, and combined with dynamic graph construction technology, a data organizational structure that conforms to professional specifications is established; then, a graph structure optimization network is used to dynamically adjust connection relationships, and important semantic links are highlighted through graph attention mechanism to filter redundant associations; finally, a hierarchical graph fusion network is used to integrate node semantics and topological structures, generating a hierarchical representation that conforms to medical professional knowledge and maintains data integrity. This embodiment is particularly suitable for strict data specification requirements and can accurately reflect the professional association relationships among various medical data.

[0123] According to one aspect of the present application, step S42 is further as follows:

[0124] S421. Based on the information hierarchical graph, calculate the context representation of each node in the graph through the self-attention mechanism; based on the information hierarchical graph, extract the local structural features of the nodes; based on the context representation and the local structural features, generate a node context feature matrix;

[0125] S422. Based on the node context feature matrix and the information hierarchical graph, construct a heterogeneous relationship graph network; based on the heterogeneous relationship graph network, use a relationship-aware message passing algorithm to calculate the direct association strength between nodes, and generate a direct association matrix;

[0126] S423. Based on the direct association matrix and the node context feature matrix, simulate the multi-step information propagation process through a multi-hop inference network, calculate the transfer dependence relationship between nodes, and generate a transfer dependence matrix;

[0127] S424. Based on the transfer dependence matrix, use a hierarchical attention pooling algorithm to perform weighted aggregation on the dependence relationships at different levels, and remove noise associations through sparse constraints to generate a hierarchical dependence matrix;

[0128] S425. Based on the hierarchical dependence matrix and the direct association matrix, construct a causal inference network, and verify the necessity and sufficiency of the association through a counterfactual inference method to generate a causal relationship matrix;

[0129] S426. Based on the causal relationship matrix and the hierarchical dependence matrix, calculate the comprehensive relationship score through the adaptive fusion network, and perform relationship strength calibration to generate the final structural relationship matrix.

[0130] In an embodiment of the present application, the process of simulating the multi-step information propagation process through the multi-hop inference network is specifically as follows: The k-hop relationship strength Rk(i, j) = ∑[λt·∏(m = 1 to k)A(path(t, m))]; where A(p) = softmax(Wa·[hp; hq] + ba) is the single-step attention score, p is the path edge, hp, hq are the node features connected by the edge; path(t, m) represents the m-th step edge of the t-th path with length k; λt is the path importance weight, satisfying ∑λt = 1; Wa is the attention weight matrix, ba is the bias vector; i, j are the starting and ending node indices; k is the hop number parameter.

[0131] In this embodiment, by constructing a deep relationship inference framework, high-precision structured parsing is achieved for the complex logical relationships and implicit dependencies in medical forms. First, the self-attention mechanism is used to calculate the context representation of each node, combined with the local structural features of the nodes, effectively capturing the contextual association between test items and relevant data; second, through the heterogeneous relationship graph network and the relationship-aware message passing algorithm, the direct association strength between different types of nodes is accurately calculated, solving the deviation problem of traditional methods in dealing with heterogeneous relationships; third, the multi-hop inference network is used to simulate multi-step information propagation, and through hierarchical attention pooling, remote dependency relationships are accurately identified, improving the understanding ability of complex medical index relationships; then, the causal inference network is used to verify the necessity and sufficiency of the association, and through the counterfactual reasoning method, the reliability and rationality of the inference result are ensured; finally, the adaptive fusion network is used to integrate various relationship information to generate a structured representation that accurately reflects the internal connection of medical data. This embodiment is particularly suitable for strict logical reasoning requirements and can accurately discover and verify potential associations between data.

[0132] As Figure 6 shown, according to one aspect of the present application, step S5 is further as follows:

[0133] S51. According to the preset professional data rules and value ranges, perform multi-dimensional verification on the data in the structured domain feature information for format, numerical range, and logical relationship, and generate a verification result matrix representing the verification result;

[0134] S52. Organize and encapsulate the data in the verification result matrix in a standardized format to generate normalized data for subsequent processing.

[0135] In this embodiment, by implementing an intelligent data verification mechanism, high-reliability data quality assurance is achieved for the data accuracy and consistency requirements in medical forms. Multidimensional verification of data is performed based on preset professional rules, which can effectively reduce the medical risks brought by data errors.

[0136] In another embodiment of the present application, step S31 can also be:

[0137] S31a. Receive the multi-resolution optimized image sequence Io, calculate the local response features through the multi-scale feature decomposition network, extract the multi-scale context information by combining the spatial pyramid pooling, and generate the hierarchical visual feature matrix Hv.

[0138] S31b. Receive the hierarchical visual feature matrix Hv, calculate the importance distribution of the feature regions using the adaptive attention module, perform feature enhancement through the dynamic weight allocation mechanism, and generate the attention-enhanced feature matrix Av.

[0139] S31c. Receive the attention-enhanced feature matrix Av, construct the cross-scale feature interaction network, realize the information exchange of features at different scales through bidirectional feature propagation, and generate the cross-scale interaction feature matrix Cv.

[0140] S31d. Receive the cross-scale interaction feature matrix Cv, capture the long-range dependence relationship using the non-local attention mechanism, calculate the global context representation through the adaptive feature aggregation, and generate the global context feature matrix Gv.

[0141] S31e. Receive the global context feature matrix Gv and the cross-scale interaction feature matrix Cv, merge the local details and the global semantic information through the residual feature fusion network, and generate the multi-scale fusion feature matrix Mv.

[0142] S31f. Receive the multi-scale fusion feature matrix Mv, perform feature distribution alignment using the feature calibration network, and generate the final visual feature tensor Fv through the adaptive normalization process.

[0143] In another embodiment of the present application, step S32 can also be:

[0144] S32a. Receive the text recognition result Rt, extract the multi-granularity language units using the hierarchical tokenization process, and generate the initial text representation matrix It through the position-aware embedding encoding.

[0145] S32b. Receive the initial text representation matrix It, construct the hierarchical semantic parsing network, extract the semantic features at different levels through the progressive semantic understanding, and generate the multi-level semantic feature matrix Ht.

[0146] S32c. Receive the multi-level semantic feature matrix Ht, establish the dependency relationship between text units using the semantic dependency analysis module, and generate the semantic dependency feature matrix Dt through the structured attention mechanism.

[0147] S32d. Receive the semantic dependency feature matrix Dt, integrate the long-distance semantic association through the context enhancement network, and generate the context enhancement feature matrix Ct using the bidirectional feature propagation mechanism.

[0148] S32e. Receive the context enhancement feature matrix Ct and the multi-level semantic feature matrix Ht, screen the key semantic information using the adaptive feature selection network, and generate the core semantic feature matrix Kt through importance weighted aggregation.

[0149] S32f. Receive the core semantic feature matrix Kt, perform distribution alignment and dimension unification through the feature normalization network, and generate the final text feature tensor Ft using the adaptive pooling operation.

[0150] According to one aspect of the present application, the medical form data recognition method based on OCR and MLLM includes the following steps:

[0151] Step 1. Collect the image to be processed, construct a prompt, and input it into the MLLM for image classification. Construct a prompt based on the image category and the text set, use the large language model to obtain the medical form data recognition information; use the MLLM to perform a finer-grained classification of the medical form to more accurately determine the information to be extracted.

[0152] Step 2. Obtain the information to be extracted according to the image category and perform preprocessing such as cropping on the image to obtain subgraphs. Dynamically adjust the list of information to be extracted according to different types of medical forms to ensure that the extracted content is more accurate and relevant; according to the information to be extracted, further crop and preprocess the original image to generate multiple subgraphs, and each subgraph corresponds to a specific extraction task.

[0153] Step 3. Use OCR to recognize the subgraph to obtain a text set. Construct a prompt based on the center point coordinates of the text lines in the third set and the OCR content recognition result, including setting the task description, providing context information or display related to the task, setting the input format, setting placeholders, setting the output format requirements, input the constructed prompt into the large language model for context understanding and position relationship analysis, extract the expense details recognition information, and output the result.

[0154] In one embodiment of the present application, the user captures an image of a medical examination form by taking a photo or using a scanner. These images can be photos or scans of paper documents. According to the content of the captured images, an appropriate prompt is constructed to describe the basic information of the images and the operations to be performed. The constructed prompt and the images are input into a pre-trained multi-modal large language model (MLLM). The MLLM can process text and image data simultaneously to classify the images. The MLLM outputs the category of the images, such as "Doppler ultrasound examination form", "epidermal cell count examination form", etc.

[0155] According to the category of the images, the system determines the key information to be extracted from this type of examination form from a pre-defined information extraction rule library. For example, for a Doppler ultrasound examination form, the information that may need to be extracted includes "ultrasound prompt", etc. According to the information to be extracted, the original image is cropped to only retain the part containing the key information. Preprocessing operations such as contrast enhancement are performed on the cropped sub-images to improve the accuracy of OCR recognition. After preprocessing, clear sub-images containing key information are obtained.

[0156] Optical character recognition (OCR) technology is used to recognize the text in the preprocessed sub-images to generate a set of text. Based on the image category, the information to be extracted, and the OCR results, a new prompt is constructed. This prompt details the specific information to be extracted from the OCR results. The new prompt and the OCR results are input into the MLLM large model. The MLLM large model performs semantic understanding and information extraction on the input text and outputs structured key information. The system analyzes the inference results of the MLLM large model, extracts the required key information, and outputs it in a structured format. This information can be further used for storage, analysis, or other subsequent processing.

[0157] In this embodiment, through image classification, the system can determine the specific type of the examination form, thereby selecting appropriate preprocessing steps and information extraction rules, which helps to improve the pertinence and accuracy of subsequent processing. Performing preprocessing operations such as cropping and contrast enhancement on the images according to the image category can remove irrelevant information, highlight key content, and improve the accuracy of OCR recognition. Using OCR technology to recognize the text in the preprocessed sub-images generates a high-quality set of text, providing a basis for subsequent semantic understanding and information extraction. Constructing a prompt and inputting the OCR results into the MLLM large model for semantic understanding and information extraction, through the powerful semantic understanding ability of the MLLM, the system can accurately extract the required key information from the OCR results. Analyzing the inference results of the MLLM and outputting structured key information to generate a structured data output, which is convenient for subsequent processing and storage.

[0158] According to one aspect of the present application, a medical form image recognition system based on OCR and MLLM includes: an image acquisition module for acquiring a medical form image to be processed; an MLLM image classification module for classifying the acquired image to determine its specific type (such as a blood routine test form, a urine routine test form, etc.); a correction processing module for performing operations such as cropping and contrast enhancement on the image according to the classification result to generate sub-images; an OCR recognition module for using OCR technology to recognize the text in the sub-images to generate a text set; an MLLM text extraction module for using MLLM to perform semantic understanding and information extraction on the text set; and an output module for generating a structured data output for subsequent processing or storage.

[0159] In this embodiment, by combining OCR technology and a multi-modal large language model (MLLM), efficient and accurate parsing of medical test forms is achieved. Through image classification and preprocessing, it is ensured that OCR recognition is carried out under optimal conditions, thereby improving the accuracy of text recognition; by utilizing the powerful semantic understanding and information extraction capabilities of MLLM, complex medical terms and structured data can be parsed more accurately; the automated processing flow reduces the workload of manual entry and improves the speed and efficiency of data processing; the system can process test forms in various formats provided by different medical institutions, and has high flexibility and versatility.

[0160] According to one aspect of the present application, a medical form data recognition system based on OCR and MLLM includes:

[0161] at least one processor; and,

[0162] a memory communicatively connected to at least one of the processors; wherein,

[0163] the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method for recognizing medical form data based on OCR and MLLM described in any one of the above embodiments.

[0164] The present invention realizes the full - process automated processing from image input to normalized data output by constructing an end - to - end intelligent medical form recognition system. First, through a multi - level image enhancement and optimization mechanism, the problem of diverse imaging quality of medical forms is solved, providing a high - quality image basis for subsequent recognition; second, a text recognition framework that integrates spatial location information accurately captures the correlation features between the text content and location in the form, breaking through the limitations of traditional OCR in processing complex tables; third, a cross - modal feature alignment and fusion mechanism based on transformer realizes the deep integration of visual layout information and text semantic information, providing strong support for accurately understanding the form content; fourth, through a hierarchical information structure parsing framework and in - depth reasoning combined with professional medical knowledge, the professionalism and logic of data association are ensured; finally, an intelligent data verification and correction mechanism is adopted to realize the automated guarantee of data quality. The present invention is particularly adapted to the characteristics of medical forms, can efficiently process different types of medical documents such as inspection reports and prescription forms, improves the efficiency and accuracy of medical data digitization, and provides reliable technical support for the construction of medical informatization.

[0165] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above - mentioned embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A method for identifying medical form data based on OCR and MLLM, characterized in that, It includes the following steps: S1. Receive the original image data of the professional field form, perform contrast enhancement to obtain an enhanced image; Based on the enhanced image, calculate the regional weights to obtain a regional importance weight matrix; based on the regional importance weight matrix, adjust the resolution to obtain a multi-resolution image sequence; Perform edge optimization processing on the multi-resolution image sequence to obtain a multi-resolution optimized image sequence; S2. Based on the multi-resolution optimized image sequence, extract feature vectors to obtain text feature vectors and spatial position feature vectors; based on the text feature vectors and spatial position feature vectors, construct an attention association to obtain a text-spatial attention map; based on the text-spatial attention map, perform text recognition to obtain an initial text recognition result and a confidence matrix; based on the initial text recognition result and the confidence matrix, extract low-confidence regions and optimize them to obtain an updated text recognition result; S3. Perform visual encoding on the multi-resolution optimized image sequence to obtain a visual feature tensor; Perform text encoding on the updated text recognition result to obtain a text feature tensor; based on the visual feature tensor and the text feature tensor, calculate feature alignment to obtain a cross-modal alignment matrix; Based on the cross-modal alignment matrix, fuse the visual feature tensor and the text feature tensor to obtain a unified feature representation; S4. Based on the unified feature representation, construct a hierarchical graph structure to obtain an information hierarchical graph; perform relationship reasoning on the information hierarchical graph to obtain a structural relationship matrix; Based on the structural relationship matrix and the pre-stored professional knowledge base, verify to obtain a relationship confidence matrix; integrate the structural relationship matrix and the relationship confidence matrix to obtain structured domain feature information; S5. Perform data verification on the structured domain feature information to obtain a verification result matrix; Generate normalized data based on the verification result matrix.

2. The medical form data recognition method based on OCR and MLLM according to claim 1, wherein Step S1 is further as follows: S11. Receive the original image data of the professional field form, calculate the gray histogram distribution characteristics of each region of the image; based on the gray histogram distribution characteristics, dynamically adjust the local contrast parameter, perform adaptive histogram equalization processing, and generate an enhanced image; S12. Based on the enhanced image, extract an image feature map through a convolutional neural network; Calculate the attention scores of each region in the image feature map; Based on the attention scores, construct a regional importance distribution to generate a regional importance weight matrix representing the importance degree of each region; S13. Based on the weight values in the regional importance weight matrix, resample different regions in the enhanced image at different resolution sampling rates to generate a multi-resolution image sequence containing a predetermined number of different resolution versions; S14. Based on the multi-resolution image sequence, use an anisotropic diffusion filtering algorithm with an adaptive threshold to suppress image noise, and at the same time use an edge enhancement algorithm based on local gradients to strengthen the edge features of the image to generate a multi-resolution optimized image sequence.

3. The method for identifying medical form data based on OCR and MLLM according to claim 2, wherein Step S2 is further as follows: S21. Based on the multi-resolution optimized image sequence, extract the visual features of the text region through a first convolutional neural network to obtain text feature vectors; at the same time, extract the spatial layout information of the text through a second convolutional neural network to obtain spatial position feature vectors; S22. Calculate the similarity matrix between the text feature vector and the spatial position feature vector; based on the similarity matrix, perform normalization processing through the softmax function to generate a text-spatial attention map representing the association strength between the text region and the spatial region; S23. Perform weighted fusion on the text-spatial attention map, the text feature vector, and the spatial position feature vector to obtain a fused feature vector; Input the fused feature vector into a pre-configured text recognition model for decoding to generate an initial text recognition result; based on the initial text recognition result, calculate the confidence score of each recognition result and construct a confidence matrix; S24. Based on the initial text recognition result, extract the image patches corresponding to the regions in the confidence matrix that are lower than the preset threshold to obtain low-confidence region image patches; based on the low-confidence region image patches, re-recognize using a pre-configured high-precision text recognition model to obtain a re-recognition result; replace the low-confidence region image patches with the re-recognition result to obtain an updated text recognition result.

4. The method for identifying medical form data based on OCR and MLLM according to claim 3, wherein Step S3 is further as follows: S31. Based on the multi-resolution optimized image sequence, extract multi-level visual features through a vision transformer network and perform feature aggregation to generate a visual feature tensor representing the image content; S32. Based on the updated text recognition result, encode the text sequence through a text transformer network to generate a text feature tensor representing the text semantics; S33. Calculate the cosine similarity between the visual feature tensor and the text feature tensor and perform normalization processing to generate a cross-modal alignment matrix representing the corresponding relationship between the visual feature and the text feature; S34. Based on the corresponding relationship in the cross-modal alignment matrix, perform weighted fusion on the visual feature tensor and the text feature tensor to generate a unified feature representation.

5. The method for identifying medical form data based on OCR and MLLM according to claim 4, wherein Step S4 is further as follows: S41. Based on the unified feature representation, use a hierarchical clustering algorithm for grouping to obtain a grouping result; based on the grouping result, construct a hierarchical graph structure including inspection items, numerical values, units, and reference value nodes to generate an information hierarchy graph representing the information hierarchical relationship; S42. Based on the information hierarchy graph, perform message passing and information aggregation on the nodes in the graph through a graph neural network to obtain updated node representations; based on the updated node representations, infer the association relationship and dependence degree between the nodes to generate a structure relationship matrix representing the structural relationship between the nodes; S43. Call a pre-stored professional knowledge base to verify whether the structural relationships in the structure relationship matrix conform to professional knowledge to obtain a verification result; Based on the verification result, calculate the credibility score of each relationship to generate a relationship confidence matrix; S44. Based on the confidence in the relationship confidence matrix, screen and optimize the structural relationships in the structure relationship matrix to obtain structural relationships that meet the requirements; Integrate the structural relationships that meet the requirements into a standardized data structure to generate structured domain feature information.

6. The method for identifying medical form data based on OCR and MLLM according to claim 5, wherein Step S5 is further as follows: S51. Perform multi-dimensional verification on the data in the structured domain feature information in terms of format, numerical range, and logical relationship according to preset professional data rules and value ranges, and generate a verification result matrix representing the verification results. S52. Organize and encapsulate the data in the verification result matrix in a standardized format to generate normalized data for subsequent processing.

7. The method for identifying medical form data based on OCR and MLLM according to claim 6, wherein Step S12 is further as follows: S121. Based on the enhanced image, extract hierarchical visual features through a multi-scale convolutional network; based on the hierarchical visual features, construct a feature pyramid to generate a multi-scale feature tensor containing different resolution information. S122. Based on the multi-scale feature tensor, use an adaptive feature aggregation algorithm to calculate the attention weights for each scale layer; perform weighted fusion on the multi-scale feature tensor and the corresponding attention weights to generate a feature aggregation tensor. S123. Based on the feature aggregation tensor, construct a spatial attention network, and generate an attention feature map through dual modeling of channel attention and position attention. S124. Based on the attention feature map, use a pre-configured probability graph model to model the spatial dependence relationship to obtain a spatial dependence relationship model. Based on the spatial dependence relationship model, use the variational inference method to calculate the conditional probability distribution between regions and generate a spatial dependence matrix. S125. Based on the spatial dependence matrix and the attention feature map, construct a region importance evaluation network, and generate a region score matrix through comprehensive calculation of multi-dimensional indicators. S126. Based on the region score matrix, use an adaptive threshold segmentation algorithm to stratify the scores and perform smoothing processing through density estimation to generate the final region importance weight matrix.

8. The method for identifying medical form data based on OCR and MLLM according to claim 6, wherein Step S23 is further as follows: S231. Use an adaptive grid division algorithm to block the text-spatial attention map to generate an attention region grid. S232. Based on the attention distribution of the attention region grid, calculate the weight coefficients of the text feature vector and the spatial position feature vector. Based on the weight coefficients, use the weighted average method for feature fusion to generate a fusion feature matrix. S233. Based on the fusion feature matrix, calculate the local statistical features of the fused features within each grid; perform template matching based on the local statistical features and a pre-stored character feature template library to generate a character similarity matrix. S234. Based on the character similarity matrix, construct a spatial constraint relationship graph between characters. Based on the spatial constraint relationship graph, calculate the optimal character combination scheme through a graph matching algorithm to generate a text combination scheme matrix. S235. Based on the text combination scheme matrix, use a pre-configured conditional random field model to perform probability modeling on the character sequence, calculate the conditional probability distribution of different text combinations, and generate a sequence probability matrix. S236. Based on the sequence probability matrix, use the dynamic programming algorithm to find the optimal text sequence path to obtain the initial text recognition result; based on the initial text recognition result, calculate the confidence score for each recognition result to generate a confidence matrix.

9. The method for identifying medical form data based on OCR and MLLM according to claim 6, wherein, Step S42 is further as follows: S421. Based on the information hierarchy graph, calculate the context representation of each node in the graph through the self-attention mechanism; based on the information hierarchy graph, extract the local structural features of the nodes; based on the context representation and the local structural features, generate the node context feature matrix; S422. Based on the node context feature matrix and the information hierarchy graph, construct a heterogeneous relationship graph network; based on the heterogeneous relationship graph network, use a relationship-aware message passing algorithm to calculate the direct association strength between nodes and generate a direct association matrix; S423. Based on the direct association matrix and the node context feature matrix, simulate the multi-step information propagation process through a multi-hop inference network, calculate the transfer dependence relationship between nodes, and generate a transfer dependence matrix; S424. Based on the transfer dependence matrix, use a hierarchical attention pooling algorithm to perform weighted aggregation on the dependence relationships at different levels, and remove noise associations through sparse constraints to generate a hierarchical dependence matrix; S425. Based on the hierarchical dependence matrix and the direct association matrix, construct a causal inference network, and verify the necessity and sufficiency of the association through a counterfactual inference method to generate a causal relationship matrix; S426. Based on the causal relationship matrix and the hierarchical dependence matrix, calculate the comprehensive relationship score through an adaptive fusion network and perform relationship strength calibration to generate the final structural relationship matrix.

10. A medical form data recognition system based on OCR and MLLM, characterized in that, Comprising: At least one processor; And A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the OCR and MLLM-based medical form data recognition method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Bill scene recognition method and device, equipment and storage medium

    CN118968537A

  • Bill identification method and device based on machine vision

    CN119068504A