A wound grading model construction method and wound self-analysis system
By constructing a wound grading model, combining wound images and clinical text, and adopting cross-modal fusion and context-aware mechanisms, the subjective differences and accuracy problems in wound assessment are solved, and efficient and accurate wound grading evaluation is achieved, supporting telemedicine and home care.
Patent Information
- Application Number
- CN202510638653.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing technology has large subjective differences, low efficiency, information omission and strong delay in wound assessment, which is difficult to meet the accuracy and timeliness requirements of telemedicine and intelligent care, and the lack of multimodal collaborative modeling mechanism, resulting in inaccurate evaluation and insufficient generalization capabilities.
A hierarchical model is constructed, acquiring wound images through a visual camera, combining clinical text description and structured background information, using feature extraction, semantic relationship prediction and graph modeling, to realize cross-modal fusion of image and text features, introduce a context perception mechanism, build a cross-modal attention fusion mechanism, and generate fusion feature vectors for grading.
It significantly improves the accuracy and generalization ability of wound grading, improves the classification accuracy by about 14.2%, ensures the accuracy of color judgment, supports telemedicine and home care, and provides reliable wound grading analysis.
Smart Images

Figure CN120221096B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wound detection, and more particularly to a method for constructing a wound grading model and a wound self-analysis system. Background Art
[0002] With the continued rise in aging and diabetes rates, the number of patients with chronic wounds is showing a significant growth trend. Chronic wounds, also known as chronic ulcers or chronic refractory wounds, are defined by the International Society for Wound Healing as wounds that cannot be restored to anatomical integrity through normal, orderly, and timely repair processes. In clinical practice, wounds that have not healed and show no tendency to heal after one month of regular treatment are generally considered chronic wounds.
[0003] Traditional wound assessment relies primarily on the physician's subjective judgment and manual recording. However, this assessment model has significant limitations. Differences in professional background, clinical experience, and personal perception often lead to significant subjective variability in the assessment of key indicators such as wound size, depth, color, and exudate. This can negatively impact diagnostic accuracy and the rationale for treatment plans. Furthermore, manual recording is inefficient and prone to omissions and errors, while also facing numerous challenges in data collation, analysis, and storage. More importantly, this traditional assessment method suffers from significant latency, making it difficult to meet the stringent timeliness and accuracy requirements of modern healthcare systems such as telemedicine and intelligent nursing. In telemedicine, physicians cannot assess patients' wounds in real time and intuitively, relying instead on the limited information provided by the patient, which significantly increases diagnostic uncertainty. In intelligent nursing, the lack of objective, accurate, and timely wound assessment data hinders the effectiveness of intelligent nursing systems and prevents them from providing precise, personalized care. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a method for constructing a wound grading model and a wound self-analysis system, which can achieve more accurate, comprehensive and explainable grading evaluation of different types of wounds.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for constructing a wound grading model comprises the following steps:
[0007] The wound surface multi-source data collection step includes obtaining a color wound surface image by photographing the wound surface with a visual camera, and obtaining clinical text description data and structured background information data corresponding to the wound surface image;
[0008] The wound surface image feature extraction step is to process the texture, color and contour edge features of the wound surface image using a feature extraction strategy to obtain an image feature vector;
[0009] The clinical text feature extraction step involves cleaning, segmenting, and labeling clinical text description data and structured background information data to obtain an entity set. The context of each individual entity is then extracted from the entity set, constructed and spliced, and input into a preset wound text language model for semantic relationship prediction. Semantic relationship groups are screened based on the prediction results, and a wound semantic evolution graph is constructed based on the entity and semantic relationship groups. The unstructured global semantic vector extracted by the wound text language model is fused with the graph structure semantic representation to obtain text embedding.
[0010] a multimodal feature fusion step, generating a wound background context vector based on the structured background information data, then dynamically weighting the image feature vector and the text feature vector based on the wound background context vector, and concatenating and fusing the weighted image feature vector and text feature vector to obtain a fused feature vector;
[0011] The hierarchical model construction step is to construct a hierarchical model according to the fused feature vector.
[0012] Furthermore, the clinical text description data reflects the wound description record of the patient's wound in natural language by medical staff, and the structured background information data includes the wound location, wound course number of days, wound type code and patient comorbidity status.
[0013] Furthermore, the clinical text feature extraction step includes an entity extraction strategy, and the entity extraction strategy includes a text preprocessing substep and an entity recognition substep;
[0014] The text preprocessing sub-step is to segment the clinical text description data into terms using a preset knowledge base, and unify entity phrases to obtain a clinical text sequence;
[0015] The entity recognition sub-step uses a preset knowledge base and a preset wound text language model to perform entity annotation on the clinical text sequence to obtain a quintuple. The quintuple includes an entity phrase, the anatomical part where the entity is located, the time point corresponding to the entity, and the diagnostic level related to the entity's estimated relationship. An entity set is constructed based on several groups of quintuples.
[0016] Furthermore, the clinical text feature extraction step includes a relationship extraction strategy, which includes a relationship type definition sub-step, a phrase splicing sub-step, and a relationship semantic screening sub-step;
[0017] The relationship type definition sub-step defines a semantic relationship category set, arbitrarily extracts two entities from the entity set, and matches the relationship type combination in the semantic relationship category set to obtain a semantic relationship group;
[0018] The phrase splicing sub-step extracts the relative context of each pair of entities in the semantic relationship group in the original text and constructs and splices them to obtain a format segment;
[0019] In the relational semantic screening sub-step, the format segment is input into a preset wound text language model to obtain a sentence-level context embedding, and then the sentence-level context embedding is input into a multi-classifier for semantic relationship prediction to obtain a prediction probability, and the semantic relationship group with a prediction probability greater than a preset threshold is recorded as a relation triple.
[0020] Furthermore, the wound semantic evolution graph includes a node set, an edge set and an attribute matrix. The node set is obtained by mapping several entities in the entity set, the edge set is obtained by converting the relationship triples, and the attribute matrix reflects the attributes of the edges in the edge set.
[0021] Furthermore, the clinical text feature extraction step includes a graph update strategy, which includes a temporal graph neural network reasoning sub-step.
[0022] The temporal graph neural network reasoning sub-step embeds the nodes in the node set into vectors, and then obtains node initialization through a layer of nonlinear mapping. The node initialization includes vector words, entity type embedding, diagnosis level embedding and time embedding. The node initialization is calculated through the temporal graph attention network through L layers of propagation to obtain a semantic representation set.
[0023] Furthermore, the graph update strategy also includes a causal semantic attention readout sub-step,
[0024] The causal semantic attention readout sub-step takes the structural embedding of each node in the semantic representation set and the corresponding diagnostic level embedding as input for attention calculation, constructs a causal attention pair, and then acts on the causal relationship pair through a trainable attention vector to obtain an attention score. The attention score is normalized by an activation function to obtain an attention weight, and all nodes are initially weighted and aggregated with the attention weight to obtain a graph structure semantic representation.
[0025] Furthermore, the multimodal feature fusion step includes a fusion strategy, which includes splicing the dynamically weighted image feature vector and text feature vector to obtain a spliced representation, and then inputting the spliced representation into a cross-modal Transformer encoder to obtain a fused feature vector through a multi-head attention mechanism.
[0026] A wound self-analysis system includes a color card, the color card includes a handheld area and a color area, the color area is provided with at least three marking points and a plurality of color blocks of different colors, and further includes
[0027] An image acquisition module acquires an image including the color card and the wound surface captured by a visual camera as an image to be processed;
[0028] a color card calibration module, which identifies a color card area in the image to be processed and deflects the angle of the color card area according to the positions of the three marking points to obtain a corrected image;
[0029] An image processing module, which segments the corrected image using an edge contour algorithm to obtain a color region, maps the color of each color block in the color region with the color of each color block in a preset standard color card to obtain a color difference, and performs color correction on the image to be processed based on the color difference to obtain an image to be analyzed;
[0030] The wound surface analysis module retrieves clinical text description data and structured background information data corresponding to the image to be analyzed, and inputs them into the grading model synchronously with the image to be analyzed to output the wound surface grade.
[0031] Furthermore, the image processing module includes a color mapping strategy, which includes constructing a color mapping relationship between each color block color and each color block color in a preset standard color card through a three-channel independent polynomial regression method to obtain color difference.
[0032] Beneficial effects of the present invention:
[0033] 1. It integrates two types of heterogeneous information, wound images and clinical texts, and realizes inter-modal semantic collaborative modeling and complementary information enhancement by introducing a context-aware cross-modal attention fusion mechanism. It significantly improves the comprehensive discrimination ability of complex wound grading features (such as color-exudation-structure-time), and solves the problem that traditional image single-modality modeling cannot explain key diagnostic information such as exudation degree and disease progression.
[0034] 2. Through the wound semantic evolution graph modeling mechanism and graph neural network reasoning, unstructured text descriptions can be converted into a clearly structured medical entity relationship graph. By modeling the propagation process of time, classification, and pathological logic in the graph, the evolution law of the wound state can be effectively captured, and the system's ability to recognize the "mild / severe" boundary state is improved, providing support for early warning of high-risk levels. Wound background context embedding vectors (such as disease duration, anatomical location, wound type, etc.) are introduced in each feature extraction and fusion stage, realizing scenario-adaptive regulation of the information processing path, enabling the model to dynamically adjust the image channel response and text semantic focus according to contextual factors such as the location of the wound and development time, significantly enhancing the system's generalization ability for different clinical scenarios.
[0035] 3. By fusing image visual features with textual clinical descriptions, a unified cross-modal diagnostic representation vector was constructed. Compared with the traditional image-based U-Net + MLP method, the overall classification accuracy in a real wound grading task was improved by approximately 14.2% (from 72.5% to 86.7%).
[0036] 4. By combining the color card with the wound surface, the analysis system can ensure accurate judgment of the color of the wound surface. Specifically, by comparing the color difference between the color card in the actual captured image and the template color card, the color deviation caused by factors such as shooting angle and lighting conditions can be effectively eliminated, ensuring the authenticity of the wound surface color information in the image, and providing a reliable data basis for subsequent wound surface analysis based on color features. In addition, the self-analysis system is designed. Whether in a professional diagnosis and treatment environment in a hospital or in a scenario where patients conduct self-monitoring at home, as long as the shooting equipment with color cards is used in accordance with the specifications, the system can effectively process and analyze the wound surface image and output reliable wound surface grade results, providing strong technical support for telemedicine, home care and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a step diagram of the method for constructing the wound grading model of the present invention;
[0038] Figure 2 This is a flow chart of wound surface image feature extraction in the present invention;
[0039] Figure 3 This is a flow chart of clinical text feature extraction in the present invention;
[0040] Figure 4 is a flow chart of color correction of an image to be processed in the present invention;
[0041] Figure 5 It is the color card structure diagram of the present invention;
[0042] Figure 6 is a first schematic diagram of the self-analysis system of the present invention;
[0043] Figure 7 is a second schematic diagram of the self-analysis system of the present invention;
[0044] Figure 8 is a third schematic diagram of the self-analysis system of the present invention;
[0045] Figure 9 It is a graphic report diagram of the wound surface in the present invention. DETAILED DESCRIPTION
[0046] The present invention will be described in further detail below with reference to the accompanying drawings and embodiments. Identical components are denoted by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the accompanying drawings, and the terms "bottom," "top," "inner," and "outer" refer to directions toward or away from the geometric center of a particular component, respectively.
[0047] Existing technologies have made some progress in wound image processing and medical record text analysis, but current research still focuses on single-modal analysis, which has the following key flaws in actual clinical applications:
[0048] ① Image segmentation methods lack the ability to perceive wound background and generalize. Specifically:
[0049] Currently widely used wound image segmentation methods, such as U-Net and its variants, mainly rely on the texture, color and edge information of a single wound image itself for feature extraction and region segmentation. These models usually divide the wound area into sub-regions such as granulation tissue, necrotic tissue, and exudate, thereby assisting in healing assessment or grade classification; however, in actual clinical environments, wound images have characteristics such as uneven lighting, complex tissue color changes, and blurred edges. Relying solely on visual features can easily lead to segmentation errors. In addition, the images often lack the expression of background information such as the duration of the disease and the location of the wound, making it difficult for the model to accurately judge the stage of the wound or the state of tissue degeneration, limiting the model's generalization ability and clinical interpretability across multiple types of wounds (such as diabetic foot and pressure ulcers).
[0050] ②Text understanding methods cannot accurately extract medical semantic features that are coordinated with images. Specifically:
[0051] Current research is beginning to explore the structural modeling of clinical text, such as nursing records and medical history descriptions, using pre-trained language models (such as BERT) to extract semantic features to assist in wound severity assessment. Although these methods can identify some subjective indicators, such as descriptions of "severe pain," "exudate," and "white edges," they are unable to effectively align descriptions with corresponding areas in the image due to the independence of text and image feature modeling and the lack of a correlation modeling mechanism between text and images, thus affecting the accuracy of the final assessment.
[0052] ③ Lack of multimodal collaborative mechanisms that integrate the wound context, specifically:
[0053] The scientificity and accuracy of wound assessment are highly dependent on the patient's wound background information, including the location of the wound, duration, underlying diseases (such as diabetes, vascular disease), etc. This information is of great guiding significance for judging the risk of wound infection and healing trend; however, existing image processing methods cannot obtain this non-visual information, and text processing methods often mix background information with subjective descriptions, and no special context modeling mechanism is established; the lack of context perception ability makes it impossible to dynamically adjust the model focus when fusing images and text, which limits the accuracy and robustness of cross-modal modeling.
[0054] In summary, the existing technologies have the following common problems in the task of automatic wound grading and assessment: the image model lacks adaptability to the wound background, the text model lacks visual collaborative semantic modeling capabilities, and there is a lack of an effective multimodal fusion mechanism between the two; therefore, the present invention designs a method for constructing a wound grading model and a wound self-analysis system, which can integrate specialized cross-modal modeling of wound image features, clinical description semantic information and background context to improve the accuracy, reliability and versatility of wound grading assessment, and accurately and quickly evaluate the grade and analyze the credibility of the wound based on the constructed grading model.
[0055] like Figure 1 As shown in the figure, it includes the steps of multi-source wound data collection, wound image feature extraction, clinical text feature extraction, multimodal feature fusion and grading model construction, and then combines the wound context-aware cross-modal attention fusion method and its automatic grading system to achieve a more accurate, comprehensive and explainable grading evaluation of different types of wounds.
[0056] The wound surface multi-source data acquisition step is to obtain a color wound surface image by taking a visual camera to capture the wound surface, which is recorded as ,in and Represent the height and width of the image respectively, and form a color input through the RGB three channels to fully record the visual information of the wound surface. In addition, all images are unified to the same size. , which can eliminate the processing difficulties caused by image size differences and improve the efficiency and accuracy of image processing;
[0057] Synchronously obtain clinical text description data and structured background information data corresponding to the wound image, where the clinical text description data reflects the wound description record of the patient's wound in natural language by medical staff ,in represents the i-th word or phrase, The length of the text is 200 words. These texts contain the subjective observations and judgments of medical staff on the wound surface, such as "exudate is high" and "redness at the edge". They are important bases for assessing the wound surface condition. The clinical text is processed by standardized word segmentation, which decomposes the natural language text into meaningful vocabulary units to facilitate subsequent text analysis and feature extraction. For example, a complex sentence can be split into individual words or phrases, such as "patient's left sole wound on the 5th day" can be segmented into "patient", "left sole", "wound", and "5th day", providing a basis for subsequent operations such as medical ontology-guided entity recognition.
[0058] Structured background information data including wound site , wound duration , wound type code and the patient's comorbidity status etc., together constitute the background variable vector: , structured background data is standardized and coded, and different types of background information are converted into formats suitable for computer processing. For example, the text description of the wound site (such as "sole" and "ankle") is converted into digital codes, and the consecutive days of the disease course are normalized, etc., so that it can be better integrated into subsequent model calculations and improve the efficiency of the model's use of background information.
[0059] like Figure 2 As shown in the figure, the wound image feature extraction step is to process the texture, color and contour edge features of the wound image through the feature extraction strategy to obtain the image feature vector; the specific steps are: ①, the collected wound image is normalized. The normalization operation can make the image data have a uniform scale and distribution, eliminate the data deviation caused by factors such as differences in image acquisition equipment and changes in lighting conditions, and ensure the accuracy and stability of subsequent feature extraction.
[0060] ② In order to accurately capture the texture and texture details of the wound area, convolution kernels of different scales (kernel sizes are ) performs convolution operations on the normalized image and, through multi-scale parallel texture response fusion, can perceive the structural differences between complex color texture areas such as "yellow, red, and white" in the wound surface, thereby improving segmentation and classification accuracy.
[0061] ③. Given that wound color is of great medical significance for distinguishing tissue types, this module is designed to enhance the model's ability to distinguish color. First, the RGB image is converted to the HSV color space. The formula is: ,in Indicates the image after converting the RGB image to the HSV color space, through Function for normalized image Convert and then extract hue , saturation , Luminance channel , the attention weight model is performed on each channel, and the formula is: , , ,in is the activation function, is a trainable weight matrix, is the global average pooling operation, and the attention weight is calculated , the color perception map after fusion modulation is expressed as follows: In this way, the model can focus on color areas with medical significance such as "granulation red" and "necrosis yellow", and improve the ability to identify wound tissue types.
[0062] ④. To address the common problems of blurred and eroded wound boundaries, the Sobel edge operator is used to extract the contour map. Then, a convolution operation is used to construct an edge attention map and perform image feature weighting. That is, by introducing an explicit edge guidance mechanism, the saliency of the wound edge is enhanced, the segmentation boundary is clearer, and the regional recognition accuracy is improved, which helps to accurately define the wound area and identify the boundary between the wound and normal skin.
[0063] ⑤. Send the fused texture, color, and edge perception features into the residual convolutional network ResNet-18 for further feature extraction to obtain the final image feature vector , image feature vector It will serve as an important input for subsequent cross-modal fusion and provide key image feature basis for automatic wound grading.
[0064] like Figure 3 As shown in the figure, the clinical text feature extraction steps include wound clinical text description, medical entity recognition, medical relationship extraction, construction of ontology knowledge graph, graph neural network reasoning and clinical text feature vector output;
[0065] These include entity extraction strategy, which includes text preprocessing sub-step and entity recognition sub-step;
[0066] The text preprocessing sub-step (clinical text description of the wound, clinical text description data and structured background information data are cleaned and segmented) is as follows: specifically, the wound description text in the medical record, such as "the patient's left sole wound is on the 5th day, the edge is red, there is a lot of yellow secretion, obvious exudation, accompanied by mild swelling, and infection is considered" is used as the original input, and it is preprocessed through a preset knowledge base. The knowledge base includes medical dictionary segmentation, which divides long text into single words or phrases, unifies entity phrases, and standardizes expressions such as "yellow secretion" and "obvious exudation", and finally obtains a standard clinical text sequence. , to prepare for subsequent analysis;
[0067] The entity recognition sub-step (medical entity recognition, which involves labeling clinical text sequences with entities using a preset knowledge base and a preset wound text language model to obtain quintuples. The quintuples include the entity phrase, the anatomical site where the entity is located, the time point corresponding to the entity, and the estimated diagnostic level associated with the entity. Entity sets are constructed based on several groups of quintuples). Specifically, using a preset knowledge base (wound medical texts such as guidelines and professional literature) and a preset wound text language model (a trained Wound-BERT model), clinically significant medical entities are identified from standard clinical text sequences. Each entity is represented by a quintuple: Indicates that It is an entity phrase, such as "red"; For entity types, like color, bleed, etc.; The anatomical part of the entity, such as "sole"; is the corresponding time point, such as “5th day of the disease course”; To estimate the diagnostic grade, all identified entities form the entity set At the same time, the Wound-BERT model is used to embed the entire standardized text sequence to obtain contextual semantic representation , generating a global semantic representation through attention-weighted aggregation ,in is the attention weight, is the context word vector of the i-th word.
[0068] The Wound-BERT model was constructed as follows: Starting with ClinicalBERT, an open-source basic model with medical semantic understanding capabilities, which has been pre-trained on a large amount of PubMed and MIMIC-III electronic medical record data, ClinicalBERT possesses sufficient medical domain knowledge and initial word vector embedding capabilities, laying the foundation for the construction of Wound-BERT. Clinical wound description corpora were collected from partner hospitals. These corpora contained descriptions such as "exudate" and "redness at the edges" recorded by real doctors. These corpora were desensitized to remove information that might involve patient privacy. The text was then cleaned to remove noise data, such as irrelevant special symbols and garbled characters. Finally, word segmentation was performed using a medical dictionary, and entity annotation was performed on the segmented text to identify the medical entities within the text, forming a training set.
[0069] A training objective is set, randomly masking some words in the training set and asking the model to predict the masked words. For example, for the text "Patient left sole wound [MASK] days," the model needs to predict the word at "[MASK]" based on the context. This approach allows the model to learn the contextual semantics of wound text, enhancing its understanding of the text. By leveraging the annotated entity information, the model learns how to identify and classify different types of medical entities, ultimately building a specialized language model, Wound-BERT, suitable for wound text.
[0070] Among them, medical relationship extraction extracts the context of each individual entity from the entity set, constructs and splices it, and inputs it into the preset wound text language model to perform semantic relationship prediction. The semantic relationship group is screened according to the prediction results, which includes a relationship extraction strategy. The relationship extraction strategy includes the relationship type definition sub-step, the phrase splicing sub-step and the relationship semantic screening sub-step. Specifically,
[0071] In the relationship type definition sub-step, since a single entity cannot fully express the logic and symptom combination meaning in the clinical text, it is necessary to extract the semantic relationship between entities. Therefore, a set of semantic relationship categories is defined first: , and matching the relationship type combination in the semantic relationship category set to obtain a semantic relationship group, such as ("redness and swelling", indication, "infection"), ("yellow exudate", location, "wound edge");
[0072] In the phrase splicing sub-step, the relative context of each pair of entities in the semantic relationship group is extracted from the original text, and the format segment is obtained by splicing. That is, two entities are randomly extracted from the entity set (each pair of candidate entities ), extract relative context from the original text and construct the spliced input in a specific format ,in, It is an entity phrase. is the context of the 5 words before and after the entity, It is a special symbol used to mark entities;
[0073] In the relation semantic screening sub-step, the format segment is input into the preset wound text language model to obtain sentence-level context embedding, and then the sentence-level context embedding is input into the multi-classifier for semantic relationship prediction to obtain the prediction probability. The semantic relationship group with a prediction probability greater than the preset threshold is recorded as a relation triple. Feed it into the Wound-BERT model to get sentence-level context embedding, and take The vector is used as the sentence meaning representation, and then input into the multi-classifier to predict the semantic relationship, and the predicted probability is greater than the threshold (in the present invention = 0.55) are recorded as relation triples ,in The probability threshold is set to filter the relationship triples. (Indicates entity pair Belonging relationship When the predicted probability of and their relationships Will be recorded as a relation triple ,This can filter out relationship predictions with low credibility and improve the reliability of the extraction results.
[0074] Among them, the ontology knowledge graph is constructed, and the wound semantic evolution graph is constructed according to the entity and semantic relationship group. The wound semantic evolution graph includes a node set, an edge set and an attribute matrix. The node set is obtained by mapping several entities in the entity set, the edge set is obtained by converting the relationship triples, and the attribute matrix reflects the attributes of the edges in the edge set. Specifically, the wound semantic evolution graph is constructed based on the previously identified entities and extracted relationship triples. ,in, For a node set, each entity Corresponding to a node , is an edge set, which is converted from the relationship triple and reflects the relationship between entities. is an attribute matrix containing the medical meaning of the edge, such as relationship type, severity label, etc., for each entity node Construct embedding vector (initial representation of node ), By entity phrase Wound-BERT encoded word vectors and entity type embeddings , diagnostic grade embedding vector and temporal embedding after temporal position encoding The final initial input vector is obtained by splicing and then undergoing nonlinear mapping.
[0075] Among them, the time series graph neural network reasoning is used to build a good graph Reasoning based on the input graph structure, including the initial representation of the node And edge attributes, the temporal graph attention network is used to update the node embedding, and the identity update formula of the i-th node in the j-th layer is: ,in is a node The neighbor set of is the linear transformation matrix, is the ReLU activation function, is the attention weight, which is propagated through L layers to obtain the final semantic representation set of all nodes , allowing each node to fuse neighbor information and evolution trajectory.
[0076] Among them, the causal semantic attention readout mechanism, after propagation through the graph neural network, different entities contribute differently to the wound grading. In order to focus on key entities, a causal semantic attention mechanism is designed to input the final representation set of all graph nodes Embed the prior diagnostic level of each node and calculate the causal semantic weight of each node First, we construct causal attention pairs, then calculate the attention score through the trainable attention vector, obtain the weight through softmax normalization, and finally, aggregate the node embeddings according to the weights to obtain the graph structure semantic representation (graph structure reasoning vector). ,This vector focuses on key entities, taking into account both structural and semantic information, and reflects the ,reasoning path and evolution trend of the disease condition.
[0077] Among them, the clinical text feature vector output is the unstructured global semantic vector extracted by the wound text language model and graph structure semantic representation Perform linear fusion to finally generate text embedding ,provides rich text semantic features for subsequent multimodal feature fusion and automatic wound classification. The specific fusion formula is: ,in is the fusion weight matrix, is the fusion bias vector.
[0078] In order to build a fusion network structure of structural specialization + semantic association + context control, and Able to efficiently align, learn complementary, and output fused embedding vectors , design a cross-modal attention fusion mechanism based on wound background context awareness, conduct deep collaborative modeling of the feature representations of the two modalities, mine their semantic associations, complementary information and common diagnostic signals, and finally generate a fused feature vector for automatic wound grading , that is, the multimodal feature fusion step, first, it is necessary to obtain the wound background context vector , this vector is a global context embedding used to guide the cross-modal attention mechanism, expressing the "wound background knowledge" of the current case. It not only provides guidance for image channel weighting and text attention, but also enhances the scene adaptability of the system.
[0079] This vector is generated based on the structured clinical fields in the patient wound case, and mainly includes the following variables:
[0080] This vector is generated based on the structured clinical fields in the patient wound case, and mainly includes the following variables:
[0081] Define the structured background input variable vector as: .
[0082] Then, each dimension of discrete variables is processed by embedding, and the continuous variables are normalized and connected to the multi-layer perceptron (MLP) for high-dimensional semantic mapping: ,in To map discrete codes (such as part, type) into continuous vectors, It is a multi-layer perceptron structure;
[0083] Wound background context vector ,Through structured medical information embedding, scene features are expressed in a semantic way, guiding the multimodal feature modulation mechanism.,The input of this module includes three parts, namely:
[0084] Image feature vector: ;
[0085] Text feature vector: ;
[0086] Wound background context vector: ;
[0087] Next, we introduce the context vector Dynamically weight image channel features:
[0088] ;
[0089] ;
[0090] in, is the Sigmoid activation function; It is an element-wise multiplication, which means channel reweighting.
[0091] This operation ensures that the image features are more focused on the structural details related to the current wound background, for example: emphasizing the identification of "yellow necrotic areas" in the late stage of diabetic foot disease;
[0092] Similarly, the context vector is used to guide the text attention weight distribution:
[0093] ;
[0094] ;
[0095] This operation can help the system improve its response sensitivity to key text fragments such as "pain description" and "exudate level". For example, if the background suggests a long course of illness, the model will automatically increase its attention to expressions such as "continuous exudate" and "granulation tissue exposure".
[0096] After adjusting the weights, concatenate the modulated image and text vectors , and finally input the cross-modal Transformer encoder to obtain the fused feature vector: , which represents the comprehensive feature embedding after context-guided modulation and cross-modal attention interaction, integrating:
[0097] Image structural elements (borders, colors, organizational levels);
[0098] clinical semantic reasoning (evolutionary clues, temporal progression, hierarchical causal pathways);
[0099] Background medical information (disease duration, anatomical location, patient type, etc.).
[0100] The fusion mechanism proposed in this paper realizes the transition from independent modality to semantic collaborative modeling. It not only retains the salient features of image and text, but also realizes dynamic guidance by introducing background context information, so that the fused feature representation is more in line with the real wound grading decision process. The output fusion vector It has the triple fusion characteristics of structural information, clinical semantics, and reasoning clues.
[0101] The grading model construction step is to construct a grading model based on the fusion feature vector, which has the ability to distinguish and medically explain, realizes the automatic grading of wounds, and outputs the multi-category diagnosis grade results. The input vector , output result (Softmax classification probability): , which represents the probability distribution of each classification level, where k is the number of classification categories, and the final classification result is the class with the largest probability: .
[0102] Data Implementation:
[0103] Wound image input: Color photos of any size, acquired using a standard digital camera or mobile phone. The image should contain the entire wound area. Auxiliary information such as skin background and color charts are permitted.
[0104] Text description input: Free-format clinical description of the wound, for example:
[0105] "The ulcer on the left plantar has lasted for 5 days. The edges are red, and there is yellow necrotic tissue in the center. There is a lot of exudate, but no obvious odor."
[0106] Structured contextual input: For example:
[0107] "location": "Left sole",
[0108] "wound_type": "diabetic foot",
[0109] "duration_days": 5,
[0110] "comorbidity": "type 2 diabetes".
[0111] Output example: {"predicted_grade": "Wagner-3", "confidence_score": 0.87}
[0112] Explanatory output:
[0113] Image Grad-CAM heat map: highlighting the necrotic edge area;
[0114] Text keyword attention: Emphasis on "yellow necrosis" and "large amounts of exudate";
[0115] Graph path reasoning diagram: showing "redness + exudate" → "infection" → "Wagner-3".
[0116] Due to the common problem of low efficiency in existing manual measurement and traditional image analysis methods, the manual measurement process is time-consuming and labor-intensive, and cannot meet the needs of efficient and batch processing, which seriously restricts the efficiency of clinical or field work. Secondly, the measurement accuracy of traditional methods is obviously insufficient. Both manual measurement and simple image processing have large subjective errors and system errors, which cannot meet the high-precision requirements of size and color measurement in the field of medical imaging. Finally, traditional positioning and size estimation methods have poor robustness to complex situations such as image rotation, tilt, scale transformation and illumination changes, which can easily lead to inaccurate positioning, which in turn seriously affects the actual application effect and reduces the credibility of diagnosis and analysis. Therefore, the present invention combines the grading model constructed by the wound grading model construction method to design a wound self-analysis system, such as Figure 5 As shown, the color card includes a handheld area and a color area. The handheld area is convenient for users to hold and operate, and prevents the color card area from being contaminated and blocked. The color area is provided with at least three marking points and a number of color blocks of different colors. The present invention takes a color card proposed in an embodiment as an example, wherein the color area is designed as a 5×5 grid layout, with a total of 25 blocks, of which 3 blocks are positioning marks, and the remaining 22 blocks are 22 standard colors. The three positioning marks are respectively located at the upper left, upper right, and lower left corners of the color card area. Each positioning mark adopts a black and white "return" pattern, which is used to locate the specific position and rotation angle of the color card in the image. It also includes
[0117] An image acquisition module acquires an image including the color card and the wound surface captured by a visual camera as an image to be processed;
[0118] The color card calibration module identifies the color card area in the image to be processed and deflects the angle of the color card area according to the positions of the three marking points to obtain a corrected image. Specifically, the coordinate center point of the image is located according to the three "Hui" coordinates, and the rotation angle of the color card is determined based on the three points to perform a rotation transformation on the image to be analyzed;
[0119] The image processing module uses an edge contour algorithm to segment the corrected image to obtain a color region, and maps the color of each color block in the color region with the color of each color block in a preset standard color card to obtain color difference. The image to be processed is color corrected based on the color difference to obtain an image to be analyzed. Specifically, based on the rotation angle and the "Hui" coordinates, the coordinates of the four vertex corners of the color card are calculated, the area of the color card is determined based on the coordinate points, and then the color of each color block is compared with the color of each color block in the preset standard color card through a three-channel independent polynomial regression method to construct a color mapping relationship to obtain color difference;
[0120] The wound surface analysis module retrieves the clinical text description data and structured background information data corresponding to the image to be analyzed, and inputs them into the grading model synchronously with the image to be analyzed to output the wound surface grade.
[0121] Combine Figure 6-9 As shown in FIG, the wound self-analysis system includes a management module, which is used to distribute and record wound images, clinical text description data, and structured background information data after being processed by a hierarchical model. Figure 9 Output the report as shown.
[0122] The wound self-analysis system can ensure that the analysis system accurately judges the color of the wound by combining the color card with the wound. Specifically, by comparing the color difference between the color card in the actual captured image and the template color card, it can effectively eliminate the color deviation caused by factors such as shooting angle and lighting conditions, ensuring the authenticity of the wound color information in the image, and providing a reliable data basis for subsequent wound analysis based on color features. In addition, whether the self-analysis system is designed in a professional diagnosis and treatment environment in a hospital or in a scenario where patients conduct self-monitoring at home, as long as the shooting equipment with color cards is used in accordance with the specifications, the system can effectively process and analyze the wound image and output reliable wound grade results, providing strong technical support for telemedicine, home care and other fields.
[0123] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that improvements and modifications that do not depart from the principles of the present invention are within the scope of protection of the present invention.
Claims
1. A method for constructing a wound grading model, characterized by: The steps include: The wound surface multi-source data acquisition step includes obtaining a color wound surface image by photographing the wound surface with a visual camera, and obtaining clinical text description data and structured background information data corresponding to the wound surface image; the structured background information data includes wound surface location, wound surface course in days, wound surface type code, and patient comorbidity status; The wound surface image feature extraction step is to process the texture, color and contour edge features of the wound surface image using a feature extraction strategy to obtain an image feature vector; The clinical text feature extraction step involves cleaning, segmenting, and labeling clinical text description data and structured background information data to obtain an entity set. The context of each individual entity is then extracted from the entity set, constructed and spliced, and input into a preset wound text language model for semantic relationship prediction. Semantic relationship groups are screened based on the prediction results, and a wound semantic evolution graph is constructed based on the entity and semantic relationship groups. The unstructured global semantic vector extracted by the wound text language model is fused with the graph structure semantic representation to obtain text embedding. a multimodal feature fusion step, generating a wound background context vector based on the structured background information data, then dynamically weighting the image feature vector and the text feature vector based on the wound background context vector, and concatenating and fusing the weighted image feature vector and text feature vector to obtain a fused feature vector; The hierarchical model construction step is to construct a hierarchical model according to the fused feature vector.
2. The method for constructing a wound grading model according to claim 1, characterized in that: The clinical text description data reflects the wound description record of the patient's wound in natural language by medical staff.
3. The method for constructing a wound grading model according to claim 2, characterized in that: The clinical text feature extraction step includes an entity extraction strategy, which includes a text preprocessing substep and an entity recognition substep; The text preprocessing sub-step is to segment the clinical text description data into terms using a preset knowledge base, and unify entity phrases to obtain a clinical text sequence; The entity recognition sub-step uses a preset knowledge base and a preset wound text language model to perform entity annotation on the clinical text sequence to obtain a quintuple. The quintuple includes an entity phrase, the anatomical part where the entity is located, the time point corresponding to the entity, and the diagnostic level related to the entity's estimated relationship. An entity set is constructed based on several groups of quintuples.
4. The method for constructing a wound grading model according to claim 3, characterized in that: The clinical text feature extraction step includes a relationship extraction strategy, which includes a relationship type definition sub-step, a phrase splicing sub-step and a relationship semantic screening sub-step; The relationship type definition sub-step defines a semantic relationship category set, arbitrarily extracts two entities from the entity set, and matches the relationship type combination in the semantic relationship category set to obtain a semantic relationship group; The phrase splicing sub-step extracts the relative context of each pair of entities in the semantic relationship group in the original text and constructs and splices them to obtain a format segment; In the relational semantic screening sub-step, the format segment is input into a preset wound text language model to obtain a sentence-level context embedding, and then the sentence-level context embedding is input into a multi-classifier for semantic relationship prediction to obtain a prediction probability, and the semantic relationship group with a prediction probability greater than a preset threshold is recorded as a relation triple.
5. The method for constructing a wound grading model according to claim 4, characterized in that: The wound semantic evolution graph includes a node set, an edge set and an attribute matrix. The node set is obtained by mapping several entities in the entity set, the edge set is obtained by converting the relationship triples, and the attribute matrix reflects the attributes of the edges in the edge set.
6. The method for constructing a wound grading model according to claim 5, characterized in that: The clinical text feature extraction step includes a graph update strategy, which includes a temporal graph neural network reasoning sub-step. The temporal graph neural network reasoning sub-step embeds the nodes in the node set into vectors, and then obtains node initialization through a layer of nonlinear mapping. The node initialization includes vector words, entity type embedding, diagnosis level embedding and time embedding. The node initialization is calculated through the temporal graph attention network through L layers of propagation to obtain a semantic representation set.
7. The method for constructing a wound grading model according to claim 6, characterized in that: The graph update strategy also includes a causal semantic attention readout sub-step, The causal semantic attention readout sub-step takes the structural embedding of each node in the semantic representation set and the corresponding diagnostic level embedding as input for attention calculation, constructs a causal attention pair, and then acts on the causal relationship pair through a trainable attention vector to obtain an attention score. The attention score is normalized by an activation function to obtain an attention weight, and all nodes are initially weighted and aggregated with the attention weight to obtain a graph structure semantic representation.
8. The method for constructing a wound grading model according to claim 1, characterized in that: The multimodal feature fusion step includes a fusion strategy, which includes splicing the dynamically weighted image feature vector and text feature vector to obtain a spliced representation, and then inputting the spliced representation into a cross-modal Transformer encoder to obtain a fused feature vector through a multi-head attention mechanism.
9. A wound surface self-analysis system, comprising a grading model constructed based on the wound surface grading model construction method according to any one of claims 1 to 8, characterized in that: The color card includes a handheld area and a color area, wherein the color area is provided with at least three marking points and a plurality of color blocks of different colors, and further includes An image acquisition module acquires an image including the color card and the wound surface captured by a visual camera as an image to be processed; a color card calibration module, which identifies a color card area in the image to be processed and deflects the angle of the color card area according to the positions of the three marking points to obtain a corrected image; An image processing module, which segments the corrected image using an edge contour algorithm to obtain a color region, maps the color of each color block in the color region with the color of each color block in a preset standard color card to obtain a color difference, and performs color correction on the image to be processed based on the color difference to obtain an image to be analyzed; The wound surface analysis module retrieves clinical text description data and structured background information data corresponding to the image to be analyzed, and inputs them into the grading model synchronously with the image to be analyzed to output the wound surface grade.
10. The wound surface self-analysis system according to claim 9, characterized in that: The image processing module includes a color mapping strategy, which includes constructing a color mapping relationship between each color block color and each color block color in a preset standard color card through a three-channel independent polynomial regression method to obtain color difference.
Citation Information
Patent Citations
Gastric ulcer diagnosis device and equipment and storage medium
CN115240847A
Medical Condition Visual Search
US20240339217A1