Calcification region segmentation method and device based on multimodal feature fusion model

The calcified region of the artificial valve is segmented through the multimodal feature fusion model, which solves the problems of low segmentation accuracy and efficiency in the prior art, and achieves higher segmentation accuracy and efficiency.

CN119624994BActive Publication Date: 2025-05-13ANHUI KUNLONG KANGXIN MEDICAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510157413.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The existing artificial valve calcification region segmentation method relies on artificial subjective judgment, has deviations and complex operations, resulting in low segmentation accuracy and efficiency.

Method used

The calcified area segmentation method based on the multimodal feature fusion model was adopted. Text features were extracted on the text diagnosis content of the target patient, and the cross-sectional image of the artificial valve was pre-processed and sectional feature extracted, and combined with the training of the convergent calcified area segmentation model, the segmentation result map was finally obtained.

Benefits of technology

It improves the accuracy and efficiency of calcified area segmentation, reduces the deviation of human subjective judgment, and enhances the comprehensiveness and rationality of the segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624994B_ABST
    Figure CN119624994B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of regional segmentation technology, and discloses a calcification region segmentation method and device based on a multimodal feature fusion model, the method comprising: determining a text feature tensor result according to the text diagnosis content of a case; determining a first image tensor result and a second image tensor result according to a cross-sectional image of an artificial valve; determining a section feature result according to the second image tensor result; inputting the text feature tensor result, the first image tensor result and the section feature result into a calcification region segmentation model for processing to obtain a segmentation result mask tensor; determining a segmentation result map according to the segmentation result mask tensor, and the segmentation result map is used to determine the artificial valve calcification region of the target patient. It can be seen that the implementation of the present invention can improve the diversity, flexibility, pertinence and comprehensiveness of the input parameters of the artificial valve calcification region segmentation model, and improve the segmentation accuracy, reliability, efficiency and convenience of the artificial valve calcification region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and in particular to a calcification region segmentation method and device based on a multimodal feature fusion model. Background Art

[0002] In the field of medical image processing, especially in the diagnosis and treatment of heart valve diseases, accurate segmentation of artificial valve calcification areas is of great significance.

[0003] However, practice shows that the existing artificial valve calcification area segmentation method is mainly for the staff to manually mark and segment the artificial valve calcification area subjectively. However, the subjective consciousness of the calcification area segmentation will be affected by many factors, such as the staff's poor mental state when performing the calcification area segmentation, different staff members have different considerations and focuses on analyzing the same calcification area, etc., which will cause deviations in the segmentation results of the same calcification area under the same conditions to a certain extent. In addition, the calcification area segmentation operation is complicated and cumbersome, which makes the segmentation accuracy and efficiency of the artificial valve calcification area low. Therefore, it is particularly important to provide a calcification area segmentation method that can improve the segmentation accuracy and efficiency of the artificial valve calcification area. Summary of the invention

[0004] The present invention provides a calcification region segmentation method and device based on a multimodal feature fusion model, which can improve the segmentation accuracy and efficiency of the calcification region.

[0005] In order to solve the above technical problems, the first aspect of the present invention discloses a calcification region segmentation method based on a multimodal feature fusion model, the method comprising:

[0006] Perform corresponding text feature extraction operations on the text diagnosis content of the case of the determined target patient to obtain a text feature tensor result;

[0007] Performing corresponding image pre-processing operations on the determined artificial valve cross-sectional image of the target patient to obtain a first image tensor result and a second image tensor result, and performing corresponding section feature extraction operations on the second image tensor result to obtain a section feature result;

[0008] Inputting the text feature tensor result, the first image tensor result and the section feature result into a trained and converged calcification region segmentation model for processing to obtain a segmentation result mask tensor;

[0009] A corresponding post-segmentation processing operation is performed on the segmentation result mask tensor to obtain a segmentation result map, and the segmentation result map is used to determine the artificial valve calcification area of ​​the target patient.

[0010] The second aspect of the present invention discloses a calcification region segmentation device based on a multimodal feature fusion model, the device comprising:

[0011] A text feature extraction module is used to perform corresponding text feature extraction operations on the text diagnosis content of the case of the determined target patient to obtain a text feature tensor result;

[0012] An image feature determination module is used to perform corresponding image pre-processing operations on the determined artificial valve cross-sectional image of the target patient to obtain a first image tensor result and a second image tensor result, and perform corresponding section feature extraction operations on the second image tensor result to obtain a section feature result;

[0013] A mask tensor determination module, used for inputting the text feature tensor result, the first image tensor result and the section feature result into a trained and converged calcification region segmentation model for processing to obtain a segmentation result mask tensor;

[0014] The segmentation result determination module is used to perform corresponding segmentation post-processing operations on the segmentation result mask tensor to obtain a segmentation result map, and the segmentation result map is used to determine the artificial valve calcification area of ​​the target patient.

[0015] The third aspect of the present invention discloses another calcification region segmentation device based on a multimodal feature fusion model, the device comprising:

[0016] A memory storing executable program code;

[0017] a processor coupled to the memory;

[0018] The processor calls the executable program code stored in the memory to execute the calcification area segmentation method based on the multimodal feature fusion model disclosed in the first aspect of the present invention.

[0019] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the calcification area segmentation method based on the multimodal feature fusion model disclosed in the first aspect of the present invention.

[0020] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0021] In an embodiment of the present invention, a corresponding text feature extraction operation is performed on the text diagnosis content of the case of the determined target patient to obtain a text feature tensor result; a corresponding image pre-processing operation is performed on the cross-sectional image of the artificial valve of the determined target patient to obtain a first image tensor result and a second image tensor result, and a corresponding section feature extraction operation is performed on the second image tensor result to obtain a section feature result; the text feature tensor result, the first image tensor result and the section feature result are input into a trained converged calcification area segmentation model for processing to obtain a segmentation result mask tensor; a corresponding segmentation post-processing operation is performed on the segmentation result mask tensor to obtain a segmentation result map, which is used to determine the artificial valve calcification area of ​​the target patient. It can be seen that the present invention can obtain text feature tensor results, image tensor results and section feature results through text feature extraction operations, image pre-processing operations and section feature extraction operations, and input the text feature tensor results, image tensor results and section feature results into the calcification region segmentation model for processing and combined with the segmentation post-processing operation to obtain a segmentation result map, which is beneficial to improving the comprehensiveness, integrity and rationality of the calcification region segmentation method, and is beneficial to improving the diversity, flexibility, pertinence and comprehensiveness of the input parameters of the artificial valve calcification region segmentation model, and further helps to improve the accuracy and reliability of the output region segmentation results, thereby helping to improve the segmentation accuracy, reliability, efficiency and convenience of the artificial valve calcification region. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 It is a flow chart of a calcification region segmentation method based on a multimodal feature fusion model disclosed in an embodiment of the present invention;

[0024] Figure 2 It is a flowchart of another calcification region segmentation method based on a multimodal feature fusion model disclosed in an embodiment of the present invention;

[0025] Figure 3 It is a structural schematic diagram of a calcification region segmentation device based on a multimodal feature fusion model disclosed in an embodiment of the present invention;

[0026] Figure 4 It is a structural schematic diagram of another calcification region segmentation device based on a multimodal feature fusion model disclosed in an embodiment of the present invention;

[0027] Figure 5 It is a structural schematic diagram of another calcification region segmentation device based on a multimodal feature fusion model disclosed in an embodiment of the present invention;

[0028] Figure 6 It is a schematic diagram of the overall process of a calcification region segmentation method based on a multimodal feature fusion model disclosed in an embodiment of the present invention;

[0029] Figure 7 It is a flowchart of a method for determining a text feature tensor result disclosed in an embodiment of the present invention;

[0030] Figure 8 It is a flowchart of a method for determining a valid text content result disclosed in an embodiment of the present invention;

[0031] Fig. 9 It is a schematic diagram of the architecture of a calcification region segmentation model disclosed in an embodiment of the present invention;

[0032] Fig.10 It is a flow chart of a single forward training scheme disclosed in an embodiment of the present invention;

[0033] Fig.11 It is a flowchart of a multiple forward and freeze training scheme disclosed in an embodiment of the present invention;

[0034] Fig.12 It is a flowchart of a multiple forward and separation training scheme disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or end including a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or ends.

[0037] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0038] The present invention discloses a calcification region segmentation method and device based on a multimodal feature fusion model, which can obtain text feature tensor results, image tensor results and section feature results through text feature extraction operations, image pre-processing operations and section feature extraction operations, input the text feature tensor results, image tensor results and section feature results into the calcification region segmentation model for processing and combine with the segmentation post-processing operation to obtain a segmentation result map, which is conducive to improving the comprehensiveness, integrity and rationality of the calcification region segmentation method, and is conducive to improving the diversity, flexibility, pertinence and comprehensiveness of the input parameters of the artificial valve calcification region segmentation model, and is conducive to improving the accuracy and reliability of the output region segmentation results, so as to improve the segmentation accuracy, reliability and segmentation efficiency and convenience of the artificial valve calcification region. The following is a detailed description.

[0039] Embodiment 1

[0040] See also Figure 1 , Figure 1 is a flow chart of a calcification region segmentation method based on a multimodal feature fusion model disclosed in an embodiment of the present invention. Figure 1 The described method can be applied to a calcification region segmentation device based on a multimodal feature fusion model, wherein the device may include a server, wherein the server includes a local server or a cloud server, which is not limited in the embodiment of the present invention. Figure 1 As shown, the calcification region segmentation method based on the multimodal feature fusion model includes the following operations:

[0041] 101. Perform corresponding text feature extraction operations on the text diagnosis content of the case of the determined target patient to obtain a text feature tensor result.

[0042] Optional, the overall flow chart of the calcification area segmentation method based on the multimodal feature fusion model can be referred to the manual Figure 6 The embodiments of the present invention are not limited thereto.

[0043] Optionally, the text diagnosis content of the case may be the text diagnosis content of the TAVR case, or may be other types of text diagnosis content of the case, which is not limited in the embodiment of the present invention.

[0044] Optional, text feature tensor result. For example, the size of the text feature tensor is , Indicates the length of the text after encoding, which is not limited in the embodiment of the present invention.

[0045] 102. Perform corresponding image pre-processing operations on the artificial valve cross-sectional image of the determined target patient to obtain a first image tensor result and a second image tensor result, and perform corresponding section feature extraction operations on the second image tensor result to obtain a section feature result.

[0046] Optionally, the artificial valve cross-sectional image may be an artificial valve cross-sectional CT image, or may be other types of artificial valve cross-sectional images, which is not limited in the embodiment of the present invention.

[0047] Optionally, the first image tensor result and the second image tensor result may be the same image tensor content or different image tensor content, which is not limited in this embodiment of the present invention.

[0048] Optionally, the size of the first image tensor result and the second image tensor result can be , the embodiments of the present invention are not limited thereto.

[0049] Optionally, the corresponding dimensions of the cut surface feature result can be , the embodiments of the present invention are not limited thereto.

[0050] 103. Input the text feature tensor result, the first image tensor result and the section feature result into the trained and converged calcification region segmentation model for processing to obtain a segmentation result mask tensor.

[0051] Optionally, the calcification area segmentation model may be a deep learning model, further, may be a ViLTransUNet deep learning network model, or may be a multimodal feature fusion model ViLTransUNet, which is not limited in the embodiments of the present invention.

[0052] Optionally, the specific method and steps for determining the segmentation result mask tensor in step 103 may refer to but not be limited to the specific relevant records in this article regarding "inputting predetermined training data into the calcification area segmentation model for analysis, and obtaining the segmentation mask tensor training results and text-related vector training results output by the calcification area segmentation model", and the embodiments of the present invention are not limited thereto.

[0053] 104. Perform corresponding post-segmentation processing operations on the segmentation result mask tensor to obtain a segmentation result map, which is used to determine the artificial valve calcification area of ​​the target patient.

[0054] It can be seen that the calcification area segmentation method based on the multimodal feature fusion model described in the embodiment of the present invention can obtain text feature tensor results, image tensor results and section feature results through text feature extraction operations, image pre-processing operations and section feature extraction operations, and input the text feature tensor results, image tensor results and section feature results into the calcification area segmentation model for processing and combined with the segmentation post-processing operation to obtain a segmentation result map, which is beneficial to improving the comprehensiveness, integrity and rationality of the calcification area segmentation method, and is beneficial to improving the diversity, flexibility, pertinence and comprehensiveness of the input parameters of the artificial valve calcification area segmentation model, and further helps to improve the accuracy and reliability of the output region segmentation results, thereby helping to improve the segmentation accuracy, reliability, efficiency and convenience of the artificial valve calcification area.

[0055] In an optional embodiment, the above-mentioned performing corresponding text feature extraction operations on the determined text diagnosis content of the case of the target patient to obtain the text feature tensor result may include:

[0056] Perform corresponding valid text interception operations on the text diagnosis content of the case of the determined target patient to obtain valid text content results;

[0057] Input the valid text content result into the text conversion model with training convergence for processing to obtain the text conversion result;

[0058] The text conversion result is input into the trained converged feature vector extraction model for processing to obtain the text feature tensor result.

[0059] Optional, the flowchart of the method for determining the text feature tensor result can be found in the manual Figure 7 The embodiments of the present invention are not limited thereto.

[0060] Optionally, the text conversion model may be a pre-trained BertTokenizer model, or other models that can achieve the same text conversion function, which is not limited in the embodiment of the present invention.

[0061] Optionally, the text conversion result may be a text Token, which is not limited in the embodiment of the present invention.

[0062] Optionally, the feature vector extraction model may be a pre-trained BertModel model, or other models that can achieve the same feature vector acquisition function, which is not limited in the embodiments of the present invention.

[0063] Optionally, the text diagnosis content of the target patient's case may include, but is not limited to, the patient's basic condition, assessment of the calcification area and severity of calcification, analysis of the symmetry of the aortic valve leaflets, and assessment of the degree of restriction of leaflet activity, etc., and the embodiments of the present invention are not limited thereto.

[0064] Optionally, the above-mentioned valid text interception operation is mainly intended to exclude text content irrelevant to the task to prevent interference with the task, which is not limited in the embodiment of the present invention.

[0065] Optionally, this solution uses the bert-base-chinese model as the pre-trained model of Bert, and inputs the truncated text of length into BertTokenizer to obtain three tensors of length, which respectively represent the index in the vocabulary, the mask value of the input data, and the segment tag index; these tensors are used as the input of BertModel, and after processing by the pre-trained BertModel model, a text feature tensor output of length can be obtained, which is not limited in the embodiment of the present invention.

[0066] It can be seen that this optional embodiment can provide a text feature extraction method, perform corresponding effective text interception operations, text conversion operations and feature vector extraction operations on the text diagnosis content of the case, and obtain a text feature tensor result, which is beneficial to improve the comprehensiveness, progressiveness and rationality of the text feature extraction method, and thus help improve the accuracy and reliability of the determined text feature tensor results.

[0067] In another optional embodiment, the above-mentioned performing corresponding valid text interception operation on the determined text diagnosis content of the case of the target patient to obtain the valid text content result may include:

[0068] Based on the preset word segmentation component library, a corresponding word segmentation operation is performed on the text diagnosis content of the case of the determined target patient to obtain a text word segmentation result, which includes one or more word segmentation text segments;

[0069] Determine the keyword screening results based on the text segmentation results and the preset keyword screening method;

[0070] According to the keyword screening results and the text segmentation results, the corresponding similarity analysis operation is performed to obtain the similarity score results corresponding to each segmented text segment;

[0071] The valid text content result is determined based on each segmented text segment and its corresponding similarity score result.

[0072] Optional, the flow chart of the method for determining the effective text content result can be referred to the instruction manual Figure 8 The embodiments of the present invention are not limited thereto.

[0073] Optionally, the word segmentation component library may be a Jieba library, or other word segmentation components that can implement the word segmentation function and can add uncommon words, which is not limited in the embodiment of the present invention.

[0074] Further optionally, the above-mentioned keyword screening results are determined based on the text segmentation results and the preset keyword screening method. For example, keywords are set, such as "calcification degree of calcification area", etc. The specific keyword setting depends on the text content that needs to be intercepted, and the embodiment of the present invention is not limited.

[0075] Further optionally, the above-mentioned similarity analysis operation is performed according to the keyword screening results and the text segmentation results to obtain the similarity score result corresponding to each segmented text segment. For example: each segmented phrase obtained in the first step is evaluated for similarity with the keyword, and the Thefuzz library is used to implement the similarity comparison of short texts. The hefuzz library calculates the similarity between two short texts by using a method based on the Levenshtein distance, and finally gives a similarity score (i.e., the similarity score result corresponding to the segmented text segment). The score is proportional to the similarity, which is not limited in the embodiments of the present invention.

[0076] Further optionally, the above determines the valid text content result based on each segmented text segment and its corresponding similarity score result. For example, determine the text segment in the area with a higher similarity score, include it in the text clipping result, and finally obtain the valid text content result. The embodiment of the present invention is not limited to this.

[0077] It can be seen that this optional embodiment can provide a method for determining effective text content results, perform corresponding word segmentation operations, keyword screening operations, similarity analysis operations and screening operations on the textual diagnosis content of the case, and obtain effective text content results, which is conducive to improving the comprehensiveness, progressiveness and rationality of the method for determining effective text content results, and further helps to improve the accuracy and reliability of the determined effective text content results.

[0078] Embodiment 2

[0079] See also Figure 2 , Figure 2 FIG. 1 is a flow chart of another calcification region segmentation method based on a multimodal feature fusion model disclosed in an embodiment of the present invention. Figure 2 The described method can be applied to a calcification region segmentation device based on a multimodal feature fusion model, wherein the device may include a server, wherein the server includes a local server or a cloud server, which is not limited in the embodiment of the present invention. Figure 2 As shown, the calcification region segmentation method based on the multimodal feature fusion model includes the following operations:

[0080] 201. Based on a preset model training strategy and predetermined training data, the calcification region segmentation model is trained to obtain a calcification region segmentation model with converged training.

[0081] Optionally, the training data may include image feature training data, section feature training data, and text feature training data, which is not limited in the embodiment of the present invention.

[0082] Optionally, the model training strategy scheme may include one or more of a single forward training scheme, a multiple forward and frozen training scheme, and a multiple forward and separated training scheme, which is not limited in the embodiment of the present invention.

[0083] Optionally, the multiple forward and frozen training schemes may include a first global training part corresponding scheme and / or a graph-text matching training part corresponding scheme, which is not limited in the embodiment of the present invention.

[0084] Optionally, the multiple forward and separate training schemes may include one or more of a corresponding scheme for the encoder training part, a corresponding scheme for the decoder training part, and a corresponding scheme for the second global part, which is not limited in the embodiment of the present invention.

[0085] Optionally, the method for selecting the model training strategy scheme is illustrated by an example: according to the requirements of training speed, training flexibility, and training comprehensiveness, the model training strategy scheme for the target application is determined from the single forward training scheme, the multiple forward and frozen training scheme, and the multiple forward and separation training scheme; further, the single forward training scheme is simpler and faster, the multiple forward and frozen training scheme is more flexible and targeted, and the multiple forward and separation training scheme is more flexible and comprehensive, which is not limited to the embodiments of the present invention.

[0086] Optionally, the method for determining training data in this scheme provides a filling method for processing missing modalities. Specifically, the data of the missing modalities are filled according to certain rules. When the image modality is missing, the data of all 1s of the same size is filled as the image feature input; and when the text modality is missing, the empty text is embedded (Embedding) as the text feature input, which is not limited to the embodiments of the present invention.

[0087] Optionally, the method for determining the training set, validation set, and test set of the present scheme, specifically, in order to support the multimodal artificial valve calcification area segmentation task, the required data set includes artificial valve calcification area segmentation samples and corresponding text diagnosis samples. Therefore, the present scheme selects artificial valve calcification area segmentation samples and text diagnosis samples from the cases, and maps the text diagnosis samples to each segmentation sample, and finally creates a multimodal artificial valve calcification area segmentation data set; the data set contains a total of 3563 images, corresponding segmentation result image labels, and the same amount of text diagnosis text data; similarly, according to the ratio of 7:1:2, this paper randomly divides the multimodal artificial valve calcification area segmentation data set into a training set, a validation set, and a test set, respectively, which is not limited to the embodiments of the present invention.

[0088] 202. Perform corresponding text feature extraction operations on the text diagnosis content of the case of the determined target patient to obtain a text feature tensor result.

[0089] 203. Perform corresponding image pre-processing operations on the determined target patient's artificial valve cross-sectional image to obtain a first image tensor result and a second image tensor result, and perform corresponding section feature extraction operations on the second image tensor result to obtain a section feature result.

[0090] 204. Input the text feature tensor result, the first image tensor result and the section feature result into the trained and converged calcification region segmentation model for processing to obtain a segmentation result mask tensor.

[0091] 205. Perform corresponding post-segmentation processing operations on the segmentation result mask tensor to obtain a segmentation result map, which is used to determine the artificial valve calcification area of ​​the target patient.

[0092] In the embodiment of the present invention, for other descriptions of step 201 to step 205, please refer to other detailed descriptions of step 101 to step 104 in embodiment 1, and the embodiment of the present invention will not be repeated here.

[0093] It can be seen that the embodiment of the present invention can obtain text feature tensor results, image tensor results and section feature results through text feature extraction operations, image pre-processing operations and section feature extraction operations, and input the text feature tensor results, image tensor results and section feature results into the calcification region segmentation model for processing and combined with the segmentation post-processing operation to obtain a segmentation result map, which is conducive to improving the comprehensiveness, integrity and rationality of the calcification region segmentation method, and is conducive to improving the diversity, flexibility, pertinence and comprehensiveness of the input parameters of the artificial valve calcification region segmentation model, thereby helping to improve the accuracy and reliability of the output region segmentation results, thereby helping to improve the segmentation accuracy, reliability and segmentation efficiency and convenience of the artificial valve calcification region; and, it can also provide a calcification region segmentation model training method, enrich a calcification region segmentation method intelligent function, train a calcification region segmentation model that fits the calcification region segmentation, which is conducive to improving the functional fit and training accuracy of the calcification region segmentation model, thereby helping to improve the segmentation accuracy and reliability of the trained calcification region segmentation model, thereby helping to improve the accuracy of the subsequent segmentation result map obtained based on the calcification region segmentation model.

[0094] In an optional embodiment, the above-mentioned training of the calcified region segmentation model based on the preset model training strategy and the predetermined training data to obtain the calcified region segmentation model with converged training may include:

[0095] Inputting the predetermined training data into the calcification region segmentation model for analysis, and obtaining the segmentation mask tensor training results and text-related vector training results output by the calcification region segmentation model;

[0096] Determine the model application loss result of the calcification area segmentation model according to the segmentation mask tensor training result, the text-related vector training result and the preset image-text matching loss determination condition;

[0097] According to the model application loss result and the preset model training strategy, it is judged whether the calcification area segmentation model meets the preset model training convergence condition;

[0098] When it is determined that the calcified region segmentation model meets the model training convergence condition, the trained calcified region segmentation model is determined to be converged, and the trained converged calcified region segmentation model is used to determine and segment the calcified region;

[0099] When it is determined that the calcified region segmentation model does not meet the model training convergence condition, the calcified region segmentation model is subjected to corresponding model training operations according to a preset model training strategy scheme until the calcified region segmentation model is trained to convergence.

[0100] Further optionally, judging whether the calcification region segmentation model satisfies a preset model training convergence condition according to the model application loss result and the preset model training strategy scheme may include:

[0101] When the preset model training strategy includes a unidirectional training scheme, determining a first loss value according to a model application loss result; judging whether the first loss value is greater than or equal to a preset first loss value threshold; when the judgment result is yes, determining that the calcified region segmentation model does not meet the preset model training convergence condition; when the judgment result is no, determining that the calcified region segmentation model meets the preset model training convergence condition;

[0102] When the preset model training strategy includes multiple forward and frozen training schemes, it is determined whether the corresponding scheme of the first global training part and the corresponding scheme of the image-text matching training part both meet the preset first scheme loss convergence condition according to the model application loss result; when the judgment result is yes, it is determined that the calcification area segmentation model meets the preset model training convergence condition; when the judgment result is no, it is determined that the calcification area segmentation model does not meet the preset model training convergence condition;

[0103] When the preset model training strategy includes multiple forward and separation training schemes, based on the model application loss result, determine whether the corresponding scheme of the encoder training part, the corresponding scheme of the decoder training part, and the corresponding scheme of the second global part all meet the preset second scheme loss convergence condition; when the judgment result is yes, determine that the calcified area segmentation model meets the preset model training convergence condition; when the judgment result is no, determine that the calcified area segmentation model does not meet the preset model training convergence condition.

[0104] Further optionally, judging whether the corresponding scheme of the first global training part and the corresponding scheme of the image-text matching training part both meet the preset first scheme loss convergence condition according to the model application loss result may include:

[0105] Determine a second loss value corresponding to the corresponding scheme of the first global training part according to the model application loss result, and determine a third loss value corresponding to the corresponding scheme of the image-text matching training part according to the model application loss result;

[0106] Determine whether the second loss value is greater than or equal to a preset second loss value threshold and determine whether the third loss value is greater than or equal to a preset third loss value threshold;

[0107] When it is determined that the second loss value is greater than or equal to the preset second loss value threshold and / or the third loss value is greater than or equal to the preset third loss value threshold, it is determined that the corresponding scheme of the first global training part and the corresponding scheme of the image-text matching training part do not both meet the preset first scheme loss convergence condition;

[0108] When it is determined that the second loss value is less than the preset second loss value threshold and the third loss value is less than the preset third loss value threshold, it is determined that the corresponding scheme of the first global training part and the corresponding scheme of the image-text matching training part both meet the preset first scheme loss convergence condition.

[0109] Further optionally, judging whether the corresponding scheme of the encoder training part, the corresponding scheme of the decoder training part, and the corresponding scheme of the second global part all meet the preset second scheme loss convergence condition according to the model application loss result may include:

[0110] Determine a fourth loss value corresponding to the corresponding scheme of the encoder training part according to the determined first loss result, determine a fifth loss value corresponding to the corresponding scheme of the decoder training part according to the model application loss result, and determine a sixth loss value corresponding to the corresponding scheme of the second global part according to the model application loss result;

[0111] Determine whether the fourth loss value is greater than or equal to a preset fourth loss value threshold, determine whether the fifth loss value is greater than or equal to a preset fifth loss value threshold, and determine whether the sixth loss value is greater than or equal to a preset sixth loss value threshold;

[0112] When it is determined that the fourth loss value is greater than or equal to the preset fourth loss value threshold and / or the fifth loss value is greater than or equal to the preset fifth loss value threshold and / or the sixth loss value is greater than or equal to the preset sixth loss value threshold, it is determined that the encoder training part corresponding scheme, the decoder training part corresponding scheme, and the second global part corresponding scheme do not all meet the preset second scheme loss convergence condition;

[0113] When it is judged that the fourth loss value is less than the preset fourth loss value threshold, the fifth loss value is judged to be less than the preset fifth loss value threshold, and the sixth loss value is judged to be less than the preset sixth loss value threshold, it is determined that the corresponding scheme of the encoder training part, the corresponding scheme of the decoder training part, and the corresponding scheme of the second global part all meet the preset second scheme loss convergence condition.

[0114] It can be seen that this optional embodiment can provide a calcification area segmentation model training method, which is conducive to improving the comprehensiveness of a calcification area segmentation method and enriching the intelligent function. The segmentation mask tensor training results and the text-related vector training results are determined according to the training data, and then the model application loss results are determined, and then the model training convergence is judged according to the model application loss results. If not, the corresponding model training operation is performed on the calcification area segmentation model according to the model training strategy scheme, which is conducive to improving the accuracy and reliability of the determined model application loss results, and then it is conducive to improving the accuracy and reliability of the determined model training convergence judgment results, so as to improve the training convergence accuracy, reliability and timeliness of the calcification area segmentation model. In addition, the corresponding model training operation is performed in combination with the preset model training strategy scheme, which is conducive to improving the model training accuracy, reliability and rationality of the calcification area segmentation model.

[0115] In another optional embodiment, the above-mentioned inputting the predetermined training data into the calcification region segmentation model for analysis to obtain the segmentation mask tensor training results and the text-related vector training results output by the calcification region segmentation model may include:

[0116] Perform corresponding splicing operations on the image feature training data and the section feature training data to obtain a first image feature vector training result;

[0117] Based on preset special marking information, corresponding splicing and filling operations are performed on the text feature training data and the first image feature vector training result to obtain the image-text modality fusion feature tensor training result, where the special marking information includes position marking information and / or modality type marking information;

[0118] Inputting the image-text modality fusion feature tensor training result into the encoder module in the calcification region segmentation model for processing to obtain the second feature tensor training result;

[0119] Performing corresponding feature extraction operations on the second feature tensor training result to obtain a third image feature tensor result, and performing corresponding reshaping operations on the third image feature tensor result to obtain a feature tensor reshaping result; inputting the feature tensor reshaping result into a decoder module in the calcification region segmentation model to perform corresponding upsampling and jump connection operations to obtain a segmentation mask tensor training result;

[0120] According to the training result of the second feature tensor, the training result of the target component tensor is determined, and based on the fully connected layer and the preset activation function in the calcified area segmentation model, the corresponding identification processing and linear projection operations are performed on the training result of the target component tensor to obtain the text-related vector training result.

[0121] Optional, as per instructions Fig. 9 As shown, this is a schematic diagram of the architecture of the calcification region segmentation model of this solution, which is not limited in the embodiment of the present invention.

[0122] Optionally, the above-mentioned corresponding splicing operations are performed on the image feature training data and the section feature training data to obtain the first image feature vector training result. For example: ViLTransUNet adopts an improved TransUNet convolutional neural network module and an artificial valve feature extraction module to splice the image features of the CT image (i.e., the image feature training data) with the artificial valve section features (i.e., the section feature training data) to obtain an image feature vector (i.e., the first image feature vector training result) with a size of 196×768. The embodiment of the present invention is not limited to this.

[0123] Optionally, based on the preset special tag information, the above-mentioned corresponding splicing and filling operations are performed on the text feature training data and the first image feature vector training results to obtain the image-text modality fusion feature tensor training results. For example: the extracted text feature vector (i.e., text feature training data) and the image feature vector (i.e., the first image feature vector training result) are spliced ​​and filled; in the splicing process, a clstoken (i.e., special tag) is added to the text feature vector and the image feature vector respectively. The purpose of this step is to use the cls token to perform feature aggregation on tokens at other positions in the feature vector, while avoiding bias towards a certain token, and finally to support downstream tasks; after this step, we get a size of In order to retain the position information of text features and image features, each token is added with a position code (i.e., a special tag). In addition, in order to help the Transformer encoder distinguish between the two modalities of image and text, each token in the feature tensor is also added with a specific modal type code (Modal-type embedding), in which all tokens representing text features are added with 0, and all tokens representing image features are added with 1, which is not limited in the embodiment of the present invention.

[0124] Optionally, the above-mentioned image-text modality fusion feature tensor training result is input into the encoder module in the calcification area segmentation model for processing to obtain the second feature tensor training result. For example, the obtained image-text modality fusion feature tensor training result is input into a 12-layer Transformer encoder. After learning through the Transformer's multi-head self-attention mechanism, the output size is also The feature tensor (i.e., the second feature tensor training result) is not limited in this embodiment of the present invention.

[0125] Optionally, the above-mentioned corresponding feature extraction operation is performed on the second feature tensor training result to obtain a third image feature tensor result, and a corresponding reshaping operation is performed on the third image feature tensor result to obtain a feature tensor reshaping result; the feature tensor reshaping result is input into the decoder module in the calcification area segmentation model for corresponding upsampling and jump connection operations to obtain a segmentation mask tensor training result. For example: a feature tensor of size (196,768) representing image features (i.e., the third image feature tensor result) is taken out, and after reshaping (i.e., the feature tensor reshaping result), it enters the decoder part, and continues to perform subsequent upsampling and jump connection, and finally obtains a segmentation mask tensor of size 2×224×224 (i.e., the segmentation mask tensor training result), which is not limited in the embodiments of the present invention.

[0126] Optionally, the above-mentioned target component tensor training result is determined according to the training result of the second feature tensor, and based on the fully connected layer and the preset activation function in the calcified area segmentation model, the target component tensor training result is subjected to corresponding identification processing and linear projection operation to obtain the text-related vector training result. For example, after the clstoken of the text part is processed by the Transformer encoder, a 1×768 tensor (i.e., the target component tensor training result) is obtained, and then after a fully connected layer and a tanh activation function, it is finally linearly projected to a 1×2 vector (i.e., the text-related vector training result), which is not limited in the embodiment of the present invention.

[0127] Optionally, the preset activation function may include but is not limited to a tanh activation function, which is not limited in this embodiment of the present invention.

[0128] It can be seen that this optional embodiment can provide a method for determining the segmentation mask tensor training results and a method for determining the text-related vector training results, and provide a specific model architecture and information analysis method, which is conducive to improving the comprehensiveness and rationality of the method for determining the segmentation mask tensor training results, and thus is conducive to improving the accuracy and reliability of the determined segmentation mask tensor training results. In addition, it is conducive to improving the comprehensiveness and rationality of the method for determining the text-related vector training results, and thus is conducive to improving the accuracy and reliability of the determined text-related vector training results.

[0129] In yet another optional embodiment, before performing corresponding splicing operations on the image feature training data and the section feature training data to obtain the first image feature vector training result, the method may further include the following operations:

[0130] When the model training strategy includes a single forward training scheme, performing a corresponding image-text pair confusion operation on the predetermined training data;

[0131] When the model training strategy includes multiple forward and frozen training schemes and the multiple forward and frozen training schemes include corresponding schemes for image-text matching training, performing corresponding image-text pair confusion operations on the predetermined training data;

[0132] When the model training strategy includes multiple forward and separate training schemes and the multiple forward and separate training schemes include a corresponding scheme for encoder training part, a corresponding image-text pair confusion operation is performed on the predetermined training data.

[0133] Optionally, the above-mentioned corresponding image-text pair obfuscation operation is performed on the predetermined training data. For example, the image-text pair is sampled and obfuscated after input, which is not limited in the embodiment of the present invention.

[0134] It can be seen that this optional embodiment can provide an image-text pair obfuscation processing method for training data, and trigger the execution of image-text pair obfuscation operations based on a model training strategy scheme, which is beneficial to improving the execution accuracy, reliability, and pertinence of image-text pair obfuscation operations, and further helps to improve the accuracy and reliability of the model training effect based on obfuscated image-text pairs.

[0135] In yet another optional embodiment, the above-mentioned performing corresponding model training operations on the calcification region segmentation model according to the preset model training strategy scheme may include:

[0136] When the model training strategy includes a single forward training scheme, a corresponding reverse gradient propagation operation is performed on the calcification region segmentation model based on the model application loss result;

[0137] When the model training strategy includes multiple forward and frozen training schemes and the multiple forward and frozen training schemes include corresponding schemes of image-text matching training, based on the model application loss result, performing corresponding reverse gradient propagation operations on the calcification region segmentation model;

[0138] When the model training strategy includes multiple forward and frozen training schemes and the multiple forward and frozen training schemes include a first global training part corresponding scheme, based on the model application loss result, performing corresponding parameter freezing operations except for the classification head on the calcification region segmentation model and performing corresponding reverse gradient propagation operations;

[0139] When the model training strategy includes multiple forward and split training schemes and the multiple forward and split training schemes include encoder training part corresponding schemes, based on the determined first loss result, performing corresponding reverse gradient propagation training encoder operations on the calcification region segmentation model;

[0140] When the model training strategy includes multiple forward and split training schemes and the multiple forward and split training schemes include a decoder training partial corresponding scheme, based on the model application loss result, performing a corresponding encoder parameter freezing operation on the calcification region segmentation model and performing a corresponding reverse gradient propagation training decoder operation;

[0141] When the model training strategy includes multiple forward and split training schemes and the multiple forward and split training schemes include a second global part corresponding scheme, a corresponding reverse gradient propagation operation is performed on the calcification region segmentation model based on the model application loss result.

[0142] Optionally, the specific process for the single forward training scheme can be referred to the instruction manual Fig.10 As shown; further, for a single forward training scheme, only one forward process and one reverse update of parameters may be performed for each training. Specifically, the image-text pair will be sampled and confused after input, and then a complete forward process is performed. In the forward process, the ITM Loss (i.e., the first loss result) of the image-text matching part and the Focal Dice Loss (i.e., the second loss result and the third loss result) of the segmentation part are calculated respectively. Finally, the total Loss is calculated and a complete reverse gradient propagation is performed, which is not limited in the embodiments of the present invention.

[0143] Optionally, the advantage of a single forward training scheme is that it is simple and fast, which is not limited in the embodiment of the present invention.

[0144] Optionally, for the specific process of multiple forward and frozen training schemes, please refer to the instructions Fig.11As shown; further, the corresponding solution for the first global training part can be to divide the entire training process into two independent stages, so as to realize purposeful model parameter training. These two stages include the training of global parameters and the stage focusing on training the image-text matching classification head, that is, linearly projecting the text cls token. Specifically, a normal forward propagation process is performed without confusing the image-text pair, and the ITM Loss (i.e., the first loss result) of the image-text matching part and the Focal Dice Loss (i.e., the second loss result and the third loss result) of the segmentation part are calculated respectively, and the total loss is used for global gradient backpropagation; for the corresponding solution for the image-text matching training part, specifically, the image-text pair is confused, and the ITM Loss (i.e., the first loss result) of the image-text matching part and the Focal Dice Loss (i.e., the second loss result and the third loss result) of the segmentation part are calculated respectively. Loss (i.e., the second loss result and the third loss result), after calculating the total loss, freeze other parameters and focus on training the linear projection layer of the classification head; further, in the actual training process, it is necessary to comprehensively consider the training progress of the two stages and adjust their training frequencies accordingly. For example, if it is found that the image-text matching part is not effective, the number of global trainings can be increased in one iteration and the number of image-text matching trainings can be reduced. This is not limited in the embodiments of the present invention.

[0145] Optionally, the main advantage of the multiple forward and frozen training schemes is their flexibility, which can train the classification head in a targeted manner and improve the performance of text modality data in the task, which is not limited in the embodiments of the present invention.

[0146] Optionally, for the specific process of multiple forward and split training schemes, please refer to the instructions Fig.12 As shown; further, the corresponding scheme for the encoder training part can be composed of three parts, namely the global training part, the encoder training part and the decoder training part. Specifically, after the input image-text pair is confused, the encoder part of the entire ViLTransUNet is used to calculate the ITM Loss (i.e., the first loss result), and then directly perform reverse gradient propagation; for the corresponding scheme for the decoder training part, specifically, the image-text pair is not confused, the parameters of the encoder part are frozen after the entire forward process, and only the parameters of the decoder part are trained by reverse gradient propagation; for the corresponding scheme for the second global part, specifically, the input image-text pair is not confused, and the ITMLoss (i.e., the first loss result) of the image-text matching part and the Focal Dice Loss (i.e., the second loss result and the third loss result) of the segmentation part are calculated respectively. Finally, the total Loss is calculated and a complete reverse gradient propagation is performed; further, in terms of the allocation of training frequency, the number of global trainings in one iteration should be greater than the number of trainings of the encoder and the decoder to ensure the segmentation accuracy of the model, which is not limited in the embodiment of the present invention.

[0147] Optionally, for multiple forward and separate training schemes, it has the advantages of flexibility and comprehensiveness. Through isolated training of the encoder part and the decoder part, the image-text matching ability of the encoder and the feature fusion and segmentation ability of the decoder part can be improved respectively. At the same time, through global training, it is ensured that the encoder part can also obtain the reverse gradient propagation of the segmentation loss, thereby ensuring the segmentation ability of the overall model, which is not limited in the embodiments of the present invention.

[0148] It can be seen that this optional embodiment can provide corresponding specific model training operations for different model training strategy schemes, which is conducive to improving the diversity, flexibility, pertinence and selectivity of model training strategy schemes, and thus is conducive to improving the execution accuracy, reliability and rationality of specific model training operations, thereby facilitating the optimization of model training effects, and is conducive to improving the timeliness, accuracy and efficiency of model training convergence.

[0149] In yet another optional embodiment, the above-mentioned determining the model application loss result of the calcification region segmentation model according to the segmentation mask tensor training result, the text-related vector training result and the preset image-text matching loss determination condition may include:

[0150] Determine a first loss result according to the text-related vector training result and the set cross entropy loss calculation method, determine a second loss result according to the segmentation mask tensor training result and the set dice loss calculation method, and determine a third loss result according to the segmentation mask tensor training result and the set focus loss calculation method;

[0151] A model application loss result of the calcification region segmentation model is determined according to the determined hyperparameter information, the first loss result, the second loss result, and the third loss result.

[0152] Optionally, the first loss result may be understood as a loss result for the image-text matching part, and the second loss result and the third loss result may be understood as loss results for the segmentation part, which is not limited in the embodiment of the present invention.

[0153] Optionally, for the determination method of the first loss result, an example is given: by intentionally replacing some samples with incorrect image-text pairs, and generating a label of the same length, the correct image-text pair is marked as 1, and the incorrect label is 0. After learning through the Transformer encoder, the class token is extracted from the text part, and a 1×2 vector is obtained through the linear projection layer. This vector is subjected to a binary classification task through the Softmax function, that is, to determine whether the image and text match. The cross entropy loss function between the binary classification result and the label is the ITM Loss. The specific formula is as follows: ;

[0154] in, N Represents the number of image-text pairs; y itm Represents the generated image-text pair Ground Truth label; Represents the binary classification prediction output result.

[0155] It can be seen that this optional embodiment can determine the model application loss result of the calcification area segmentation model through the cross entropy loss calculation method, the dice loss calculation method and the focal loss calculation method, which is conducive to improving the comprehensiveness and rationality of the method for determining the model application loss result, and then it is conducive to improving the diversity, flexibility, comprehensiveness and pertinence of the sub-loss results used to determine the model application loss result, so as to improve the accuracy and reliability of the determined model application loss result, and further help to improve the subsequent model training convergence accuracy, reliability and timeliness determined based on the model application loss result.

[0156] In yet another optional embodiment, the model application loss result can be calculated by the following formula:

[0157] ;

[0158] in, is the hyperparameter information, is the first loss result; The second loss result; The result is the third loss.

[0159] Optional, The value may be 0.2, or may be set to other values ​​according to actual needs, which is not limited in the embodiment of the present invention.

[0160] It can be seen that this optional embodiment can provide a calculation formula for the model application loss result, which is conducive to improving the scientificity, pertinence and creativity of the method for determining the model application loss result, and further helps to improve the rationality and effectiveness of the determined model application loss result.

[0161] Embodiment 3

[0162] See also Figure 3 , Figure 3 is a schematic diagram of the structure of a calcification region segmentation device based on a multimodal feature fusion model disclosed in an embodiment of the present invention. Figure 3 The described device may include a server, wherein the server includes a local server or a cloud server, which is not limited in the embodiment of the present invention. Figure 3 As shown, the calcification region segmentation device based on the multimodal feature fusion model may include:

[0163] The text feature extraction module 301 is used to perform corresponding text feature extraction operations on the text diagnosis content of the case of the determined target patient to obtain a text feature tensor result.

[0164] The image feature determination module 302 is used to perform corresponding image pre-processing operations on the determined target patient's artificial valve cross-sectional image to obtain a first image tensor result and a second image tensor result, and to perform corresponding section feature extraction operations on the second image tensor result to obtain a section feature result.

[0165] The mask tensor determination module 303 is used to input the text feature tensor result, the first image tensor result and the section feature result into the trained converged calcification region segmentation model for processing to obtain a segmentation result mask tensor.

[0166] The segmentation result determination module 304 is used to perform corresponding segmentation post-processing operations on the segmentation result mask tensor to obtain a segmentation result map, which is used to determine the artificial valve calcification area of ​​the target patient.

[0167] It can be seen that implementation Figure 3 The described calcification region segmentation device based on the multimodal feature fusion model can obtain text feature tensor results, image tensor results and section feature results through text feature extraction operations, image pre-processing operations and section feature extraction operations, and input the text feature tensor results, image tensor results and section feature results into the calcification region segmentation model for processing and combined with the segmentation post-processing operation to obtain a segmentation result map, which is conducive to improving the comprehensiveness, integrity and rationality of the calcification region segmentation method, and is conducive to improving the diversity, flexibility, pertinence and comprehensiveness of the input parameters of the artificial valve calcification region segmentation model, and further helps to improve the accuracy and reliability of the output region segmentation results, thereby helping to improve the segmentation accuracy, reliability, efficiency and convenience of the artificial valve calcification region.

[0168] In an optional embodiment, if Figure 4 As shown, the device may also include:

[0169] The segmentation model training module 305 is used to train the calcification region segmentation model based on a preset model training strategy and predetermined training data to obtain a calcification region segmentation model with converged training, wherein the training data includes image feature training data, section feature training data, and text feature training data;

[0170] Among them, the model training strategy scheme includes one or more of a single forward training scheme, multiple forward and frozen training schemes, and multiple forward and separate training schemes; the multiple forward and frozen training scheme includes a first global training part corresponding scheme and / or a graphic matching training part corresponding scheme; the multiple forward and separate training scheme includes one or more of an encoder training part corresponding scheme, a decoder training part corresponding scheme, and a second global part corresponding scheme.

[0171] It can be seen that implementation Figure 4 The described device can provide a calcification region segmentation model training method, enrich the intelligent function of a calcification region segmentation method, and train a calcification region segmentation model that fits the calcification region segmentation, which is beneficial to improving the functional fit and training accuracy of the calcification region segmentation model, and further beneficial to improving the segmentation accuracy and reliability of the trained calcification region segmentation model, thereby facilitating improving the accuracy of the subsequent segmentation result map obtained based on the calcification region segmentation model.

[0172] In another optional embodiment, the segmentation model training module 305 trains the calcification region segmentation model based on a preset model training strategy and predetermined training data, and the method of obtaining the calcification region segmentation model with converged training specifically includes:

[0173] Inputting the predetermined training data into the calcification region segmentation model for analysis, and obtaining the segmentation mask tensor training results and text-related vector training results output by the calcification region segmentation model;

[0174] Determine the model application loss result of the calcification area segmentation model according to the segmentation mask tensor training result, the text-related vector training result and the preset image-text matching loss determination condition;

[0175] According to the model application loss result and the preset model training strategy, it is judged whether the calcification area segmentation model meets the preset model training convergence condition;

[0176] When it is determined that the calcified region segmentation model meets the model training convergence condition, the trained calcified region segmentation model is determined to be converged, and the trained converged calcified region segmentation model is used to determine and segment the calcified region;

[0177] When it is determined that the calcified region segmentation model does not meet the model training convergence condition, the calcified region segmentation model is subjected to corresponding model training operations according to a preset model training strategy scheme until the calcified region segmentation model is trained to convergence.

[0178] It can be seen that implementation Figure 4The described device can also provide a calcification area segmentation model training method, which is conducive to improving the comprehensiveness of a calcification area segmentation method and enriching intelligent functions. The segmentation mask tensor training results and text-related vector training results are determined according to the training data, and then the model application loss results are determined, and then the model training convergence is judged according to the model application loss results. If not, the corresponding model training operation is performed on the calcification area segmentation model according to the model training strategy scheme, which is conducive to improving the accuracy and reliability of the determined model application loss results, and then it is conducive to improving the accuracy and reliability of the determined model training convergence judgment results, so as to improve the training convergence accuracy, reliability and timeliness of the calcification area segmentation model. In addition, the corresponding model training operation is performed in combination with the preset model training strategy scheme, which is conducive to improving the model training accuracy, reliability and rationality of the calcification area segmentation model.

[0179] In another optional embodiment, the segmentation model training module inputs the predetermined training data into the calcification region segmentation model for analysis, and obtains the segmentation mask tensor training results and the text-related vector training results output by the calcification region segmentation model in a manner that specifically includes:

[0180] Perform corresponding splicing operations on the image feature training data and the section feature training data to obtain a first image feature vector training result;

[0181] Based on preset special marking information, corresponding splicing and filling operations are performed on the text feature training data and the first image feature vector training result to obtain the image-text modality fusion feature tensor training result, where the special marking information includes position marking information and / or modality type marking information;

[0182] Inputting the image-text modality fusion feature tensor training result into the encoder module in the calcification region segmentation model for processing to obtain the second feature tensor training result;

[0183] Performing corresponding feature extraction operations on the second feature tensor training result to obtain a third image feature tensor result, and performing corresponding reshaping operations on the third image feature tensor result to obtain a feature tensor reshaping result; inputting the feature tensor reshaping result into a decoder module in the calcification region segmentation model to perform corresponding upsampling and jump connection operations to obtain a segmentation mask tensor training result;

[0184] According to the training result of the second feature tensor, the training result of the target component tensor is determined, and based on the fully connected layer and the preset activation function in the calcified area segmentation model, the corresponding identification processing and linear projection operations are performed on the training result of the target component tensor to obtain the text-related vector training result.

[0185] It can be seen that implementation Figure 4The described device can also provide a method for determining the segmentation mask tensor training results and a method for determining the text-related vector training results, and provide a specific model architecture and information analysis method, which is conducive to improving the comprehensiveness and rationality of the method for determining the segmentation mask tensor training results, and thus is conducive to improving the accuracy and reliability of the determined segmentation mask tensor training results. In addition, it is conducive to improving the comprehensiveness and rationality of the method for determining the text-related vector training results, and thus is conducive to improving the accuracy and reliability of the determined text-related vector training results.

[0186] In another optional embodiment, the segmentation model training module 305 is also used to perform corresponding splicing operations on the image feature training data and the section feature training data before obtaining the first image feature vector training result. When the model training strategy includes a single forward training scheme, the corresponding image-text pair confusion operation is performed on the predetermined training data; when the model training strategy includes multiple forward and frozen training schemes and the multiple forward and frozen training schemes include corresponding schemes for image-text matching training parts, the corresponding image-text pair confusion operation is performed on the predetermined training data; when the model training strategy includes multiple forward and separation training schemes and the multiple forward and separation training schemes include corresponding schemes for encoder training parts, the corresponding image-text pair confusion operation is performed on the predetermined training data.

[0187] It can be seen that implementation Figure 4 The described device can also provide an image-text pair obfuscation processing method for training data, and trigger the execution of image-text pair obfuscation operations based on a model training strategy scheme, which is beneficial to improving the execution accuracy, reliability, and pertinence of image-text pair obfuscation operations, and thus is beneficial to improving the accuracy and reliability of the model training effect based on obfuscated image-text pairs.

[0188] In another optional embodiment, the segmentation model training module 305 performs corresponding model training operations on the calcification region segmentation model according to a preset model training strategy scheme, specifically including:

[0189] When the model training strategy includes a single forward training scheme, a corresponding reverse gradient propagation operation is performed on the calcification region segmentation model based on the model application loss result;

[0190] When the model training strategy includes multiple forward and frozen training schemes and the multiple forward and frozen training schemes include corresponding schemes of image-text matching training, based on the model application loss result, performing corresponding reverse gradient propagation operations on the calcification region segmentation model;

[0191] When the model training strategy includes multiple forward and frozen training schemes and the multiple forward and frozen training schemes include a first global training part corresponding scheme, based on the model application loss result, performing corresponding parameter freezing operations except for the classification head on the calcification region segmentation model and performing corresponding reverse gradient propagation operations;

[0192] When the model training strategy includes multiple forward and split training schemes and the multiple forward and split training schemes include encoder training partial corresponding schemes, based on the determined first loss result, performing corresponding reverse gradient propagation training encoder operations on the calcification region segmentation model;

[0193] When the model training strategy includes multiple forward and split training schemes and the multiple forward and split training schemes include a decoder training partial corresponding scheme, based on the model application loss result, performing a corresponding encoder parameter freezing operation on the calcification region segmentation model and performing a corresponding reverse gradient propagation training decoder operation;

[0194] When the model training strategy includes multiple forward and split training schemes and the multiple forward and split training schemes include a second global part corresponding scheme, a corresponding reverse gradient propagation operation is performed on the calcification region segmentation model based on the model application loss result.

[0195] It can be seen that implementation Figure 4 The described device can also provide corresponding specific model training operations for different model training strategy schemes, which is conducive to improving the diversity, flexibility, pertinence and selectivity of model training strategy schemes, and thus is conducive to improving the execution accuracy, reliability and rationality of specific model training operations, thereby facilitating the optimization of model training effects, and is conducive to improving the timeliness, accuracy and efficiency of model training convergence.

[0196] In another optional embodiment, the segmentation model training module 305 determines the model application loss result of the calcification area segmentation model according to the segmentation mask tensor training result, the text-related vector training result and the preset image-text matching loss determination condition, specifically including:

[0197] Determine a first loss result according to the text-related vector training result and the set cross entropy loss calculation method, determine a second loss result according to the segmentation mask tensor training result and the set dice loss calculation method, and determine a third loss result according to the segmentation mask tensor training result and the set focus loss calculation method;

[0198] A model application loss result of the calcification region segmentation model is determined according to the determined hyperparameter information, the first loss result, the second loss result, and the third loss result.

[0199] It can be seen that implementation Figure 4The described device can also determine the model application loss result of the calcification area segmentation model through the cross entropy loss calculation method, the dice loss calculation method and the focal loss calculation method, which is conducive to improving the comprehensiveness and rationality of the method for determining the model application loss result, and then it is conducive to improving the diversity, flexibility, comprehensiveness and pertinence of the sub-loss results used to determine the model application loss result, so as to improve the accuracy and reliability of the determined model application loss result, and further help to improve the subsequent model training convergence accuracy, reliability and timeliness determined based on the model application loss result.

[0200] In yet another optional embodiment, the model application loss result is calculated by the following formula:

[0201] ;

[0202] in, is the hyperparameter information, is the first loss result; The second loss result; The result is the third loss.

[0203] It can be seen that implementation Figure 4 The described device can also provide a formula for calculating the model application loss result, which is conducive to improving the scientificity, pertinence and creativity of the method for determining the model application loss result, and further conducive to improving the rationality and effectiveness of the determined model application loss result.

[0204] In another optional embodiment, the text feature extraction module 301 performs corresponding text feature extraction operations on the determined text diagnosis content of the case of the target patient, and the method of obtaining the text feature tensor result specifically includes:

[0205] Perform corresponding valid text interception operations on the text diagnosis content of the case of the determined target patient to obtain valid text content results;

[0206] Input the valid text content result into the text conversion model with training convergence for processing to obtain the text conversion result;

[0207] The text conversion result is input into the trained converged feature vector extraction model for processing to obtain the text feature tensor result.

[0208] It can be seen that implementation Figure 4 The described device can also provide a text feature extraction method, which performs corresponding effective text interception operations, text conversion operations and feature vector extraction operations on the text diagnosis content of the case to obtain a text feature tensor result, which is beneficial to improving the comprehensiveness, progressiveness and rationality of the text feature extraction method, and thus helps to improve the accuracy and reliability of the determined text feature tensor results.

[0209] In another optional embodiment, the text feature extraction module 301 performs a corresponding valid text interception operation on the determined text diagnosis content of the case of the target patient, and the method of obtaining the valid text content result specifically includes:

[0210] Based on the preset word segmentation component library, a corresponding word segmentation operation is performed on the text diagnosis content of the case of the determined target patient to obtain a text word segmentation result, which includes one or more word segmentation text segments;

[0211] Determine the keyword screening results based on the text segmentation results and the preset keyword screening method;

[0212] According to the keyword screening results and the text segmentation results, the corresponding similarity analysis operation is performed to obtain the similarity score results corresponding to each segmented text segment;

[0213] The valid text content result is determined based on each segmented text segment and its corresponding similarity score result.

[0214] It can be seen that implementation Figure 4 The described device can also provide a method for determining effective text content results, perform corresponding word segmentation operations, keyword screening operations, similarity analysis operations and screening operations on the textual diagnosis content of the case, and obtain effective text content results, which is conducive to improving the comprehensiveness, progressiveness and rationality of the method for determining effective text content results, and thus is conducive to improving the accuracy and reliability of the determined effective text content results.

[0215] Embodiment 4

[0216] See also Figure 5 , Figure 5 is a structural schematic diagram of another calcification region segmentation device based on a multimodal feature fusion model disclosed in an embodiment of the present invention. Figure 5 The described device may include a server, wherein the server includes a local server or a cloud server, which is not limited in the embodiment of the present invention. Figure 5 As shown, the device may include:

[0217] A memory 401 storing executable program codes;

[0218] a processor 402 coupled to the memory 401;

[0219] Furthermore, it may also include an input interface 403 and an output interface 404 coupled to the processor 402;

[0220] The processor 402 calls the executable program code stored in the memory 401 to execute the steps in the calcification region segmentation method based on the multimodal feature fusion model described in the first or second embodiment.

[0221] Embodiment 5

[0222] An embodiment of the present invention discloses a computer storage medium storing a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the calcification region segmentation method based on a multimodal feature fusion model described in Embodiment 1 or Embodiment 2.

[0223] Embodiment 6

[0224] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps in the calcification area segmentation method based on a multimodal feature fusion model described in Embodiment 1 or Embodiment 2.

[0225] The device embodiments described above are only illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, i.e., they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.

[0226] Through the specific description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution can be essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0227] Finally, it should be noted that the calcification area segmentation method and device based on the multimodal feature fusion model disclosed in the embodiment of the present invention discloses only the preferred embodiments of the present invention, which are only used to illustrate the technical solution of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A calcification region segmentation method based on a multimodal feature fusion model, characterized in that: The method comprises: Perform corresponding text feature extraction operations on the text diagnosis content of the case of the determined target patient to obtain a text feature tensor result; Performing corresponding image pre-processing operations on the determined artificial valve cross-sectional image of the target patient to obtain a first image tensor result and a second image tensor result, and performing corresponding section feature extraction operations on the second image tensor result to obtain a section feature result; Inputting the text feature tensor result, the first image tensor result and the section feature result into a trained and converged calcification region segmentation model for processing to obtain a segmentation result mask tensor; Performing corresponding post-segmentation processing operations on the segmentation result mask tensor to obtain a segmentation result map, wherein the segmentation result map is used to determine the artificial valve calcification area of ​​the target patient; And, the method further comprises: Based on a preset model training strategy and predetermined training data, the calcification region segmentation model is trained to obtain the calcification region segmentation model with converged training, wherein the training data includes image feature training data, section feature training data, and text feature training data; The model training strategy includes one or more of a single forward training scheme, a multiple forward and frozen training scheme, and a multiple forward and separated training scheme; the multiple forward and frozen training scheme includes a first global training part corresponding scheme and / or a picture-text matching training part corresponding scheme; the multiple forward and separated training scheme includes one or more of an encoder training part corresponding scheme, a decoder training part corresponding scheme, and a second global part corresponding scheme; Furthermore, the calcification region segmentation model is trained based on a preset model training strategy and predetermined training data to obtain the calcification region segmentation model with converged training, including: Inputting the predetermined training data into the calcification region segmentation model for analysis, and obtaining the segmentation mask tensor training results and the text-related vector training results output by the calcification region segmentation model; Determining a model application loss result of the calcification region segmentation model according to the segmentation mask tensor training result, the text-related vector training result and a preset image-text matching loss determination condition; According to the model application loss result and the preset model training strategy scheme, judging whether the calcification region segmentation model meets the preset model training convergence condition; When it is determined that the calcified region segmentation model meets the model training convergence condition, determining that the trained calcified region segmentation model is converged, and the trained converged calcified region segmentation model is used to determine and segment the calcified region; When it is determined that the calcified region segmentation model does not satisfy the model training convergence condition, performing corresponding model training operations on the calcified region segmentation model according to a preset model training strategy scheme until the calcified region segmentation model is trained to convergence; And, performing corresponding model training operations on the calcification region segmentation model according to a preset model training strategy, including: When the preset model training strategy includes the single forward training scheme, based on the model application loss result, performing a corresponding reverse gradient propagation operation on the calcification region segmentation model; When the preset model training strategy includes the multiple forward and frozen training schemes and the multiple forward and frozen training schemes include the corresponding schemes of the image-text matching training part, based on the model application loss result, performing a corresponding reverse gradient propagation operation on the calcification region segmentation model; When the preset model training strategy includes the multiple forward and freeze training schemes and the multiple forward and freeze training schemes include the first global training part corresponding scheme, based on the model application loss result, performing corresponding parameter freezing operations except for the classification head on the calcification area segmentation model and performing corresponding reverse gradient propagation operations; When the preset model training strategy includes multiple forward and split training schemes and the multiple forward and split training schemes include the encoder training part corresponding scheme, based on the determined first loss result, performing a corresponding reverse gradient propagation training encoder operation on the calcified region segmentation model; When the preset model training strategy includes multiple forward and separate training schemes and the multiple forward and separate training schemes include the decoder training part corresponding scheme, based on the model application loss result, performing a corresponding encoder parameter freezing operation on the calcified area segmentation model and performing a corresponding reverse gradient propagation training decoder operation; When the preset model training strategy includes multiple forward and separate training schemes and the multiple forward and separate training schemes include the second global part corresponding scheme, based on the model application loss result, a corresponding reverse gradient propagation operation is performed on the calcification area segmentation model.

2. The calcification region segmentation method based on the multimodal feature fusion model according to claim 1, characterized in that: The step of inputting the predetermined training data into the calcified region segmentation model for analysis to obtain the segmentation mask tensor training results and text-related vector training results output by the calcified region segmentation model includes: Performing corresponding splicing operations on the image feature training data and the section feature training data to obtain a first image feature vector training result; Based on preset special marking information, corresponding splicing and filling operations are performed on the text feature training data and the first image feature vector training result to obtain a text-image modality fusion feature tensor training result, wherein the special marking information includes position marking information and / or modality type marking information; Inputting the image-text modality fusion feature tensor training result into the encoder module in the calcification region segmentation model for processing to obtain a second feature tensor training result; Performing a corresponding feature extraction operation on the second feature tensor training result to obtain a third image feature tensor result, and performing a corresponding reshaping operation on the third image feature tensor result to obtain a feature tensor reshaping result; inputting the feature tensor reshaping result into a decoder module in the calcification region segmentation model to perform corresponding upsampling and skip connection operations to obtain a segmentation mask tensor training result; According to the second feature tensor training result, determine the target component tensor training result, and based on the fully connected layer and the preset activation function in the calcified area segmentation model, perform corresponding identification processing and linear projection operations on the target component tensor training result to obtain the text-related vector training result; Furthermore, before performing corresponding splicing operations on the image feature training data and the section feature training data to obtain the first image feature vector training result, the method further includes: When the preset model training strategy includes the single forward training scheme, performing a corresponding image-text pair obfuscation operation on the predetermined training data; When the preset model training strategy includes the multiple forward and frozen training schemes and the multiple forward and frozen training schemes include the corresponding schemes of the image-text matching training part, performing the corresponding image-text pair obfuscation operation on the predetermined training data; When the preset model training strategy includes multiple forward and separate training schemes and the multiple forward and separate training schemes include the encoder training part corresponding scheme, the corresponding image-text pair confusion operation is performed on the predetermined training data.

3. The calcification region segmentation method based on the multimodal feature fusion model according to claim 1, characterized in that: The step of determining the model application loss result of the calcified area segmentation model according to the segmentation mask tensor training result, the text-related vector training result and a preset image-text matching loss determination condition includes: Determine a first loss result according to the text-related vector training result and the set cross entropy loss calculation method, determine a second loss result according to the segmentation mask tensor training result and the set dice loss calculation method, and determine a third loss result according to the segmentation mask tensor training result and the set focus loss calculation method; Determine a model application loss result of the calcification region segmentation model according to the determined hyperparameter information, the first loss result, the second loss result, and the third loss result; And, the model application loss result is calculated by the following formula: ; in, is the hyperparameter information, is the first loss result; The second loss result; This is the third loss result.

4. The calcification region segmentation method based on a multimodal feature fusion model according to any one of claims 1 to 3, characterized in that: The corresponding text feature extraction operation is performed on the determined text diagnosis content of the target patient's case to obtain a text feature tensor result, including: Perform corresponding valid text interception operations on the text diagnosis content of the case of the determined target patient to obtain valid text content results; Inputting the valid text content result into the trained and converged text conversion model for processing to obtain a text conversion result; The text conversion result is input into a trained and converged feature vector extraction model for processing to obtain a text feature tensor result.

5. The calcification region segmentation method based on the multimodal feature fusion model according to claim 4, characterized in that: The corresponding valid text interception operation is performed on the determined target patient's case text diagnosis content to obtain a valid text content result, including: Based on a preset word segmentation component library, a corresponding word segmentation operation is performed on the text diagnosis content of the determined target patient's case to obtain a text word segmentation result, wherein the text word segmentation result includes one or more word segmentation text segments; Determine the keyword screening result according to the text segmentation result and the preset keyword screening method; According to the keyword screening results and the text segmentation results, a corresponding similarity analysis operation is performed to obtain a similarity score result corresponding to each segmented text segment; A valid text content result is determined based on each of the segmented text segments and their corresponding similarity score results.

6. A calcification region segmentation device based on a multimodal feature fusion model, characterized in that: The device comprises: A text feature extraction module is used to perform corresponding text feature extraction operations on the text diagnosis content of the case of the determined target patient to obtain a text feature tensor result; An image feature determination module is used to perform corresponding image pre-processing operations on the determined artificial valve cross-sectional image of the target patient to obtain a first image tensor result and a second image tensor result, and perform corresponding section feature extraction operations on the second image tensor result to obtain a section feature result; A mask tensor determination module, used for inputting the text feature tensor result, the first image tensor result and the section feature result into a trained and converged calcification region segmentation model for processing to obtain a segmentation result mask tensor; A segmentation result determination module, used for performing corresponding segmentation post-processing operations on the segmentation result mask tensor to obtain a segmentation result map, wherein the segmentation result map is used for determining the artificial valve calcification area of ​​the target patient; And, the device also includes: A segmentation model training module is used to train the calcification region segmentation model based on a preset model training strategy and predetermined training data to obtain the calcification region segmentation model with converged training, wherein the training data includes image feature training data, section feature training data and text feature training data; The model training strategy includes one or more of a single forward training scheme, a multiple forward and frozen training scheme, and a multiple forward and separated training scheme; the multiple forward and frozen training scheme includes a first global training part corresponding scheme and / or a picture-text matching training part corresponding scheme; the multiple forward and separated training scheme includes one or more of an encoder training part corresponding scheme, a decoder training part corresponding scheme, and a second global part corresponding scheme; Furthermore, the segmentation model training module trains the calcification region segmentation model based on a preset model training strategy and predetermined training data, and the method of obtaining the calcification region segmentation model with converged training specifically includes: Inputting the predetermined training data into the calcification region segmentation model for analysis, and obtaining the segmentation mask tensor training results and the text-related vector training results output by the calcification region segmentation model; Determining a model application loss result of the calcification region segmentation model according to the segmentation mask tensor training result, the text-related vector training result and a preset image-text matching loss determination condition; According to the model application loss result and the preset model training strategy scheme, judging whether the calcification region segmentation model meets the preset model training convergence condition; When it is determined that the calcified region segmentation model meets the model training convergence condition, determining that the trained calcified region segmentation model is converged, and the trained converged calcified region segmentation model is used to determine and segment the calcified region; When it is determined that the calcified region segmentation model does not satisfy the model training convergence condition, performing corresponding model training operations on the calcified region segmentation model according to a preset model training strategy scheme until the calcified region segmentation model is trained to convergence; And, the segmentation model training module performs corresponding model training operations on the calcification region segmentation model according to the preset model training strategy scheme, specifically including: When the preset model training strategy includes the single forward training scheme, based on the model application loss result, performing a corresponding reverse gradient propagation operation on the calcification region segmentation model; When the preset model training strategy includes the multiple forward and frozen training schemes and the multiple forward and frozen training schemes include the corresponding schemes of the image-text matching training part, based on the model application loss result, performing a corresponding reverse gradient propagation operation on the calcification region segmentation model; When the preset model training strategy includes the multiple forward and freeze training schemes and the multiple forward and freeze training schemes include the first global training part corresponding scheme, based on the model application loss result, performing corresponding parameter freezing operations except for the classification head on the calcification area segmentation model and performing corresponding reverse gradient propagation operations; When the preset model training strategy includes multiple forward and split training schemes and the multiple forward and split training schemes include the encoder training part corresponding scheme, based on the determined first loss result, performing a corresponding reverse gradient propagation training encoder operation on the calcified region segmentation model; When the preset model training strategy includes multiple forward and separate training schemes and the multiple forward and separate training schemes include the decoder training part corresponding scheme, based on the model application loss result, performing a corresponding encoder parameter freezing operation on the calcified area segmentation model and performing a corresponding reverse gradient propagation training decoder operation; When the preset model training strategy includes multiple forward and separate training schemes and the multiple forward and separate training schemes include the second global part corresponding scheme, based on the model application loss result, a corresponding reverse gradient propagation operation is performed on the calcification area segmentation model.

7. A calcification region segmentation device based on a multimodal feature fusion model, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the calcification area segmentation method based on the multimodal feature fusion model according to any one of claims 1 to 5.

8. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the calcification area segmentation method based on the multimodal feature fusion model according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-modal information fusion AI platform and application method

    CN119293726A