A structured sign-based multi-modal joint anchor training method

CN122491416BActive Publication Date: 2026-09-18WANLIYUN MEDICAL INFORMATION TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610991835.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-18
Estimated Expiration
2046-07-06

AI Technical Summary

Technical Problem

[0005]但是,通过上述方式训练出的视觉编码器或文本编码器,由于未在病理类别维度对视觉-文本语义空间施加显式结构约束,特征仅在样本层面被拉近而缺乏类别判别方向的一致性保障,在用于例如病灶定位的任务时,往往难以准确精细地定位出医学影像图像中病灶的位置

Benefits of technology

[0012]In this embodiment, the computing device first extracts information from the image report sample, generating structured semantic units corresponding to each sign involved in the image report sample. Then, using a text encoding model, semantic codes for the corresponding signs are generated based on the structured semantic units. Visual vectors for each image region of the medical image sample are generated using a visual encoding model. Subsequently, based on the visual vectors corresponding to each image region and the semantic vectors of the corresponding signs, visual evidence retrieval is performed to generate visual evidence vectors corresponding to each sign. Finally, the text encoding model and the visual encoding model are jointly optimized and trained based on the semantic vectors and the visual evidence vectors. Compared to existing technologies, this method generates structured semantic units for signs and combines these units with an active visual evidence retrieval mechanism to jointly optimize the text encoding model and the visual encoding model, enabling the establishment of a fine correspondence between image regions and text descriptions at the sign level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491416B_ABST
    Figure CN122491416B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal joint anchor training method based on structured signs. Information extraction is performed on image report samples to generate structured semantic units corresponding to each sign involved in the image report samples. Then, a text encoding model is used to generate semantic codes of the corresponding signs according to the structured semantic units, and a visual encoding model is used to generate visual vectors of image regions of the medical image samples. Subsequently, visual evidence vectors are generated through visual evidence retrieval based on the visual vectors and the semantic vectors of the corresponding signs. Finally, the text encoding model and the visual encoding model are optimized according to the semantic vectors and the visual evidence vectors. The method generates structured semantic units of signs, and combines the structured semantic units with active visual evidence retrieval to jointly optimize the text encoding model and the visual encoding model, so that a fine corresponding relationship between image regions and text descriptions can be established at the sign level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical image processing technology, and in particular to a multimodal joint anchoring training method based on structured features. Background Technology

[0002] Chest DR (Digital Radiography) is one of the most commonly used medical imaging methods in clinical practice, and it is widely used for screening and auxiliary diagnosis of lung diseases, pleural diseases, and cardiopulmonary structural abnormalities.

[0003] Currently, visual-language models can be pre-trained based on paired data of medical images and image reports to train visual encoders for encoding medical images and text encoders for encoding image reports. The trained visual encoders and text encoders can be used to support tasks such as disease classification, lesion localization, report generation, and cross-modal retrieval.

[0004] In existing technologies, the medical image is typically encoded by a visual encoder, and the image report text is encoded by a text encoder. Then, the visual encoder and the text encoder are pre-trained through contrastive learning.

[0005] However, visual encoders or text encoders trained in the above manner often fail to accurately and precisely locate lesions in medical images when used for tasks such as lesion localization because they do not impose explicit structural constraints on the visual-text semantic space in the pathological category dimension. The features are only brought closer at the sample level and lack consistency in the category discrimination direction.

[0006] There is currently no effective solution to the technical problem that existing visual encoders or text encoders often have difficulty accurately and precisely locating lesions in medical images when used for tasks such as lesion localization. Summary of the Invention

[0007] The embodiments of this disclosure provide a multimodal joint anchoring training method based on structured features, which at least solves the technical problem that existing visual encoders or text encoders often have difficulty accurately and precisely locating the position of lesions in medical images when used for tasks such as lesion localization.

[0008] According to one aspect of the present disclosure, a multimodal joint anchoring training method based on structured features is provided, comprising: acquiring image report samples and medical image samples corresponding to the image report samples; determining structured semantic units corresponding to each feature involved in the image report samples, wherein the structured semantic units include at least a feature field and a location field; for each feature, generating a semantic vector corresponding to the corresponding feature based on the structured semantic units using a text encoding model; dividing the medical image samples into multiple image regions, generating a visual vector corresponding to each image region using a visual encoding model; for each feature, determining the association weight between each visual vector and the semantic vector of the corresponding feature based on the semantic vector corresponding to the corresponding feature and the visual vector corresponding to each image region, and weighting and summing the visual vectors corresponding to each image region according to the corresponding association weights to generate a visual evidence vector corresponding to the corresponding feature; determining the joint loss of the text encoding model and the visual encoding model based on the semantic vector and the visual evidence vector; jointly training the text encoding model and the visual encoding model based on the joint loss; and using the trained text encoding model and / or visual encoding model to perform downstream tasks related to medical images.

[0009] According to another aspect of the present disclosure, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0010] According to another aspect of the present disclosure, a multimodal joint anchoring training device based on structured features is also provided, comprising: an acquisition module for acquiring image report samples and medical image samples corresponding to the image report samples; a structured unit determination module for determining structured semantic units corresponding to each feature involved in the image report samples, wherein the structured semantic units include at least a feature field and a location field; a semantic encoding module for generating a semantic vector corresponding to each feature based on the structured semantic units using a text encoding model; and a visual encoding module for dividing the medical image samples into multiple image regions and generating a semantic vector corresponding to each image region using a visual encoding model. The system comprises: a visual evidence generation module, which, for each sign, determines the association weight between each visual vector and the semantic vector of the corresponding sign based on the semantic vector corresponding to the sign and the visual vector corresponding to each image region, and weights and sums the visual vectors corresponding to each image region according to the corresponding association weights to generate a visual evidence vector corresponding to the corresponding sign; and a training module, which determines the joint loss of the text encoding model and the visual encoding model based on the semantic vector and the visual evidence vector, and jointly trains the text encoding model and the visual encoding model based on the joint loss. The trained text encoding model and / or visual encoding model are used to perform downstream tasks related to medical imaging.

[0011] According to another aspect of the present disclosure, a multimodal joint anchoring training device based on structured features is also provided, comprising: a processor; and a memory connected to the processor, configured to provide the processor with instructions for processing the following steps: acquiring image report samples and medical image samples corresponding to the image report samples; determining structured semantic units corresponding to each feature involved in the image report samples, wherein the structured semantic units include at least a feature field and a location field; for each feature, generating a semantic vector corresponding to the corresponding feature based on the structured semantic units using a text encoding model; dividing the medical image samples into multiple image regions, and using a visual encoding model, For each sign, a visual vector corresponding to each image region is generated. Based on the semantic vector corresponding to the sign and the visual vector corresponding to each image region, the association weight between each visual vector and the semantic vector of the sign is determined. The visual vectors corresponding to each image region are then weighted and summed according to their respective association weights to generate a visual evidence vector corresponding to the sign. Based on the semantic vector and the visual evidence vector, the joint loss of the text encoding model and the visual encoding model is determined. Based on the joint loss, the text encoding model and the visual encoding model are jointly trained. The trained text encoding model and / or visual encoding model are used to perform downstream tasks related to medical imaging.

[0012] In this embodiment, the computing device first extracts information from the image report sample, generating structured semantic units corresponding to each sign involved in the image report sample. Then, using a text encoding model, semantic codes for the corresponding signs are generated based on the structured semantic units. Visual vectors for each image region of the medical image sample are generated using a visual encoding model. Subsequently, based on the visual vectors corresponding to each image region and the semantic vectors of the corresponding signs, visual evidence retrieval is performed to generate visual evidence vectors corresponding to each sign. Finally, the text encoding model and the visual encoding model are jointly optimized and trained based on the semantic vectors and the visual evidence vectors. Compared to existing technologies, this method generates structured semantic units for signs and combines these units with an active visual evidence retrieval mechanism to jointly optimize the text encoding model and the visual encoding model, enabling the establishment of a fine correspondence between image regions and text descriptions at the sign level. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure. In the drawings: Figure 1 This is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of this disclosure; Figure 2 This is a flowchart illustrating the multimodal joint anchoring training method based on structured features according to the first aspect of Embodiment 1 of this disclosure; Figure 3 This is a schematic diagram of the overall process of generating structured semantic units and training the model according to the first aspect of Embodiment 1 of this disclosure; Figure 4 This is a schematic diagram of the process of active visual evidence retrieval according to the first aspect of Embodiment 1 of this disclosure; Figure 5 This is a schematic diagram of the process of generating structured semantic units according to the first aspect of Embodiment 1 of this disclosure; Figure 6 This is a schematic diagram illustrating the joint optimization of the text encoding model and the visual encoding model according to the first aspect of Embodiment 1 of this disclosure; Figure 7 This is a schematic diagram of a heat map according to the first aspect of Embodiment 1 of this disclosure; Figure 8 This is a schematic diagram of a multimodal joint anchoring training device based on structured features according to the first aspect of Embodiment 2 of this disclosure; Figure 9This is a schematic diagram of a multimodal joint anchoring training device based on structured features, according to the first aspect of Embodiment 3 of this disclosure. Detailed Implementation

[0014] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this disclosure.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Example 1

[0017] According to this embodiment, a multimodal joint anchoring training method based on structured features is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0018] The method embodiments provided in this example can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Figure 1 A hardware block diagram of a computing device for implementing a multimodal joint anchoring training method based on structured features is shown. Figure 1As shown, a computing device may include one or more processors (processors may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include a display, keyboard, and cursor control device connected to the input / output interface. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, a computing device may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0019] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element in a computing device. As involved in the embodiments of this disclosure, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).

[0020] The memory can be used to store software programs and modules of application software, such as the program instruction / data storage device corresponding to the multimodal joint anchoring training method based on structured features in the embodiments of this disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the above-mentioned application's multimodal joint anchoring training method based on structured features. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computing device via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0021] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the computing device's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0022] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of the computing device.

[0023] It should be noted here that, in some optional embodiments, the above... Figure 1 The computing device shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computing devices.

[0024] Under the aforementioned operating environment, according to the first aspect of this embodiment, a multimodal joint anchoring training method based on structured features is provided. This method consists of... Figure 1 The computing device shown is implemented. Figure 2 A flowchart illustrating the method is shown below. (Refer to...) Figure 2 As shown, the method includes: S202: Obtain an image report sample and the corresponding medical image sample; S204: Determine the structured semantic units corresponding to each feature involved in the image report sample, wherein the structured semantic units include at least a feature field and a location field; S206: For each symptom, a semantic vector corresponding to the corresponding symptom is generated based on the structured semantic units using a text encoding model; S208: Divide the medical image sample into multiple image regions and generate a visual vector corresponding to each image region through a visual coding model; S210: For each feature, based on the semantic vector corresponding to the corresponding feature and the visual vector corresponding to each image region, determine the association weight between each visual vector and the semantic vector of the corresponding feature, and then sum the visual vectors corresponding to each image region according to the corresponding association weights to generate a visual evidence vector corresponding to the corresponding feature; and S212: Determine the joint loss of the text encoding model and the visual encoding model based on the semantic vector and the visual evidence vector. Based on the joint loss, jointly train the text encoding model and the visual encoding model. The trained text encoding model and / or visual encoding model are used to perform downstream tasks related to medical images.

[0025] In this embodiment, a method for jointly pre-training a medical-related text encoding model and a visual encoding model is provided. The text encoding model can be used to extract sign-related features from image report text, and the visual encoding model can be used to extract features from medical images.

[0026] First, the computing device can acquire image report samples and corresponding medical image samples (S202). The image report samples and corresponding medical image samples are used to jointly train the text encoding model and the visual encoding model, which will be mentioned later.

[0027] Then, refer to Figure 3 As shown, the computing device can determine the structured semantic units corresponding to each feature involved in the image report sample, wherein the structured semantic units include at least a feature field and a location field (S204).

[0028] In other words, the computing device can extract the sign information of each sign involved in the image report sample, as well as the information related to each sign (such as location information, nature information, degree information, and genus information), thereby determining the structured semantic unit corresponding to each sign. The structured semantic unit corresponding to a sign includes at least a sign field and a location field. That is, the sign information corresponding to the sign (e.g., "patchy shadow") can be used as the sign field in the structured semantic unit of that sign, and the location information related to the sign (e.g., "both lungs" or "lower right lung field") can be used as the location field in the structured semantic unit. The structured semantic unit may also include a nature field, a degree field, and an attribute field. Examples of structured semantic units will be provided later.

[0029] After the computing device determines the structured semantic unit corresponding to each feature, it can generate a semantic vector corresponding to the corresponding feature based on the structured semantic unit through a text encoding model for each feature (S206).

[0030] Continue to refer to Figure 3 As shown, after the computing device determines the structured semantic unit corresponding to each feature, it can generate a semantic vector corresponding to each feature based on the structured semantic unit using a text encoding model. The semantic vector expression for each feature is as follows:

[0031] in, Indicates the first The semantic vector corresponding to each feature. For the number of signs, This represents the dimension of the semantic vector.

[0032] The computing device then divides the medical image sample into multiple image regions and generates a visual vector corresponding to each image region through a visual coding model (S208).

[0033] Continue to refer to Figure 3 As shown, the computing device can divide the medical image corresponding to the medical image sample into multiple image regions. ,in, Indicates the first Each image region is then divided into several regions. A visual encoding model is then used to generate a visual vector corresponding to each region. The expression for the visual vector corresponding to each image region is shown below:

[0034] in, Indicates the first The visual vector corresponding to each image region The number of image regions. is the dimension of the visual vector.

[0035] The visual encoding model can be a Transformer visual encoding network, a convolutional neural network, or a CNN-Transformer hybrid network. Preferably, the visual token (i.e., the visual vector) retains spatial location information to support subsequent region localization tasks. For example, if the visual vector is generated through a Transformer visual encoding network, the location encoding output by the Transformer visual encoding network can be retained.

[0036] Continue to refer to Figure 3 As shown, the computing device performs visual evidence retrieval based on the visual vector corresponding to each image region and the semantic vector corresponding to the corresponding feature, and generates a visual evidence vector corresponding to the corresponding feature.

[0037] That is, after the computing device determines the semantic vector corresponding to the sign and the visual vector corresponding to each image region, for each sign, it determines the association weight between each visual vector and the semantic vector of the corresponding sign based on the semantic vector corresponding to the corresponding sign and the visual vector corresponding to each image region, and then weights and sums the visual vectors corresponding to each image region according to the corresponding association weights to generate the visual evidence vector corresponding to the corresponding sign (S210).

[0038] Specifically, refer to Figure 4 As shown, the process by which a computing device generates a visual evidence vector corresponding to a given sign can be called active visual evidence retrieval.

[0039] First, the computing device generates the semantic vector corresponding to the symptom using the following formula (1). Association weights between each visual vector ~ The association weight between the semantic vector and the visual vector corresponding to a sign is equivalent to the association weight between the sign and the image region corresponding to the visual vector: (1) in, =1~ , Used to indicate the first The first sign and the first Association weights between image regions.

[0040] In other words, proactive visual evidence retrieval leverages pathological features to actively drive visual region selection by pointing to clearly defined, structured signs. For each sign's semantic vector... Use it as a query condition for the visual token set Active visual evidence retrieval is performed using the set of visual vectors corresponding to each image region. The association weight is calculated from the dynamic relationship between the semantic vector of the feature and the visual token representation (i.e., the visual vector), and is used to characterize the degree of attention that the corresponding feature pays to different visual regions (image regions).

[0041] Then, the computing device weights and aggregates the visual vectors corresponding to each image region according to their respective association weights to generate a result similar to the first image region. Visual evidence vector corresponding to each sign : (2) Through the above mechanism, a fine-grained cross-modal correspondence is established between the structured semantic units of each sign and the local regions of the chest DR image. It should be noted that the active visual evidence retrieval mechanism in this embodiment differs from the cross-attention fusion operation in conventional multimodal models: conventional cross-attention is usually used as an internal component of the modality fusion module, and its output is used for subsequent classification or generation tasks, rather than directly as an independent supervisory signal aligned with the sign semantic vector; furthermore, in this embodiment, the visual evidence vector obtained through retrieval aggregation is regarded as the sign semantic vector. The model learns to actively learn "what combinations of visual regions can constitute evidence supporting the semantics of structured features" by mapping the corresponding regions in the visual space and using the high similarity between the two regions in the embedding space as an explicit pre-training optimization target.

[0042] Finally, the computing device can generate visual evidence vectors corresponding to each feature, expressed as:

[0043] The aforementioned correlation weights are not only used to form visual evidence vectors for corresponding signs, but also as a measure of regional response intensity for heatmap generation in subsequent sign localization tasks.

[0044] Finally, the computing device determines the joint loss of the text encoding model and the visual encoding model based on the semantic vector and the visual evidence vector. Based on the joint loss, the text encoding model and the visual encoding model are jointly trained. The trained text encoding model and / or visual encoding model are used to perform downstream tasks related to medical images (S212).

[0045] Specifically, after determining the visual evidence vector corresponding to each feature, the text encoding model and the visual encoding model can be jointly trained (pre-trained). The computing device can determine the joint loss for both the text encoding model and the visual encoding model based on the semantic vector and the visual evidence vector. By minimizing this joint loss, the text encoding model and the visual encoding model are jointly trained. The joint loss includes joint anchoring loss, multimodal alignment loss, and regularization loss. The regularization loss is optional, and each loss included in the joint loss will be explained in detail later.

[0046] As described in the background section, in existing technologies, a visual encoder typically encodes the entire medical image, while a text encoder encodes the image report text. Both the visual and text encoders are then pre-trained using contrastive learning. However, the visual or text encoder trained in this way lacks explicit structural constraints on the visual-text semantic space at the pathological category dimension. Features are only brought together at the sample level, lacking consistency in category discrimination direction. Therefore, when used for tasks such as lesion localization, it often struggles to accurately and precisely locate lesions in medical images.

[0047] In view of this, in this embodiment, the computing device first extracts information from the image report sample, generating structured semantic units corresponding to each sign involved in the image report sample. Then, through a text encoding model, semantic codes (semantic vectors) for the corresponding signs are generated based on the structured semantic units corresponding to the signs. Visual vectors for each image region of the medical image sample are generated through a visual encoding model. Subsequently, based on the visual vectors corresponding to each image region and the semantic vectors of the corresponding signs, visual evidence retrieval is performed to generate visual evidence vectors. Finally, the text encoding model and the visual encoding model are jointly optimized and trained based on the semantic vectors and the visual evidence vectors. Compared with existing technologies, this method generates structured semantic units for signs and combines structured semantic units with active visual evidence retrieval to jointly optimize the text encoding model and the visual encoding model, enabling the establishment of a fine correspondence between image regions and text descriptions at the sign level.

[0048] The computing device needs to acquire image report samples and corresponding medical image samples in advance. In this embodiment, medical images can refer to chest DR images.

[0049] First, the computing device constructs a paired dataset of chest DR images and radiological diagnostic reports. After acquiring the chest DR images and their corresponding radiological diagnostic reports, the "Imaging Findings" text paragraph is extracted from the reports as input for subsequent structured parsing. To ensure the integrity of subtle signs and anatomical orientation information in the chest DR images, the chest DR images undergo standardized preprocessing: image format standardization, grayscale value normalization, window width and level (or grayscale range) standardization, resolution standardization, and necessary de-identification processing.

[0050] The computing device can use the preprocessed chest DR image as a medical image sample and the paired "Imaging Findings" text paragraph as an image report sample corresponding to the corresponding medical image sample.

[0051] To avoid disrupting medical anatomical information and the spatial relationships of lesions, strong data augmentation strategies sensitive to spatial orientation, such as horizontal flipping, large-angle rotation, or excessive random cropping, are not employed during the preprocessing stage. These constraints ensure that the model effectively maintains its ability to perceive spatial relationships and subtle lesion signs in chest DR images.

[0052] Optionally, the operation of determining the structured semantic units corresponding to each sign involved in the image report sample includes: extracting information from the image report sample to determine the sign information contained in the report text corresponding to the image report sample, as well as the location information, nature information, degree information, and attribute information corresponding to the sign information; generating structured semantic units corresponding to the corresponding signs based on the sign information, the location information, nature information, degree information, and attribute information corresponding to the sign information, wherein the structured semantic units include sign fields, location fields, nature fields, degree fields, and attribute fields corresponding to the corresponding signs.

[0053] Specifically, the process of generating structured semantic units of features is as follows: Figure 5 As shown, an image report sample may involve multiple features, and the computing device can extract structured semantic units corresponding to multiple features from the report text ("Imagery Views") of the image report sample. A structured semantic unit corresponding to a feature can be a quintuple. .in: E represents the Entity field, such as "patchy shadow"; P represents the Position field, such as "both lungs"; N represents the nature field, such as "high density"; D represents the degree field, such as "frequent". A represents an attribute field, such as "companion nodules".

[0054] Therefore, when generating structured semantic units, the computing device can extract information from the image report sample, determine the symptom information contained in the report text corresponding to the image report sample, as well as the location information, nature information, degree information and attribute information corresponding to the symptom information. Then, based on the symptom information and the location information, nature information, degree information and attribute information corresponding to the symptom information, a structured semantic unit corresponding to the symptom is generated. The generated structured semantic unit includes the symptom field, location field, nature field, degree field and attribute field corresponding to the symptom.

[0055] In other words, generating structured semantic units is equivalent to the computing device extracting structured information from the image report sample, extracting the sign information, location information, nature information, degree information and attribute information related to each sign in the image report sample, and combining this information into structured semantic units corresponding to the corresponding signs.

[0056] Structured parsing (i.e., extracting sign information, location information, nature information, degree information, and attribute information for each sign from image report samples) can employ large language models, medical named entity recognition models, rule-based terminology parsing models, or combinations of the aforementioned models. Medical lexicon constraints and structured rule constraints can be introduced to improve the stability and consistency of the extraction.

[0057] Optionally, for each symptom, the operation of generating a semantic vector corresponding to the corresponding symptom through a text encoding model based on structured semantic units includes: converting the structured semantic unit corresponding to the corresponding symptom into a natural language description sentence corresponding to the corresponding symptom; inputting the natural language description sentence into the text encoding model to generate a semantic vector corresponding to the corresponding symptom.

[0058] Specifically, to enhance the semantic robustness of the text encoding model and eliminate the rigidity of expression caused by structured parsing, the computing device can convert the structured semantic units corresponding to the symptoms into natural language descriptions with different styles and complete information structures. For example, the structured semantic unit {E: "consolidation"; P: "right lower lung field"; D: "single"; N: "high density, blurred boundaries, patchy"} can be converted into natural language descriptions such as: "Consolidation occurs in the right lower lung field, characterized by patchy, high density, blurred boundaries, and single lesion," or "A single patchy high-density consolidation is visible in the right lower lung field, with unclear boundaries." Then, the computing device can input the converted natural language descriptions into the text encoding model to generate semantic vectors corresponding to the respective symptoms.

[0059] If the computing device generates multiple different natural language descriptions for the same symptom's structured semantic unit, then the pairing between each natural language description and the medical image sample can be used as a separate training sample.

[0060] Text encoding models can employ Transformer text encoding models, medical text pre-training models, or general language models.

[0061] The joint losses mentioned above include: joint anchoring loss, multimodal alignment loss, and regularization loss. That is, refer to... Figure 6 As shown, the computing device jointly trains the text encoding model and the visual encoding model with the optimization objectives of minimizing the joint anchoring loss, multimodal alignment loss, and regularization loss.

[0062] The aforementioned joint anchoring loss and multimodal alignment loss are complementary and work together to address the problem of poor clinical interpretability and integrate global alignment goals. They construct a semantically consistent shared space through bidirectional constraints, enabling consistent learning between semantic features and visual evidence. Regularization loss plays a supplementary role, enhancing the distinguishability between each feature and its corresponding region, thereby improving the discriminative power of the weight distribution and retrieval stability.

[0063] Optionally, the process of jointly optimizing the text encoding model and the visual encoding model through the joint anchoring loss in the joint loss is as follows: by using a preset shared classifier, a first classification result is generated based on the semantic vector of the corresponding feature, and a second classification result is generated based on the visual evidence vector of the corresponding feature. The text encoding model and the visual encoding model are jointly optimized with the optimization objective of minimizing the difference between the first classification result and the feature field and minimizing the difference between the second classification result and the feature field. Furthermore, the optimization of joint training of the text encoding model and the visual encoding model through the multimodal alignment loss in the joint loss is to optimize the text encoding model and the visual encoding model by maximizing the similarity between the semantic vector and the visual evidence vector corresponding to the same feature and minimizing the similarity between the semantic vector and the visual evidence vector corresponding to different features.

[0064] In the joint anchoring loss, a pre-defined shared classifier generates classification results corresponding to the two modalities based on the semantic vectors and visual evidence vectors of the corresponding features. The text encoding model and the visual encoding model are jointly optimized with the objective of minimizing the difference between the two classification results and the feature fields. In other words, this pre-defined shared classifier is shared between the text modality and the image modality, and it can receive visual evidence vectors. semantic vector of or phenomenon (After projection) these features are used as input, and classification losses are calculated separately. Since the two branches share the same set of classification weights, the output features of the two branches are forced to be arranged according to the same class discrimination direction, thereby jointly anchoring the semantic spaces of the two modalities to a consistent structure. The class labels for classification-supervised fine-tuning use sign fields obtained in the construction of structured semantic units, such as "nodular shadow" and "mass".

[0065] The cross-modal alignment loss aims to maximize the similarity between the semantic vector and visual evidence vector corresponding to the same feature, while minimizing the similarity between the semantic vector and visual evidence vector corresponding to different features. This is achieved through joint optimization of the text encoding model and the visual encoding model. In other words, the goal of the cross-modal alignment loss is to maximize the similarity between the semantic vectors s corresponding to the same feature and the visual evidence vectors.j With visual evidence vector z j They exhibit high similarity in the embedding space, while maintaining low similarity with visual evidence vectors of non-corresponding features.

[0066] The expression for cross-modal alignment loss can be given as:

[0067] Where sim(·,·) is a similarity function (such as cosine similarity, dot product similarity, or Euclidean distance transformation). This refers to the temperature parameter.

[0068] Optionally, the process of jointly optimizing the text encoding model and the visual encoding model using the regularization loss in the joint loss is as follows: The text encoding model and the visual encoding model are jointly optimized by minimizing the entropy constraint and the regional difference constraint, where the expression for the entropy constraint is:

[0069] For the number of signs, The number of image regions. This represents the association weight between the semantic vector corresponding to the j-th feature and the visual vector corresponding to the i-th image region; The expression for the regional difference constraint is:

[0070] Indicates the relationship with the first The regional correlation weight distribution corresponding to each symptom.

[0071] Specifically, regularization loss This can include: entropy constraints and regional differences constraints Regularization loss is used to improve the discriminativeness of the weight distribution and the stability of retrieval, so as to avoid different features from degenerating into the same region.

[0072] Among them, entropy constraint The expression is:

[0073] For the number of signs, The number of image regions. Let represent the association weight between the semantic vector corresponding to the j-th feature and the visual vector corresponding to the i-th image region. The purpose of minimizing the entropy constraint is to make the association weight distribution of each feature more deterministic (i.e., to make the shape of the association weight distribution sharper), thereby indirectly increasing the possibility of forming differentiated distributions among different features.

[0074] Regional differences constraints The expression is:

[0075] This represents the region association weight distribution corresponding to the j-th feature. Wherein, include .

[0076] In summary, the total loss function (i.e., the joint loss function) used for joint training of the text encoding model and the visual encoding model is: .in, , as well as These represent the weights of the joint anchoring loss, cross-modal alignment loss, and regularization loss, respectively. The computing device can use this total loss function, combined with the joint anchoring loss, cross-modal alignment loss, and regularization loss, to jointly optimize the text encoding model and the visual encoding model. This bidirectional constraint mechanism ensures that the shared latent space possesses both clear semantic directionality and accurate cross-modal correspondences.

[0077] Furthermore, before jointly training the text encoding model and the visual encoding model using the aforementioned total loss function, the computing device can perform independent, lightweight classification-supervised fine-tuning on each model separately. This step aims to enable the text encoding model and the visual encoding model to independently learn and extract discriminative features related to the semantics of the features without relying on the other modality, providing a better starting point for subsequent accurate alignment.

[0078] The category labels for classification-supervised fine-tuning utilize the feature fields obtained in the structured semantic unit construction step, such as "nodular shadow" and "mass." Specifically, a first predicted category can be generated based on the semantic vector corresponding to the feature using a classifier corresponding to the text encoding model, and a second predicted category can be generated based on the visual evidence vector corresponding to the feature using a classifier corresponding to the visual encoding model. The text encoding model is trained separately with the optimization objective of minimizing the difference between the first predicted category and the category label. Similarly, the visual encoding model is trained separately with the optimization objective of minimizing the difference between the second predicted category and the category label.

[0079] Optionally, the method further includes: acquiring the image report to be processed and the medical image corresponding to the image report; using a trained text encoding model, generating semantic vectors corresponding to the structured semantic units corresponding to the signs involved in the image report; using a trained visual encoding model, generating visual vectors corresponding to multiple image regions of the medical image; generating association weights between the corresponding signs and multiple image regions of the medical image based on the visual vectors corresponding to the corresponding signs and multiple image regions of the medical image; and generating a lesion sign heatmap based on the association weights between the semantic vectors corresponding to the corresponding signs and multiple image regions of the medical image.

[0080] Specifically, the text encoding model and visual encoding model provided in this application can be used to generate, for example... Figure 7 The heatmap shown is shown below.

[0081] The computing device can acquire the image report to be processed and the corresponding medical images. Subsequently, based on the image report and the corresponding medical images, a heatmap of lesion signs will be generated, such as... Figure 7 The heatmap of lesion signs shown is used to indicate the location of the signs involved in the image report to be processed in the corresponding medical images.

[0082] The computing device can determine the text seen in the image report and extract information from it to identify the structured semantic units corresponding to the signs involved. Then, using a trained text encoding model, the device can generate semantic vectors for the corresponding signs, and using a trained visual encoding model, it can generate visual vectors corresponding to multiple image regions of the medical image. Subsequently, using a method for generating visual evidence vectors, it can generate association weights between the corresponding signs and multiple image regions of the medical image, and based on these association weights, generate a heatmap of lesion signs. In other words, the magnitude of the association weights directly determines the color of the corresponding location in the heatmap. Figure 7 As can be seen, areas of high attention tend to be red, while areas of low attention tend to be blue. The higher the correlation weight, the higher the attention given to the corresponding area in the heatmap. Interpolation can be used to restore the correlation weights to the pixel level, thereby generating a corresponding heatmap of lesion signs.

[0083] Figure 7The example shown uses the query input "a nodule is visible in the left lung and upper field, with smooth edges and a size of about 2.4 × 2.4 cm". The regional response heatmap generated by the conventional basic pre-trained model is usually widely distributed in the left lung region with low focus. The regional response heatmap (lesion sign heatmap) generated by the present invention based on structured sign semantic active visual evidence retrieval can focus more concentratedly on the vicinity of the actual nodule location, achieving more accurate fine-grained lesion localization.

[0084] In addition to generating heatmaps for feature localization as mentioned above, the trained text encoding model and visual encoding model in this application can also be used for downstream tasks such as disease classification, radiological report generation, and cross-modal retrieval.

[0085] Disease classification refers to classifying medical images for diseases based on visual vectors generated by a visual encoding model (e.g., classifying chest diseases or identifying multi-label anomalies in chest DR images). Radiology report generation refers to generating structured medical descriptions or complete radiology reports based on visual vectors generated by a visual encoding model. Cross-modal retrieval refers to retrieving images from image report text or retrieving descriptions of signs from medical images.

[0086] Furthermore, the structured semantic units of features are not limited to quintuples; quadtuples or graph structures can be used. Visual encoding models are not limited to ViT; ResNet or Swin Transformer can be used. In active visual evidence retrieval, the method for generating association weights between features and image regions is not limited to dot-product attention; multi-head attention can be used.

[0087] The method of this invention was pre-trained and tested on a publicly available chest DR dataset. Preliminary results show that the invention can effectively establish the correspondence between signs and image regions, verifying the feasibility and effectiveness of the technical solution provided in this embodiment.

[0088] In addition, refer to Figure 1 As shown, according to a second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0089] Thus, according to this embodiment, (1) fine-grained cross-modal semantic alignment is achieved: compared with existing methods that use only disease entities or complete reports for image-text alignment, this embodiment can establish a fine correspondence between image regions and text descriptions at the sign level through joint optimization of structured sign semantic units and active visual evidence retrieval.

[0090] (2) No anatomical detector required: Compared with existing methods that rely on pre-trained anatomical region detectors or lung region annotations (such as anatomical region alignment-based technical solutions), this embodiment can achieve implicit correspondence learning of signs and regions without additional annotations or pre-trained models.

[0091] (3) Enhance the ability to express complex medical semantics: Compared with existing methods that only extract disease categories, the five-tuple structure of the present invention includes key clinical dimensions such as properties, degree, and attributes, which can represent more complex symptom semantics.

[0092] (4) Fine-grained interpretability: Compared with the black-box image-text alignment model, this invention actively retrieves visual evidence through structured sign semantics, outputs the regional association weight distribution corresponding to the sign, and can generate a regional response heatmap (lesion sign heatmap) corresponding to the specific sign semantics. At the same time, through the joint semantic anchoring mechanism (shared classifier + contrast loss bidirectional constraint), visual evidence and sign semantics are arranged in the same discriminative direction in the shared space, so that the model reasoning process has dual interpretability of regional visual interpretation and semantic direction interpretability, which significantly improves clinical credibility.

[0093] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0095] Example 2

[0096] Figure 8 A multimodal joint anchoring training apparatus based on structured features according to a first aspect of this embodiment is shown, which corresponds to the method described according to the first aspect of Embodiment 1. Reference Figure 8As shown, the device includes: an acquisition module 810 for acquiring image report samples and medical image samples corresponding to the image report samples; a structured unit determination module 820 for determining structured semantic units corresponding to each sign involved in the image report samples, wherein the structured semantic units include at least a sign field and a location field; a semantic encoding module 830 for generating a semantic vector corresponding to each sign based on the structured semantic units using a text encoding model; a visual encoding module 840 for dividing the medical image samples into multiple image regions and generating a visual vector corresponding to each image region using a visual encoding model; and a visual evidence generation module 850 for... For each sign, based on the semantic vector corresponding to the corresponding sign and the visual vector corresponding to each image region, the association weight between each visual vector and the semantic vector of the corresponding sign is determined, and the visual vectors corresponding to each image region are weighted and summed according to the corresponding association weights to generate a visual evidence vector corresponding to the corresponding sign; the training module 860 is used to determine the joint loss of the text encoding model and the visual encoding model based on the semantic vector and the visual evidence vector, and to jointly train the text encoding model and the visual encoding model based on the joint loss. The trained text encoding model and / or visual encoding model is used to perform downstream tasks related to medical imaging.

[0097] Optionally, the structured unit determination module 820 is used to extract information from the image report sample, determine the sign information contained in the report text corresponding to the image report sample, as well as the location information, nature information, degree information and attribute information corresponding to the sign information; and generate a structured semantic unit corresponding to the corresponding sign based on the sign information, the location information, nature information, degree information and attribute information corresponding to the sign information, wherein the structured semantic unit includes the sign field, location field, nature field, degree field and attribute field corresponding to the corresponding sign.

[0098] Optionally, the semantic encoding module 830 is used to convert the structured semantic units corresponding to the corresponding features into natural language description sentences corresponding to the corresponding features; input the natural language description sentences into the text encoding model to generate semantic vectors corresponding to the corresponding features.

[0099] Optionally, the joint loss includes: joint anchoring loss, multimodal alignment loss, and regularization loss; the training module 860 is used to jointly optimize the text encoding model and the visual encoding model through the joint anchoring loss in the joint loss, the process of which is as follows: by using a preset shared classifier, a first classification result is generated based on the semantic vector of the corresponding feature, and a second classification result is generated based on the visual evidence vector of the corresponding feature, and the text encoding model and the visual encoding model are jointly optimized with the optimization objective of minimizing the difference between the first classification result and the feature field and minimizing the difference between the second classification result and the feature field.

[0100] Optionally, the training module 860 is used to jointly optimize the text encoding model and the visual encoding model through the multimodal alignment loss in the joint loss. The process is as follows: the text encoding model and the visual encoding model are jointly optimized with the optimization objective of maximizing the similarity between the semantic vector and the visual evidence vector corresponding to the same feature and minimizing the similarity between the semantic vector and the visual evidence vector corresponding to different features.

[0101] Optionally, the training module 860 is used to jointly optimize the text encoding model and the visual encoding model using the regularization loss in the joint loss. The process involves training the text encoding model and the visual encoding model to minimize entropy constraints and regional difference constraints, wherein the expression for the entropy constraint is:

[0102] For the number of signs, The number of image regions. This represents the association weight between the semantic vector corresponding to the j-th feature and the visual vector corresponding to the i-th image region; The expression for the regional difference constraint is:

[0103] This represents the regional association weight distribution corresponding to the j-th symptom.

[0104] Optionally, the visual evidence generation module 850 is further configured to acquire the image report to be processed and the medical image corresponding to the image report; generate semantic vectors corresponding to the corresponding signs based on the structured semantic units corresponding to the signs involved in the image report using a trained text encoding model; generate visual vectors corresponding to multiple image regions of the medical image using a trained visual encoding model; generate association weights between the corresponding signs and multiple image regions of the medical image based on the semantic vectors corresponding to the corresponding signs and the visual vectors corresponding to the multiple image regions of the medical image; and generate a heatmap of lesion signs based on the association weights between the corresponding signs and multiple image regions of the medical image.

[0105] Thus, according to this embodiment, (1) fine-grained cross-modal semantic alignment is achieved: compared with existing methods that use only disease entities or complete reports for image-text alignment, this embodiment can establish a fine correspondence between image regions and text descriptions at the sign level through joint optimization of structured sign semantic units and active visual evidence retrieval.

[0106] (2) No anatomical detector required: Compared with existing methods that rely on pre-trained anatomical region detectors or lung region annotations (such as anatomical region alignment-based technical solutions), this embodiment can achieve implicit correspondence learning of signs and regions without additional annotations or pre-trained models.

[0107] (3) Enhance the ability to express complex medical semantics: Compared with existing methods that only extract disease categories, the five-tuple structure of the present invention includes key clinical dimensions such as properties, degree, and attributes, which can represent more complex symptom semantics.

[0108] (4) Fine-grained interpretability: Compared with the black-box image-text alignment model, this invention actively retrieves visual evidence through structured sign semantics, outputs the regional association weight distribution corresponding to the sign, and can generate a regional response heatmap (lesion sign heatmap) corresponding to the specific sign semantics. At the same time, through the joint semantic anchoring mechanism (shared classifier + contrast loss bidirectional constraint), visual evidence and sign semantics are arranged in the same discriminative direction in the shared space, so that the model reasoning process has dual interpretability of regional visual interpretation and semantic direction interpretability, which significantly improves clinical credibility.

[0109] Example 3

[0110] Figure 9 A multimodal joint anchoring training apparatus based on structured features according to a first aspect of this embodiment is shown, which corresponds to the method described according to the first aspect of Embodiment 1. Reference Figure 9As shown, the device includes: a processor 910; and a memory 920 connected to the processor 910, used to provide the processor with instructions to perform the following processing steps: acquiring an image report sample and a medical image sample corresponding to the image report sample; determining the structured semantic units corresponding to each sign involved in the image report sample, wherein the structured semantic units include at least a sign field and a location field; for each sign, generating a semantic vector corresponding to the corresponding sign based on the structured semantic units using a text encoding model; dividing the medical image sample into multiple image regions, and generating a visual encoding model corresponding to each image region. For each sign, based on the semantic vector corresponding to the sign and the visual vector corresponding to each image region, the association weight between each visual vector and the semantic vector of the corresponding sign is determined. The visual vectors corresponding to each image region are then weighted and summed according to their respective association weights to generate a visual evidence vector corresponding to the sign. Based on the semantic vector and the visual evidence vector, the joint loss of the text encoding model and the visual encoding model is determined. Based on the joint loss, the text encoding model and the visual encoding model are jointly trained. The trained text encoding model and / or visual encoding model are used to perform downstream tasks related to medical imaging.

[0111] Optionally, the operation of determining the structured semantic units corresponding to each sign involved in the image report sample includes: extracting information from the image report sample to determine the sign information contained in the report text corresponding to the image report sample, as well as the location information, nature information, degree information, and attribute information corresponding to the sign information; generating structured semantic units corresponding to the corresponding signs based on the sign information, the location information, nature information, degree information, and attribute information corresponding to the sign information, wherein the structured semantic units include sign fields, location fields, nature fields, degree fields, and attribute fields corresponding to the corresponding signs.

[0112] Optionally, for each symptom, the operation of generating a semantic vector corresponding to the corresponding symptom through a text encoding model based on structured semantic units includes: converting the structured semantic unit corresponding to the corresponding symptom into a natural language description sentence corresponding to the corresponding symptom; inputting the natural language description sentence into the text encoding model to generate a semantic vector corresponding to the corresponding symptom.

[0113] Optionally, the joint loss includes: joint anchoring loss, multimodal alignment loss, and regularization loss; and the process of jointly optimizing the text encoding model and the visual encoding model through the joint anchoring loss in the joint loss is as follows: by using a preset shared classifier, a first classification result is generated based on the semantic vector of the corresponding feature, and a second classification result is generated based on the visual evidence vector of the corresponding feature. The text encoding model and the visual encoding model are jointly optimized with the optimization objective of minimizing the difference between the first classification result and the feature field and minimizing the difference between the second classification result and the feature field.

[0114] Optionally, the process of jointly optimizing the text encoding model and the visual encoding model through the multimodal alignment loss in the joint loss is as follows: the optimization objective is to maximize the similarity between the semantic vector and the visual evidence vector corresponding to the same feature, and minimize the similarity between the semantic vector and the visual evidence vector corresponding to different features, and then jointly optimize the text encoding model and the visual encoding model.

[0115] Optionally, the process of jointly optimizing the text encoding model and the visual encoding model using the regularization loss in the joint loss is as follows: The text encoding model and the visual encoding model are jointly optimized to minimize the entropy constraint and the regional difference constraint, wherein the expression for the entropy constraint is:

[0116] For the number of signs, The number of image regions. This represents the association weight between the semantic vector corresponding to the j-th feature and the visual vector corresponding to the i-th image region; The expression for the regional difference constraint is:

[0117] This represents the regional association weight distribution corresponding to the j-th symptom.

[0118] Optionally, the memory 920 is also configured to provide the processor with instructions for processing the following steps: acquiring an image report to be processed and the medical image corresponding to the image report; using a trained text encoding model, generating semantic vectors corresponding to the structured semantic units corresponding to the signs involved in the image report; using a trained visual encoding model, generating visual vectors corresponding to multiple image regions of the medical image; generating association weights between the corresponding signs and multiple image regions of the medical image based on the semantic vectors corresponding to the corresponding signs and the visual vectors corresponding to the multiple image regions of the medical image; and generating a lesion sign heatmap based on the association weights between the corresponding signs and multiple image regions of the medical image.

[0119] Thus, according to this embodiment, (1) fine-grained cross-modal semantic alignment is achieved: compared with existing methods that use only disease entities or complete reports for image-text alignment, this embodiment can establish a fine correspondence between image regions and text descriptions at the sign level through joint optimization of structured sign semantic units and active visual evidence retrieval.

[0120] (2) No anatomical detector required: Compared with existing methods that rely on pre-trained anatomical region detectors or lung region annotations (such as anatomical region alignment-based technical solutions), this embodiment can achieve implicit correspondence learning of signs and regions without additional annotations or pre-trained models.

[0121] (3) Enhance the ability to express complex medical semantics: Compared with existing methods that only extract disease categories, the five-tuple structure of the present invention includes key clinical dimensions such as properties, degree, and attributes, which can represent more complex symptom semantics.

[0122] (4) Fine-grained interpretability: Compared with the black-box image-text alignment model, this invention actively retrieves visual evidence through structured sign semantics, outputs the regional association weight distribution corresponding to the sign, and can generate a regional response heatmap (lesion sign heatmap) corresponding to the specific sign semantics. At the same time, through the joint semantic anchoring mechanism (shared classifier + contrast loss bidirectional constraint), visual evidence and sign semantics are arranged in the same discriminative direction in the shared space, so that the model reasoning process has dual interpretability of regional visual interpretation and semantic direction interpretability, which significantly improves clinical credibility.

[0123] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0124] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0125] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0126] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0127] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0128] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0129] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multimodal joint anchoring training method based on structured features, characterized in that, include: Obtain image report samples and corresponding medical image samples; The structured semantic units corresponding to each feature involved in the image report sample are determined, wherein the structured semantic units include at least a feature field and a location field; For each symptom, a semantic vector corresponding to the symptom is generated based on the structured semantic unit using a text encoding model. The medical image sample is divided into multiple image regions, and a visual vector corresponding to each image region is generated through a visual coding model. For each sign, based on the semantic vector corresponding to the corresponding sign and the visual vector corresponding to each image region, the association weight between each visual vector and the semantic vector of the corresponding sign is determined, and the visual vectors corresponding to each image region are weighted and summed according to the corresponding association weights to generate a visual evidence vector corresponding to the corresponding sign. Based on the semantic vector and the visual evidence vector, a joint loss for the text encoding model and the visual encoding model is determined. Based on the joint loss, the text encoding model and the visual encoding model are jointly trained. The trained text encoding model and / or visual encoding model are then used to perform downstream tasks related to medical images. The joint loss includes: joint anchoring loss, multimodal alignment loss, and regularization loss; and wherein... The process of jointly optimizing the text encoding model and the visual encoding model using the joint anchoring loss in the joint loss is as follows: by using a preset shared classifier, a first classification result is generated based on the semantic vector of the corresponding feature, and a second classification result is generated based on the visual evidence vector of the corresponding feature. The text encoding model and the visual encoding model are jointly optimized with the optimization objective of minimizing the difference between the first classification result and the feature field and minimizing the difference between the second classification result and the feature field.

2. The method according to claim 1, characterized in that, The operation of determining the structured semantic units corresponding to each feature involved in the image report sample includes: Information is extracted from the image report sample to determine the symptom information contained in the report text corresponding to the image report sample, as well as the location information, nature information, degree information and attribute information corresponding to the symptom information; Based on the symptom information, the location information, nature information, degree information, and attribute information corresponding to the symptom information, a structured semantic unit corresponding to the corresponding symptom is generated. The structured semantic unit includes a symptom field, a location field, a nature field, a degree field, and an attribute field corresponding to the corresponding symptom.

3. The method according to claim 1, characterized in that, For each symptom, the operation of generating a semantic vector corresponding to the corresponding symptom using a text encoding model based on the structured semantic unit includes: The structured semantic units corresponding to the corresponding features are converted into natural language description sentences corresponding to the corresponding features; The natural language description is input into the text encoding model to generate a semantic vector corresponding to the corresponding feature.

4. The method according to claim 1, characterized in that, in, The process of jointly optimizing the text encoding model and the visual encoding model using the multimodal alignment loss in the joint loss is as follows: the optimization objective is to maximize the similarity between the semantic vector and the visual evidence vector corresponding to the same feature, and minimize the similarity between the semantic vector and the visual evidence vector corresponding to different features, and the text encoding model and the visual encoding model are jointly optimized.

5. The method according to claim 1, characterized in that, in, The process of jointly optimizing the text encoding model and the visual encoding model using the regularization loss in the joint loss is as follows: The text encoding model and the visual encoding model are jointly optimized to minimize the entropy constraint and the regional difference constraint, wherein the expression for the entropy constraint is: For the number of signs, The number of image regions. This represents the association weight between the semantic vector corresponding to the j-th feature and the visual vector corresponding to the i-th image region; The expression for the regional difference constraint is: This represents the regional association weight distribution corresponding to the j-th symptom.

6. The method according to claim 1, characterized in that, Also includes: Obtain the image report to be processed and the corresponding medical images; Using the trained text encoding model, semantic vectors corresponding to the corresponding features are generated based on the structured semantic units corresponding to the features involved in the image report to be processed. Using the trained visual coding model, visual vectors corresponding to multiple image regions of the medical image are generated; Based on the semantic vector corresponding to the corresponding sign and the visual vector corresponding to multiple image regions of the medical image, the association weight between the corresponding sign and multiple image regions of the medical image is generated. A heatmap of lesion signs is generated based on the correlation weights between the corresponding signs and multiple image regions of the medical image.

7. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, a processor performs the method according to any one of claims 1 to 6.

8. A multimodal joint anchoring training device based on structured features, characterized in that, include: The acquisition module is used to acquire image report samples and medical image samples corresponding to the image report samples; The structured unit determination module is used to determine the structured semantic units corresponding to each feature involved in the image report sample, wherein the structured semantic units include at least a feature field and a location field; The semantic encoding module is used to generate a semantic vector corresponding to each feature based on the structured semantic unit using a text encoding model. The visual encoding module is used to divide the medical image sample into multiple image regions and generate a visual vector corresponding to each image region through a visual encoding model. The visual evidence generation module is used to determine the association weight between each visual vector and the semantic vector of the corresponding sign based on the semantic vector corresponding to the corresponding sign and the visual vector corresponding to each image region for each sign, and to generate a visual evidence vector corresponding to the corresponding sign by weighted summation of the visual vectors corresponding to each image region according to the corresponding association weight. The training module is used to determine the joint loss of the text encoding model and the visual encoding model based on the semantic vector and the visual evidence vector, and to jointly train the text encoding model and the visual encoding model based on the joint loss. The trained text encoding model and / or visual encoding model is then used to perform downstream tasks related to medical images. The joint loss includes: joint anchoring loss, multimodal alignment loss, and regularization loss; and wherein... The process of jointly optimizing the text encoding model and the visual encoding model using the joint anchoring loss in the joint loss is as follows: by using a preset shared classifier, a first classification result is generated based on the semantic vector of the corresponding feature, and a second classification result is generated based on the visual evidence vector of the corresponding feature. The text encoding model and the visual encoding model are jointly optimized with the optimization objective of minimizing the difference between the first classification result and the feature field and minimizing the difference between the second classification result and the feature field.

9. A multimodal joint anchoring training device based on structured features, characterized in that, include: processor; as well as A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: Obtain image report samples and corresponding medical image samples; The structured semantic units corresponding to each feature involved in the image report sample are determined, wherein the structured semantic units include at least a feature field and a location field; For each symptom, a semantic vector corresponding to the symptom is generated based on the structured semantic unit using a text encoding model. The medical image sample is divided into multiple image regions, and a visual vector corresponding to each image region is generated through a visual coding model. For each sign, based on the semantic vector corresponding to the corresponding sign and the visual vector corresponding to each image region, the association weight between each visual vector and the semantic vector of the corresponding sign is determined, and the visual vectors corresponding to each image region are weighted and summed according to the corresponding association weights to generate a visual evidence vector corresponding to the corresponding sign. Based on the semantic vector and the visual evidence vector, a joint loss for the text encoding model and the visual encoding model is determined. Based on the joint loss, the text encoding model and the visual encoding model are jointly trained. The trained text encoding model and / or visual encoding model are then used to perform downstream tasks related to medical images. The joint loss includes: joint anchoring loss, multimodal alignment loss, and regularization loss; and wherein... The process of jointly optimizing the text encoding model and the visual encoding model using the joint anchoring loss in the joint loss is as follows: by using a preset shared classifier, a first classification result is generated based on the semantic vector of the corresponding feature, and a second classification result is generated based on the visual evidence vector of the corresponding feature. The text encoding model and the visual encoding model are jointly optimized with the optimization objective of minimizing the difference between the first classification result and the feature field and minimizing the difference between the second classification result and the feature field.

Citation Information

Patent Citations

  • Method and device for generating medical image report

    CN117352121A

  • Anatomical region guided medical vision-language pre-training system

    CN119227831A