Multi-path intersection detection model training method

By cropping and extracting features from medical image sample slices, constructing a topological structure for key feature representation, and using a multi-way cross-detection model for cross-attention processing, the problem of low prediction accuracy of medical image classification detection models in existing technologies is solved, and higher prediction accuracy is achieved.

CN118298421BActive Publication Date: 2025-10-21SOUTH CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410298095.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-10-21
Estimated Expiration
2044-03-15

AI Technical Summary

Technical Problem

Existing technologies ignore the spatial relationship between medical image samples in the training of medical image classification detection models, resulting in low prediction accuracy.

Method used

By cropping medical image sample slices, the feature representations of multiple medical image samples are extracted, and the topological structure of key feature representations is constructed. The multi-way cross detection model is used to perform cross attention processing and obtain the attention weight matrix to improve the prediction accuracy.

Benefits of technology

The prediction accuracy of the medical image classification detection model is improved, and the classification ability of the model is enhanced by considering the spatial relationship of image samples and the probabilistic relationship of feature representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118298421B_ABST
    Figure CN118298421B_ABST
Patent Text Reader

Abstract

The application provides a multi-path cross detection model training method, including: medical image sample slices are cropped to obtain a plurality of medical image samples; the plurality of medical image samples are input to a preset encoder for feature extraction processing to obtain a plurality of first feature representations and a package level feature representation; according to the probability of the first feature representation corresponding to each image category of a preset, key feature representations corresponding to each image category are obtained; local structure information between the key feature representations is obtained; according to the key feature representations, the package level feature representation and the local structure information, a first attention weight matrix and a second attention weight matrix are obtained to combine the package level feature representation and the key feature representation to obtain a prediction result of the medical image sample slice; according to the prediction result and the sample image category, model parameters of a to-be-trained multi-path cross detection model are trained to obtain a multi-path cross detection model used for predicting the image category of the medical image sample slice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image classification detection, and in particular to a multi-path cross-detection model training method. Background Art

[0002] Whole-slide imaging refers to the process of scanning an entire microscope slide and converting it into a digital whole-slide image. However, a single digital whole-slide image can introduce significant redundancy and irrelevant input. Therefore, using a single digital whole-slide image for classification detection model training can compromise training accuracy. Existing techniques segment medical image samples into multiple medical image samples for classification detection model training. However, this technique ignores the spatial relationships between medical image samples, resulting in low prediction accuracy in the trained classification detection models. Summary of the Invention

[0003] The purpose of this application is to overcome the shortcomings and deficiencies in the prior art and to provide a multi-path cross-detection model training method. The trained image classification detection model can improve the prediction accuracy of image classification.

[0004] A first aspect of an embodiment of the present application provides a multi-path cross-detection model training method, comprising:

[0005] Cropping the medical image sample slices to obtain a plurality of medical image samples, wherein the medical image samples correspond to sample image categories;

[0006] Inputting the multiple medical image samples into a preset encoder for feature extraction processing to obtain a first feature representation of each medical image sample and a packet-level feature representation corresponding to the medical image sample slice; the packet-level feature representation is obtained by concatenating multiple first feature representations of the same medical image sample slice;

[0007] According to the probability that the first feature representation corresponds to each preset image category, obtaining the key feature representation corresponding to each image category mined based on the first feature representation;

[0008] Determining each medical image sample corresponding to the key feature representation as a key sample, and constructing a topological structure of the key samples corresponding to the same image category to obtain local structural information between the key feature representations;

[0009] Inputting the key feature representation, the packet-level feature representation, and the local structure information into a multi-way cross detection model to be trained, obtaining a first attention weight matrix for performing cross attention processing on the key feature representation and the packet-level feature representation, and a second attention weight matrix for performing cross attention processing on the packet-level feature representation and the local structure information input;

[0010] Obtaining a prediction result of the medical image sample slice according to the first attention weight matrix, the second attention weight matrix, the packet-level feature representation, and the key feature representation;

[0011] The model parameters of the multi-way cross detection model to be trained are trained according to the prediction result and the sample image category to obtain a multi-way cross detection model for predicting the image category of the medical image sample slice.

[0012] Compared with the related art, the present application performs feature extraction processing on multiple medical image samples obtained by cropping medical image sample slices to obtain a first feature representation of each medical image sample and a package-level feature representation of the medical image sample slices; obtains a key feature representation based on the probability of each image category corresponding to the first feature representation, establishes a topological structure of the key samples corresponding to the key feature representation, and obtains local structural information between the key feature representations; then inputs the key feature representation, the package-level feature representation, and the local structural information into the multi-way cross detection model to be trained to obtain a first attention weight matrix and a second attention weight matrix obtained by cross-attention processing; then obtains a prediction result of the medical image sample slice based on the first attention weight matrix, the second attention weight matrix, the package-level feature representation, and the key feature representation, so as to train the multi-way cross detection model based on the prediction result and the sample image category of the medical image sample. The present application can train the multi-way cross detection model based on the probabilistic relationship between the image category of the medical image sample and the feature representation of the segmented medical image sample, as well as the spatial relationship between the feature representations of the medical image sample, and can improve the prediction accuracy of the image classification output by the multi-way cross detection model trained as an image classification model.

[0013] In order to provide a clearer understanding of the present application, the specific implementation methods of the present application will be described below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a flowchart of a multi-path cross-detection model training method according to an embodiment of the present application.

[0015] Figure 2 This is a flowchart of steps S31-S33 of the multi-path cross detection model training method according to one embodiment of the present application.

[0016] Figure 3 This is a flowchart of steps S61-S64 of the multi-path cross detection model training method according to one embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be described in further detail below with reference to the accompanying drawings.

[0018] It should be clear that the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the embodiments of the present application.

[0019] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. In the description of this application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to the specific circumstances. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates other meanings. The words "if" / "if" used herein can be interpreted as "at the time of" or "when" or "in response to determination".

[0020] In addition, in this application, unless otherwise specified, "plurality" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0021] See also Figure 1 , which is a flowchart of a multi-path cross detection model training method according to an embodiment of the present application, including:

[0022] S1: Crop the medical image sample slices to obtain a plurality of medical image samples, where the medical image samples correspond to sample image categories.

[0023] The medical image sample slice is an image processed by RGB thresholding, which may be an image of a sliced ​​tissue region, and then cut into multiple non-overlapping medical image samples of 512×512 pixels at a 20× magnification.

[0024] The sample image category refers to the category of the tissue corresponding to the medical image sample, for example, the category of the tissue corresponding to the medical image sample is in a normal state, the category of the tissue corresponding to the medical image sample is in an abnormal state, etc.

[0025] S2: Input the multiple medical image samples into a preset encoder for feature extraction processing to obtain a first feature representation of each medical image sample and a packet-level feature representation corresponding to the medical image sample slice; the packet-level feature representation is obtained by splicing multiple first feature representations of the same medical image sample slice.

[0026] The encoder refers to a feature encoder, which can be used to extract feature representations of medical image samples. The encoder can adopt a Resnet50 model.

[0027] S3: According to the probability that the first feature representation corresponds to each preset image category, obtain the key feature representation corresponding to each image category mined based on the first feature representation.

[0028] See also Figure 2 Step S3 can be implemented by a built-in image anchoring module, wherein the built-in image anchoring module may include an index branch and two module branches. Specifically, step S2 includes the step of obtaining key feature representations corresponding to each image category mined based on the first feature representation according to the probability that the first feature representation corresponds to each preset image category, including:

[0029] S31: Inputting each of the first feature representations into a probability index branch to obtain the probability that each of the first feature representations corresponds to each preset image category.

[0030] Among them, the probability index branch module branch includes two fully connected layers, a ReLu activation function and a Softmax function connected in sequence. This module branch can obtain the probability of each first feature representation corresponding to each preset image category.

[0031] S32: Input each of the first feature representations into a feature mining branch to obtain a plurality of second feature representations of the first feature representations.

[0032] The feature mining branch includes two sequentially connected fully connected layers, a normalization layer, and a ReLu activation function. This module branch projects the first feature representation into a relatively compact cluster in the feature space, thereby facilitating the feature mining capabilities of the subsequent attention mechanism. Step S32 and step S31 are parallel steps and have no logical precedence.

[0033] S33: According to the relationship between the probability and the first feature representation, and the relationship between the first feature representation and the second feature representation, determine the second feature representation whose probability is greater than a preset probability threshold as a key feature representation.

[0034] Among them, S3 can use the probability of each first feature representation corresponding to each preset image category as an index from large to small, that is, from the feature representation output by the second module branch, select multiple key feature representations with higher probabilities. The higher probability can be a number of feature representations output by the second module branch corresponding to a preset probability threshold. The probability threshold can be set by the user. In other feasible embodiments, it can also be a number of feature representations output by the second module branch corresponding to a relatively high probability. Among them, in order to facilitate the display of the key feature representation, the second feature representation corresponding to the same first feature representation can be arranged from high to low according to the probability of each first feature representation corresponding to each preset image category, so as to obtain the key feature representation therefrom.

[0035] S4: Determine each medical image sample corresponding to the key feature representation as a key sample, and construct a topological structure of the key samples corresponding to the same image category to obtain local structural information between the key feature representations.

[0036] Among them, the topological structure of key samples can be constructed by determining nodes and using the nearest neighbor algorithm, and the local structural information between key feature representations can be run through the topological structure.

[0037] S5: Input the key feature representation, the packet-level feature representation and the local structure information into the multi-way cross-detection model to be trained, and obtain a first attention weight matrix for cross-attention processing based on the key feature representation and the packet-level feature representation, and a second attention weight matrix for cross-attention processing of the packet-level feature representation and the local structure information input.

[0038] Among them, the multi-way cross detection model is an image classification detection model that performs feature detection based on the first attention weight matrix and the second attention weight matrix obtained based on key feature representation and packet-level feature representation, and packet-level feature representation and local structure information, which can improve the accuracy of image classification prediction.

[0039] S6: Obtain a prediction result of the medical image sample slice according to the first attention weight matrix, the second attention weight matrix, the packet-level feature representation, and the key feature representation.

[0040] Among them, the first attention weight matrix, the second attention weight matrix, the package-level feature representation and the key feature representation can be subjected to feature fusion and feature summation to obtain medical image sample slices and feature representations of medical image sample slices, thereby predicting the category of the medical image sample slices and obtaining prediction results of the medical image sample slices. For example, the prediction results may refer to the category of the tissue corresponding to the predicted medical image sample being in a normal state, the category of the tissue corresponding to the medical image sample being in an abnormal state, and other results.

[0041] S7: Training the model parameters of the to-be-trained multi-way cross detection model according to the prediction result and the sample image category, to obtain a multi-way cross detection model for predicting the image category of the medical image sample slice.

[0042] The training can be guided by a function. For example, a loss function constructed by using prediction results and sample image categories can be used to guide the training of model parameters of the multi-way cross detection model to be trained until the value of the loss function is less than a preset function threshold. The training of model parameters of the multi-way cross detection model to be trained can also be guided by setting the number of training times, for example, 100 times, 200 times, 300 times, etc.

[0043] Compared with the related art, the present application performs feature extraction processing on multiple medical image samples obtained by cropping medical image sample slices to obtain a first feature representation of each medical image sample and a package-level feature representation of the medical image sample slices; obtains a key feature representation based on the probability of each image category corresponding to the first feature representation, establishes a topological structure of the key samples corresponding to the key feature representation, and obtains local structural information between the key feature representations; then inputs the key feature representation, the package-level feature representation, and the local structural information into the multi-way cross detection model to be trained to obtain a first attention weight matrix and a second attention weight matrix obtained by cross-attention processing; then obtains a prediction result of the medical image sample slice based on the first attention weight matrix, the second attention weight matrix, the package-level feature representation, and the key feature representation, so as to train the multi-way cross detection model based on the prediction result and the sample image category of the medical image sample. The present application can train the multi-way cross detection model based on the probabilistic relationship between the image category of the medical image sample and the feature representation of the segmented medical image sample, as well as the spatial relationship between the feature representations of the medical image sample, and can improve the prediction accuracy of the image classification output by the multi-way cross detection model trained as an image classification model.

[0044] In a feasible embodiment, the step of S5: determining each medical image sample corresponding to the key feature representation as a key sample, and constructing a topological structure of the key samples corresponding to the same image category to obtain local structural information between the key feature representations, includes:

[0045] S51: The key samples corresponding to the same image category are used as nodes, the key feature representations are used as node features, and the topological structure is obtained by using a nearest neighbor algorithm and each node feature.

[0046] Specifically, S51 includes:

[0047] S511: Obtain the similarity of node features of two nearest neighbor nodes.

[0048] The two nearest neighbor nodes are the two key samples with the closest physical distance. The physical distance of the key samples can be obtained from the clipping in step S1. That is, the two key samples with no other key samples in the straight-line distance are the two nearest neighbor nodes. The similarity of the node features of the two nearest neighbor nodes can be calculated using cosine similarity.

[0049] S512: If the similarity is greater than or equal to a preset similarity threshold, a connection line is established between the two nodes.

[0050] The connecting line is used to represent the spatial relationship between two nodes. The two nodes connected by the connecting line are two nodes associated with the spatial relationship.

[0051] S513: Obtain the topological structure according to each node and the connection line.

[0052] Based on each node and connection line, the topological structure between each key sample can be generated.

[0053] S52: Input the topological structure into the graph convolution layer to obtain local structural information between the key feature representations.

[0054] The graph convolution layer can be multiple graph convolution layers connected in sequence. After the outputs of each graph convolution layer are spliced, the spliced ​​features are input into the 1×1 convolution layer for feature dimension alignment to obtain the local structural information between the key feature representations.

[0055] See also Figure 3 In a feasible embodiment, the step S6 of obtaining a prediction result of the medical image sample slice according to the first attention weight matrix, the second attention weight matrix, the packet-level feature representation, and the key feature representation includes:

[0056] S61: Perform weighted sum processing on the first attention weight matrix and the second attention weight matrix to obtain a weighted sum feature representation.

[0057] The first attention weight matrix and the second attention weight matrix can be weighted and summed using the same weight coefficient, for example, the weight coefficient is also 0.5.

[0058] S62: Fusing the weighted sum feature representation and the packet-level feature representation to obtain a fused feature representation.

[0059] Among them, the fusion processing refers to the process of multiplying the weighted sum feature representation and the package-level feature representation, and the product result is the fused feature representation.

[0060] S63: summing the fused feature representation and the key feature representation to obtain a feature summation representation.

[0061] Among them, the fusion feature representation and the key feature representation are added together, and the sum obtained is the feature sum representation.

[0062] S64: Performing self-attention processing and prediction processing on the feature sum representation to obtain a prediction result of the medical image sample slice.

[0063] Among them, self-attention processing can effectively capture the medium and long-distance dependencies represented by feature summation, and can easily expand the length of features; and prediction processing can be implemented using prediction functions, such as the softmax function, etc. For example, the features obtained by self-attention processing can be classified and predicted by the softmax function, thereby obtaining the prediction results of medical image sample slices.

[0064] In this embodiment, the first attention weight matrix, the second attention weight matrix, the packet-level feature representation, and the key feature representation can be processed through steps S61-S64 to accurately obtain the prediction results of the medical image sample slices.

[0065] In a feasible embodiment, after the step of training the model parameters of the to-be-trained multi-way cross detection model based on the prediction result and the sample image category to obtain the multi-way cross detection model for predicting the image category of the medical image sample slice, the method further comprises:

[0066] S81: Perform data enhancement processing on the medical image sample slice to obtain a first medical image sample and a second medical image sample.

[0067] The data enhancement processing may include brightness processing, contrast processing, flip processing, symmetry processing, and the like on the medical image sample slices.

[0068] S82: Determine the image category as a pseudo-label category based on the highest value of the probability of each image category represented by the first feature, and obtain the similarities and differences between the image category corresponding to the medical image sample slice and the pseudo-label category.

[0069] The highest probability value of step S82 can be obtained from step S3. Since the image category corresponding to the medical image sample slice is a known image category, the similarities and differences between the image category corresponding to the medical image sample slice and the pseudo label category can be directly obtained.

[0070] S83: Acquire a first encoder and a second encoder having the same structure as the encoder.

[0071] The structures of the first encoder and the second encoder are the same as those of the encoder in step S2, but the parameters of the first encoder and the second encoder are different.

[0072] S84: Input the first medical image sample into the first encoder and the first mapper to obtain first sample features after feature extraction and depth mapping.

[0073] The first encoder is used for feature extraction, and the first mapper is used for performing depth mapping on the output of the first encoder to obtain a first sample feature.

[0074] S85: Input the first medical image sample into the second encoder and the second mapper to obtain second sample features after feature extraction and depth mapping.

[0075] The second encoder is used for feature extraction, and the second mapper is used for performing depth mapping on the output of the second encoder to obtain second sample features.

[0076] Since the parameters of the first encoder and the second encoder are different, the first sample features and the second sample features obtained in steps S84 and S85 are different.

[0077] S86: Divide the second sample features into a first queue and a second queue according to the similarity and difference relationship;

[0078] Among them, the first queue and the second queue respectively accommodate the second sample features based on the difference and similarity relationship. For example, if the difference and similarity relationship is the same, the second sample features can be divided into the first queue; if the difference and similarity relationship is different, the second sample features can be divided into the second queue.

[0079] S87: Train the first encoder according to the first sample features, the first queue, and the second queue to obtain a trained optimized encoder.

[0080] Among them, step S87 includes training the first encoder according to a preset number of training times, using the first sample feature, the first queue and the second queue as inputs of a weakly supervised contrastive learning loss function, and obtaining a trained optimized encoder.

[0081] The training of the first encoder may be performed according to a preset number of training times, that is, when the preset number of training times is reached, an optimized encoder may be obtained.

[0082] The weakly supervised contrastive learning loss function is shown in the following formula:

[0083]

[0084] Among them, L WSCL is the output of the weakly supervised contrastive learning loss function; log() is the logarithmic operation; exp() is the exponential operation; Τ is the transposition operation; S(q i,j ) is the same as the current feature representation q i,j The set of queues with the same pseudo-label, for example, the first queue; D(q i,j ) is the same as the current feature representation q i,j A set of queues with different pseudo labels, for example, the second queue; k + is the current feature representation; q i,j For the same feature representation as the pseudo label, k - is the same as the current feature representation q i,j is the different feature representations of pseudo labels; τ is the temperature coefficient, which is a constant.

[0085] S88: Use the optimized encoder as the encoder to continue training the multi-path cross detection model.

[0086] Step S88 is to use the output of the optimized encoder to participate in the secondary training of the multi-path cross detection model, that is, after updating the encoder, it can be executed from step S2 until the secondary training obtains a multi-path cross detection model that meets the preset training times or is less than the preset loss function threshold.

[0087] In a feasible embodiment, before the step of inputting the first medical image sample into the second encoder and the second mapper to obtain the second sample features after feature extraction and depth mapping, the method includes:

[0088] S8501: Update the parameters of the second encoder according to the first encoding parameters of the first encoder and the second encoding parameters of the second encoder before updating to obtain the updated second encoder.

[0089] Specifically, the updated parameters of the second encoder can be obtained by the following formula:

[0090] E2=m·E′2+(1-m)E1

[0091] Wherein, E2 is the parameter of the second encoder after updating; m is the momentum coefficient, which is a constant; E′2 is the second encoding parameter of the second encoder before updating; and E1 is the first encoding parameter of the first encoder.

[0092] S8502: Update the parameters of the second mapper according to the first mapping parameters of the first mapper and the second mapping parameters of the second mapper before updating, to obtain the updated second mapper.

[0093] Specifically, the updated parameters of the second mapper can be obtained by the following formula:

[0094] P2=m·P′2+(1-m)P1

[0095] Wherein, P2 is the parameter of the second mapper after update; m is the momentum coefficient, which is a constant; P′2 is the second mapping parameter of the second mapper before update; and P1 is the first mapping parameter of the first mapper.

[0096] In this embodiment, the second encoder and the second mapper can be updated according to steps S8501 and S8502, which is beneficial to iterative update of the second encoder and the second mapper, so as to output the second sample features through the updated second encoder and the second mapper.

[0097] The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application. Those of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0098] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0099] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the function selected in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 function selected in a box or multiple boxes.

[0100] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 steps for the function selected in a box or multiple boxes.

[0101] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0102] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0103] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0104] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0105] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A multi-path cross-detection model training method, characterized in that: include: Cropping the medical image sample slices to obtain a plurality of medical image samples, wherein the medical image samples correspond to sample image categories; Inputting the multiple medical image samples into a preset encoder for feature extraction processing to obtain a first feature representation of each medical image sample and a packet-level feature representation corresponding to the medical image sample slice; the packet-level feature representation is obtained by concatenating multiple first feature representations of the same medical image sample slice; According to the probability that the first feature representation corresponds to each preset image category, obtaining the key feature representation corresponding to each image category mined based on the first feature representation; Determining each medical image sample corresponding to the key feature representation as a key sample, and constructing a topological structure of the key samples corresponding to the same image category to obtain local structural information between the key feature representations; Inputting the key feature representation, the packet-level feature representation, and the local structure information into a multi-way cross detection model to be trained, obtaining a first attention weight matrix for performing cross attention processing on the key feature representation and the packet-level feature representation, and a second attention weight matrix for performing cross attention processing on the packet-level feature representation and the local structure information input; Obtaining a prediction result of the medical image sample slice according to the first attention weight matrix, the second attention weight matrix, the packet-level feature representation, and the key feature representation; The model parameters of the multi-way cross detection model to be trained are trained according to the prediction result and the sample image category to obtain a multi-way cross detection model for predicting the image category of the medical image sample slice.

2. The multi-path cross detection model training method according to claim 1, characterized in that: The step of obtaining, based on the probability that the first feature representation corresponds to each preset image category, a key feature representation corresponding to each image category mined based on the first feature representation includes: Inputting each of the first feature representations into a probability index branch to obtain a probability that each of the first feature representations corresponds to each preset image category; Inputting each of the first feature representations into a feature mining branch to obtain a plurality of second feature representations of the first feature representations; According to the relationship between the probability and the first feature representation, and the relationship between the first feature representation and the second feature representation, the second feature representation having a probability greater than a preset probability threshold is determined as the key feature representation.

3. The multi-path cross detection model training method according to claim 1, characterized in that: The step of determining each medical image sample corresponding to the key feature representation as a key sample, and constructing a topological structure of the key samples corresponding to the same image category to obtain local structural information between the key feature representations includes: The key samples corresponding to the same image category are used as nodes, the key feature representations are used as node features, and the topological structure is obtained by using a nearest neighbor algorithm and each node feature; The topological structure is input into the graph convolution layer to obtain the local structural information between the key feature representations.

4. The multi-path cross detection model training method according to claim 3, characterized in that: The step of obtaining the topological structure by using the nearest neighbor algorithm and the characteristics of each node includes: Get the similarity of node features of the two nearest neighbor nodes; If the similarity is greater than or equal to a preset similarity threshold, establishing a connection line between the two nodes; The topological structure is obtained according to each node and the connection line.

5. The multi-path cross detection model training method according to claim 1, characterized in that: The step of obtaining a prediction result of the medical image sample slice according to the first attention weight matrix, the second attention weight matrix, the packet-level feature representation, and the key feature representation includes: Performing weighted sum processing on the first attention weight matrix and the second attention weight matrix to obtain a weighted sum feature representation; Fusing the weighted sum feature representation and the packet-level feature representation to obtain a fused feature representation; Summing the fused feature representation and the key feature representation to obtain a feature summation representation; Self-attention processing and prediction processing are performed on the feature sum representation to obtain a prediction result of the medical image sample slice.

6. The multi-path cross detection model training method according to any one of claims 1 to 5, characterized in that: After the step of training the model parameters of the to-be-trained multi-way cross detection model according to the prediction result and the sample image category to obtain the multi-way cross detection model for predicting the image category of the medical image sample slice, the method includes: performing data enhancement processing on the medical image sample slices to obtain a first medical image sample and a second medical image sample; Determining the image category as a pseudo-label category based on the highest value of the probability of each image category represented by the first feature, and obtaining a similarity and difference relationship between the image category corresponding to the medical image sample slice and the pseudo-label category; Acquire a first encoder and a second encoder having the same structure as the encoder; Inputting the first medical image sample into the first encoder and the first mapper to obtain first sample features after feature extraction and depth mapping; Inputting the first medical image sample into the second encoder and the second mapper to obtain second sample features after feature extraction and depth mapping; Dividing the second sample features into a first cohort and a second cohort according to the similarity and difference relationship; training the first encoder according to the first sample feature, the first queue, and the second queue to obtain a trained optimized encoder; The optimized encoder is used as the encoder to continue training the multi-way cross detection model.

7. The multi-path cross detection model training method according to claim 6, characterized in that: Before the step of inputting the first medical image sample into the second encoder and the second mapper to obtain the second sample features after feature extraction and depth mapping, the method includes: Updating parameters of the second encoder according to the first encoding parameters of the first encoder and the second encoding parameters of the second encoder before updating to obtain an updated second encoder; The parameters of the second mapper are updated according to the first mapping parameters of the first mapper and the second mapping parameters of the second mapper before updating, to obtain the updated second mapper.

8. The multi-path cross detection model training method according to claim 7, characterized in that: The step of updating the parameters of the second encoder according to the first encoding parameters of the first encoder and the second encoding parameters of the second encoder before updating to obtain the updated second encoder includes: The updated parameters of the second encoder are obtained by the following formula: E2=m·E′2+(1-m)E1 Wherein, E2 is the parameter of the second encoder after updating; m is the momentum coefficient, which is a constant; E′2 is the second encoding parameter of the second encoder before updating; and E1 is the first encoding parameter of the first encoder.

9. The multi-path cross detection model training method according to claim 7, characterized in that: The step of updating the parameters of the second mapper according to the first mapping parameters of the first mapper and the second mapping parameters of the second mapper before updating to obtain the updated second mapper includes: The updated parameters of the second mapper are obtained by the following formula: P2=m·P′2+(1-m)P1 Wherein, P2 is the parameter of the second mapper after the update; m is the momentum coefficient, which is a constant; P′2 is the second mapping parameter of the second mapper before the update; P i is the first mapping parameter of the first mapper.

10. The multi-path cross detection model training method according to claim 6, characterized in that: The step of training the first encoder according to the first sample feature, the first queue, and the second queue to obtain a trained optimized encoder includes: According to a preset number of training times, the first sample feature, the first queue, and the second queue are used as inputs of a weakly supervised contrastive learning loss function to train the first encoder and obtain a trained optimized encoder.

Citation Information

Patent Citations

  • Primary tumor staging multi-example learning method in pathological image based on hierarchical graph, framework, equipment and medium

    CN116012332A

  • Fine granularity classification method based on multi-granularity interaction and feature recombination network

    CN116883748A