Tumor tissue metastasis risk identification method and device

Through deep learning processing of video frames and lymph node image features in gastric cancer intraoperative surgery, combined with preset distribution characteristics, real-time accurate prediction of lymph node metastasis risks are achieved, and the accuracy and reliability of lymph node metastasis prediction in the prior art are solved, and surgical efficiency and safety are improved.

CN120279331APending Publication Date: 2025-07-08INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432125.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has low accuracy and poor reliability in predicting lymph node metastasis in gastric cancer. It is difficult to accurately determine whether lymph nodes in each group are metastasized, resulting in the possibility of missing lesional lymph nodes during the operation or excessive non-lesional lymph nodes being removed, increasing the risk of surgery.

Method used

By acquiring intraoperative video frames and lymph node image features, using recognition models for feature extraction and fusion, and combining preset preoperative and intraoperative distribution features, real-time accurate prediction of lymph node metastasis risks, including feature filtering and correction, using deep learning and generative adversarial networks and other technologies.

Benefits of technology

Real-time and accurate prediction of lymph node metastasis in various groups of gastric cancer patients has been achieved, assisting doctors to efficiently remove lesion lymph nodes, reduce non-lesion lymph node resection, and improve medical efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279331A_ABST
    Figure CN120279331A_ABST
Patent Text Reader

Abstract

The invention provides a tumor tissue metastasis risk identification method and device, which can be applied to the technical field of image processing. The method comprises the following steps: acquiring an intra-operative video frame representing surgical operation for tumor tissue of a target object and lymph node image features related to a target area of the target object, wherein the lymph node image features represent lymph node attributes before the surgical operation is performed on the target object; feature extraction is conducted on the intra-operative video frame, target image features corresponding to the target area are obtained, and the target image features represent the lymph node attributes of the target object in the surgical operation process; and processing the lymph node image features corresponding to the target area and the target image features by using an identification model to obtain an identification result, the identification result representing a risk level of metastasis of the tumor tissue to the target area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and particularly to a method and device for identifying the risk of tumor tissue metastasis. Background Art

[0002] Gastric cancer is one of the most common malignant tumors, with relatively high incidence and mortality rates globally. Perigastric lymph nodes are very prone to metastasis in gastric cancer. If the metastasized lymph nodes are not completely resected during the operation, it will cause the recurrence of gastric cancer and lead to the death of the patient. Therefore, accurately predicting whether tumor tissue metastasis has occurred in the lymph node region has a crucial impact on the prognosis of gastric cancer patients.

[0003] However, due to the complexity of the vascular anatomy of the stomach and the surrounding lymphatic drainage, the existing identification methods have low accuracy and poor reliability in real-time metastasis prediction. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides a method and device for identifying the risk of tumor tissue metastasis.

[0005] According to a first aspect of the present disclosure, there is provided a method for identifying the risk of tumor tissue metastasis, including: obtaining intraoperative video frames representing surgical operations on tumor tissue of a target object, and lymph node image features related to a target region of the target object, where the lymph node image features represent the lymph node attributes before performing the surgical operation on the target object; extracting features from the intraoperative video frames to obtain target image features corresponding to the target region, where the target image features represent the lymph node attributes of the target object during the surgical operation; using an identification model to process the lymph node image features and the target image features corresponding to the target region to obtain an identification result, where the identification result represents the risk level of tumor tissue metastasis to the target region.

[0006] According to an embodiment of the present disclosure, extracting features from the intraoperative video frames to obtain target image features corresponding to the target region includes: extracting initial image features from the intraoperative video frames; filtering the initial image features based on the mutual information between a preset preoperative distribution feature and the initial image features to obtain intraoperative filtered features, where the preset preoperative distribution feature represents a lymph node image feature distribution library related to the target region constructed based on sample objects before performing the surgical operation; correcting the intraoperative filtered features based on a preset intraoperative distribution feature to obtain target image features, where the preset intraoperative distribution feature represents a target image feature distribution library related to the target region constructed based on sample objects during the surgical operation.

[0007] According to an embodiment of the present disclosure, the initial image features are obtained by performing feature extraction based on a first feature extraction model, and the first feature extraction model includes a feature extraction layer and a long short-term memory network layer; wherein, performing feature extraction on the intraoperative video frames to obtain the initial image features includes: processing the intraoperative video frames by using the feature extraction layer to obtain an image fusion feature corresponding to the target region; processing the image fusion feature corresponding to the target region by using the long short-term memory network layer to obtain the initial image features corresponding to the target region.

[0008] According to an embodiment of the present disclosure, the preset preoperative distribution feature and the preset intraoperative distribution feature are determined based on the following operations: performing feature extraction on the sample preoperative tissue section image of the sample object to obtain the sample lymph node image features; performing feature extraction on the sample intraoperative video frames of the sample object to obtain the sample target image features; processing the sample lymph node image features and the target image features by using a conditional generative adversarial network to obtain the sample preoperative feature distribution corresponding to the sample lymph node image features and the sample intraoperative feature distribution corresponding to the sample target image features; respectively performing filtering processing on the sample preoperative feature distribution and the sample intraoperative feature distribution based on the mutual information between the sample preoperative feature distribution and the sample intraoperative feature distribution to obtain the sample preoperative corrected feature distribution and the sample intraoperative corrected feature distribution; respectively modeling the sample preoperative corrected feature distribution and the sample intraoperative corrected feature distribution based on the kernel density estimation algorithm to obtain the preset preoperative distribution feature and the preset intraoperative distribution feature.

[0009] According to an embodiment of the present disclosure, performing feature extraction on the sample intraoperative video frames of the sample object to obtain the sample target image features includes: processing the sample intraoperative video frames by using a convolutional neural network to obtain the sample lymph node static features; processing the sample intraoperative video frames by using an image coding network to obtain the sample lymph node dynamic features; obtaining the sample target image features according to the sample lymph node static features and the sample lymph node dynamic features.

[0010] According to an embodiment of the present disclosure, the lymph node image features related to the target region are obtained based on the following operations: performing feature extraction on the preoperative tissue section image of the target object to obtain the initial lymph node features; filtering the initial lymph node features based on the mutual information between the preset intraoperative distribution feature and the initial lymph node features to obtain the preoperative filtered features; correcting the preoperative filtered features based on the preset preoperative distribution feature to obtain the intermediate lymph node features; performing fusion processing on the intermediate lymph node features and the preoperative tissue section image to obtain the lymph node image features related to the target region.

[0011] According to an embodiment of the present disclosure, the lymph node image features are obtained by performing fusion processing based on a second feature extraction model. The second feature extraction model includes a multi-head attention layer and a weighted fusion layer, and the target region includes a lymph node region. Among them, performing fusion processing on the intermediate lymph node features and the preoperative tissue section image to obtain lymph node image features related to the target region includes: determining the lymph node region, the gastric cancer primary focus region, and the visceral fat region with a mapping relationship from the preoperative tissue section image of the target object; using the multi-head attention layer to respectively extract features from the gastric cancer primary focus region and the visceral fat region that have a mapping relationship with the lymph node region, to obtain the gastric cancer primary focus region features and the visceral fat region features; using the weighted fusion layer to process the intermediate lymph node features, the gastric cancer primary focus region features, and the visceral fat region features to obtain the lymph node image features.

[0012] According to an embodiment of the present disclosure, the second feature extraction model is trained based on the following operations: obtaining training samples, where the training samples include the sample preoperative tissue section image and the sample label of the sample object. The sample preoperative tissue section image includes the target region, the gastric cancer primary focus region, and the visceral fat region, and the sample label represents that the tumor tissue has metastasized to the target region or the tumor tissue has not metastasized to the target region; using the second feature extraction model to extract features from the sample preoperative tissue section image to obtain sample lymph node image features related to the target region; identifying the sample lymph node image features to obtain a sample recognition result; using a loss function to process the sample recognition result, the sample label, the area of the visceral fat region, and the distance between the gastric cancer primary focus region and the lymph node region to obtain a loss value; training the second feature extraction model according to the loss value to obtain the trained second feature extraction model.

[0013] According to an embodiment of the present disclosure, the recognition model includes a cross-attention layer and a decoding layer. Among them, using the recognition model to process the lymph node image features corresponding to the target region and the target image features to obtain the recognition result includes: performing an adaptive convolution operation on the lymph node image features corresponding to the target region to obtain preoperative key features; performing an adaptive convolution operation on the target image features corresponding to the target region to obtain intraoperative key features; using the cross-attention layer to process the preoperative key features and the intraoperative key features to obtain an initial fusion feature; splicing the initial fusion feature, the preoperative key features, the intraoperative key features, the lymph node image features, and the target image features to obtain a target fusion feature; using the decoding layer to process the target fusion feature to obtain the recognition result corresponding to the target region.

[0014] According to a second aspect of the present disclosure, there is provided a device for identifying the risk of tumor tissue metastasis, including: an acquisition module, configured to acquire intraoperative video frames representing surgical operations on tumor tissue of a target object, and lymph node image features related to a target region of the target object, where the lymph node image features represent the lymph node attributes before performing the surgical operation on the target object; an extraction module, configured to extract features from the intraoperative video frames to obtain target image features corresponding to the target region, where the target image features represent the lymph node attributes of the target object during the surgical operation; and an identification module, configured to process the lymph node image features and the target image features corresponding to the target region by using an identification model to obtain an identification result, where the identification result represents the risk level of tumor tissue metastasis to the target region.

[0015] According to the method and device for identifying the risk of tumor tissue metastasis provided by the present disclosure, by acquiring intraoperative video frames representing surgical operations on tumor tissue of a target object, and lymph node image features related to a target region of the target object; extracting features from the intraoperative video frames to obtain target image features corresponding to the target region; and processing the lymph node image features and the target image features corresponding to the target region by using an identification model to obtain an identification result. Since static lymph node image features highly related to the tumor tissue metastasis state of each group of lymph nodes have been extracted before performing the surgical operation on the target object, and then based on the dynamic target image features highly related to the tumor tissue metastasis state of each group of lymph nodes during the surgical operation obtained in real time, the identification model can fully consider the complementarity and fusion between before surgery and during surgery, so as to realize real-time and accurate prediction of whether tumor tissue metastasis occurs in the lymph node region where each group of lymph nodes is located during the operation, and can assist the doctor in excising the metastasized lymph node region during the operation, improving the medical efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become clearer.

[0017] Figure 1 The flowchart of the method for identifying the risk of tumor tissue metastasis according to an embodiment of the present disclosure is shown.

[0018] Figure 2 The schematic diagram of an example of training the first feature extraction model according to an embodiment of the present disclosure is shown.

[0019] Figure 3 The schematic diagram of an example of the second feature extraction model according to an embodiment of the present disclosure is shown.

[0020] Figure 4 The schematic diagram of an example of the identification model according to an embodiment of the present disclosure is shown.

[0021] Figure 5 Shows an exemplary schematic diagram of obtaining an identification result according to an embodiment of the present disclosure.

[0022] Figure 6 Shows a structural block diagram of a tumor tissue metastasis risk identification device according to an embodiment of the present disclosure. Detailed implementation manners

[0023] Hereinafter, embodiments according to the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.

[0024] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0026] In cases where expressions such as "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0027] In the process of implementing the present disclosure, it is found that there are up to a dozen groups of perigastric lymph nodes, with the number ranging from dozens to hundreds. Against the background of the complex perigastric structure, the accuracy of manual detection is significantly limited, and the operation is time-consuming and laborious, prone to misjudgment, and requires high experience of diagnostic physicians. At the same time, due to problems such as insufficient specificity of these signs, high dependence on doctors' experience, and poor repeatability, the accuracy of judging lymph node metastasis by CT images is only 50% - 70%. Currently, in the research work on preoperative CT lymph node metastasis prediction for gastric cancer at home and abroad, most use radiomics and deep learning models to indirectly predict lymph node metastasis based on the primary focus of gastric cancer, with an accuracy rate of 70% - 75%, and mainly focus on judging whether patients have lymph node metastasis, lacking research on judging the metastasis situation of each group of lymph nodes in gastric cancer patients. Therefore, how to accurately judge whether each group of lymph nodes in gastric cancer patients has metastasized, so as to assist doctors in efficiently and accurately removing as many diseased lymph nodes as possible while reducing the removal of non-diseased lymph nodes and reducing surgery-related complications is an urgent problem to be solved.

[0028] In view of this, according to an embodiment of the present disclosure, a method and device for identifying the risk of tumor tissue metastasis are provided. The method includes: obtaining intraoperative video frames representing surgical operations on the tumor tissue of a target object, and lymph node image features related to the target area of the target object, where the lymph node image features represent the lymph node attributes before performing the surgical operation on the target object; extracting features from the intraoperative video frames to obtain target image features corresponding to the target area, where the target image features represent the lymph node attributes of the target object during the surgical operation; using an identification model to process the lymph node image features and the target image features corresponding to the target area to obtain an identification result, where the identification result represents the risk level of tumor tissue metastasis to the target area.

[0029] In the technical solution of the present disclosure, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. And the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0030] It should be noted that the sequence numbers of each operation in the following method are only used for representing the operation for description, and should not be regarded as indicating the execution sequence of each operation. Unless explicitly stated, the method does not need to be executed exactly in the order shown.

[0031] Figure 1The figure shows a flowchart of a method for identifying the risk of tumor tissue metastasis according to an embodiment of the present disclosure.

[0032] As Figure 1 shown, the method 100 for identifying the risk of tumor tissue metastasis includes operations S110 to S130.

[0033] In operation S110, intraoperative video frames representing surgical operations on tumor tissue of a target object are obtained, as well as lymph node image features related to the target area of the target object.

[0034] In operation S120, feature extraction is performed on the intraoperative video frames to obtain target image features corresponding to the target area.

[0035] In operation S130, the recognition model processes the lymph node image features and target image features corresponding to the target area to obtain a recognition result.

[0036] According to an embodiment of the present disclosure, the target object is an object to be identified with tumor tissue in the gastric region, and the intraoperative video frames are real-time image data of each group of lymph node regions around the stomach obtained during the operation of cleaning the tumor tissue in the gastric region.

[0037] According to an embodiment of the present disclosure, the target area can be one lymph node region or multiple lymph node regions around the stomach of the target object.

[0038] According to an embodiment of the present disclosure, before performing a surgical operation on the target object, CT image data of the target object is first obtained, and lymph node image features related to the target area are extracted from the CT image data. The lymph node image features represent the static lymph node attributes before performing the surgical operation on the target object.

[0039] According to an embodiment of the present disclosure, preprocessing operations such as noise removal and image enhancement are performed on the real-time obtained intraoperative video frames.

[0040] According to an embodiment of the present disclosure, a convolutional neural network or the like can be used to perform feature extraction on the preprocessed intraoperative video frames to obtain target image features corresponding to the target area.

[0041] According to an embodiment of the present disclosure, the target image features represent the dynamic lymph node attributes of the target object during the surgical operation. The target image features include lymph node characteristic information, spatio-temporal information, and temporal information of the target area.

[0042] According to an embodiment of the present disclosure, the recognition model can be constructed based on a cross-attention mechanism.

[0043] According to an embodiment of the present disclosure, the lymph node image features and target image features corresponding to the target area are input into the recognition model together, and a recognition result corresponding to the target area is output.

[0044] According to an embodiment of the present disclosure, the recognition result characterizes the risk level of tumor tissue metastasis to the target area. For example, a recognition result of 1 represents that the tumor tissue has metastasized to the target area, and a recognition result of 0 represents that the tumor tissue has not metastasized to the target area.

[0045] According to an embodiment of the present disclosure, there are usually 16 groups of lymph nodes distributed in 4 stations around the stomach, and one or more lymph nodes in each group gather together to form the lymph node area where this group of lymph nodes is located. When the target area is preset as a single lymph node area around the stomach of the target object, the recognition model is used to process the lymph node image features and target image features corresponding to each lymph node area in the target area to obtain the recognition result corresponding to each lymph node area.

[0046] According to an embodiment of the present disclosure, the recognition result corresponding to the target area includes the recognition results corresponding to the lymph node areas where each group of lymph nodes is located.

[0047] According to an embodiment of the present disclosure, since the static lymph node image features highly related to the tumor tissue metastasis state of each group of lymph nodes have been extracted before performing the surgical operation on the target object, and then based on the dynamic target image features highly related to the tumor tissue metastasis state of each group of lymph nodes during the surgical operation obtained in real time, the recognition model can fully consider the complementarity and fusion between before surgery and during surgery, so as to realize the real-time and accurate prediction of whether tumor tissue metastasis occurs in the lymph node area where each group of lymph nodes is located during the operation, which can assist the doctor to excise the metastasized lymph node area during the operation and improve the medical efficiency.

[0048] According to an embodiment of the present disclosure, feature extraction is performed on the intraoperative video frame to obtain the target image features corresponding to the target area, including: feature extraction is performed on the intraoperative video frame to obtain initial image features; based on the mutual information between the preset preoperative distribution features and the initial image features, the initial image features are filtered to obtain intraoperative filtered features, where the preset preoperative distribution features characterize the lymph node image feature distribution library related to the target area constructed based on the sample object before the surgical operation; the intraoperative filtered features are corrected based on the preset intraoperative distribution features to obtain the target image features, where the preset intraoperative distribution features characterize the target image feature distribution library related to the target area constructed based on the sample object during the surgical operation.

[0049] According to an embodiment of the present disclosure, a deep learning model can be used to perform feature extraction on the intraoperative video frame to obtain initial image features, and the initial image features characterize the lymph node characteristic information, spatio-temporal information, and temporal sequence information initially extracted during the surgical operation.

[0050] According to an embodiment of the present disclosure, the sample object is an object with tumor cells in the gastric region, and the preset preoperative distribution feature is a standardized lymph node image feature distribution library related to the target region pre-constructed for the sample object before the surgical operation. The lymph node image feature distribution library includes multiple lymph node image features.

[0051] According to an embodiment of the present disclosure, the intraoperative distribution feature is a standardized target image feature distribution library related to the target region constructed for the sample object during the surgical operation. The target image feature distribution library includes multiple target image features.

[0052] According to an embodiment of the present disclosure, the mutual information between the preset preoperative distribution feature and the initial image feature can be calculated using the kernel density estimation algorithm. The mutual information can measure the shared information amount between the same variables in the preset preoperative distribution feature and the initial image feature. The larger the mutual information, the more redundant information is contained. The variables in the initial image feature with mutual information higher than the first threshold are filtered to obtain the intraoperative filtered feature. The intraoperative filtered feature represents the feature after removing redundancy from the preset preoperative distribution feature.

[0053] According to an embodiment of the present disclosure, the maximum mean discrepancy algorithm can be used to process the preset intraoperative distribution feature and the intraoperative filtered feature, calculate the distribution difference between the variables in the preset intraoperative distribution feature and the intraoperative filtered feature, and align and correct the data of the variables with the difference value higher than the second threshold to obtain the target image feature.

[0054] According to an embodiment of the present disclosure, aligning and correcting the data of the variables with the difference value higher than the second threshold may include: for this variable, standardizing or normalizing the data of the variable in the intraoperative filtered feature using the data mean or standard deviation of the variable in the preset intraoperative distribution feature for alignment and correction.

[0055] According to an embodiment of the present disclosure, the target image feature is the feature after feature distribution alignment and correction with the preset intraoperative distribution feature.

[0056] According to an embodiment of the present disclosure, by adopting the pre-constructed standardized preset preoperative distribution feature and preset intraoperative distribution feature, feature redundancy removal and feature distribution alignment and correction are performed on the extracted initial image feature, which can not only enable the initial image feature to retain features with strong complementarity and low redundancy between the preoperative modalities, but also significantly reduce the domain shift caused by data source and imaging differences, provide a more accurate and efficient feature representation for the recognition model, and ultimately improve the lymph node metastasis prediction performance.

[0057] According to an embodiment of the present disclosure, the initial image features are obtained by performing feature extraction based on a first feature extraction model, and the first feature extraction model includes a feature extraction layer and a long short-term memory network layer; wherein, performing feature extraction on intraoperative video frames to obtain the initial image features includes: processing the intraoperative video frames by using the feature extraction layer to obtain an image fusion feature corresponding to a target region; and processing the image fusion feature corresponding to the target region by using the long short-term memory network layer to obtain an initial image feature corresponding to the target region.

[0058] According to an embodiment of the present disclosure, the feature extraction layer may be constructed based on a self-supervised network (Distillation with No Labels, DINO), and the long short-term memory network layer may be constructed based on Long Short-Term Memory (LSTM).

[0059] According to an embodiment of the present disclosure, inputting the intraoperative video frames into the feature extraction layer to obtain an image fusion feature corresponding to the target region, and the image fusion feature is a fusion of the local detail information and the global semantic information of the target region in the intraoperative video frames.

[0060] According to an embodiment of the present disclosure, inputting the image fusion feature into the long short-term memory network layer to obtain an initial image feature corresponding to the target region.

[0061] According to an embodiment of the present disclosure, compared with the image fusion feature, the initial image feature represents a richer temporal dependence relationship.

[0062] According to an embodiment of the present disclosure, during the training process of the first feature extraction model, an intraoperative video frame sample set of 2000 gastric cancer patients is obtained, the intraoperative video frame sample set is preprocessed, one frame is sampled every 5 frames, and it is divided into a pre-training data set and a downstream fine-tuning data set according to a set ratio. The target regions in the downstream fine-tuning data set are manually labeled, and each target region aggregates a single group of lymph nodes. Each group of lymph nodes is separately outlined. The radiologists manually outline the target regions where each group of lymph nodes is located in the downstream fine-tuning data set, and obtain the sample class labels of whether the tumor tissue has metastasized to the target region.

[0063] Figure 2 Shows an example schematic diagram of training the first feature extraction model according to an embodiment of the present disclosure.

[0064] As Figure 2As shown, the first feature extraction model includes a feature extraction layer 210, a long short-term memory network layer 220, a fully connected layer 230, and an activation layer 240. Among them, the fully connected layer and the activation layer are used for classification output during the training process. Train the first feature extraction model: The target regions are not labeled in the pre-training dataset. The feature extraction layer 210 is pre-trained using the pre-training dataset, enabling the feature extraction layer 210 to self-supervised learn preliminary tissue features; then the downstream fine-tuning dataset is input into the pre-trained feature extraction layer 210 to obtain image fusion features corresponding to the target regions; the image fusion features are input into the long short-term memory network layer 220 to obtain initial image features; the initial image features are input into the fully connected layer 230 to obtain output features, and the output features are input into the activation layer 240 to obtain classification output results. The classification output results predict whether the tumor tissue has metastasized to the target regions where each group of lymph nodes is located. 10-fold cross-validation can be used to evaluate the first feature extraction model. Under the condition that the cross-entropy loss value between the classification output results and the sample category labels does not meet the training end condition, the parameters of the first feature extraction model are fine-tuned for iterative training until the cross-entropy loss value meets the training end condition, and a trained first feature extraction model is obtained.

[0065] According to an embodiment of the present disclosure, the preset preoperative distribution feature and the preset intraoperative distribution feature are determined based on the following operations: Feature extraction is performed on the sample preoperative tissue section image of the sample object to obtain sample lymph node image features; feature extraction is performed on the sample intraoperative video frames of the sample object to obtain sample target image features; a conditional generative adversarial network is used to process the sample lymph node image features and the target image features to obtain a sample preoperative feature distribution corresponding to the sample lymph node image features and a sample intraoperative feature distribution corresponding to the sample target image features; based on the mutual information between the sample preoperative feature distribution and the sample intraoperative feature distribution, the sample preoperative feature distribution and the sample intraoperative feature distribution are respectively filtered to obtain a sample preoperative corrected feature distribution and a sample intraoperative corrected feature distribution; based on the kernel density estimation algorithm, the sample preoperative corrected feature distribution and the sample intraoperative corrected feature distribution are respectively modeled to obtain the preset preoperative distribution feature and the preset intraoperative distribution feature.

[0066] According to an embodiment of the present disclosure, for the sample preoperative tissue section images (CT image data) obtained before performing surgical operations on ten thousand sample objects, a segmentation software is used to outline the target regions where each group of lymph nodes is located and perform manual verification and correction to obtain corrected sample preoperative tissue section images.

[0067] According to an embodiment of the present disclosure, the sample intraoperative video frames for cleaning each group of lymph node regions of the sample object are screened, and one frame is sampled from every 15 frames of the sample intraoperative video frames to obtain sample intraoperative video frames. A segmentation software is used to outline each group of lymph nodes and perform manual verification and correction to obtain corrected sample intraoperative video frames.

[0068] According to an embodiment of the present disclosure, a convolutional neural network and a Vision Transformer can be used to extract features from the corrected preoperative tissue section image of the sample object and the corrected intraoperative video frame of the sample object respectively, so as to obtain the sample lymph node image features and the sample target image features.

[0069] According to an embodiment of the present disclosure, the sample lymph node image features are the lymph node attributes related to the target area before performing a surgical operation on the sample object, and the sample target image features are the lymph node attributes of the target area during the surgical operation on the sample object.

[0070] According to an embodiment of the present disclosure, taking the sample lymph node image features and the target image features as conditional information, a conditional generative adversarial network is used to process the sample lymph node image features and the target image features, so as to obtain a continuous sample preoperative feature distribution consistent with the distribution of the sample lymph node image features and a continuous sample intraoperative feature distribution consistent with the distribution of the sample target image features.

[0071] According to an embodiment of the present disclosure, a sample estimation algorithm can be used to calculate the mutual information between the sample preoperative feature distribution and the sample intraoperative feature distribution, and high-redundancy information filtering processing is respectively performed on the variables with high mutual information values in the sample preoperative feature distribution and the sample intraoperative feature distribution, so as to obtain the sample preoperative corrected feature distribution and the sample intraoperative corrected feature distribution.

[0072] According to an embodiment of the present disclosure, the sample preoperative corrected feature distribution is the sample preoperative features after removing redundant features. The sample intraoperative corrected feature distribution is the sample intraoperative features after removing redundant features.

[0073] According to an embodiment of the present disclosure, the sample preoperative corrected feature distribution is modeled based on the kernel density estimation algorithm, distribution parameters are extracted, and a preset preoperative distribution feature is formed based on the distribution parameters.

[0074] According to an embodiment of the present disclosure, the sample intraoperative corrected feature distribution is modeled based on the kernel density estimation algorithm, distribution parameters are extracted, and a preset intraoperative distribution feature is formed based on the distribution parameters.

[0075] According to an embodiment of the present disclosure, the characteristic distributions of each group of lymph nodes in the preset preoperative distribution characteristics can be directly established and completed in the preoperative stage, including the information of all preoperative lymph node groups. The characteristic distributions of each group of lymph nodes in the preset intraoperative distribution characteristics are established based on the intraoperative video frames of the samples in the regions of each group of lymph nodes dissected around the stomach during the surgical procedure of the sample object, including the information of all intraoperative lymph node groups. The preset preoperative distribution characteristics and the preset intraoperative distribution characteristics are constructed based on ten thousand case sample data and can provide a stable prior distribution. When the number of samples in the current task is limited or the data distribution fluctuates greatly, through distribution alignment and redundancy removal, the reliability, comparability, and transferability of feature expression can be improved to assist in constructing a more discriminative feature space.

[0076] According to an embodiment of the present disclosure, feature extraction is performed on the intraoperative video frames of the sample object to obtain sample target image features, including: processing the intraoperative video frames of the sample using a convolutional neural network to obtain sample lymph node static features; processing the intraoperative video frames of the sample using an image encoding network to obtain sample lymph node dynamic features; and obtaining sample target image features based on the sample lymph node static features and the sample lymph node dynamic features.

[0077] According to an embodiment of the present disclosure, the convolutional neural network is a Convolutional Neural Network (CNN).

[0078] According to an embodiment of the present disclosure, the sample lymph node static features represent the visual information in the intraoperative video frames of the sample, for example, image color information, tissue texture information, lymph node shape information, etc.

[0079] According to an embodiment of the present disclosure, the image encoding network is constructed based on a TimeSformer.

[0080] According to an embodiment of the present disclosure, the sample lymph node dynamic features represent the dynamic transformation information in the intraoperative video frames of the sample, for example, motion information, temporal change information, etc.

[0081] According to an embodiment of the present disclosure, splicing processing is performed on the sample lymph node static features and the sample lymph node dynamic features to obtain sample target image features.

[0082] According to an embodiment of the present disclosure, the sample lymph node static features and the sample lymph node dynamic features are extracted through a convolutional neural network and an image encoding network, so as to more effectively analyze and learn the lymph node attribute information in the intraoperative video frames of the sample.

[0083] According to embodiments of the present disclosure, the lymph node image features related to the target region are obtained based on the following operations: extracting features from the preoperative tissue section image of the target object to obtain initial lymph node features; filtering the initial lymph node features based on the mutual information between the preset intraoperative distribution features and the initial lymph node features to obtain preoperative filtered features; correcting the preoperative filtered features based on the preset preoperative distribution features to obtain intermediate lymph node features; and performing a fusion process on the intermediate lymph node features and the preoperative tissue section image to obtain the lymph node image features related to the target region.

[0084] According to embodiments of the present disclosure, the preoperative tissue section image is the CT three-dimensional image data before performing a surgical operation on the target object. Preprocessing the preoperative tissue section image may include: standardizing and denoising the preoperative tissue section image to ensure that the resolution and size of the image are consistent and improve the image quality.

[0085] According to embodiments of the present disclosure, a multi-head attention network may be used to extract features from the preprocessed preoperative tissue section image to obtain initial lymph node features, and the initial lymph node features represent the lymph node attribute information of the preoperatively extracted preoperative tissue section image.

[0086] According to embodiments of the present disclosure, the mutual information between the preset intraoperative distribution features and the initial lymph node features may be calculated using a kernel density estimation algorithm. The mutual information can measure the shared information amount between the same variables in the preset intraoperative distribution features and the initial lymph node features. The larger the mutual information, the more redundant information is contained. Variables in the initial lymph node features with mutual information higher than the first threshold are filtered to obtain preoperative filtered features. The preoperative filtered features represent the features after removing redundancy from the preset intraoperative distribution features.

[0087] According to embodiments of the present disclosure, the maximum mean discrepancy algorithm may be used to process the preset preoperative distribution features and the preoperative filtered features, calculate the distribution difference between the variables in the preset preoperative distribution features and the preoperative filtered features, and align and correct the data of the variables with the difference value higher than the second threshold to obtain intermediate lymph node features.

[0088] According to embodiments of the present disclosure, the intermediate lymph node features are the features after aligning the feature distributions with the preset preoperative distribution features.

[0089] According to embodiments of the present disclosure, an image feature fusion algorithm may be used to fuse the intermediate lymph node features and the preoperative tissue section image to obtain the lymph node image features related to the target region.

[0090] According to an embodiment of the present disclosure, by adopting pre-constructed standardized preset preoperative distribution features and preset intraoperative distribution features, and performing feature redundancy removal and feature distribution alignment correction on the extracted initial lymph node features, not only can the initial lymph node features retain features with strong complementarity and low redundancy between intraoperative modalities, but also can significantly reduce the domain shift caused by data sources and imaging differences, provide a more accurate and efficient feature representation for the recognition model, and ultimately improve the lymph node metastasis prediction performance.

[0091] According to an embodiment of the present disclosure, the lymph node image features are obtained through fusion processing based on a second feature extraction model. The second feature extraction model includes a multi-head attention layer and a weighted fusion layer, and the target region includes the lymph node region. Among them, the fusion processing of the intermediate lymph node features and the preoperative tissue section image to obtain the lymph node image features related to the target region includes: determining the lymph node region, gastric cancer primary focus region, and visceral fat region with a mapping relationship from the preoperative tissue section image of the target object; using the multi-head attention layer to extract features from the gastric cancer primary focus region and the visceral fat region that have a mapping relationship with the lymph node region respectively to obtain lymph node region features, gastric cancer primary focus region features, and visceral fat region features; using the weighted fusion layer to process the intermediate lymph node features, gastric cancer primary focus region features, and visceral fat region features to obtain the lymph node image features.

[0092] According to an embodiment of the present disclosure, the second feature extraction model includes two separate multi-head attention layers and a weighted fusion layer, and the multi-head attention layer is constructed based on the multi-head attention mechanism.

[0093] According to an embodiment of the present disclosure, the preoperative tissue section image is manually annotated to outline the target regions, gastric cancer primary focus regions, and visceral fat regions where each group of lymph nodes are located.

[0094] According to an embodiment of the present disclosure, a one-to-many mapping relationship is constructed from the preoperative tissue section image: the association relationship between the gastric cancer primary focus region, the visceral fat region, and the lymph node regions where each group of lymph nodes are located.

[0095] According to an embodiment of the present disclosure, for the lymph node region where a certain group of lymph nodes are located, two parallel multi-head attention layers are used to extract features from the gastric cancer primary focus region and the visceral fat region respectively to obtain gastric cancer primary focus region features and visceral fat region features.

[0096] According to an embodiment of the present disclosure, the intermediate lymph node features represent the lymph node morphology, size, and boundary information in the corrected lymph node region. The gastric cancer primary focus region features represent the tissue morphology, size, and boundary information in the gastric cancer primary focus region. The visceral fat region features represent the tissue morphology, size, and boundary information in the visceral fat region.

[0097] According to an embodiment of the present disclosure, a weighted sum processing is performed on the intermediate lymph node features, the gastric cancer primary focus region features, and the visceral fat region features by using a weighted fusion layer to obtain lymph node image features corresponding to the lymph node region where this group of lymph nodes is located. Thus, lymph node image features corresponding to each of the lymph node regions where each group of lymph nodes is located are obtained.

[0098] According to an embodiment of the present disclosure, the location, size, shape of the primary focus, and its relationship with surrounding tissues can directly affect the probability and pattern of lymph node lesions. The depth of invasion, pathological type of the gastric cancer primary focus, and the anatomical relationship with the lymphatic drainage path are all important factors determining whether lymph nodes metastasize. Therefore, extracting the gastric cancer primary focus region features can not only provide key clues for the spatial association between tumor tissues and lymph nodes, but also provide more comprehensive and accurate support for predicting the metastasis risk of lymph nodes by integrating the feature information of the primary focus (such as tumor boundary irregularity, depth of invasion, etc.).

[0099] According to an embodiment of the present disclosure, visceral fat is not only an energy storage site, but also secretes a variety of cytokines and adipokines, which can promote chronic inflammation and tumor-related inflammation, thereby providing a suitable microenvironment for the progression and metastasis of cancer. The proportion of visceral fat can reflect the metabolic state of the patient, and further affect the ability of tumor cells to metastasize and the degree of lymph node involvement.

[0100] According to an embodiment of the present disclosure, in the process of extracting lymph node image features corresponding to each group of lymph nodes from preoperative tissue section images, based on the gastric cancer primary focus region features and the visceral fat region features as auxiliary features, and using the intermediate lymph node features as the main features, the boundary information, cell change information, etc. of each group of lymph nodes of the target object can be learned more comprehensively, so as to more accurately predict lymph node metastasis.

[0101] According to an embodiment of the present disclosure, the second feature extraction model is trained based on the following operations: obtaining training samples, where the training samples include sample preoperative tissue section images of sample objects and sample labels, the sample preoperative tissue section images include target regions, gastric cancer primary focus regions, and visceral fat regions, and the sample labels represent that tumor tissues metastasize to the target region or tumor tissues do not metastasize to the target region; using the second feature extraction model to perform feature extraction on the sample preoperative tissue section images to obtain sample lymph node image features related to the target region; performing recognition on the sample lymph node image features to obtain a sample recognition result; using a loss function to process the sample recognition result, the sample label, the area of the visceral fat region, and the distance between the gastric cancer primary focus region and the lymph node region to obtain a loss value; training the second feature extraction model according to the loss value to obtain the trained second feature extraction model.

[0102] According to an embodiment of the present disclosure, the target region may be the lymph node regions where each group of lymph nodes is located. In the preoperative tissue section image of the sample, the lymph node regions where each group of lymph nodes is located, the primary gastric cancer focus region, and the visceral fat region are marked. Each lymph node region where a group of lymph nodes is located corresponds to a sample label, and the sample label indicates whether the tumor tissue has metastasized to this lymph node region or not.

[0103] According to an embodiment of the present disclosure, the sample lymph node image features related to the target region are the lymph node features corresponding to each of the lymph node regions where each group of lymph nodes is located.

[0104] According to an embodiment of the present disclosure, the sample lymph node image features are input into the fully connected layer, and then classified through the activation function softmax to obtain the sample recognition result.

[0105] According to an embodiment of the present disclosure, each lymph node region where a group of lymph nodes is located corresponds to a sample recognition result, and the sample recognition result represents the prediction result of whether the tumor tissue has metastasized to the lymph node region where this group of lymph nodes is located.

[0106] According to an embodiment of the present disclosure, 10-fold cross-validation can be used for model evaluation. If the loss function value of the sample recognition result and the sample label does not meet the training end condition, the parameters are fine-tuned and the training is iterated until the loss function value meets the training end condition, and a trained second feature extraction model is obtained.

[0107] In one embodiment, the loss function is as shown in formula (1):

[0108] (1);

[0109] Where: B represents the number of preoperative tissue section images of the sample in each round of the training process, represents the number of groups of lymph nodes labeled in the j-th preoperative tissue section image of the sample, represents the sample label of the i-th group of lymph nodes in the j-th preoperative tissue section image of the sample, represents the sample recognition result of the i-th group of lymph nodes in the j-th preoperative tissue section image of the sample, represents the distance between the lymph node region where the i-th group of lymph nodes is located and the primary gastric cancer focus region in the j-th preoperative tissue section image of the sample, is the average distance between the lymph node regions where each group of lymph nodes is located and the primary gastric cancer focus region in B preoperative tissue section images of the sample, represents the area of the visceral fat region in the j-th preoperative tissue section image of the sample where the i-th group of lymph nodes is located, is the average value of the areas of the visceral fat regions in B preoperative tissue section images of the sample, is a hyperparameter for adjusting the influence of the fat region, and are coefficients used to balance different loss terms.

[0110] According to an embodiment of the present disclosure, the distance between the primary gastric cancer focus region and the lymph node region is added to the loss optimization, thereby constraining the distance distribution between the primary focus and the lymph nodes, making the second feature extraction model closer to the actual clinical scenario, paying attention to the spatial relationship with consistency or statistical significance, and improving the reliability and interpretability of the prediction; the area of the visceral fat region is added to the loss optimization, thereby constraining the second feature extraction model to pay attention to the impact of the change in fat distribution on the prediction result. By combining the visceral fat area index, the model can more comprehensively reflect the overall health status of the patient, such as nutritional status, obesity-related inflammation, etc., so as to more accurately predict lymph node metastasis.

[0111] Figure 3 Shows an example schematic diagram of the second feature extraction model according to an embodiment of the present disclosure.

[0112] As Figure 3 shown, the second feature extraction model includes a multi-head attention layer 310, a multi-head attention layer 320, and a weighted fusion layer 330. For a lymph node region where a certain group of lymph nodes is located, the multi-head attention layer 310 is used to extract features from the primary gastric cancer focus region to obtain the primary gastric cancer focus region features; the multi-head attention layer 320 is used to extract features from the visceral fat region to obtain the visceral fat region features; the weighted fusion layer 330 is used to perform weighted summation or weighted average processing on the intermediate lymph node features, the primary gastric cancer focus region features, and the visceral fat region features to obtain the lymph node image features corresponding to the lymph node region where this group of lymph nodes is located.

[0113] According to an embodiment of the present disclosure, the recognition model includes a cross-attention layer and a decoding layer; wherein, using the recognition model to process the lymph node image features and the target image features corresponding to the target region, the obtained recognition results include: performing an adaptive convolution operation on the lymph node image features corresponding to the target region to obtain preoperative key features; performing an adaptive convolution operation on the target image features corresponding to the target region to obtain intraoperative key features; using the cross-attention layer to process the preoperative key features and the intraoperative key features to obtain initial fusion features; splicing the initial fusion features, the preoperative key features, the intraoperative key features, the lymph node image features, and the target image features to obtain target fusion features; using the decoding layer to process the target fusion features to obtain the recognition result corresponding to the target region.

[0114] According to an embodiment of the present disclosure, the cross-attention layer can be constructed based on the cross-attention mechanism (CrossAttention). The decoding layer can be constructed based on the mixture-of-experts decoder (Mixture-of-Experts Decoder).

[0115] According to an embodiment of the present disclosure, the recognition model further includes an adaptive convolutional layer. An adaptive convolution operation is performed on the lymph node image features corresponding to the target region by using the adaptive convolutional layer to obtain preoperative key features. The preoperative key features are the key features after enhancing the lymph node image features.

[0116] According to an embodiment of the present disclosure, an adaptive convolution operation is performed on the target image features corresponding to the target region by using the adaptive convolutional layer to obtain intraoperative key features. The intraoperative key features are the key features after enhancing the target image features.

[0117] According to an embodiment of the present disclosure, the initial fusion feature is the information after enhancing the highly correlated features between the preoperative key features and the intraoperative key features.

[0118] According to an embodiment of the present disclosure, the initial fusion feature, the preoperative key feature, the intraoperative key feature, the lymph node image feature, and the target image feature are spliced to obtain a target fusion feature; then the decoding layer is used to perform classification prediction on the target fusion feature to obtain a recognition result corresponding to the target region.

[0119] According to an embodiment of the present disclosure, during the training process of the recognition model, the area under the ROC curve (Receiver Operating Characteristic), accuracy, sensitivity, and specificity can be used to verify and evaluate the classification performance of the recognition model, and the hyperparameters of model training, such as the learning rate and batch size, are adjusted when the accuracy and robustness of the recognition model are low.

[0120] Figure 4 An exemplary schematic diagram of the recognition model according to an embodiment of the present disclosure is shown.

[0121] As Figure 4As shown, the recognition model includes an adaptive convolutional layer 410, an adaptive convolutional layer 420, a cross-attention layer 430, a splicing layer 440, and a decoding layer 450. The adaptive convolutional layer 410 performs an adaptive convolution operation on the lymph node image features corresponding to the target region to obtain preoperative key features; the adaptive convolutional layer 420 performs an adaptive convolution operation on the target image features corresponding to the target region to obtain intraoperative key features; the cross-attention layer 430 processes the preoperative key features and the intraoperative key features to obtain initial fusion features; the splicing layer 440 processes the initial fusion features, the preoperative key features, the intraoperative key features, the lymph node image features, and the target image features to obtain target fusion features; and the decoding layer 450 processes the target fusion features to obtain a recognition result corresponding to the target region.

[0122] Figure 5 FIG. shows an exemplary schematic diagram of obtaining a recognition result according to an embodiment of the present disclosure.

[0123] As Figure 5 shown, feature extraction is performed on the intraoperative video frame 501 to obtain initial image features 502; the initial image features 502 are corrected based on a preset preoperative distribution feature 503 and a preset intraoperative distribution feature 504 to obtain target image features 505; feature extraction is performed on the preoperative tissue section image 506 to obtain initial lymph node features 507; the initial lymph node features 507 are corrected based on the preset preoperative distribution feature 503 and the preset intraoperative distribution feature 504 to obtain intermediate lymph node features 508; feature extraction is performed on the preoperative tissue section image 506 to obtain a gastric cancer primary focus region feature 509 and a visceral fat region feature 510; the intermediate lymph node features 508, the gastric cancer primary focus region feature 509, and the visceral fat region feature 510 are fused to obtain lymph node image features 511 related to the target region; and the recognition model processes the lymph node image features 511 and the target image features 505 to obtain a recognition result 512.

[0124] Based on the above tumor tissue metastasis risk recognition method, the present disclosure also provides a tumor tissue metastasis risk recognition device. The following will be combined with Figure 6 to describe this device in detail.

[0125] Figure 6 FIG. shows a structural block diagram of a tumor tissue metastasis risk recognition device according to an embodiment of the present disclosure.

[0126] As Figure 6 shown, the tumor tissue metastasis risk recognition device 600 of this embodiment includes an acquisition module 610, an extraction module 620, and a recognition module 630.

[0127] An acquisition module 610 is configured to acquire intraoperative video frames representing surgical operations on tumor tissues of a target object, and lymph node image features related to a target region of the target object, where the lymph node image features represent the lymph node attributes before performing the surgical operation on the target object. In an embodiment, the acquisition module 610 may be configured to perform the operation S110 described above, which will not be elaborated herein.

[0128] An extraction module 620 is configured to extract features from the intraoperative video frames to obtain target image features corresponding to the target region, where the target image features represent the lymph node attributes of the target object during the surgical operation. In an embodiment, the extraction module 620 may be configured to perform the operation S120 described above, which will not be elaborated herein.

[0129] An identification module 630 is configured to process the lymph node image features and the target image features corresponding to the target region by using an identification model to obtain an identification result, where the identification result represents the risk level of tumor tissue metastasis to the target region. In an embodiment, the identification module 630 may be configured to perform the operation S130 described above, which will not be elaborated herein.

[0130] According to an embodiment of the present disclosure, the extraction module 620 includes a first extraction sub-module, a second extraction sub-module, and a third extraction sub-module.

[0131] The first extraction sub-module is configured to extract features from the intraoperative video frames to obtain initial image features.

[0132] The second extraction sub-module is configured to filter the initial image features based on the mutual information between the preset preoperative distribution features and the initial image features to obtain intraoperative filtered features, where the preset preoperative distribution features represent a lymph node image feature distribution library related to the target region constructed based on sample objects before the surgical operation.

[0133] The third extraction sub-module is configured to correct the intraoperative filtered features based on the preset intraoperative distribution features to obtain target image features, where the preset intraoperative distribution features represent a target image feature distribution library related to the target region constructed based on sample objects during the surgical operation.

[0134] According to an embodiment of the present disclosure, the first extraction sub-module includes a first extraction unit and a second extraction unit.

[0135] The first extraction unit block is configured to process the intraoperative video frames by using a feature extraction layer to obtain image fusion features corresponding to the target region.

[0136] The second extraction unit block is configured to process the image fusion features corresponding to the target region by using a long short-term memory network layer to obtain initial image features corresponding to the target region.

[0137] According to an embodiment of the present disclosure, the tumor tissue metastasis risk identification device 600 further includes a first construction module, a second construction module, a third construction module, a fourth construction module, and a fifth construction module.

[0138] The first construction module is configured to extract features from the preoperative tissue section image of the sample object to obtain sample lymph node image features.

[0139] The second construction module is configured to extract features from the intraoperative video frames of the sample object to obtain sample target image features.

[0140] The third construction module is configured to use a conditional generative adversarial network to process the sample lymph node image features and the target image features to obtain a sample preoperative feature distribution corresponding to the sample lymph node image features and a sample intraoperative feature distribution corresponding to the sample target image features.

[0141] The fourth construction module is configured to perform filtering processing on the sample preoperative feature distribution and the sample intraoperative feature distribution respectively based on the mutual information between the sample preoperative feature distribution and the sample intraoperative feature distribution to obtain a sample preoperative corrected feature distribution and a sample intraoperative corrected feature distribution.

[0142] The fifth construction module is configured to respectively model the sample preoperative corrected feature distribution and the sample intraoperative corrected feature distribution based on the kernel density estimation algorithm to obtain a preset preoperative distribution feature and a preset intraoperative distribution feature.

[0143] According to an embodiment of the present disclosure, the second construction module includes a first construction sub-module, a second construction sub-module, and a third construction sub-module.

[0144] The first construction sub-module is configured to use a convolutional neural network to process the intraoperative video frames of the sample to obtain sample lymph node static features.

[0145] The second construction sub-module is configured to use an image coding network to process the intraoperative video frames of the sample to obtain sample lymph node dynamic features.

[0146] The third construction sub-module is configured to obtain sample target image features according to the sample lymph node static features and the sample lymph node dynamic features.

[0147] According to an embodiment of the present disclosure, the tumor tissue metastasis risk identification device 600 further includes a first preoperative feature processing module, a second preoperative feature processing module, a third preoperative feature processing module, and a fourth preoperative feature processing module.

[0148] The first preoperative feature processing module is configured to extract features from the preoperative tissue section image of the target object to obtain initial lymph node features.

[0149] The second pre-operative feature processing module is used to filter the initial lymph node features based on the mutual information between the preset intra-operative distribution features and the initial lymph node features, so as to obtain pre-operative filtered features.

[0150] The third pre-operative feature processing module is used to correct the pre-operative filtered features based on the preset pre-operative distribution features, so as to obtain intermediate lymph node features.

[0151] The fourth pre-operative feature processing module is used to perform fusion processing on the intermediate lymph node features and the pre-operative tissue section images, so as to obtain lymph node image features related to the target area.

[0152] According to an embodiment of the present disclosure, the first pre-operative feature processing module includes a first processing sub-module, a second processing sub-module, and a third processing sub-module.

[0153] The first processing sub-module is used to determine the lymph node area, gastric cancer primary focus area, and visceral fat area with a mapping relationship from the pre-operative tissue section images of the target object.

[0154] The second processing sub-module is used to respectively extract features of the gastric cancer primary focus area and the visceral fat area with a mapping relationship with the lymph node area by using a multi-head attention layer, so as to obtain the gastric cancer primary focus area features and the visceral fat area features.

[0155] The third processing sub-module is used to process the intermediate lymph node features, the gastric cancer primary focus area features, and the visceral fat area features by using a weighted fusion layer, so as to obtain lymph node image features.

[0156] According to an embodiment of the present disclosure, the recognition module 630 includes a first recognition sub-module, a second recognition sub-module, a third recognition sub-module, a fourth recognition sub-module, and a fifth recognition sub-module.

[0157] The first recognition sub-module is used to perform an adaptive convolution operation on the lymph node image features corresponding to the target area, so as to obtain pre-operative key features.

[0158] The second recognition sub-module is used to perform an adaptive convolution operation on the target image features corresponding to the target area, so as to obtain intra-operative key features.

[0159] The third recognition sub-module is used to process the pre-operative key features and the intra-operative key features by using a cross-attention layer, so as to obtain initial fusion features.

[0160] The fourth recognition sub-module is used to splice the initial fusion features, the pre-operative key features, the intra-operative key features, the lymph node image features, and the target image features, so as to obtain target fusion features.

[0161] The fifth recognition sub-module is used to process the target fusion features by using a decoding layer, so as to obtain a recognition result corresponding to the target area.

[0162] According to embodiments of the present disclosure, any plurality of modules among modules, sub-modules, units, and sub-units may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present disclosure, at least one of modules, sub-modules, units, and sub-units may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), programmable logic array (PLA), system on a chip, system on a substrate, system in a package, application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware for integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of modules, sub-modules, units, and sub-units may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.

[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0164] The embodiments according to the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all of these substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A method for identifying the risk of tumor tissue metastasis, characterized in that, The method includes: Obtaining intraoperative video frames representing a surgical operation on a tumor tissue of a target object, and lymph node image features related to a target region of the target object, where the lymph node image features represent the lymph node attributes before performing the surgical operation on the target object; Performing feature extraction on the intraoperative video frames to obtain target image features corresponding to the target region, where the target image features represent the lymph node attributes of the target object during the surgical operation; Processing the lymph node image features and the target image features corresponding to the target region by using an identification model to obtain an identification result, where the identification result represents the risk level of tumor tissue metastasis to the target region.

2. The method according to claim 1, characterized in that, The performing feature extraction on the intraoperative video frames to obtain target image features corresponding to the target region includes: Performing feature extraction on the intraoperative video frames to obtain initial image features; Filtering the initial image features based on the mutual information between a preset preoperative distribution feature and the initial image features to obtain intraoperative filtered features, where the preset preoperative distribution feature represents a lymph node image feature distribution library related to the target region constructed based on sample objects before performing the surgical operation; Correcting the intraoperative filtered features based on a preset intraoperative distribution feature to obtain target image features, where the preset intraoperative distribution feature represents a target image feature distribution library related to the target region constructed based on the sample objects during the surgical operation.

3. The method according to claim 2, wherein The initial image features are obtained by performing feature extraction based on a first feature extraction model, and the first feature extraction model includes a feature extraction layer and a long short-term memory network layer; Wherein, the performing feature extraction on the intraoperative video frames to obtain initial image features includes: Processing the intraoperative video frames by using the feature extraction layer to obtain an image fusion feature corresponding to the target region; Processing the image fusion feature corresponding to the target region by using the long short-term memory network layer to obtain initial image features corresponding to the target region.

4. The method according to claim 2, wherein The preset preoperative distribution feature and the preset intraoperative distribution feature are determined based on the following operations: Performing feature extraction on sample preoperative tissue section images of the sample objects to obtain sample lymph node image features; Performing feature extraction on sample intraoperative video frames of the sample objects to obtain sample target image features; Processing the sample lymph node image features and the target image features by using a conditional generative adversarial network to obtain a sample preoperative feature distribution corresponding to the sample lymph node image features and a sample intraoperative feature distribution corresponding to the sample target image features; Performing filtering processing on the sample preoperative feature distribution and the sample intraoperative feature distribution respectively based on the mutual information between the sample preoperative feature distribution and the sample intraoperative feature distribution to obtain a sample preoperative corrected feature distribution and a sample intraoperative corrected feature distribution; Based on the kernel density estimation algorithm, respectively modeling the sample preoperative corrected feature distribution and the sample intraoperative corrected feature distribution to obtain the preset preoperative distribution feature and the preset intraoperative distribution feature.

5. The method according to claim 4, wherein Performing feature extraction on the intraoperative video frames of the sample object to obtain sample target image features, including: Processing the intraoperative video frames of the sample using a convolutional neural network to obtain static features of the sample lymph nodes; Processing the intraoperative video frames of the sample using an image encoding network to obtain dynamic features of the sample lymph nodes; Obtaining the sample target image features based on the static features of the sample lymph nodes and the dynamic features of the sample lymph nodes.

6. The method according to claim 2, wherein The lymph node image features related to the target region are obtained based on the following operations: Performing feature extraction on the preoperative tissue section image of the target object to obtain initial lymph node features; Filtering the initial lymph node features based on the mutual information between the preset intraoperative distribution features and the initial lymph node features to obtain preoperative filtered features; Correcting the preoperative filtered features based on the preset preoperative distribution features to obtain intermediate lymph node features; Performing fusion processing on the intermediate lymph node features and the preoperative tissue section image to obtain the lymph node image features related to the target region.

7. The method according to claim 6, wherein The lymph node image features are obtained through fusion processing based on a second feature extraction model, the second feature extraction model includes a multi-head attention layer and a weighted fusion layer, and the target region includes a lymph node region; Among them, the performing fusion processing on the intermediate lymph node features and the preoperative tissue section image to obtain the lymph node image features related to the target region includes: Determining the lymph node region, gastric cancer primary focus region, and visceral fat region with a mapping relationship from the preoperative tissue section image of the target object; Using the multi-head attention layer to perform feature extraction on the gastric cancer primary focus region and the visceral fat region with a mapping relationship to the lymph node region respectively to obtain the gastric cancer primary focus region features and the visceral fat region features; Using the weighted fusion layer to process the intermediate lymph node features, the gastric cancer primary focus region features, and the visceral fat region features to obtain the lymph node image features related to the target region.

8. The method according to claim 7, wherein The second feature extraction model is trained based on the following operations: Obtaining training samples, the training samples include the preoperative tissue section image of the sample object, sample labels, the preoperative tissue section image of the sample includes a target region, a gastric cancer primary focus region, and a visceral fat region, and the sample labels indicate that the tumor tissue has metastasized to the target region or the tumor tissue has not metastasized to the target region; Using the second feature extraction model to perform feature extraction on the preoperative tissue section image of the sample to obtain sample lymph node image features related to the target region; Identifying the sample lymph node image features to obtain a sample identification result; Using a loss function to process the sample identification result, the sample labels, the area of the visceral fat region, and the distance between the gastric cancer primary focus region and the lymph node region to obtain a loss value; Training the second feature extraction model according to the loss value to obtain the trained second feature extraction model.

9. The method according to claim 1, wherein The identification model includes a cross-attention layer and a decoding layer; Among them, the process of using the recognition model to process the lymph node image features and the target image features corresponding to the target area to obtain the recognition result includes: Performing an adaptive convolution operation on the lymph node image features corresponding to the target area to obtain preoperative key features; Performing an adaptive convolution operation on the target image features corresponding to the target area to obtain intraoperative key features; Using the cross-attention layer to process the preoperative key features and the intraoperative key features to obtain initial fusion features; Concatenating the initial fusion features, the preoperative key features, the intraoperative key features, the lymph node image features, and the target image features to obtain target fusion features; Using the decoding layer to process the target fusion features to obtain the recognition result corresponding to the target area.

10. A tumor tissue metastasis risk identification device, characterized in that, The device includes: An acquisition module, configured to acquire an intraoperative video frame representing a surgical operation on a tumor tissue of a target object, and lymph node image features related to a target area of the target object, where the lymph node image features represent the lymph node attributes before performing the surgical operation on the target object; An extraction module, configured to perform feature extraction on the intraoperative video frame to obtain target image features corresponding to the target area, where the target image features represent the lymph node attributes of the target object during the surgical operation; A recognition module, configured to use a recognition model to process the lymph node image features and the target image features corresponding to the target area to obtain a recognition result, where the recognition result represents the risk level of tumor tissue metastasis to the target area.