Lymph node classification method, device and equipment

Through self-supervised feature extraction and spatiotemporal feature processing of lymph node areas, combined with self-supervised learning and convolutional neural network, efficient and accurate diagnosis of lymph node lesion types is achieved, and poor diagnosis accuracy and delay problems caused by relying on experience in existing methods are solved, improving medical efficiency and real-time diagnosis.

CN120279332APending Publication Date: 2025-07-08INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432127.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing lymph node lesion type identification methods rely on the experience of medical technicians, resulting in poor diagnostic accuracy and diagnostic delays, affecting the real-time treatment.

Method used

By self-supervised feature extraction of target videos within the preset time, combined with spatiotemporal feature processing of target frame images and adjacent frame images, the classification module is used to predict the lesion type of lymph node area, including feature buffering units and self-supervised feature extraction units, and feature extraction and classification are combined with self-attention mechanisms and convolutional neural networks.

Benefits of technology

It improves the accuracy of lymph node regional diagnosis, reduces the need for manual labeling and pathological examination, saves labor costs, and realizes non-invasive real-time diagnosis and rapid judgment, avoids missed or incorrect cleaning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279332A_ABST
    Figure CN120279332A_ABST
Patent Text Reader

Abstract

The invention provides a lymph node classification method, device and equipment, which can be applied to the technical field of image processing. The method comprises the following steps: carrying out self-supervised feature extraction on a target video within a preset time to obtain an initial feature, and marking a lymph node region in a target frame image; performing organization category identification on each pixel point of the target frame image to obtain an organization category of each pixel point in the target frame image; performing organization category identification on each pixel point of each adjacent frame image to obtain an organization category of each pixel point in each adjacent frame image; performing spatial-temporal feature processing on the initial feature, pixel points of a target frame image corresponding to the tissue category and pixel points of an adjacent frame image by using feature extraction modules corresponding to the tissue categories to obtain processed spatial-temporal features; and performing lesion type prediction on the processed spatial-temporal characteristics by using a classification module to obtain a type result corresponding to the lymph node region in the target frame image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and particularly to a lymph node classification method, apparatus, and device. Background Art

[0002] Gastric cancer is one of the high-incidence cancer types. Due to its inconspicuous early symptoms and easy occurrence of lymph node metastasis, the survival period of patients is significantly reduced. In order to improve the survival period of gastric cancer patients, how to accurately identify whether tumor tissue metastasis has occurred in the lymph node area for cleaning has become an important research direction.

[0003] However, the existing methods for identifying lymph node lesion types rely on the experience and technical level of medical technicians. The diagnostic accuracy deviation caused by the experience differences of different medical technicians is relatively large, and visual judgment may take several days, which may lead to the problem of diagnostic delay and poor real-time treatment effect. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides a lymph node classification method, apparatus, and device.

[0005] According to a first aspect of the present disclosure, a lymph node classification method is provided, including: performing self-supervised feature extraction on a target video within a preset time to obtain initial features, where the target video includes a target frame image and a plurality of adjacent frame images having an adjacent time sequence relationship with the target frame image, and the lymph node area is marked in the target frame image; performing tissue category recognition on each pixel point of the target frame image to obtain the tissue category of each pixel point in the target frame image; performing tissue category recognition on each pixel point of each adjacent frame image to obtain the tissue category of each pixel point in each adjacent frame image; using a feature extraction module corresponding to each tissue category to perform spatio-temporal feature processing on the initial features, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame images to obtain processed spatio-temporal features; and using a classification module to perform lesion type prediction on the processed spatio-temporal features to obtain a type result corresponding to the lymph node area in the target frame image.

[0006] According to an embodiment of the present disclosure, the feature extraction module includes a feature buffer unit and a self-supervised feature extraction unit. The feature extraction module has a one-to-one correspondence with the tissue category. The number of feature extraction modules and the number of tissue categories are both N. The processed spatio-temporal features include buffer features and self-supervised features, 1 N; wherein, by using the feature extraction modules respectively corresponding to the tissue categories, spatio-temporal feature processing is performed on the initial features, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame images, and the processed spatio-temporal features obtained include: using the (n + 1)-th self-supervised feature extraction unit to perform self-supervised feature processing on the pixel points of the target frame image corresponding to the (n + 1)-th tissue category, the pixel points of the adjacent frame images, the n-th self-supervised feature, and the n-th buffered feature to obtain the (n + 1)-th self-supervised feature, wherein the n-th buffered feature is obtained by using the n-th buffer unit to perform temporal dynamic adjustment on the (n - 1)-th self-supervised feature. In the case of n = 1, the 1st self-supervised feature is obtained by performing self-supervised feature processing on the initial features, the pixel points of the target frame image corresponding to the 1st tissue category, and the pixel points of the adjacent frame images, and the 1st buffered feature is obtained by using the 1st feature buffer unit to perform temporal dynamic adjustment on the initial features; using the (n + 1)-th buffered feature unit to perform temporal dynamic adjustment processing on the n-th self-supervised feature to obtain the (n + 1)-th buffered feature.

[0007] According to an embodiment of the present disclosure, using the classification module to perform lesion type prediction on the processed spatio-temporal features, and the type result corresponding to the lymph node region in the target frame image obtained includes: performing self-supervised feature extraction on the N-th self-supervised feature and the N-th buffered feature to obtain target features; using the classification module to perform lesion type prediction on the target features to obtain the type result corresponding to the lymph node region in the target frame image.

[0008] According to an embodiment of the present disclosure, the classification module includes a fully connected unit and an activation unit. Using the classification module to perform lesion type prediction on the target features, and the type result corresponding to the lymph node region in the target frame image obtained includes: using the fully connected unit to process the target features to obtain fully connected features; using the activation unit to process the fully connected features to obtain the type result corresponding to the lymph node region in the target frame image.

[0009] According to an embodiment of the present disclosure, the tissue category of each pixel point in the target frame image is obtained by using a tissue category model to perform tissue category recognition on each pixel point in the target frame image. The tissue category model includes an embedding module, an encoding module, a feature pyramid module, a decoding module, and a segmentation module; wherein, performing tissue category recognition on each pixel point of the target frame image to obtain the tissue category of each pixel point in the target frame image includes: using the embedding module to process each pixel point of the target frame image to obtain mapped features; using the encoding module to process the mapped features to obtain encoded features; using the feature pyramid module to process the encoded features to obtain multi-scale features; using the decoding module to process the multi-scale features and pixel point position information to obtain decoded features; using the segmentation module to process the decoded features to obtain the tissue category of the pixel points.

[0010] According to an embodiment of the present disclosure, the encoding module includes a self-attention unit, a feed-forward neural unit, and a normalization unit; wherein, processing the mapped features by the encoding module to obtain encoded features includes: processing the mapped features by the self-attention unit to obtain attention features; processing the attention features by the feed-forward neural unit to obtain feed-forward neural features; performing a residual connection on the attention features and the feed-forward neural features to obtain residual features; processing the residual features by the normalization unit to obtain encoded features.

[0011] According to an embodiment of the present disclosure, the classification model includes a classification module and a feature extraction module corresponding to each tissue category respectively. The classification model is trained based on the following operations: obtaining training samples, where the training samples include a sample target video and a sample label within a preset time. The sample target video includes a sample target frame image and a plurality of sample adjacent frame images having an adjacent time sequence relationship with the sample target frame image, and the sample label represents the lesion type of the lymph node region in the sample target frame image; performing self-supervised feature extraction on the sample target video to obtain sample initial features; respectively performing tissue category recognition on each pixel point of the sample target frame image and the sample adjacent frame images to obtain the tissue category of each pixel point in the sample target frame image and the tissue category of each pixel point in the sample adjacent frame images; processing the sample initial features, the pixel points of the sample target frame image corresponding to the tissue category, and the pixel points of the sample adjacent frame images by the classification model to be trained to obtain a sample type result corresponding to the lymph node region in the sample target frame image; training the classification model to be trained according to the sample initial features, the sample type result, and the sample label to obtain a trained classification model.

[0012] According to an embodiment of the present disclosure, the sample initial features include a plurality of sub-features; wherein, training the classification model to be trained according to the sample initial features, the sample type result, and the sample label to obtain a trained classification model includes: calculating a loss value between the sample type result and the sample label by using a classification loss function to obtain a classification loss value; calculating a loss value between every two adjacent time-sequence sub-features by using a temporal dependence loss function to obtain a temporal dependence loss value; obtaining a target loss value according to the classification loss value and the temporal dependence loss value; training the classification model to be trained according to the target loss value to obtain a trained classification model.

[0013] The second aspect of the present disclosure provides a lymph node classification device, including: a first feature extraction module configured to perform self-supervised feature extraction on a target video within a preset time to obtain initial features, where the target video includes a target frame image and a plurality of adjacent frame images having an adjacent temporal relationship with the target frame image, and the lymph node region is marked in the target frame image; a first recognition module configured to perform tissue category recognition on each pixel point of the target frame image to obtain the tissue category of each pixel point in the target frame image; a second recognition module configured to perform tissue category recognition on each pixel point of each adjacent frame image to obtain the tissue category of each pixel point in each adjacent frame image; a second feature extraction module configured to perform spatio-temporal feature processing on the initial features, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame images by using the feature extraction module corresponding to each tissue category to obtain processed spatio-temporal features; and a prediction module configured to perform lesion type prediction on the processed spatio-temporal features by using a classification module to obtain a type result corresponding to the lymph node region in the target frame image.

[0014] The third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned lymph node classification method.

[0015] According to the lymph node classification method, device, and equipment provided by the present disclosure, initial features are obtained by performing self-supervised feature extraction on a target video within a preset time; tissue category recognition is performed on the pixel points of the target frame image and each adjacent frame image to obtain the tissue category of each pixel point in the target frame image and the adjacent frame images; spatio-temporal feature processing is performed on the initial features, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame images by using the feature extraction module corresponding to each tissue category to obtain processed spatio-temporal features; and lesion type prediction is performed on the processed spatio-temporal features to obtain a type result. Since auxiliary diagnosis can be performed by combining adjacent frame images when diagnosing the lymph node region in the target frame image extracted from intraoperative image data, the diagnostic accuracy is improved. First, the tissue category of each pixel point in the target frame image and the adjacent frame images is recognized, and then the pixel points are input into the feature extraction module corresponding to the tissue category for efficient feature extraction. The hierarchical extraction of the feature extraction module by tissue category improves the accuracy of lymph node region diagnosis, reduces the need for manual annotation and pathological examination, thereby saving labor costs and improving medical efficiency; in addition, by non-invasively predicting the type result and providing real-time feedback to assist doctors in quickly judging the benign and malignant nature of lymph nodes, it provides a basis for surgical decision-making and avoids missed dissection or misdissection. Description of the Drawings

[0016] The above and other objects, features, and advantages of the present disclosure will become more apparent from the following description of the embodiments of the present disclosure with reference to the accompanying drawings.

[0017] Figure 1 The flowchart of the lymph node classification method according to an embodiment of the present disclosure is shown.

[0018] Figure 2 The exemplary schematic diagram of the classification model according to an embodiment of the present disclosure is shown.

[0019] Figure 3 The exemplary schematic diagram of the tissue category model according to an embodiment of the present disclosure is shown.

[0020] Figure 4 The exemplary schematic diagram of obtaining the type result corresponding to the lymph node region according to an embodiment of the present disclosure is shown.

[0021] Figure 5 The structural block diagram of the lymph node classification device according to an embodiment of the present disclosure is shown.

[0022] Figure 6 The block diagram of the electronic device suitable for implementing the lymph node classification method according to an embodiment of the present disclosure is schematically shown. Detailed Embodiments

[0023] Hereinafter, embodiments according to the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.

[0024] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0026] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning that a person skilled in the art usually understands this expression (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0027] Embodiments of the present disclosure provide a lymph node classification method, device, and equipment. The method includes: performing self-supervised feature extraction on a target video within a preset time to obtain initial features, where the lymph node region is marked in the target frame image; performing tissue category recognition on each pixel point of the target frame image to obtain the tissue category of each pixel point in the target frame image; performing tissue category recognition on each pixel point of each adjacent frame image to obtain the tissue category of each pixel point in each adjacent frame image; using a feature extraction module corresponding to each tissue category to perform spatio-temporal feature processing on the initial features, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame images to obtain processed spatio-temporal features; and using a classification module to perform lesion type prediction on the processed spatio-temporal features to obtain a type result corresponding to the lymph node region in the target frame image.

[0028] In the technical solution of the present disclosure, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. And the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0029] It should be noted that the sequence numbers of each operation in the following methods are only used as representations of the operation for description, and should not be regarded as indicating the execution order of each operation. Unless explicitly stated, the method does not need to be executed exactly in the order shown.

[0030] Figure 1 A flowchart of a lymph node classification method according to an embodiment of the present disclosure is shown.

[0031] As Figure 1 shown, the method 100 includes operations S110 to S150.

[0032] In operation S110, perform self-supervised feature extraction on a target video within a preset time to obtain initial features.

[0033] In operation S120, tissue category recognition is performed on each pixel point of the target frame image to obtain the tissue category of each pixel point in the target frame image.

[0034] In operation S130, tissue category recognition is performed on each pixel point in each adjacent frame image to obtain the tissue category of each pixel point in each adjacent frame image.

[0035] In operation S140, feature extraction modules corresponding to the tissue categories are used to perform spatiotemporal feature processing on the initial features, the pixels of the target frame image corresponding to the tissue category, and the pixels of the adjacent frame images to obtain processed spatiotemporal features.

[0036] In operation S150, a classification module is used to predict the lesion type based on the processed spatiotemporal features to obtain a type result corresponding to the lymph node region in the target frame image.

[0037] According to an embodiment of the present disclosure, the initial video may be video data captured during the exploration of perigastric lymph nodes during laparoscopic surgery for gastric cancer, such as obtaining a 30-second initial video. The video laparoscope camera should be located in an area where the perigastric lymph nodes can be clearly observed, and the camera angle needs to be kept stable to avoid excessive movement or lens defocusing.

[0038] According to an embodiment of the present disclosure, the initial video includes multiple frames of images, and the center points of the lymph nodes of interest in the multiple frames of images are marked, and the geometric center position of the lymph node area is clicked, that is, the symmetrical midpoint of its outline or the center of the largest volume area. If the shape of the lymph node is irregular, the visually uniform and unobstructed area should be selected as the center point, and click on the boundary or adjacent structure should be avoided. The principles for processing the covering connective tissue on the surface of each organ are as follows: for thin layers of connective tissue membranes (such as subserous layers or loose connective tissue), they can be classified as organ parenchyma areas; for thicker connective tissue layers (such as fascia) or connective tissue aggregated in blocks, when selecting the center of the lymph node, attention should be paid to the position of the covering connective tissue in the frame image, and avoid selecting the area where the connective tissue and other adjacent structures overlap, and finally select the image frame with the most complete display of the lymph node structure as the target frame image.

[0039] According to an embodiment of the present disclosure, a target frame image is first determined from an initial video according to a lymph node region marker, and a plurality of adjacent frame images having an adjacent temporal relationship with the target frame image are acquired based on a preset time and the moment of the target frame image.

[0040] For example, the preset time is 15 seconds, and based on the time of the target frame image as the center, multiple adjacent frame images within 15 seconds before and after the target frame image are obtained from the initial video, and the multiple adjacent frame images have an adjacent time sequence relationship.

[0041] According to an embodiment of the present disclosure, a target video is obtained according to a target frame image and a plurality of adjacent frame images having an adjacent temporal relationship with the target frame image.

[0042] According to an embodiment of the present disclosure, self-supervised feature extraction is performed on a target video within a preset time to obtain initial features. The initial features are temporal and spatial information of the target video that is preliminarily extracted and encoded.

[0043] According to an embodiment of the present disclosure, the target frame image has multiple adjacent frame images, and tissue category recognition is performed on each pixel point in each adjacent frame image to obtain the tissue category to which each pixel point in each adjacent frame image belongs.

[0044] According to an embodiment of the present disclosure, the tissue category represents the tissue classification of the area around the stomach, such as stomach, gallbladder, liver, peritoneum, intestine, and connective tissue.

[0045] According to an embodiment of the present disclosure, an encoder network may be used to perform tissue category recognition on each pixel in a target frame image to determine the tissue category to which each pixel in the target frame image belongs.

[0046] According to an embodiment of the present disclosure, the feature extraction module can be constructed based on the self-supervised attention mechanism, each feature extraction module corresponds to a tissue category one by one, and there is a cascade relationship between the feature extraction modules. For example, the preset tissue categories are six tissue categories of stomach, gallbladder, liver, peritoneum, intestine and connective tissue, and there are six pre-trained feature extraction modules corresponding to the six tissue categories.

[0047] According to an embodiment of the present disclosure, six feature extraction modules are sorted according to preset rules, the initial features are input into the first feature extraction module, and the pixel points of the target frame image corresponding to the tissue category and the pixel points of the adjacent frame image are input into the feature extraction module corresponding to the tissue category. After step-by-step spatiotemporal feature extraction, the processed spatiotemporal features are output in the six feature extraction modules.

[0048] For example, the pixel points of the target frame image corresponding to the tissue category of the stomach and the pixel points of the adjacent frame images are input into the feature extraction module corresponding to the stomach.

[0049] According to the embodiments of the present disclosure, the classification module can be constructed based on the activation function, and the classification module is used to perform probability prediction of the lesion type on the processed spatiotemporal features to obtain the type result corresponding to the lymph node area in the target frame image. The real-time classification result will be displayed on the surgical display screen.

[0050] According to an embodiment of the present disclosure, the type result may be the type of lesion corresponding to the lymph node region, such as a benign category or a malignant category.

[0051] According to an embodiment of the present disclosure, when diagnosing in the lymph node region of the target frame image extracted based on intraoperative image data, adjacent frame images can be combined for auxiliary diagnosis, which improves the diagnostic accuracy. First, the tissue categories of each pixel point in the target frame image and the adjacent frame images are identified, and then the pixel points are input into the feature extraction module corresponding to the tissue category for efficient feature extraction. The hierarchical extraction of the feature extraction module by tissue category improves the accuracy of lymph node region diagnosis, reduces the need for manual annotation and pathological examination, thus saving labor costs and improving medical efficiency. In addition, by non-invasively predicting the type result and providing real-time feedback to assist the doctor in quickly judging the benign and malignant nature of the lymph nodes, and providing a basis for surgical decision-making, it avoids missed dissection or misdissection.

[0052] According to an embodiment of the present disclosure, the feature extraction module includes a feature buffer unit and a self-supervised feature extraction unit. The feature extraction module has a one-to-one correspondence with the tissue category. The number of feature extraction modules and the number of tissue categories are both N. The processed spatio-temporal features include buffer features and self-supervised features, 1 N; wherein, using the feature extraction module corresponding to each tissue category respectively, spatio-temporal feature processing is performed on the initial feature, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame image. The obtained processed spatio-temporal features include: using the (n + 1)-th self-supervised feature extraction unit to perform self-supervised feature processing on the pixel points of the target frame image corresponding to the (n + 1)-th tissue category, the pixel points of the adjacent frame image, the n-th self-supervised feature, and the n-th buffer feature to obtain the (n + 1)-th self-supervised feature. Wherein, the n-th buffer feature is obtained by the n-th buffer unit performing temporal dynamic adjustment on the (n - 1)-th self-supervised feature. In the case of n = 1, the first self-supervised feature is obtained by performing self-supervised feature processing on the initial feature, the pixel points of the target frame image corresponding to the first tissue category, and the pixel points of the adjacent frame image. The first buffer feature is obtained by the first feature buffer unit performing temporal dynamic adjustment on the initial feature; using the (n + 1)-th buffer feature unit to perform temporal dynamic adjustment processing on the n-th self-supervised feature to obtain the (n + 1)-th buffer feature.

[0053] According to an embodiment of the present disclosure, the number of feature extraction modules and the number of tissue categories are both N. Each feature extraction module corresponds to a tissue category, and each feature extraction module includes a feature buffer unit and a self-supervised feature extraction unit.

[0054] According to an embodiment of the present disclosure, the self-supervised feature extraction unit is responsible for extracting effective spatio-temporal features from the input data. Without manually labeled data, the self-supervised feature extraction unit learns the internal structure and rules of the target video by designing self-supervised tasks. The self-supervised feature extraction unit is constructed based on a self-supervised learning framework, which combines an encoder (Transformer) and three Convolutional Neural Networks (CNNs).

[0055] According to an embodiment of the present disclosure, the feature buffer unit can solve the problem of information forgetting in target video processing. The feature buffer unit dynamically stores and updates the features of the feature buffer unit by introducing the long short-term memory mechanism of 4 recurrent neural networks. The long short-term memory mechanism dynamically adjusts the information in the feature buffer unit by processing the output features of the previous self-supervised feature extraction unit, ensuring that the features at each moment in the target video can be updated and supplemented in a timely manner according to the time change.

[0056] According to an embodiment of the present disclosure, when n = 1, the initial feature, the pixel points of the target frame image corresponding to the first tissue category, and the pixel points of the adjacent frame image are input into the first self-supervised feature extraction unit, and the first self-supervised feature is output. The first self-supervised feature is the spatio-temporal information that fuses the target video and the pixel points corresponding to the first tissue category.

[0057] According to an embodiment of the present disclosure, when n = 1, the initial feature is input into the first feature buffer unit, and the first buffer feature is output. The first buffer feature is the long-term temporal dependence information of the target video that is dynamically cached and updated.

[0058] According to an embodiment of the present disclosure, the nth self-supervised feature, the nth buffer feature, the pixel points of the target frame image corresponding to the n + 1th tissue category, and the pixel points of the adjacent frame image are input into the n + 1th self-supervised feature extraction unit, and the n + 1th self-supervised feature is output. The n + 1th self-supervised feature is the spatio-temporal information that fuses the target video and the pixel points corresponding to the 1st to n + 1th tissue categories.

[0059] According to an embodiment of the present disclosure, the nth self-supervised feature is input into the n + 1th feature buffer unit, and the n + 1th buffer feature is output. The n + 1th buffer feature is the long-term temporal dependence information of the nth self-supervised feature that is dynamically cached and updated.

[0060] According to an embodiment of the present disclosure, for the Nth feature extraction module, the (N - 1)th self-supervised feature, the (N - 1)th buffered feature, the pixel points of the target frame image corresponding to the Nth tissue category, and the pixel points of the adjacent frame image are input into the Nth self-supervised feature extraction unit, and the Nth self-supervised feature is output. The Nth self-supervised feature is the spatio-temporal information that fuses the target video and the pixel points corresponding to the 1st to Nth tissue categories.

[0061] According to an embodiment of the present disclosure, the (N - 1)th self-supervised feature is input into the Nth feature buffer unit, and the Nth buffered feature is output. The Nth buffered feature is the long-term temporal dependence information of the dynamically cached and updated (N - 1)th self-supervised feature.

[0062] According to an embodiment of the present disclosure, each current feature extraction module receives the self-supervised feature and the buffered feature output by the previous feature extraction module: the self-supervised feature and the buffered feature output by the previous feature extraction module are input into the current self-supervised feature extraction unit together, and at the same time, the self-supervised feature output by the previous feature extraction module is separately input into the current feature buffer unit. The feature buffer unit dynamically updates and stores the self-supervised feature through the Long Short Term Memory (LSTM) mechanism to ensure the effective transmission of temporal information.

[0063] According to an embodiment of the present disclosure, the encoder serves as the backbone network of the self-supervised feature extraction unit and encodes the features of the target video using the global attention mechanism; a convolutional neural network is used as an auxiliary network to extract local spatial features from the input target video. Through local convolution operations, the CNN can efficiently capture the local texture information of the target video and combine it with the global information of the Transformer to obtain a more diverse feature representation; whenever new pixel point information is input, the LSTM network in the feature buffer unit will update its internal state to better reflect the changes occurring in the target video. Thus, the feature cache unit can effectively maintain and transmit the long-term temporal dependence information in the target video, thereby enhancing the long-term memory and temporal modeling ability of the entire feature extraction module for the target video.

[0064] According to an embodiment of the present disclosure, the classification module is used to predict the lesion type of the processed spatio-temporal features, and the type result corresponding to the lymph node region in the target frame image is obtained as follows: self-supervised feature extraction is performed on the Nth self-supervised feature and the Nth buffered feature to obtain the target feature; the classification module is used to predict the lesion type of the target feature to obtain the type result corresponding to the lymph node region in the target frame image.

[0065] According to an embodiment of the present disclosure, self-supervised feature extraction is performed on the Nth self-supervised feature and the Nth buffered feature to obtain the target feature.

[0066] According to an embodiment of the present disclosure, the target feature integrates the spatio-temporal information of the target video, the pixel points corresponding to the first to Nth tissue categories, and the temporal dependence information.

[0067] According to an embodiment of the present disclosure, the classification module is a classification head, which is responsible for outputting the final type result, and further determining the benignity and malignancy of the lymph nodes. The classification module is used to predict the lesion type of the target feature, and obtain the type result corresponding to the lymph node region in the target frame image.

[0068] According to an embodiment of the present disclosure, the type result may be the lesion type corresponding to the lymph node region, such as a benign category or a malignant category.

[0069] According to an embodiment of the present disclosure, the classification module includes a fully connected unit and an activation unit. Using the classification module to predict the lesion type of the target feature, obtaining the type result corresponding to the lymph node region in the target frame image includes: processing the target feature using the fully connected unit to obtain a fully connected feature; processing the fully connected feature using the activation unit to obtain the type result corresponding to the lymph node region in the target frame image.

[0070] According to an embodiment of the present disclosure, the classification module includes 6 fully connected units. The first 5 fully connected units are connected to a non-linear activation unit (Rectified Linear Unit, ReLU), and the activation unit can enhance the ability of non-linear feature representation. The last fully connected layer is connected to a classification activation unit (Sigmoid), which maps the fully connected feature to the category space and generates the probability of each lesion type.

[0071] According to an embodiment of the present disclosure, the target feature is input into the first fully connected unit to obtain the first fully connected feature; the first fully connected feature is input into the first non-linear activation unit to obtain the first activation feature.

[0072] According to an embodiment of the present disclosure, the first activation feature is input into the second fully connected unit to obtain the second fully connected feature; the second fully connected feature is input into the second non-linear activation unit to obtain the second activation feature.

[0073] According to an embodiment of the present disclosure, the activation unit of the upper layer is connected to the fully connected unit of the lower layer. The fifth activation feature is input into the sixth fully connected unit to obtain the sixth fully connected feature; the sixth fully connected feature is input into the classification activation unit to obtain the type result corresponding to the lymph node region in the target frame image.

[0074] According to an embodiment of the present disclosure, the classification model includes a classification module and a feature extraction module corresponding to each tissue category. The classification model is trained based on the following operations: obtaining training samples, where the training samples include sample target videos and sample labels within a preset time. Among them, the sample target videos include sample target frame images and a plurality of sample adjacent frame images having an adjacent temporal relationship with the sample target frame images, and the sample labels represent the lesion types of the lymph node regions in the sample target frame images; performing self-supervised feature extraction on the sample target videos to obtain sample initial features; respectively performing tissue category recognition on each pixel point of the sample target frame images and the sample adjacent frame images to obtain the tissue categories of the respective pixel points in the sample target frame images and the tissue categories of the respective pixel points in the sample adjacent frame images; using the to-be-trained classification model to process the sample initial features, the pixel points of the sample target frame images corresponding to the tissue categories, and the pixel points of the sample adjacent frame images to obtain the sample type results corresponding to the lymph node regions in the sample target frame images; training the to-be-trained classification model according to the sample initial features, the sample type results, and the sample labels to obtain the trained classification model.

[0075] According to an embodiment of the present disclosure, the training samples include a large number of sample target videos within a preset time. The sample target videos are extracted from the annotated intraoperative gastric cancer videos in the database. Each sample target video contains annotation information of different tissue categories (such as stomach, gallbladder, liver, peritoneum, intestine, connective tissue) and lymph node regions. The sample target frame images and the sample adjacent frame images are uniformly adjusted to a unified size, and the sample target frame images are annotated to mark the lymph node regions and the sample labels, and the sample labels are the lymph node benign and malignant information.

[0076] According to an embodiment of the present disclosure, in order to improve the robustness and generalization ability of the classification model, data augmentation is performed on the sample target frame images and the sample adjacent frame images during the training process, including operations such as random rotation, flipping, cropping, color jittering, scaling, and translation.

[0077] According to an embodiment of the present disclosure, self-supervised feature extraction is performed on the sample target videos to obtain sample initial features; respectively performing tissue category recognition on each pixel point of the sample target frame images and the sample adjacent frame images to obtain the tissue categories of the respective pixel points in the sample target frame images and the tissue categories of the respective pixel points in the sample adjacent frame images.

[0078] According to an embodiment of the present disclosure, the feature extraction modules in the classification model to be trained are sorted according to a preset rule. During training, they are respectively input into different feature extraction modules layer by layer in the order of the number of occurrences of sample pixel points corresponding to each tissue category such as the stomach, gallbladder, liver, peritoneum, intestine, and connective tissue from more to less. Spatiotemporal features are extracted through the self-supervised feature extraction unit, and the spatiotemporal features are transmitted to the feature cache unit for dynamic update, retaining the temporal information, and finally the updated features are passed into the classification module for lesion type classification.

[0079] According to an embodiment of the present disclosure, the sample initial features, sample type results, and sample labels are processed according to the loss function to obtain a loss value, and the parameters of the classification model to be trained are updated by backpropagation based on the loss value and a preset performance threshold. In addition, accuracy, recall rate, and precision can also be used as evaluation indicators, especially paying attention to the confusion matrix in the classification task to analyze the classification effect of the classification model on lymph nodes with different lesion types.

[0080] According to an embodiment of the present disclosure, during the training process, an Adaptive Moment Estimation (Adam) optimizer is used as the optimizer of the classification model to accelerate convergence and avoid the problems of gradient vanishing or explosion. The learning rate uses an adaptive adjustment strategy, with the initial learning rate set to 1e-4 and gradually decaying to 1e-6 as the training progresses, which is achieved through a learning rate scheduler. The weight decay factor is set to 3e-5 to reduce the overfitting phenomenon. During the optimization process, a batch size of 8 is used, and the number of training times per round is adjusted according to the size of the training samples. Usually, 120 rounds of training are performed to ensure the convergence of the classification model.

[0081] According to an embodiment of the present disclosure, the sample initial features include multiple sub-features; among them, according to the sample initial features, sample type results, and sample labels, training the classification model to be trained to obtain a trained classification model includes: calculating the loss value between the sample type result and the sample label using the classification loss function to obtain the classification loss value; calculating the loss value between sub-features of every two adjacent time series using the temporal dependence loss function to obtain the temporal dependence loss value; obtaining the target loss value according to the classification loss value and the temporal dependence loss value; and training the classification model to be trained according to the target loss value to obtain a trained classification model.

[0082] According to an embodiment of the present disclosure, the sample initial features include multiple sub-features, and the multiple sub-features correspond to the spatiotemporal features of the target frame image and multiple sample adjacent frame images respectively.

[0083] According to an embodiment of the present disclosure, the classification loss function is used to guide the classification model to perform benign and malignant classification of the lesion types in the lymph node region based on the difference between the sample label and the sample type result.

[0084] In one embodiment, the classification loss value is as shown in formula (1):

[0085] (1);

[0086] where K represents the total number of training samples, represents the sample label of the k-th sample, represents the sample type result of the k-th sample, represents the logarithmic function.

[0087] According to an embodiment of the present disclosure, the temporal dependence loss function is used to perform contrastive learning based on the features of the front and rear frames. The smaller the difference between the features, the smaller the temporal dependence loss value, thereby ensuring that the initial features of self-supervised feature extraction can retain long-term temporal information and consistency, and avoiding information loss.

[0088] In one embodiment, the temporal dependence loss value is as shown in formula (2):

[0089] (2);

[0090] where T represents the total number of frame image pairs of the target frame image and the adjacent frame image in the target video, represents the sub-feature corresponding to the t-th frame image, represents the sub-feature corresponding to the (t + 1)-th frame image, represents the similarity function.

[0091] In one embodiment, the target loss value is as shown in formula (3):

[0092] (3);

[0093] where represents the classification loss value, represents the temporal dependence loss value, represents a hyperparameter used to control the balance between the classification loss and the temporal dependence loss.

[0094] According to an embodiment of the present disclosure, the obtained target loss value can be used to adjust the parameters of the classification model to be trained until the target loss value reaches the convergence condition, stop training, and obtain the trained classification model.

[0095] Figure 2 Shows an example schematic diagram of a classification model according to an embodiment of the present disclosure.

[0096] As Figure 2As shown, the classification model includes a first self-supervised feature extraction module, a second self-supervised feature extraction module, feature extraction modules corresponding to each of the 6 tissue categories, and a classification module. Each feature extraction module corresponding to a tissue category includes a feature buffer unit and a self-supervised feature extraction unit.

[0097] According to an embodiment of the present disclosure, the target video is input into the first self-supervised feature extraction module to obtain initial features; in the first feature extraction module, the initial features, the pixel points of the target frame image corresponding to the first tissue category, and the pixel points of the adjacent frame images are input into the first self-supervised feature extraction unit to output the first self-supervised feature, and at the same time, the initial features are input into the first feature buffer unit to output the first buffered feature; in the second feature extraction module, the first self-supervised feature, the first buffered feature, the pixel points of the target frame image corresponding to the second tissue category, and the pixel points of the adjacent frame images are input into the second self-supervised feature extraction unit to output the second self-supervised feature, and at the same time, the first self-supervised feature is input into the second feature buffer unit to output the second buffered feature. The feature extraction processes in the 3rd, 4th, 5th, and 6th feature extraction modules are not described in detail here.

[0098] According to an embodiment of the present disclosure, finally, the sixth self-supervised feature and the sixth buffered feature are input into the second self-supervised feature extraction module to obtain target features, and then the classification module is used to predict the lesion type of the target features to obtain the type result corresponding to the lymph node region in the target frame image.

[0099] According to an embodiment of the present disclosure, the tissue category of each pixel point in the target frame image is obtained by identifying the tissue category of each pixel point in the target frame image based on a tissue category model. The tissue category model includes an embedding module, an encoding module, a feature pyramid module, a decoding module, and a segmentation module; wherein, identifying the tissue category of each pixel point in the target frame image to obtain the tissue category of each pixel point in the target frame image includes: using the embedding module to process each pixel point of the target frame image to obtain mapping features; using the encoding module to process the mapping features to obtain encoded features; using the feature pyramid module to process the encoded features to obtain multi-scale features; using the decoding module to process the multi-scale features and pixel point position information to obtain decoded features; using the segmentation module to process the decoded features to obtain the tissue category of the pixel points.

[0100] According to an embodiment of the present disclosure, the embedding module is responsible for performing position encoding and pixel encoding on each pixel point in the target frame image, and then mapping the features to a high-dimensional space to obtain mapping features.

[0101] According to an embodiment of the present disclosure, the encoding module can be constructed based on an encoder (Transformer).

[0102] According to an embodiment of the present disclosure, the mapping features are processed by an encoding module to obtain encoded features, and the encoded features are low-dimensional information.

[0103] According to an embodiment of the present disclosure, a Feature Pyramid Network (FPN) consists of multiple levels. The lower levels extract more detailed information, while the higher levels extract more global semantic information. The detailed information and global semantic information are fused layer by layer to form multi-scale features. Each level is composed of a convolutional layer, and deformable convolutions can be used in the convolutional layer.

[0104] According to an embodiment of the present disclosure, the encoded features are input into the feature pyramid module to obtain multi-scale features, which fuse the pixel detail information and global semantic information of the target frame image.

[0105] According to an embodiment of the present disclosure, the decoding module can be constructed based on a decoder, and the structures of the encoding module and the decoding module are symmetric.

[0106] According to an embodiment of the present disclosure, the decoding module gradually decodes the multi-scale features and pixel position information to obtain decoded features, which are high-dimensional mask information after restoring the dimension.

[0107] According to an embodiment of the present disclosure, the segmentation module can be constructed based on a convolutional network. First, it processes the decoded features to output mask features with the same size as the target frame image, and then based on a 3x3 convolutional kernel, it performs pixel-level prediction to generate the probability value of each pixel point, where the probability value indicates whether the pixel point belongs to a certain organ or tissue, thereby obtaining the tissue category of each pixel point.

[0108] According to an embodiment of the present disclosure, the feature pyramid module performs multi-scale fusion on the encoded features output by the encoding module. Through the top-down path and the bottom-up path, the low-level detailed features and high-level semantic features are fused to better process organs and tissues at different scales, especially having advantages when dealing with the gastric perigastric lymph nodes with large anatomical structure changes.

[0109] According to an embodiment of the present disclosure, the encoding module includes a self-attention unit, a feed-forward neural unit, and a normalization unit. Among them, processing the mapping features by the encoding module to obtain encoded features includes: processing the mapping features by the self-attention unit to obtain attention features; processing the attention features by the feed-forward neural unit to obtain feed-forward neural features; performing residual connection on the attention features and the feed-forward neural features to obtain residual features; processing the residual features by the normalization unit to obtain encoded features.

[0110] According to an embodiment of the present disclosure, the encoding module may include six blocks, each of which includes a self-attention unit, a feed-forward neural unit, and a normalization unit (Layer Normalization). A multi-head self-attention mechanism is embedded in the self-attention unit, and each feed-forward neural unit includes two linear layers and a non-linear activation function ReLU.

[0111] According to an embodiment of the present disclosure, the self-attention unit is used to process the mapped features to obtain attention features. The attention features are the initially captured spatio-temporal information.

[0112] According to an embodiment of the present disclosure, in the feed-forward neural unit, the attention features are sequentially input into two linear layers and a non-linear activation function to obtain feed-forward neural features. The feed-forward neural features are the enhanced spatio-temporal information.

[0113] According to an embodiment of the present disclosure, residual connection is performed on the attention features and the feed-forward neural features to obtain residual features. The normalization unit is used to process the residual features to obtain encoded features.

[0114] According to an embodiment of the present disclosure, the above-mentioned tissue category model can be used to perform tissue category recognition on each pixel point of each adjacent frame image, so as to obtain the tissue category of each pixel point in each adjacent frame image. In the tissue category model, the dependency relationships of different regions of the target frame image and the adjacent frame images can be calculated in parallel through the multi-head self-attention mechanism, so that the features at each position can integrate information from other positions, and then processed by the feed-forward neural unit to further enhance the feature representation.

[0115] Figure 3 An example schematic diagram of the tissue category model according to an embodiment of the present disclosure is shown.

[0116] As Figure 3 shown, the tissue category model includes an embedding module, an encoding module, a feature pyramid module, a decoding module, and a segmentation module. The embedding module is used to process each pixel point of the target frame image to obtain mapped features; the encoding module (Transformer) is used to process the mapped features to obtain encoded features; the feature pyramid module (Feature Pyramid Network, FPN) is used to process the encoded features to obtain multi-scale features; the decoding module (Transformer) is used to process the multi-scale features and the pixel point position information to obtain decoded features; the segmentation module is used to process the decoded features to obtain the tissue category of the pixel points.

[0117] Figure 4 An example schematic diagram of obtaining the type result corresponding to the lymph node region according to an embodiment of the present disclosure is shown.

[0118] As Figure 4As shown, the target video includes a target frame image and a plurality of adjacent frame images having an adjacent temporal relationship with the target frame image. The target frame image and the plurality of adjacent frame images are processed using an organ category model to obtain the organ categories of the respective pixel points in the target frame image and each adjacent frame image; then, based on the organ categories of the pixel points, the pixel points of the target frame image and each adjacent frame image are input into a classification model to obtain the type result corresponding to the lymph node region in the target frame image.

[0119] According to an embodiment of the present disclosure, to train the organ category model: a computer-aided manual annotation dataset including organ categories such as stomach, gallbladder, liver, peritoneum, intestine, and connective tissue is constructed. Each frame image in the dataset is annotated by a professional doctor to ensure the data quality and annotation accuracy. To improve the generalization ability of the organ category model and avoid overfitting, the sizes of all frame images are uniformly adjusted to 1024×1024. At the same time, data augmentation strategies such as random flipping, rotation, and cropping are adopted to increase the data diversity, so that the organ category model has stronger robustness when processing frame images at different angles, sizes, or positions.

[0120] According to an embodiment of the present disclosure, before starting to train the organ category model, to ensure that the organ category model can converge effectively, Xavier initialization is used to initialize the weights. Xavier initialization initializes the weights using a uniform distribution by considering the number of input and output nodes of each module. In addition, He initialization can be used for the ReLU activation function to further improve the training effect of the feedforward neural units.

[0121] According to an embodiment of the present disclosure, the training of the organ category model will be carried out within 60 epochs, and the training method with a batch size of 8 is adopted in each epoch. The Dice loss is selected as the loss function to ensure the stability and efficiency during the training process.

[0122] In one embodiment, the Dice loss function is as shown in (4):

[0123] (4);

[0124] where X represents the organ category of the pixel points output by the organ category model during training, Y represents the sample category label of the pixel points, |X∩Y| represents the intersection of the organ category of the pixel points and the sample category label in the training dataset, and |·| is the operation of calculating the number of samples in the calculation object.

[0125] According to an embodiment of the present disclosure, the Adam optimizer is used to adaptively adjust the learning rate of each parameter. The initial learning rate is set to 1e-4, and the weight decay factor is 4e-5, thereby preventing overfitting through regularization.

[0126] Based on the above lymph node classification method, the present disclosure also provides a lymph node classification device. The following will be combined with Figure 5 to describe this device in detail.

[0127] Figure 5 Fig. shows a structural block diagram of a lymph node classification device according to an embodiment of the present disclosure.

[0128] As Figure 5 shown, the lymph node classification device 500 of this embodiment includes a first feature extraction module 510, a first recognition module 520, a second recognition module 530, a second feature extraction module 540, and a prediction module 550.

[0129] The first feature extraction module 510 is configured to perform self-supervised feature extraction on a target video within a preset time to obtain initial features, where the target video includes a target frame image and a plurality of adjacent frame images having an adjacent temporal relationship with the target frame image, and the lymph node region is marked in the target frame image. In one embodiment, the first feature extraction module 510 may be configured to perform the operation S110 described above, which will not be elaborated here.

[0130] The first recognition module 520 is configured to perform tissue category recognition on each pixel point of the target frame image to obtain the tissue category of each pixel point in the target frame image. In one embodiment, the first recognition module 520 may be configured to perform the operation S120 described above, which will not be elaborated here.

[0131] The second recognition module 530 is configured to perform tissue category recognition on each pixel point of each adjacent frame image to obtain the tissue category of each pixel point in each adjacent frame image. In one embodiment, the second recognition module 530 may be configured to perform the operation S130 described above, which will not be elaborated here.

[0132] The second feature extraction module 540 is configured to perform spatio-temporal feature processing on the initial features, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame images by using the feature extraction module corresponding to each tissue category to obtain the processed spatio-temporal features. In one embodiment, the second feature extraction module 540 may be configured to perform the operation S140 described above, which will not be elaborated here.

[0133] The prediction module 550 is configured to use a classification module to perform lesion type prediction on the processed spatio-temporal features to obtain a type result corresponding to the lymph node region in the target frame image. In one embodiment, the prediction module 550 may be configured to perform the operation S150 described above, which will not be elaborated here.

[0134] According to an embodiment of the present disclosure, the second feature extraction module 540 includes a first feature extraction sub-module and a second feature extraction sub-module.

[0135] The first feature extraction sub-module is configured to perform self-supervised feature processing on the pixel points of the target frame image corresponding to the (n + 1)-th tissue category, the pixel points of the adjacent frame image, the n-th self-supervised feature, and the n-th buffered feature by using the (n + 1)-th self-supervised feature extraction unit, so as to obtain the (n + 1)-th self-supervised feature. Wherein, the n-th buffered feature is obtained by performing temporal dynamic adjustment on the (n - 1)-th self-supervised feature by using the n-th buffer unit. In the case of n = 1, the first self-supervised feature is obtained by performing self-supervised feature processing on the initial feature, the pixel points of the target frame image corresponding to the first tissue category, and the pixel points of the adjacent frame image, and the first buffered feature is obtained by performing temporal dynamic adjustment on the initial feature by using the first feature buffer unit.

[0136] The second feature extraction sub-module is configured to perform temporal dynamic adjustment processing on the n-th self-supervised feature by using the (n + 1)-th buffered feature unit, so as to obtain the (n + 1)-th buffered feature.

[0137] According to an embodiment of the present disclosure, the prediction module 550 includes a first prediction sub-module and a second prediction sub-module.

[0138] The first prediction sub-module is configured to perform self-supervised feature extraction on the N-th self-supervised feature and the N-th buffered feature to obtain a target feature.

[0139] The second prediction sub-module is configured to use the classification module to perform lesion type prediction on the target feature, so as to obtain a type result corresponding to the lymph node region in the target frame image.

[0140] According to an embodiment of the present disclosure, the second prediction sub-module includes a first prediction unit and a second prediction unit.

[0141] The first prediction unit is configured to process the target feature by using a fully connected unit to obtain a fully connected feature.

[0142] The second prediction unit is configured to process the fully connected feature by using an activation unit to obtain a type result corresponding to the lymph node region in the target frame image.

[0143] According to an embodiment of the present disclosure, the first recognition module includes a first recognition sub-module, a second recognition sub-module, a third recognition sub-module, a fourth recognition sub-module, and a fifth recognition sub-module.

[0144] The first recognition sub-module is configured to process each pixel point of the target frame image by using an embedding module to obtain a mapped feature.

[0145] The second recognition sub-module is configured to process the mapped feature by using an encoding module to obtain an encoded feature.

[0146] The third recognition sub-module is configured to process the encoded feature by using a feature pyramid module to obtain multi-scale features.

[0147] The fourth recognition sub-module is configured to use the decoding module to process the multi-scale features and the pixel position information to obtain decoded features.

[0148] The fifth recognition sub-module is configured to use the segmentation module to process the decoded features to obtain the tissue categories of the pixels.

[0149] According to an embodiment of the present disclosure, the second recognition sub-module includes a first recognition unit, a second recognition unit, a third recognition unit, and a fourth recognition unit.

[0150] The first recognition unit is configured to use the self-attention unit to process the mapped features to obtain attention features.

[0151] The second recognition unit is configured to use the feed-forward neural unit to process the attention features to obtain feed-forward neural features.

[0152] The third recognition unit is configured to perform a residual connection on the attention features and the feed-forward neural features to obtain residual features.

[0153] The fourth recognition unit is configured to use the normalization unit to process the residual features to obtain encoded features.

[0154] According to an embodiment of the present disclosure, any plurality of modules among the modules, sub-modules, units, and sub-units may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the modules, sub-modules, units, and sub-units may be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware by integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware or in any appropriate combination of several of them. Alternatively, at least one of the modules, sub-modules, units, and sub-units may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0155] Figure 6 A block diagram of an electronic device suitable for implementing the lymph node classification method according to an embodiment of the present disclosure is schematically shown.

[0156] Figure 6 The electronic device shown is only an example and should not impose any limitation on the functions and the scope of use of the embodiments of the present disclosure.

[0157] AsFigure 6 As shown, the computer electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 can include, for example, a general-purpose microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), and so on. The processor 601 can also include on-board memory for caching purposes. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0158] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in one or more memories.

[0159] According to an embodiment of the present disclosure, the electronic device 600 can further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 can further include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.

[0160] According to an embodiment of the present disclosure, the method flow according to the embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0161] The present disclosure also provides a computer-readable storage medium, which can be included in the device / device / system described in the above embodiment; or can exist alone without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, a lymph node classification method according to an embodiment of the present disclosure is implemented.

[0162] According to an embodiment of the present disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. For example, it can include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device.

[0163] For example, according to an embodiment of the present disclosure, the computer-readable storage medium can include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.

[0164] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method provided by the embodiment of the present disclosure. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the lymph node classification method provided by the embodiment of the present disclosure.

[0165] When the computer program is executed by the processor 601, the above functions defined in the system / device of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, devices, modules, units, etc. can be implemented by computer program modules.

[0166] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 609, and / or installed from the removable medium 611. The program code included in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0167] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions. Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0169] The above describes embodiments according to the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A method for classifying lymph nodes, characterized in that, The method includes: Performing self-supervised feature extraction on a target video within a preset time to obtain initial features, where the target video includes a target frame image and a plurality of adjacent frame images having an adjacent temporal relationship with the target frame image, and the lymph node region is marked in the target frame image; Performing tissue category recognition on each pixel point of the target frame image to obtain the tissue category of each pixel point in the target frame image; Performing tissue category recognition on each pixel point of each adjacent frame image to obtain the tissue category of each pixel point in each adjacent frame image; Using a feature extraction module corresponding to each tissue category to perform spatio-temporal feature processing on the initial features, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame images to obtain processed spatio-temporal features; Using a classification module to perform lesion type prediction on the processed spatio-temporal features to obtain a type result corresponding to the lymph node region in the target frame image.

2. The method according to claim 1, wherein The feature extraction module includes a feature buffer unit and a self-supervised feature extraction unit. The feature extraction module has a one-to-one correspondence with the tissue category. The number of the feature extraction modules and the number of the tissue categories are both N. The processed spatio-temporal features include buffered features and self-supervised features, 1 N; Wherein, the using a feature extraction module corresponding to each tissue category to perform spatio-temporal feature processing on the initial features, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame images to obtain processed spatio-temporal features includes: Using the (n + 1)th self-supervised feature extraction unit to perform self-supervised feature processing on the pixel points of the target frame image and the adjacent frame images corresponding to the (n + 1)th tissue category, the nth self-supervised feature, and the nth buffered feature to obtain the (n + 1)th self-supervised feature, where the nth buffered feature is obtained by using the nth buffer unit to perform temporal dynamic adjustment on the (n - 1)th self-supervised feature. In the case of n = 1, the 1st self-supervised feature is obtained by performing self-supervised feature processing on the initial features, the pixel points of the target frame image corresponding to the 1st tissue category, and the adjacent frame images, and the 1st buffered feature is obtained by using the 1st feature buffer unit to perform temporal dynamic adjustment on the initial features; Using the (n + 1)th buffered feature unit to perform temporal dynamic adjustment processing on the nth self-supervised feature to obtain the (n + 1)th buffered feature.

3. The method according to claim 2, wherein The using a classification module to perform lesion type prediction on the processed spatio-temporal features to obtain a type result corresponding to the lymph node region in the target frame image includes: Performing self-supervised feature extraction on the Nth self-supervised feature and the Nth buffered feature to obtain target features; Using the classification module to perform lesion type prediction on the target features to obtain a type result corresponding to the lymph node region in the target frame image.

4. The method according to claim 3, wherein The classification module includes a fully connected unit and an activation unit. The using the classification module to perform lesion type prediction on the target features to obtain a type result corresponding to the lymph node region in the target frame image includes: Using the fully connected unit to process the target features to obtain fully connected features; Using the activation unit to process the fully connected features to obtain a type result corresponding to the lymph node region in the target frame image.

5. The method according to claim 1, wherein The tissue category of each pixel point in the target frame image is obtained by performing tissue category recognition on each pixel point in the target frame image based on a tissue category model, and the tissue category model includes an embedding module, an encoding module, a feature pyramid module, a decoding module, and a segmentation module; Among them, performing tissue category recognition on each pixel point of the target frame image to obtain the tissue category of each pixel point in the target frame image includes: Processing each pixel point of the target frame image by using the embedding module to obtain a mapped feature; Processing the mapped feature by using the encoding module to obtain an encoded feature; Processing the encoded feature by using the feature pyramid module to obtain multi-scale features; Processing the multi-scale features and pixel point position information by using the decoding module to obtain a decoded feature; Processing the decoded feature by using the segmentation module to obtain the tissue category of the pixel point.

6. The method according to claim 5, characterized in that The encoding module includes a self-attention unit, a feed-forward neural unit, and a normalization unit; Among them, processing the mapped feature by using the encoding module to obtain an encoded feature includes: Processing the mapped feature by using the self-attention unit to obtain an attention feature; Processing the attention feature by using the feed-forward neural unit to obtain a feed-forward neural feature; Performing a residual connection on the attention feature and the feed-forward neural feature to obtain a residual feature; Processing the residual feature by using the normalization unit to obtain the encoded feature.

7. The method according to claim 1, characterized in that, The classification model includes the classification module and a feature extraction module corresponding to each tissue category, and the classification model is trained based on the following operations: Obtaining a training sample, where the training sample includes a sample target video and a sample label within a preset time. Among them, the sample target video includes a sample target frame image and a plurality of sample adjacent frame images having an adjacent time sequence relationship with the sample target frame image, and the sample label represents the lesion type of the lymph node region in the sample target frame image; Performing self-supervised feature extraction on the sample target video to obtain a sample initial feature; Performing tissue category recognition on each pixel point of the sample target frame image and the sample adjacent frame images respectively to obtain the tissue category of each pixel point in the sample target frame image and the tissue category of each pixel point in the sample adjacent frame images; Processing the sample initial feature, the pixel points of the sample target frame image corresponding to the tissue category, and the pixel points of the sample adjacent frame images by using the classification model to be trained to obtain a sample type result corresponding to the lymph node region in the sample target frame image; Training the classification model to be trained according to the sample initial feature, the sample type result, and the sample label to obtain a trained classification model.

8. The method according to claim 7, wherein The sample initial feature includes a plurality of sub-features; Among them, training the classification model to be trained according to the sample initial feature, the sample type result, and the sample label to obtain a trained classification model includes: Calculating a loss value between the sample type result and the sample label by using a classification loss function to obtain a classification loss value; Calculate the loss value between the sub-features of every two adjacent time series by using a time series dependent loss function to obtain a time series dependent loss value; Obtain a target loss value according to the classification loss value and the time series dependent loss value; Train the classification model to be trained according to the target loss value to obtain a trained classification model.

9. A lymph node classification device, characterized in that The device includes: A first feature extraction module, configured to perform self-supervised feature extraction on a target video within a preset time to obtain initial features, where the target video includes a target frame image and a plurality of adjacent frame images having an adjacent time series relationship with the target frame image, and the lymph node area is marked in the target frame image; A first recognition module, configured to perform tissue category recognition on each pixel point of the target frame image to obtain the tissue category of each pixel point in the target frame image; A second recognition module, configured to perform tissue category recognition on each pixel point of each adjacent frame image to obtain the tissue category of each pixel point in each adjacent frame image; A second feature extraction module, configured to perform spatio-temporal feature processing on the initial features, the pixel points of the target frame image corresponding to the tissue category, and the pixel points of the adjacent frame images by using a feature extraction module corresponding to each tissue category to obtain processed spatio-temporal features; A prediction module, configured to use a classification module to predict a lesion type of the processed spatio-temporal features to obtain a type result corresponding to the lymph node area in the target frame image.

10. An electronic device, comprising: One or more processors; A memory, configured to store one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.