Method and system for intraoperative identification of blood vessels, lymph nodes and nerves

By constructing a recognition model using an FPN neural network and a multi-feature extraction network, blood vessels, lymph nodes, and nerves in surgical videos can be identified and visualized in real time, solving the problem of structural recognition during surgery, reducing the risk of injury, and improving surgical outcomes.

CN116612411BActive Publication Date: 2026-06-02CHENGDU WITHAI INNOVATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU WITHAI INNOVATION TECH CO LTD
Filing Date
2023-05-05
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

During surgery, the removal and damage of blood vessels, lymph nodes, and nerves are common, leading to postoperative complications that affect the patient's recovery and quality of life.

Method used

An initial recognition model was constructed using an FPN neural network and a multi-feature extraction neural network. The model was then trained and optimized using a database of blood vessel, lymph node, and nerve samples. Surgical videos were acquired in real time and imported into the recognition model to identify and delineate the contours of blood vessels, lymph nodes, and nerves. Combined with visualization, this helped the surgeon identify and remove the target structures.

Benefits of technology

It enables precise identification of blood vessels, lymph nodes, and nerves in the laparoscopic surgical field, reduces surgical risks, assists surgeons in the rational removal and avoidance of unnecessary structures, and improves the success rate of surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612411B_ABST
    Figure CN116612411B_ABST
Patent Text Reader

Abstract

The application discloses a method and system for identification of blood vessels, lymph nodes and nerves in surgery, and relates to the technical field of computers, which comprises the following steps: S1, acquiring a surgery video, marking and establishing a blood vessel sample database, a lymph sample database and a nerve sample database; S2, constructing an initial identification model; S3, training the initial identification model to obtain a blood vessel identification model, a lymph identification model and a nerve identification model; S4, collecting the surgery video, and identifying and outlining the contours of blood vessels, lymph and nerves; and S5, visually displaying the surgery video and the contours to a surgeon; the system comprises a collection module, a central processing unit and a display module; the blood vessels, lymph and nerve structures under the laparoscopic surgery visual field are identified in real time by using an artificial intelligence computer model, and the visual display function is combined, so that the surgeon is provided with guidance for resection and avoidance of corresponding blood vessels, lymph nodes and nerves, the surgeon is guided to reasonably identify and resect blood vessels and lymph nodes, the blood vessels and lymph nodes that do not need to be resected are avoided, and the surgery is smoothly assisted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method and system for identifying blood vessels, lymph nodes and nerves during surgery. Background Technology

[0002] During surgery, due to the complexity of anatomical structures and the influence of the surgeon's subjective factors, the resection and damage of blood vessels, lymph nodes, and nerves are relatively common. Improper damage or resection errors are highly likely to lead to postoperative complications, hindering patient recovery and affecting their quality of life. Summary of the Invention

[0003] The purpose of this invention is to design a method and system for identifying blood vessels, lymph nodes and nerves during surgery in order to solve the above problems.

[0004] The present invention achieves the above objectives through the following technical solutions:

[0005] Methods for identifying blood vessels, lymph nodes, and nerves during surgery include:

[0006] S1. Acquire surgical videos of various procedures, mark the surgical stages and the blood vessels, lymph nodes and nerves in the field of vision, and establish a blood vessel sample database, a lymph node sample database and a nerve sample database;

[0007] S2. Construct three initial recognition models. Each initial recognition model includes an FPN neural network and a multi-feature extraction neural network. The FPN neural network is a bottom-up, top-down, and horizontally connected network structure. The multi-feature extraction neural network includes four pooling attention layers. The four pooling attention layers are connected to form a bottom-up downsampling structure. The FPN neural network has four bottom-up layers. One pooling attention layer corresponds to one bottom-up downsampling layer. The outputs of the previous pooling attention layer and the previous downsampling layer are used as the inputs of the next downsampling layer.

[0008] S3. Import the initial recognition model into the vascular sample database and train and optimize it to obtain the vascular recognition model; import the initial recognition model into the lymph node sample database and train and optimize it to obtain the lymph node recognition model; import the initial recognition model into the neural sample database and train and optimize it to obtain the neural recognition model.

[0009] S4. Real-time acquisition of surgical videos, and import them into the vascular recognition model, lymph node recognition model and nerve recognition model respectively, to obtain the vascular recognition results, lymph node recognition results and nerve recognition results, and outline the contours of the vascular recognition results, lymph node recognition results and nerve recognition results;

[0010] S5. Visualize and display surgical videos and outlines to the surgeon.

[0011] A system for identifying blood vessels, lymph nodes, and nerves during surgery, including:

[0012] Acquisition module for real-time acquisition of surgical videos;

[0013] Central Processing Unit (CPU); The CPU is used to analyze surgical videos to identify blood vessels, lymph nodes, and nerves, and to annotate the results of blood vessel identification, lymph node identification, and nerve identification in the surgical videos.

[0014] Display module; The display module is used to display the annotated surgical video to the surgeon.

[0015] The beneficial effects of this invention are as follows: It uses an artificial intelligence computer model to identify blood vessels, lymph nodes and nerve structures in the laparoscopic surgical field in real time. At the same time, it combines visualization display function to provide surgeons with guidance on the removal and avoidance of corresponding blood vessels, lymph nodes and nerves. This guides surgeons to reasonably identify and remove blood vessels and lymph nodes and avoid blood vessels and lymph nodes that do not need to be removed, thus assisting the smooth progress of the operation. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the method for identifying blood vessels, lymph nodes, and nerves during surgery according to the present invention.

[0017] Figure 2 This is a schematic diagram of the structure of the recognition model in this invention;

[0018] Figure 3 This is a schematic diagram of the structure of the multi-feature extraction neural network in this invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0020] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0021] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0022] In the description of this invention, it should be understood that the terms "upper," "lower," "inner," "outer," "left," "right," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used to facilitate the description of this invention and to simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0023] Furthermore, the terms "first," "second," etc., are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0024] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, terms such as "set" and "connection" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0025] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0026] like Figure 1 , Figure 2 , Figure 3 As shown, the method for identifying blood vessels, lymph nodes, and nerves during surgery includes:

[0027] S1. Acquire surgical videos of various procedures, mark the surgical stages and the blood vessels, lymph nodes and nerves in the field of vision, and establish a blood vessel sample database, a lymph node sample database and a nerve sample database;

[0028] S2. Construct three initial recognition models. Each initial recognition model includes an FPN neural network and a multi-feature extraction neural network. The FPN neural network is a bottom-up, top-down, and horizontally connected network structure. The multi-feature extraction neural network includes four pooling attention layers. The four pooling attention layers are connected to form a bottom-up downsampling structure. The FPN neural network has four bottom-up layers. One pooling attention layer corresponds to one bottom-up downsampling layer. The outputs of the previous pooling attention layer and the previous downsampling layer are used as the inputs of the next downsampling layer.

[0029] S3. Import the initial recognition model into the vascular sample database and train and optimize it to obtain the vascular recognition model; import the initial recognition model into the lymph node sample database and train and optimize it to obtain the lymph node recognition model; import the initial recognition model into the neural sample database and train and optimize it to obtain the neural recognition model.

[0030] S4. Real-time acquisition of surgical videos, and import them into the vascular recognition model, lymph node recognition model and nerve recognition model respectively, to obtain the vascular recognition results, lymph node recognition results and nerve recognition results, and outline the contours of the vascular recognition results, lymph node recognition results and nerve recognition results;

[0031] S5. Visualize and display surgical videos and outlines to the surgeon.

[0032] Each pooling attention layer consists of a pooling layer and a first self-attention layer. The output of the pooling layer serves as the input to the first self-attention layer, which downsamples the surgical image through local aggregation and calculates global self-attention.

[0033] The multi-feature extraction neural network also includes four second self-attention layers. The second self-attention layers divide the input into non-overlapping windows and compute local self-attention within each window. The output of the second self-attention layers is used as the input of the FPN neural network.

[0034] The multi-feature extraction neural network also includes three hybrid window layers. One hybrid window layer is used to calculate the local attention within a window. Except for the last block of the hybrid window layer, the output of the hybrid window layer is used as the input of the FPN neural network.

[0035] For any input sequence, the pooling attention layer performs a linear projection, resulting in three tensors Q, K, and V, represented as follows: Pooling and attention calculations are performed on the three tensors sequentially, incorporating relative position information into the attention calculation. The attention calculation is represented as follows: The distance calculation between elements is decomposed along the spacetime axis as follows: , where h and w represent the vertical and horizontal directions, respectively.

[0036] A system for identifying blood vessels, lymph nodes, and nerves during surgery, including:

[0037] Acquisition module for real-time acquisition of surgical videos;

[0038] Central Processing Unit (CPU); The CPU is used to analyze surgical videos to identify blood vessels, lymph nodes, and nerves, and to annotate the results of blood vessel identification, lymph node identification, and nerve identification in the surgical videos.

[0039] Display module; The display module is used to display the annotated surgical video to the surgeon.

[0040] The working principle of the method and system for identifying blood vessels, lymph nodes, and nerves during surgery in this invention is as follows:

[0041] Generally, the diameter of arteries and veins does not exceed 3 cm, the diameter of lymph nodes is usually 0.5-2 cm, and the width of nerves does not exceed 1 cm. These three types of targets that need to be identified and detected are all relatively small individuals, usually with a pixel area of ​​no more than 32x32 pixels in the laparoscopic lens. They themselves lack sufficient information and contain insufficient discriminative information. In addition, there is also the problem of imbalanced datasets and difficulty in matching reference boxes.

[0042] To better address these issues, the FPN neural network introduces a bottom-up, top-down network structure. It achieves feature enhancement by fusing features from adjacent layers, demonstrating good performance for targets with insufficient pixels and limited feature richness. A strong baseline is created, focusing attention is enhanced along two axes, and positional information is injected into the pooling attention layer using decomposed positional distances. Pooling residual connections compensate for the impact of the pooling stride on attention calculation. Combined with the FPN neural network, it is used for object detection and instance segmentation.

[0043] The key idea of ​​this algorithm is to achieve the construction of different stages of high-dimensional and low-dimensional visual modeling by expanding the channel width while reducing the resolution, rather than using single-scale blocks. Because upsampling and downsampling are needed to extract and fuse features at different scales, this algorithm proposes pooling attention, such as... Figure 3 As shown. For any input sequence, a linear projection onto it yields three tensors: Q (query), K (key), and V (value). ,

[0044] Q represents the query vector, K represents the vector of the relevance of the queried information to other information, V represents the vector of the queried information, W is the linear projection matrix, X is the self-attention weight, and P is the pooling operator.

[0045] Q, K, and V then undergo pooling, primarily to shorten the length of the K and V sequences. Attention is then calculated based on this pooled sequence. ,

[0046] Where K T Let K be the transpose of matrix K, and D be the vector dimension of the multi-head pooling self-attention processing.

[0047] The first self-attention layer of the multi-feature extraction neural network can be pooled at each step, which can greatly reduce the memory cost and computational load of QKV calculation. This will greatly reduce the hardware and computing power requirements for the system to perform recognition, enabling more medical institutions to operate at low cost.

[0048] Because absolute position encoding only provides location information but ignores the translation invariance of features, changes in absolute position will alter the dependencies between them, even though the relative positions of the two regions remain unchanged. To address this issue, this model incorporates relative position information into the self-attention calculation of the pooling attention layer, which depends only on the relative distance of K: ,in ,

[0049] Where i represents the i-th token in the time dimension, j represents the j-th token in the spatial dimension, R is the relative position code, d is the absolute value of the range of R, and P represents the spatiotemporal position of elements i and j.

[0050] To simplify the calculation and reduce complexity, the algorithm decomposes the distance calculation between elements along the spacetime axis into: ,

[0051] Where h and w represent the vertical and horizontal directions respectively, since this project uses single-frame images of the surgery rather than surgical videos, and t is the time dimension, the t dimension is 0, i.e. .

[0052] After the aforementioned improvements, the structure of the multi-feature extraction neural network can be integrated into the FPN neural network. Structurally, the multi-feature extraction neural network generates multi-scale feature maps in four stages, thus naturally integrating into the FPN neural network used for object detection tasks. The top-down pyramid with lateral connections in the FPN neural network constructs semantically powerful feature maps at all scales, such as... Figure 2 As shown.

[0053] In addition to the pooling attention layer, this algorithm proposes a second self-attention layer to significantly reduce computational and memory complexity. The pooling attention layer is characterized by downsampling through local aggregation while maintaining global self-attention computation, while the second self-attention layer maintains tensor resolution but performs self-attention locally by dividing the input into non-overlapping windows and then computing only the local self-attention within each window. This inherent difference between the two approaches leads to their potential to perform complementary object detection tasks. A hybrid window layer is also proposed to add cross-window connections. The hybrid window layer computes local attention within a window, except for the last block of the last three stages, all of which are fed into the FPN neural network, thus mapping information that includes global context.

[0054] Two attention mechanisms enable the blood vessel recognition model, lymph node recognition model, and neural recognition model to accurately identify targets at various scales. Regardless of how the laparoscope moves during the operation, as long as its features are exposed under the lens, it can be accurately identified. At the same time, it does not require a large amount of computation. The model has only 42.1G floating-point operations per second (FLOPs) and only 56M computer functions (Params). This allows the system to segment and label in real time during the operation, helping medical workers to quickly find blood vessels, lymph nodes, and nerves and reduce surgical risks.

[0055] Artificial intelligence computer models are used to identify blood vessels, lymph nodes, and nerve structures in the laparoscopic surgical field in real time. Combined with visualization display functions, it provides surgeons with guidance on the removal and avoidance of corresponding blood vessels, lymph nodes, and nerves. This guides surgeons to rationally identify and remove blood vessels and lymph nodes, as well as avoid blood vessels and lymph nodes that do not need to be removed, thus assisting in the smooth progress of the operation.

[0056] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.

Claims

1. A method for identifying blood vessels, lymph nodes, and nerves during surgery, characterized in that, include: S1. Acquire surgical videos of various procedures, mark the surgical stages and the blood vessels, lymph nodes and nerves in the field of vision, and establish a blood vessel sample database, a lymph node sample database and a nerve sample database; S2. Construct three initial recognition models. Each initial recognition model includes an FPN neural network and a multi-feature extraction neural network. The FPN neural network has a bottom-up, top-down, and laterally connected network structure. The multi-feature extraction neural network includes four pooling attention layers. These four pooling attention layers are connected to form a bottom-up downsampling structure. The FPN neural network has four bottom-up layers, with one pooling attention layer corresponding to one bottom-up downsampling layer. The outputs of the previous pooling attention layer and the previous downsampling layer are used as the inputs of the next downsampling layer. Each pooling attention layer includes a pooling layer and a first self-attention layer. The output of the first self-attention layer is used as the input of the first self-attention layer. The first self-attention layer downsamples the surgical image through local aggregation and calculates global self-attention. The multi-feature extraction neural network also includes four second self-attention layers. The second self-attention layers divide the input into non-overlapping windows and calculate the local self-attention within each window. The output of the second self-attention layer is used as the input of the FPN neural network. The multi-feature extraction neural network also includes three hybrid window layers. One hybrid window layer is used to calculate the local attention within a window. Except for the last block of the hybrid window layer, the output of the hybrid window layer is used as the input of the FPN neural network. For any input sequence, the pooling attention layer performs a linear projection, resulting in three tensors Q, K, and V, represented as follows: Where Q represents the query vector, K represents the vector representing the relevance of the queried information to other information, V represents the vector of the queried information, W is the linear projection matrix, X is the self-attention weight, and P is the pooling operator; pooling and attention calculations are performed on the three tensors sequentially, and the information of relative position is incorporated into the attention calculation. The attention calculation is expressed as follows: The distance calculation between elements is decomposed along the spacetime axis as follows: Where h and w represent the vertical and horizontal directions, respectively. Where i represents the i-th token in the time dimension, j represents the j-th token in the spatial dimension, R is the relative position code, d is the absolute value of the range of R, p represents the spatiotemporal position of elements i and j, and t is the time dimension, where t is 0. ; S3. Import the initial recognition model into the vascular sample database and train and optimize it to obtain the vascular recognition model; import the initial recognition model into the lymphatic sample database and train and optimize it to obtain the lymphatic recognition model; import the initial recognition model into the neural sample database and train and optimize it to obtain the neural recognition model. S4. Real-time acquisition of surgical videos, and import them into the vascular recognition model, lymph node recognition model and nerve recognition model respectively, to obtain the vascular recognition results, lymph node recognition results and nerve recognition results, and outline the contours of the vascular recognition results, lymph node recognition results and nerve recognition results; S5. Visualize and display surgical videos and outlines to the surgeon.

2. A system for identifying blood vessels, lymph nodes, and nerves during surgery, used to implement the method for identifying blood vessels, lymph nodes, and nerves during surgery as described in claim 1, characterized in that, include: Acquisition module for real-time acquisition of surgical videos; Central Processing Unit (CPU); The CPU is used to analyze surgical videos to identify blood vessels, lymph nodes, and nerves, and to annotate the results of blood vessel identification, lymph node identification, and nerve identification in the surgical videos. Display module; The display module is used to display the annotated surgical video to the surgeon.