Nodule risk stratification method, apparatus, electronic device, and storage medium

CN122223450BActive Publication Date: 2026-09-22脉得智能科技(无锡)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610661325.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-14
Publication Date
2026-09-22
Estimated Expiration
2046-05-14

AI Technical Summary

Technical Problem

[0003]然而,现有人工智能方法普遍依赖分类网络输出的单一恶性概率值,并通过固定阈值判别良恶性

Benefits of technology

[0018]本发明实施例提供的一种结节风险分层方法、装置、电子设备及存储介质,通过将待处理结节超声图像输入预训练编码网络获得第一归一化特征,并基于预构建的类别平衡特征流形库计算其第一局部离散度与第一邻域恶性占比,进而映射至预构建的二维风险图谱以获取对应网格单元的风险层级标签,由此实现:输出结果的空间可定位性,风险层级标签由查询特征与特征流形库中参考样本的距离统计量(第一局部离散度)及标签统计量(第一邻域恶性占比)共同确定,且该二维统计量可映射至二维风险图谱的确定坐标;风险层级标签的获取不依赖于对分类概率设置任何人为阈值,而是通过查询特征在特征流形中的邻域统计关系与预校准图谱的网格映射关系直接导出。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223450B_ABST
    Figure CN122223450B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a nodule risk stratification method and device, electronic equipment and storage medium, relating to the field of image processing. The method inputs the obtained to-be-processed nodule ultrasound image into a pre-trained encoding network to obtain first normalized features, calculates first local dispersion and first neighborhood malignant proportion of the first normalized features based on a pre-constructed feature manifold library, locates a corresponding first grid unit in a pre-constructed two-dimensional risk map according to the first local dispersion and the first neighborhood malignant proportion, and takes a risk level label of the first grid unit as a risk stratification result of the nodule ultrasound image. Thus, a risk stratification mechanism based on feature space geometry and neighborhood semantic distribution is constructed, so that the output result has spatial locatability and hierarchical interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to a method, apparatus, electronic device, and storage medium for nodule risk stratification. Background Technology

[0002] Ultrasound screening for breast nodules is an important means of early detection of breast malignancies in clinical practice. Physicians need to comprehensively assess the BI-RADS risk based on morphology, borders, echogenicity, and other signs to guide follow-up or biopsy decisions.

[0003] However, existing artificial intelligence methods generally rely on a single malignancy probability value output by a classification network and use a fixed threshold to distinguish between benign and malignant cases. This approach fails to reflect the relative position and neighborhood stability of the sample in the historical case feature space, and therefore cannot generate hierarchical risk conclusions that match the BI-RADS clinical classification. Specifically, probability values ​​alone cannot distinguish between "low-risk images located in stable benign areas," "high-risk images located in stable malignant areas," and "medium-risk images located in areas with ambiguous boundaries." Summary of the Invention

[0004] In view of this, the object of the present invention is to provide a nodule risk stratification method, apparatus, electronic device and storage medium to at least partially improve the above-mentioned problems.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, embodiments of the present invention provide a nodule risk stratification method, including: The obtained ultrasound images of the nodules to be processed are input into a pre-trained encoding network to obtain the first normalized features; Based on a pre-constructed feature manifold library, the first local dispersion and the first neighborhood malignancy ratio of the first normalized feature are calculated; wherein, the feature manifold library is a set of normalized reference features constructed by extracting a preset proportion of samples from the training set used to train the encoding network. Based on the first local dispersion and the first neighborhood malignancy percentage, the corresponding first grid cell is located in the pre-constructed two-dimensional risk map, and the risk level label of the first grid cell is used as the risk stratification result of the nodule ultrasound image; wherein, the two-dimensional risk map is obtained by discretizing each sample in the test set for training the coding network with the first local dispersion and the first neighborhood malignancy percentage as the axis.

[0006] Optionally, the construction steps of the feature manifold library include: Equal amounts of benign and malignant samples are extracted from the training set according to a preset ratio and used as reference samples; Each of the reference samples is input into the coding network to obtain normalized reference features; The normalized reference features and the benign / malignant labels corresponding to the reference samples are used to construct a category-balanced feature manifold library.

[0007] Optionally, calculating the first local discreteness and the proportion of malignancy in the first neighborhood of the first normalized feature based on the pre-built feature manifold library includes: Calculate the distance between the first normalized feature and each normalized reference feature in the feature manifold library; The nearest neighbors with the smallest distance are used as the neighborhood set; Calculate the average distance between each nearest neighbor in the neighborhood set to obtain the first local dispersion of the first normalized feature; Calculate the proportion of malignant samples corresponding to each nearest neighbor in the neighborhood set to obtain the first neighborhood malignant proportion of the first normalized feature.

[0008] Optionally, the steps for constructing the two-dimensional risk map include: Calculate the second local dispersion and the proportion of malignancy in the second neighborhood for each sample in the test set; Based on the preset local discreteness step size and neighborhood malignancy percentage step size, the two-dimensional plane formed by each second local discreteness and each second neighborhood malignancy percentage is discretized into multiple grid cells. For each sample in the test set, its grid cell is determined based on its second local dispersion and the second neighborhood malignancy ratio; Count the number of samples falling into each grid cell to obtain the sample density of each grid cell; Based on the samples in each grid cell, calculate the classification accuracy, benign classification accuracy, and malignant classification accuracy for each grid cell. Based on the sample density, classification accuracy, benign classification accuracy, and malignant classification accuracy, a risk level label and a basis coefficient are calculated for each grid cell; wherein, the risk level label represents the clinically significant risk stratification result of the nodule, and the basis coefficient represents the reliability of the grid cell as the final basis for nodule-level determination.

[0009] Optionally, the formula for calculating the risk level label is:

[0010] in, These represent the index numbers of the two-dimensional risk map in the direction of dispersion and the direction of malignancy, respectively; The second risk level in the two-dimensional risk map represents the first risk level in the two-dimensional risk map. One grid cell; Represents grid cells The corresponding risk level labels are: LR for low risk, MLR for low to medium risk, MR for medium risk, MHR for medium to high risk, and HR for high risk. Represents grid cells The average percentage of malignant cells in the second neighborhood of each sample; Represents grid cells The accuracy of benign classification; Represents grid cells The accuracy of malignancy classification; This represents the threshold for segmenting malignancy proportions into low-malignancy, intermediate, and high-malignancy regions. ; This represents the accuracy threshold parameter used to determine stable low-risk and stable high-risk areas.

[0011] Optionally, the formula for calculating the coefficient is:

[0012]

[0013] in, Indicates the basic basis coefficient; Represents grid cells The basis coefficient; This represents the normalized result of the grid cell index number; This represents the normalized result of the grid index number; This represents the horizontal position weighting coefficient, used to adjust the intensity of the influence of the dispersion direction on the basic coefficient; This represents the vertical position weighting coefficient, used to adjust the intensity of the influence of the malignancy proportion direction on the base coefficient; Represents grid cells Normalized sample density; Represents grid cells The accuracy of classification; This represents the density compensation weighting coefficient; This represents the weighting coefficient for classification accuracy compensation. Indicates risk level label The corresponding weights.

[0014] Optionally, the method further includes: Multiple ultrasound images of the same nodule are input into the coding network to obtain their respective second normalized features and malignancy probability. Based on the feature manifold library, calculate the third local dispersion and the third neighborhood malignancy ratio of each second normalized feature; Based on the third local dispersion and the malignancy ratio of each third neighborhood, the corresponding second grid cell is located in the two-dimensional risk map, and the basis coefficient of each second grid cell is used as the basis coefficient of each nodule ultrasound image. The ultrasound image of the nodule with the smallest criterion coefficient was selected as the ultrasound image of the target nodule. The risk level label corresponding to the ultrasound image of the target nodule is used as the risk stratification result of the nodule, and the corresponding malignancy probability is used as the benign or malignant classification of the nodule.

[0015] Secondly, embodiments of the present invention provide a nodule risk stratification device, comprising: The encoding unit is used to input the obtained ultrasound image of the nodule to be processed into the pre-trained encoding network to obtain the first normalized feature; The parameter calculation unit is used to calculate the first local dispersion and the first neighborhood malignancy ratio of the first normalized feature based on the pre-constructed feature manifold library; wherein, the feature manifold library is a set of normalized reference features constructed by extracting a preset proportion of samples from the training set used to train the encoding network. A risk stratification unit is used to locate the corresponding first grid cell in a pre-constructed two-dimensional risk map based on the first local dispersion and the first neighborhood malignancy ratio, and to use the risk level label of the first grid cell as the risk stratification result of the nodule ultrasound image; wherein, the two-dimensional risk map is obtained by discretizing the first local dispersion and the first neighborhood malignancy ratio of each sample in the test set for training the encoding network as the axis.

[0016] Thirdly, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the method described in any of the above-mentioned embodiments.

[0017] Fourthly, embodiments of the present invention provide a storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in any of the preceding claims.

[0018] This invention provides a nodule risk stratification method, apparatus, electronic device, and storage medium. The method involves inputting an ultrasound image of the nodule to be processed into a pre-trained coding network to obtain a first normalized feature. Based on a pre-constructed category-balanced feature manifold library, the method calculates the first local dispersion and the proportion of malignant cells in the first neighborhood. This information is then mapped to a pre-constructed two-dimensional risk map to obtain the risk level label for the corresponding grid cell. This achieves: spatial localization of the output result; the risk level label is jointly determined by the distance statistics (first local dispersion) between the query feature and the reference samples in the feature manifold library, and the label statistics (proportion of malignant cells in the first neighborhood), and this two-dimensional statistic can be mapped to the defined coordinates of the two-dimensional risk map; the acquisition of the risk level label does not depend on setting any artificial threshold for the classification probability, but is directly derived through the neighborhood statistical relationship of the query feature in the feature manifold and the grid mapping relationship of the pre-calibrated map.

[0019] This makes the output results spatially locatable. The risk level label is jointly determined by the local geometric position (dispersion) of the query feature in the feature manifold and the neighborhood semantic distribution (malignancy rate). Doctors can trace the basis for judgment through this two-dimensional coordinate. Moreover, the risk judgment is not based on a manually set classification probability threshold, but is dynamically defined by the real distribution density and classification reliability of historical cases in the feature space, thus improving the robustness of the model.

[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A schematic structural block diagram of an electronic device provided in an embodiment of the present invention; Figure 2 A flowchart illustrating a nodule risk stratification method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a process for constructing a feature manifold library according to an embodiment of the present invention; Figure 4 A flowchart illustrating step S220 provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of a process for constructing a two-dimensional risk map, provided by an embodiment of the present invention. Figure 6 A schematic diagram of a two-dimensional risk map provided in an embodiment of the present invention; Figure 7 This is another flowchart illustrating a nodule risk stratification method provided in an embodiment of the present invention; Figure 8 This is a schematic structural block diagram of a nodule risk stratification device provided in an embodiment of the present invention.

[0023] Icons: 100 - Electronic device; 101 - Memory; 102 - Communication interface; 103 - Processor; 104 - Communication bus; 600 - Nodule risk stratification device; 610 - Encoding unit; 620 - Parameter calculation unit; 630 - Risk stratification unit. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0025] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0026] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0027] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0028] For ultrasound screening of breast nodules, existing AI solutions can be broadly categorized into two types: the first type mimics the BI-RADS rating or risk range provided by doctors; the second type directly outputs a binary classification result of benign or malignant nodules. The former is essentially closer to refitting existing rules. For example, based on the malignancy probability of the nodule, it divides it into multiple risk ranges based on several preset thresholds, making it difficult to fully utilize the structural relationships of samples in a high-dimensional feature space. The latter, while able to provide pathological attribute judgments, is mostly based on the output probability of a single image, without explicitly describing the sample's position in the historical case feature space, the density of its neighborhood, or the consistency of neighborhood label distribution. Therefore, it is difficult to form interpretable risk stratification results.

[0029] Based on the above, embodiments of the present invention provide a nodule risk stratification method, device, electronic device, and storage medium. It abandons the threshold dependence on a single classification probability and instead utilizes normalized features extracted by a pre-trained coding network, combined with a pre-constructed category-balanced feature manifold library to perform local neighborhood statistics, and obtains a two-dimensional statistic that represents geometric stability and semantic consistency. Then, it uses this two-dimensional statistic to locate grid cells in a pre-calibrated two-dimensional risk map and directly outputs risk level labels that are aligned with clinical risk semantics.

[0030] To implement the process steps and functions of this invention, please refer to [link / reference]. Figure 1 , Figure 1 This is a schematic structural block diagram of an electronic device 100 provided in an embodiment of the present invention. The electronic device 100 includes a memory 101 and a processor 103, which are electrically connected directly or indirectly to each other to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses 104 or signal lines. The memory 101 can be used to store software programs and modules, and the processor 103 executes the software programs and modules stored in the memory 101, thereby performing various functional applications and data processing.

[0031] Electronic device 100 can be, but is not limited to, a personal computer (PC), a server, a distributed computer, etc. It is understood that electronic device 100 is not limited to a physical server, but can also be a virtual machine on a physical server, a virtual machine built on a cloud platform, or any other computer that can provide the same functionality as the server or virtual machine. The operating system of electronic device 100 can be, but is not limited to, Windows, Linux, etc.

[0032] The memory 101 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0033] The communication connection between the electronic device 100 and external devices is achieved through at least one communication interface 102 (which can be wired or wireless).

[0034] Processor 103 may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of this embodiment can be completed by integrated logic circuits in the hardware of processor 103 or by instructions in software form. Processor 103 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0035] Understandable. Figure 1 The structure shown is for illustrative purposes only; the electronic device 100 may also include components that are more advanced than those shown. Figure 1 The more or fewer components shown, or having the same Figure 1The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.

[0036] The nodule risk stratification method provided in this embodiment of the invention will be described below by way of example. See also Figure 2 The subject executing this method can be one of the above. Figure 1 The electronic device 100 shown, the method includes as follows Figure 2 The following steps are described: S210: Input the obtained ultrasound image of the nodule to be processed into the pre-trained encoding network to obtain the first normalized feature.

[0037] S220: Based on the pre-built feature manifold library, calculate the first local dispersion and the first neighborhood malignancy ratio of the first normalized feature; wherein, the feature manifold library is a set of normalized reference features constructed by extracting a preset proportion of samples from the training set of the training encoding network.

[0038] S230: Based on the first local dispersion and the proportion of malignancy in the first neighborhood, locate the corresponding first grid cell in the pre-constructed two-dimensional risk map, and use the risk level label of the first grid cell as the risk stratification result of the nodule ultrasound image; wherein, the two-dimensional risk map is obtained by discretizing the first local dispersion and the proportion of malignancy in the first neighborhood of each sample in the test set of the training coding network as the axis.

[0039] This method does not rely on classification probabilities or their artificial thresholds. Instead, it uses normalized features as a medium to obtain local geometric and semantic two-dimensional statistics by querying a pre-built feature manifold library. Based on these two-dimensional statistics, it directly maps and outputs clinically interpretable risk level labels in a pre-calibrated two-dimensional risk map.

[0040] In one alternative implementation, the obtained ultrasound images of the nodules to be processed are directly stratified by risk. First, step S210 is performed, in which the ultrasound images of the nodules to be processed are input into a pre-trained encoding network to obtain the first normalized features.

[0041] This encoding network is used to learn mid-to-high-dimensional semantic representations with pathological discrimination from raw ultrasound images of breast nodules. For example, let the input image be... The corresponding tag is Where 0 represents benign and 1 represents malignant. Through the backbone network Features are extracted and then subjected to L2 normalization to obtain discriminative representations. ; at the same time, it can also be classified by head Output the probability of malignancy The processing procedure is shown in the following formula:

[0042]

[0043] in, Input image; Represents the input image The corresponding pathology label, among which Indicates benign. Indicates malignancy; The parameter is The backbone network; Represents the set of trainable parameters of the encoding network; Indicates the input image The original feature representation; This represents the normalization operator, typically... Normalization operator; This represents the normalized feature after normalization; The parameter is The category header; Represents the set of trainable parameters for the classification head; Indicates the input image The probability of malign prediction; This represents the Sigmoid activation function. The backbone network can be ResNet, EfficientNet, ConvNeXt, SwinTransformer, or other convolutional networks or visual Transformer networks capable of outputting discriminative features.

[0044] To implement this method, the encoding network needs to be trained first, and then the training set and test set used to train the encoding network are used to construct a feature manifold library and a two-dimensional risk map.

[0045] First, the encoding network undergoes end-to-end supervised training. In this embodiment of the invention, the encoding network adopts the EfficientNet-B3 architecture, and its training process is as follows: Data preparation: The breast ultrasound image database was used as the original training set, with labels of binary pathological ground truth (0: benign; 1: malignant); images were uniformly resampled to 512×512.

[0046] To improve the inter-class separation in the manifold, in this embodiment of the invention, a prototype margin constraint is superimposed on the weighted binary cross-entropy loss, making similar samples more compact in the feature space and dissimilar samples easier to separate.

[0047] The loss function is designed as follows:

[0048]

[0049]

[0050] in, The loss for image-level good / bad classification is represented by a weighted binary cross-entropy loss. This represents the prototype margin constraint loss, used to enhance intra-class compactness and inter-class separation in the feature space; This represents the total loss during model training; Indicates the number of samples participating in the current training batch; Indicates the relationship with sample labels The prototype center of the corresponding category; Indicates the relationship with sample labels The prototype center of the opposite category; This represents the hyperparameter indicating the interval between prototype centers; The weighting coefficients representing the prototype interval constraint loss; This represents the distance metric function in the feature space, which can be cosine distance, Euclidean distance, or Mahalanobis distance.

[0051] The prototype center represents the average of the feature vectors of all samples in the same category.

[0052] Prototype Spacing Constraint Loss The purpose of this is to pull each benign sample toward the prototype center of benign samples and away from the prototype center of malignant samples during the training process, and to pull each malignant sample toward the prototype center of malignant samples and away from the prototype center of benign samples.

[0053] Then, use the training set used to train this encoding network to construct a feature manifold library. In one optional implementation, see [link to implementation details]. Figure 3 The steps for constructing a feature manifold library may include: S310: Extract equal amounts of benign and malignant samples from the training set according to a preset ratio, and use them as reference samples.

[0054] S320: Input each reference sample into the encoding network to obtain normalized reference features.

[0055] S330: Construct a class-balanced feature manifold library by combining each normalized reference feature with the benign / malignant label corresponding to each reference sample.

[0056] Steps S201 to S203 are performed during the offline calibration phase after the encoding network training is completed, and do not participate in the backpropagation or parameter update of the encoding network.

[0057] For example, the training set used in this embodiment of the invention contains 10,000 samples, including 8,000 benign nodules and 2,000 malignant nodules, with a benign-to-malignant ratio of 4:1. To eliminate the interference of class imbalance on the feature space structure, this step performs strict equal-quantity sampling: 1,000 cases are randomly selected from the benign samples and 1,000 cases are randomly selected from the malignant samples, for a total of 2,000 cases as reference samples. This sampling ratio is a preset fixed value and is not adjusted with changes in the total amount of the training set.

[0058] For each of the 2000 reference samples obtained from S201, forward propagation is performed to extract their feature vectors, and L2 normalization is immediately performed to obtain the normalized reference features.

[0059] Each normalized reference feature and its corresponding benign / malignant label are used to construct a class-balanced feature manifold library:

[0060] in, Represents a labeled, category-balanced feature manifold library; Indicates the first Normalized feature vectors of reference samples; Indicates reference sample Corresponding benign or malignant labels; This represents the total number of reference samples in the manifold library.

[0061] This feature manifold library is used only for non-parametric neighborhood retrieval during the online inference phase and does not participate in any gradient calculation, model fine-tuning, or loss function construction.

[0062] Then, step S220 is executed. There are multiple ways to calculate the first local dispersion and the proportion of malignancy in the first neighborhood. In one optional implementation, see [link to implementation details]. Figure 4 S220 may include the following sub-steps: S221: Calculate the distance between the first normalized feature and each normalized reference feature in the feature manifold library.

[0063] The distance can be expressed as a cosine distance:

[0064] in, This represents the first normalized feature; Represents the first in the feature manifold library One reference feature; express With reference features The distance between them; and Representing vectors respectively with vector of Norm.

[0065] S222: Select the nearest neighbors with the smallest distance as the neighborhood set.

[0066] Set a preset quantity K, for example, K = 20, and select... The smallest 20 Forming a neighborhood set:

[0067] in, This indicates sorting by distance and taking the first few items. Operations on each element, for each reference feature in the set. All carry their original labels (0 or 1) to provide a basis for subsequent statistics.

[0068] S223: Calculate the average distance between each nearest neighbor in the neighborhood set to obtain the first local dispersion of the first normalized feature.

[0069] Calculate each of the neighborhood sets corresponding distance The average value is used to obtain the first local dispersion:

[0070] in, This represents the average dispersion between the query sample and its local reference neighborhood. The smaller the value, the closer the sample is to a known stable manifold region.

[0071] S224: Calculate the proportion of malignant samples corresponding to each nearest neighbor in the neighborhood set to obtain the first neighborhood malignant proportion of the first normalized feature.

[0072] Calculate the proportion of malignant samples in the neighborhood set to obtain the malignant proportion of the first neighborhood:

[0073] in, This indicates the proportion of malignant cells in the local neighborhood, reflecting the pathological predisposition of the queried sample's neighborhood. Indicates reference feature The corresponding benign or malignant labels.

[0074] Finally, a two-dimensional risk map is constructed using the test set used to train the encoding network. In one alternative implementation, see [link to implementation details]. Figure 5 The steps for constructing a two-dimensional risk map may include: S410: Calculate the second local dispersion and the proportion of malignancy in the second neighborhood for each sample in the test set.

[0075] S420: Based on the preset local discreteness step size and neighborhood malignancy percentage step size, the two-dimensional plane formed by each second local discreteness and each second neighborhood malignancy percentage is discretized into multiple grid cells.

[0076] S430: For each sample in the test set, determine its grid cell based on its second local dispersion and the proportion of malignancy in its second neighborhood.

[0077] S440: Count the number of samples falling into each grid cell to obtain the sample density of each grid cell.

[0078] S450: Based on the samples in each grid cell, calculate the classification accuracy, benign classification accuracy, and malignant classification accuracy for each grid cell.

[0079] S460: Calculate the risk level label and basis coefficient for each grid cell based on sample density, classification accuracy, benign classification accuracy, and malignant classification accuracy; whereby the risk level label represents the clinically significant risk stratification result of the nodule, and the basis coefficient represents the reliability of the grid cell as the final basis for nodule-level determination.

[0080] The test set is a dataset independent of the training set, for example, 1000 samples.

[0081] For each sample in the test set The same process as S220 is executed, ultimately resulting in 1000 pairs. The data points form the original distribution of the two-dimensional statistical plane.

[0082] Set the local dispersion step size and the neighborhood malignancy percentage step size. For example, if the range of both local dispersion and neighborhood malignancy percentage is [0.0, 1.0], and the step size is 0.1, then the resulting mesh cells are:

[0083] in, Represents the first in the two-dimensional statistical plane One grid cell, For the first test set Each sample has one feature, u = 1 to 11; v = 1 to 11; a total of 121 grid cells.

[0084] Understandably, this is a simple way of dividing the grid. However, for the sake of explanation, adaptive methods or other methods can also be used for grid division.

[0085] For each sample Calculate its grid coordinates:

[0086]

[0087] in, They represent Index numbering in the two-dimensional grid along the direction of discreteness and the direction of malignancy percentage; This indicates the grid step size in the direction of discreteness; Indicates the grid step size in the direction of malignancy percentage; This indicates the floor function.

[0088] Next, calculate the sample density for each grid cell:

[0089] in, Represents grid cells The sample density; This indicates an indicator function that takes the value 1 when the condition within the parentheses is true, and 0 otherwise.

[0090] For example, There are 45 samples, then .

[0091] For example, the number of samples obtained from the grid cells is as follows Figure 6 As shown, there are a total of 5424 samples, with each grid cell corresponding to a certain number of samples.

[0092] For each non-empty grid Using the real labels of the samples With model prediction labels Calculate classification accuracy Accuracy of benign classification And the accuracy of malignant classification :

[0093]

[0094]

[0095] in, , These represent the number of real benign and malignant samples within the grid, respectively.

[0096] Finally, based on the sample density, classification accuracy, benign classification accuracy, and malignant classification accuracy calculated for each grid, the risk level label and basis coefficient for each grid unit are calculated.

[0097] Alternatively, the formula for calculating the risk level label can be:

[0098] in, These represent the index numbers of the two-dimensional risk map in the direction of dispersion and the direction of malignancy, respectively; Represents the first in a two-dimensional risk map One grid cell; Represents grid cells The corresponding risk level labels are: LR for low risk, MLR for low to medium risk, MR for medium risk, MHR for medium to high risk, and HR for high risk. Represents grid cells The average percentage of malignant cells in the second neighborhood of each sample; Represents grid cells The accuracy of benign classification; Represents grid cells The accuracy of malignancy classification; This represents the threshold for segmenting malignancy proportions into low-malignancy, intermediate, and high-malignancy regions. ; This represents the accuracy threshold parameter used to determine stable low-risk and stable high-risk areas.

[0099] In low-risk areas, the proportion of malignant samples is very low (benign samples are absolutely dominant), and the model has a high accuracy rate in classifying benign samples. Therefore, the area is highly safe, and the model rarely misclassifies benign samples as malignant, so it can directly output "low risk".

[0100] In high-risk areas, the proportion of malignant samples is very high (malignant samples are absolutely dominant), and the model has a high accuracy rate in classifying malignant samples. This area is highly dangerous, and the model can reliably identify malignant samples and directly output "high risk".

[0101] In the low-to-medium risk area, the proportion of malignant cells in the grid is in the benign range (below 50% but above the low threshold), and the accuracy of benign cell classification is higher than that of malignant cell classification. This area is still predominantly benign, and the model has a stronger ability to identify benign cells. Overall, it tends to favor the low-risk side and is defined as "low-to-medium risk".

[0102] In the medium-to-high risk grids, the proportion of malignant cells is in the pre-malignant range (≥50% but below the high threshold), and the accuracy of malignant cell classification is not lower than that of benign cells. The region is already biased towards malignant cells, and the model's ability to identify malignant cells is no worse than that of benign cells. Overall, it tends to be on the high-risk side and is classified as "medium-to-high risk".

[0103] Other situations indicate insufficient evidence in the current area or ambiguous model performance, making it unreliable to bias towards low or high risk, and are classified as "medium risk," with further investigation recommended.

[0104] The basis coefficient essentially represents the reliability of the current grid cell as the final criterion for nodule-level determination. When a sample is located in a stable benign region at both ends or a stable malignant region at both ends, although the risk levels are different, both may have high decision certainty, and therefore can be assigned a smaller basis coefficient. Conversely, images located in the intermediate blur region, sparse sample region, or region with low reliability should have a larger basis coefficient. Therefore, the formula for calculating the basis coefficient can be:

[0105]

[0106] in, Indicates the basic basis coefficient; Represents grid cells The basis coefficient; This represents the normalized result of the grid cell index number; This represents the normalized result of the grid index number; This represents the horizontal position weighting coefficient, used to adjust the intensity of the influence of the dispersion direction on the basic coefficient; This represents the vertical position weighting coefficient, used to adjust the intensity of the influence of the malignancy proportion direction on the base coefficient; Represents grid cells Normalized sample density; Represents grid cells The accuracy of classification; This represents the density compensation weighting coefficient; This represents the weighting coefficient for classification accuracy compensation. Indicates risk level label The corresponding weights.

[0107] Explain the normalization indices in the parameters, such as Figure 6 middle, This refers to numbering the grid cells by index. Normalize 1 to 11 to 0 to 1. This refers to numbering the grid cells by index. Normalize 1 to 11 to 0 to 1. This refers to normalizing the number of samples in a grid cell to between 0 and 1.

[0108] Basic basis coefficient It is determined solely by the position of the grid in a two-dimensional plane, reflecting the "inherent instability of the region".

[0109] The factor for the direction of dispersion is , It is a normalized local dispersion index; the greater the dispersion, the more unstable the feature. Therefore... The larger the value, the larger the fundamental coefficient, and the more unreliable the region is considered.

[0110] The factor in the direction of malignancy proportion is ,when When the value is 0.5 (the median percentage of malignant cases), 1 2 | 0.5 0.5|=1, this term is at its maximum. When When the value is 0 or 1 (at both ends of the malignancy rate), this term has its minimum value. The effect is that the closer the malignancy rate is to 50%, the larger the base coefficient, and the less reliable the intermediate transition zone becomes.

[0111] Multiplying the two factors indicates that the region with high dispersion and a middle proportion has the largest basic coefficient, while the region with low dispersion and a proportion at both ends has the smallest basic coefficient.

[0112] Based on weight Can be taken =1、 =3、 =5、 =3、 =1, thus forming a U-shaped structure with the highest weight in the middle fuzzy region and the lowest weight in the stable regions at both ends.

[0113] Density and accuracy compensation items middle, The larger the value, the sparser the sample, the less stable the statistics, and thus the larger the coefficient. The larger the coefficient, the worse the model performs in that region. This compensation term serves to add an extra coefficient to indicate unreliability when a grid sample is scarce or the model accuracy is low.

[0114] Furthermore, clinical nodules often contain multiple static images or multiple key sections. Existing methods typically use simple averaging, maximum probability, or subjective image selection for comprehensive nodule-level judgment, lacking a unified image-based mechanism. That is, existing technologies can often provide the probability of a single image, but cannot answer which image among multiple images of the same nodule best represents its true risk and pathological attributes, directly limiting its clinical applicability and interpretability. Therefore, in one alternative implementation method, see [link to relevant documentation]. Figure 7 The method may also include: S510: Input multiple ultrasound images of the same nodule into the coding network to obtain their respective second normalized features and malignancy probability.

[0115] S520: Based on the feature manifold library, calculate the third local dispersion and the proportion of malignancy in the third neighborhood for each second normalized feature.

[0116] S530: Based on the dispersion of each third local area and the proportion of malignancy in each third neighborhood, locate the corresponding second grid cell in the two-dimensional risk map, and use the basis coefficient of each second grid cell as the basis coefficient of each nodule ultrasound image.

[0117] S540: Select the nodule ultrasound image with the smallest criterion coefficient as the target nodule ultrasound image.

[0118] S550: The risk level label corresponding to the ultrasound image of the target nodule is used as the risk stratification result of the nodule, and the corresponding malignancy probability is used as the benign or malignant classification of the nodule.

[0119] For example, N standard-section ultrasound images of the same thyroid / breast nodule are acquired, and each image is independently input into a pre-trained encoding network to obtain a normalized feature vector. and the corresponding probability of malignancy For each normalized feature Based on a pre-built and frozen feature manifold library, K-nearest neighbors are used to retrieve the neighborhood set of each feature manifold in the feature space, and the third local discreteness is calculated. and the proportion of malignant cases in the third neighborhood ; each group Project the data onto the calibrated two-dimensional risk map, locate the corresponding grid cell, and extract the pre-stored basis coefficients for that grid cell. Among N criteria coefficients, the ultrasound image corresponding to the minimum value is selected as the ultrasound image of the target nodule. The risk level label of the grid cell where the target image is located is used as the final risk stratification result of the nodule, and its corresponding malignancy probability is also determined. As a binary classification of benign and malignant, the output is shown.

[0120] Optionally, when multiple images have the same criterion coefficient, the image with a higher proportion of local malignancy is selected to improve the sensitivity to malignant nodules.

[0121] Furthermore, embodiments of the present invention also provide a nodule risk stratification device, see [link to relevant documentation]. Figure 8 The nodule risk stratification device 600 includes: The encoding unit 610 is used to input the obtained ultrasound image of the nodule to be processed into a pre-trained encoding network to obtain the first normalized feature; The parameter calculation unit 620 is used to calculate the first local dispersion and the first neighborhood malignancy ratio of the first normalized feature based on the pre-built feature manifold library; wherein, the feature manifold library is a set of normalized reference features constructed by extracting a preset proportion of samples from the training set of the training encoding network. Risk stratification unit 630 is used to locate the corresponding first grid cell in the pre-constructed two-dimensional risk map based on the first local dispersion and the proportion of malignancy in the first neighborhood, and to use the risk level label of the first grid cell as the risk stratification result of the nodule ultrasound image; wherein, the two-dimensional risk map is obtained by discretizing the first local dispersion and the proportion of malignancy in the first neighborhood of each sample in the test set of the training coding network as the axis.

[0122] In summary, the nodule risk stratification method, apparatus, electronic device, and storage medium provided in this invention transform the risk stratification of nodule ultrasound images into a joint discrimination process based on their local geometric stability and neighborhood semantic consistency in a high-dimensional feature space by constructing a category-balanced feature manifold library and a pre-calibrated two-dimensional risk map. This method eliminates reliance on a single malignancy probability and arbitrary thresholds, enabling the output risk stratification labels to possess clear spatial localization, clinical interpretability, and decision traceability. Furthermore, the introduction of a criterion coefficient mechanism not only achieves automatic optimization of multi-section images at the nodule level but also quantifies the reliability of each risk assessment.

[0123] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0124] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0125] If the functionality is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0126] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0127] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A method for risk stratification of nodules, characterized in that, include: The obtained ultrasound images of the nodules to be processed are input into a pre-trained encoding network to obtain the first normalized features; Based on a pre-constructed feature manifold library, the first local dispersion and the proportion of malignant features in the first neighborhood of the first normalized feature are calculated; wherein, the feature manifold library is a set of normalized reference features constructed by extracting a preset proportion of samples from the training set used to train the encoding network; the construction steps of the feature manifold library include: extracting equal amounts of benign and malignant samples from the training set according to a preset proportion as reference samples; inputting each of the reference samples into the encoding network to obtain normalized reference features; and constructing a class-balanced feature manifold library by combining each of the normalized reference features and the benign / malignant labels corresponding to each of the reference samples; Based on the first local dispersion and the first neighborhood malignancy ratio, the corresponding first grid cell is located in the pre-constructed two-dimensional risk map, and the risk level label of the first grid cell is used as the risk stratification result of the nodule ultrasound image; wherein, the two-dimensional risk map is obtained by discretizing the first local dispersion and the first neighborhood malignancy ratio of each sample in the test set for training the coding network as the axis. The steps for constructing the two-dimensional risk map include: Calculate the second local dispersion and the proportion of malignancy in the second neighborhood for each sample in the test set; Based on the preset local discreteness step size and neighborhood malignancy percentage step size, the two-dimensional plane formed by each second local discreteness and each second neighborhood malignancy percentage is discretized into multiple grid cells. For each sample in the test set, its grid cell is determined based on its second local dispersion and the second neighborhood malignancy ratio; Count the number of samples falling into each grid cell to obtain the sample density of each grid cell; Based on the samples in each grid cell, calculate the classification accuracy, benign classification accuracy, and malignant classification accuracy for each grid cell. Based on the sample density, classification accuracy, benign classification accuracy, and malignant classification accuracy, a risk level label and a basis coefficient are calculated for each grid cell; wherein, the risk level label represents the clinically significant risk stratification result of the nodule, and the basis coefficient represents the reliability of the grid cell as the final basis for nodule-level determination.

2. The method according to claim 1, characterized in that, The calculation of the first local dispersion and the first neighborhood malignancy ratio of the first normalized feature based on the pre-constructed feature manifold library includes: Calculate the distance between the first normalized feature and each normalized reference feature in the feature manifold library; The nearest neighbors with the smallest distance are used as the neighborhood set; Calculate the average distance between each nearest neighbor in the neighborhood set to obtain the first local dispersion of the first normalized feature; Calculate the proportion of malignant samples corresponding to each nearest neighbor in the neighborhood set to obtain the first neighborhood malignant proportion of the first normalized feature.

3. The method according to claim 1, characterized in that, The formula for calculating the risk level label is: in, These represent the index numbers of the two-dimensional risk map in the direction of dispersion and the direction of malignancy, respectively; The second risk level in the two-dimensional risk map represents the first risk level in the two-dimensional risk map. One grid cell; Represents grid cells The corresponding risk level labels are: LR for low risk, MLR for low to medium risk, MR for medium risk, MHR for medium to high risk, and HR for high risk. Represents grid cells The average percentage of malignant cells in the second neighborhood of each sample; Represents grid cells The accuracy of benign classification; Represents grid cells The accuracy of malignancy classification; This represents the threshold for segmenting malignancy proportions into low-malignancy, intermediate, and high-malignancy regions. ; This represents the accuracy threshold parameter used to determine stable low-risk and stable high-risk areas.

4. The method according to claim 3, characterized in that, The formula for calculating the coefficient is as follows: in, Indicates the basic basis coefficient; Represents grid cells The basis coefficient; This represents the normalized result of the grid cell index number; This represents the normalized result of the grid index number; This represents the horizontal position weighting coefficient, used to adjust the intensity of the influence of the dispersion direction on the basic coefficient; This represents the vertical position weighting coefficient, used to adjust the intensity of the influence of the malignancy proportion direction on the base coefficient; Represents grid cells Normalized sample density; Represents grid cells The accuracy of classification; This represents the density compensation weighting coefficient; This represents the weighting coefficient for classification accuracy compensation. Indicates risk level label The corresponding weights.

5. The method according to claim 1, characterized in that, The method further includes: Multiple ultrasound images of the same nodule are input into the coding network to obtain their respective second normalized features and malignancy probability. Based on the feature manifold library, calculate the third local dispersion and the third neighborhood malignancy ratio of each second normalized feature; Based on the third local dispersion and the malignancy ratio of each third neighborhood, the corresponding second grid cell is located in the two-dimensional risk map, and the basis coefficient of each second grid cell is used as the basis coefficient of each nodule ultrasound image. The ultrasound image of the nodule with the smallest criterion coefficient was selected as the ultrasound image of the target nodule. The risk level label corresponding to the ultrasound image of the target nodule is used as the risk stratification result of the nodule, and the corresponding malignancy probability is used as the benign or malignant classification of the nodule.

6. A nodule risk stratification device, characterized in that, include: The encoding unit is used to input the obtained ultrasound image of the nodule to be processed into the pre-trained encoding network to obtain the first normalized feature; A parameter calculation unit is used to calculate the first local dispersion and the proportion of malignant features in the first neighborhood of the first normalized feature based on a pre-constructed feature manifold library. The feature manifold library is a set of normalized reference features constructed by extracting a preset proportion of samples from the training set used to train the encoding network. The construction steps of the feature manifold library include: extracting equal amounts of benign and malignant samples from the training set according to a preset proportion as reference samples; inputting each reference sample into the encoding network to obtain normalized reference features; and constructing a class-balanced feature manifold library by combining each normalized reference feature with the corresponding benign / malignant label of each reference sample. A risk stratification unit is used to locate the corresponding first grid cell in a pre-constructed two-dimensional risk map based on the first local dispersion and the first neighborhood malignancy percentage, and to use the risk level label of the first grid cell as the risk stratification result of the nodule ultrasound image; wherein, the two-dimensional risk map is obtained by discretizing the first local dispersion and the first neighborhood malignancy percentage of each sample in the test set used to train the encoding network as axes; the construction steps of the two-dimensional risk map include: calculating the second local dispersion and the second neighborhood malignancy percentage of each sample in the test set; and discretizing the two-dimensional plane formed by each second local dispersion and each second neighborhood malignancy percentage into multiple grids according to the preset local dispersion step size and neighborhood malignancy percentage step size. The unit; for each sample in the test set, its grid cell is determined according to its second local dispersion and the proportion of malignant nodules in its second neighborhood; the number of samples falling into each grid cell is counted to obtain the sample density of each grid cell; based on the samples in each grid cell, the classification accuracy, benign classification accuracy, and malignant classification accuracy of each grid cell are calculated respectively; based on the sample density, the classification accuracy, the benign classification accuracy, and the malignant classification accuracy, the risk level label and the basis coefficient of each grid cell are calculated; wherein, the risk level label represents the risk stratification result of the nodule in a clinical sense, and the basis coefficient represents the reliability of the grid cell as the final basis for nodule-level determination.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 5.

8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Neural network-based thyroid nodule preoperative malignant risk intelligent assessment system

    CN117174315A

  • Benign and malignant nodule grading evaluation system based on large model fusion ultrasonic imaging and thyroid gene marker

    CN120452757A