Liver slice recognition method and system based on feature alignment and priori reasoning

By employing feature alignment and prior reasoning methods, noise in liver ultrasound images is suppressed, anisotropic features are extracted, and hierarchical matching is performed in conjunction with clinical logic. This solves the problems of noise interference and feature misalignment in liver ultrasound imaging, achieving highly accurate and safe liver section recognition.

CN122493495APending Publication Date: 2026-07-31HUAQIAO UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAQIAO UNIVERSITY
Filing Date
2026-06-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for liver ultrasound imaging suffer from speckle noise, difficulty in modeling the geometric anisotropy of anatomical structures, misalignment of classification and localization task features, and lack of clinical interpretability, resulting in insufficient diagnostic accuracy and safety.

Method used

A method based on feature alignment and prior reasoning is adopted. Noise is suppressed by feature alignment network, anisotropic features are extracted, and logical matching is performed by hierarchical anatomical reasoning engine to realize the recognition of standard ultrasound sections of the liver.

Benefits of technology

It improves the accuracy and safety of liver ultrasound diagnosis, reduces the risk of misdiagnosis and missed diagnosis, increases diagnostic efficiency, and meets the needs of real-time auxiliary diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493495A_ABST
    Figure CN122493495A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for liver cross-section recognition based on feature alignment and prior reasoning, belonging to the field of image processing technology. The method includes: acquiring liver ultrasound images and performing masking and cropping preprocessing to construct a dataset; constructing and training a deep learning model containing a context and feature alignment network and a hierarchical anatomical inference engine to obtain an automatic liver cross-section recognition model; inputting the image to be recognized into the model, outputting key anatomical structure detection results through the context and feature alignment network, and then using the hierarchical anatomical inference engine to infer based on clinical prior logic to determine the standard cross-section category. This invention, by constructing a deep learning model containing a context and feature alignment network and a hierarchical anatomical inference engine, achieves the recognition and determination of standard cross-sections in liver ultrasound images through multi-level feature alignment and clinical prior logic reasoning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to a method and system for liver cross-section recognition based on feature alignment and prior reasoning. Background Technology

[0002] Liver ultrasound imaging is a core imaging modality for clinical screening and follow-up of liver diseases such as fatty liver, cirrhosis, and liver tumors. In standardized ultrasound scanning procedures, a key prerequisite for high-quality clinical screening is the accurate acquisition and identification of standard liver sections. However, due to inherent speckle noise, beam attenuation, and complex acoustic environments in ultrasound images, the judgment and measurement of traditional 2D ultrasound anatomical structures heavily rely on the operator's experience. This not only leads to significant inter-observer variability but also greatly restricts the widespread application of high-quality ultrasound screening in primary healthcare settings or resource-constrained environments.

[0003] To standardize ultrasound diagnosis, clinical practice typically defines standard sections based on authoritative guidelines (such as the ACR-AIUM-SPR-SRU abdominal ultrasound operation guidelines). This essentially relies on the spatial relationships of a set of key anatomical structures (such as the hepatic vein, portal vein, and inferior vena cava). However, in actual clinical screening, physicians must visually capture and confirm numerous anatomical landmarks during dynamic scanning, which is not only time-consuming and laborious but also prone to inconsistent diagnostic results due to visual fatigue. In recent years, deep learning methods have been introduced into this field, but existing methods mostly rely on purely data-driven global image classification, lacking clinical interpretability (i.e., "black box" characteristics). Directly introducing general object detection models would face complex cross-level "feature misalignment" dilemmas.

[0004] With the development of technology, computer-aided diagnostic techniques for ultrasound sections have gradually matured. Current research on liver ultrasound standard section recognition mainly faces the following problems: First, during low-level feature extraction, speckle noise exhibits a multiplicative distribution, making convolutional downsampling prone to disrupting fine-grained tissue boundaries. Second, tubular anatomical structures such as hepatic veins exhibit significant geometric anisotropy, making traditional isotropic convolution difficult to model effectively. Third, classification and localization tasks suffer from representational conflicts in the shared feature space. Finally, the lack of logical reasoning constraints incorporating clinical prior guidelines results in insufficient diagnostic safety for high-risk cases. To address these issues, this paper proposes a liver ultrasound standard section recognition method and system that combines contextual feature alignment and clinical prior reasoning. This aims to shift from purely data-driven "visual perception" to knowledge-guided "logical reasoning," suppressing noise interference, resolving feature misalignment issues, and equipping the system with an active rejection mechanism based on missing core biomarkers, thereby improving diagnostic accuracy and clinical safety. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a liver section recognition method and system based on feature alignment and prior reasoning. By using a feature alignment network to suppress noise, extract anisotropic features, and alleviate representational conflicts, and then matching clinical logic through a hierarchical prior reasoning engine, the recognition of standard ultrasound sections of the liver is achieved.

[0006] On the one hand, a liver section recognition method based on feature alignment and prior reasoning includes:

[0007] S1. Acquire liver ultrasound images and perform masking and cropping preprocessing to construct the input dataset;

[0008] S2, Construct a deep learning model, which includes a context and feature alignment network and a hierarchical anatomical inference engine; Train the deep learning model using the input dataset to obtain an automatic liver section recognition model;

[0009] S3. Input the ultrasound image of the liver to be identified into the liver section automatic recognition model. Output the key anatomical structure detection results through the context and feature alignment network. The hierarchical anatomical reasoning engine infers the detection results based on clinical prior logic and determines the standard section category.

[0010] The context and feature alignment network includes a low-level feature extraction module, a mid-level semantic modeling module, and a high-level detection head. The low-level feature extraction module suppresses speckle noise in liver ultrasound images while preserving boundary information to obtain low-level features. The mid-level semantic modeling module extracts anisotropic features of tubular anatomical structures from the low-level features and generates mid-level features based on these anisotropic features. The high-level detection head alleviates the representational conflict between classification and localization tasks based on the mid-level features and outputs the detection results of key anatomical structures.

[0011] The hierarchical anatomical reasoning engine uses a hierarchical anatomical reasoning algorithm to transform clinical guidelines into a three-layer logical structure including a basic set of markers, an auxiliary feature set, and an exclusionary constraint set. Based on the three-layer logical structure, it performs hierarchical matching on the detection results of key anatomical structures to obtain the marker matching status corresponding to each candidate section. Finally, based on the marker matching status, it determines the standard section category to which the liver ultrasound image belongs.

[0012] Furthermore, in S1, the mask cropping preprocessing specifically includes: locating the effective imaging region of the ultrasound fan-shaped area corresponding to the liver ultrasound image, constructing the corresponding binary mask to remove the background area and additional information outside the fan-shaped area, retaining the effective acoustic area, and unifying the image into a square structure through zero padding. Finally, the liver ultrasound image is scaled to a specified resolution to construct the input dataset.

[0013] Furthermore, in S3, the low-level feature extraction module suppresses speckle noise in the liver ultrasound image while preserving boundary information to obtain the low-level features. The calculation formula is as follows:

[0014] ;

[0015] ;

[0016] ;

[0017] ;

[0018] ;

[0019] ;

[0020] in, This represents a characteristic ultrasound image of the liver. The pixel value at position (i, j) in channel c of the liver ultrasound feature image; Represents the row index of the feature map; Column index representing the feature map; Indicates the height of the feature map; The width of the feature map is represented by z; the channel descriptor is represented by z. This represents the c-th element in the channel descriptor; This represents the first fully connected layer; This indicates the second fully connected layer; This represents a convolution transformation with a 7×7 kernel; Represents the Sigmoid function; Represents the channel attention map; Represents a spatial attention map; Represents the ReLU activation function; Represents the intermediate feature map; This indicates element-wise multiplication; and Indicates the middle feature map The channel dimension is used to apply average pooling and max pooling to obtain a two-dimensional spatial graph; Indicates a splicing operation; Indicates spatial alignment features; Represents underlying features; and All are learnable 1×1 convolutional projections.

[0021] Furthermore, in S3, the mid-level semantic modeling module extracts the anisotropic features of the tubular anatomical structure from the low-level features to obtain the mid-level features. The calculation formula is as follows:

[0022] ;

[0023] ;

[0024] ;

[0025] in, Indicates a splicing operation; This represents the 5×5 average pooling operator; This represents a deep feature map generated based on low-level features; This represents a local context. Indicates directional characteristics; and This represents an asymmetric convolutional transform with both horizontal and vertical receptive fields. Represents the Sigmoid function; This represents a learnable 1×1 convolutional projection; This represents mid-level features.

[0026] Furthermore, in S3, the high-level detection head mitigates the representational conflict between classification and localization tasks based on mid-level features, and outputs the detection results of key anatomical structures. The calculation formula is as follows:

[0027] ;

[0028] ;

[0029] ;

[0030] in, This represents the sequence after flattening the spatial feature map; Indicates linear projection; This represents the index of the linear transformation corresponding to the generated key vector k; This represents the index of the linear transformation corresponding to the generated value vector v; The encoding represents the global context representation; Q represents the learnable query vector; K and V represent the keys and values ​​generated based on mid-level features; Indicates the scaling factor; Represents the original spatial characteristics. Indicates a space broadcasting mechanism; This indicates the results of key anatomical structure examinations; represents the normalization exponential function, used to transform the calculated attention weights into a probability distribution; q represents the global descriptor; Represents a learnable linear transformation; Expand represents spatial broadcasting, used to match original spatial features. The resolution.

[0031] Furthermore, in S3, hierarchical matching of key anatomical structure detection results is performed based on a three-layer logical structure, specifically including:

[0032] The set D of credible anatomical structures in the current image is selected based on the detection confidence level.

[0033] Determine whether set D contains all the core structures in the basic label set; if no key structure is satisfied, then the aspect is deemed invalid.

[0034] Determine if set D intersects with the set of exclusionary constraints; if spatially mutually exclusive structures exist, exclude the corresponding candidate sections.

[0035] The number of matches between the statistical set D and the auxiliary feature set is counted; when both the basic label constraint and the exclusion constraint are satisfied, and the number of auxiliary feature matches reaches a preset threshold, the image is determined to belong to the corresponding standard section.

[0036] On the other hand, a liver section recognition system based on feature alignment and prior reasoning includes:

[0037] The dataset construction module is used to acquire liver ultrasound images and perform mask cropping preprocessing to construct the input dataset;

[0038] The model building module is used to build a deep learning model, which includes a context and feature alignment network and a hierarchical anatomical inference engine; the deep learning model is trained using the input dataset to obtain an automatic liver section recognition model;

[0039] The identification and judgment module is used to input the ultrasound image of the liver to be identified into the liver section automatic identification model, output the key anatomical structure detection results through the context and feature alignment network, and infer the detection results based on clinical prior logic through the hierarchical anatomical reasoning engine to determine the standard section category.

[0040] The context and feature alignment network includes a low-level feature extraction module, a mid-level semantic modeling module, and a high-level detection head. The low-level feature extraction module suppresses speckle noise in liver ultrasound images while preserving boundary information to obtain low-level features. The mid-level semantic modeling module extracts anisotropic features of tubular anatomical structures from the low-level features and generates mid-level features based on these anisotropic features. The high-level detection head alleviates the representational conflict between classification and localization tasks based on the mid-level features and outputs the detection results of key anatomical structures.

[0041] The hierarchical anatomical reasoning engine uses a hierarchical anatomical reasoning algorithm to transform clinical guidelines into a three-layer logical structure including a basic set of markers, an auxiliary feature set, and an exclusionary constraint set. Based on the three-layer logical structure, it performs hierarchical matching on the detection results of key anatomical structures to obtain the marker matching status corresponding to each candidate section. Finally, based on the marker matching status, it determines the standard section category to which the liver ultrasound image belongs.

[0042] The present invention adopts the above technical solution and has the following beneficial effects:

[0043] (1) The present invention effectively suppresses multiplicative speckle noise in ultrasound images and preserves tissue boundary information by using channel-space attention and learnable projection in the mask cropping preprocessing and low-level feature extraction module;

[0044] (2) This invention extracts anisotropic features based on axial pyramid pooling and asymmetric convolution in the mid-level semantic modeling module, and combines learnable queries and global context encoding in the high-level detection head to alleviate the representation conflict between classification and localization tasks.

[0045] (3) This invention transforms clinical guidelines into a three-layer logical structure of basic markers, auxiliary features, and exclusionary constraints based on a hierarchical anatomical reasoning engine, thereby achieving hierarchical matching and active rejection that conforms to clinical priors. Attached Figure Description

[0046] Figure 1 This is a flowchart of liver section recognition based on feature alignment and prior reasoning according to an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of the overall architecture of the model in an embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of the MedSutur module structure according to an embodiment of the present invention;

[0049] Figure 4 This is a schematic diagram of the AxiPyS operator structure according to an embodiment of the present invention;

[0050] Figure 5 This is a schematic diagram illustrating the bounding box regression optimization principle of an embodiment of the present invention;

[0051] Figure 6 This is an example diagram showing the standard cross-section of the liver and the annotation of key anatomical structures in an embodiment of the present invention;

[0052] Figure 7 This is a flowchart of the KG-HAR hierarchical anatomical reasoning process according to an embodiment of the present invention;

[0053] Figure 8 This is a schematic diagram of the organ detection output results based on the liver according to an embodiment of the present invention;

[0054] Figure 9 This is a diagram of a liver section recognition system based on feature alignment and prior reasoning, according to an embodiment of the present invention. Detailed Implementation

[0055] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0056] like Figure 1 As shown, the present invention provides a large model transfer method based on sparse activation of image representations, comprising:

[0057] S1. Acquire liver ultrasound images and perform masking and cropping preprocessing to construct the input dataset.

[0058] Specifically, the mask cropping preprocessing includes: locating the effective imaging region of the ultrasound fan-shaped area corresponding to the liver ultrasound image, constructing the corresponding binary mask to remove the background area and additional information outside the fan-shaped area, retaining the effective acoustic area, and unifying the image into a square structure through zero padding. Finally, the liver ultrasound image is scaled to a specified resolution to construct the input dataset.

[0059] Specifically, this embodiment performs automated mask cropping preprocessing on the original DICOM ultrasound images. First, the effective imaging region of the ultrasound fan shape is located, and a corresponding binary mask is constructed to remove the background region, device parameters, and additional information outside the fan shape. Then, only the effective acoustic region is retained, and the image is unified into a square structure through zero-padding. Finally, it is scaled to 640×640 resolution and input into the network to improve the model's generalization ability across devices and complex clinical scenarios.

[0060] S2, Construct a deep learning model, which includes a context and feature alignment network and a hierarchical anatomical inference engine; Train the deep learning model using the input dataset to obtain an automatic liver section recognition model.

[0061] Specifically, the automatic identification model includes CFA-Net in the perception stage and KG-HAR in the inference stage. CFA-Net addresses speckle noise, geometric anisotropy, and multi-task representation conflicts in ultrasound imaging by introducing the MedSutur module at the bottom layer, designing the AxiPyS spatial operator at the deep layer, and constructing the Query Guided Detection Head (QG-Head) at the high layer.

[0062] Specifically, such as Figure 2As shown, CFA-Net (Context and Feature Alignment Network) first extracts multi-level features from the input image through a backbone network containing multiple MedSutur modules, followed by final AxiPyS pooling layer processing. Then, it utilizes a Neck (FPN / PAN) structure for bidirectional fusion of multi-scale features to enhance the detection capability for targets of different sizes. Its core innovation lies in the query-guided detection head, which achieves accurate localization and classification of target regions through the interaction of query vectors and feature maps. During training, the model combines Focal Loss and MPDIoU loss functions to address class imbalance and improve bounding box regression accuracy. Figure 3 ( Figure 2 As shown in the dashed box in the lower left corner, this module splits the input features into two paths: the main path extracts features through repeated MedSutur basic units, and the other path performs residual connections. Internally, the module incorporates a GAM attention mechanism, calculating channel attention and spatial attention separately, and achieving adaptive alignment and enhancement of features in both channel and spatial dimensions through element-wise multiplication and addition operations, ultimately outputting refined features. Figure 4 ( Figure 2 As shown in the lower middle dashed box, this operator first performs multi-scale average pooling on the input features and then concatenates them, before entering the axial attention branch. It uses 1×7 and 7×1 orthogonal asymmetric convolution kernels to extract contextual information in the horizontal and vertical directions, respectively. After element-wise addition and sigmoid activation, it is then multiplied element-wise with the original features, thereby effectively reconstructing the receptive field shape in the deep semantic modeling stage to adapt to complex anatomical structures such as tubular structures.

[0063] S3. Input the ultrasound image of the liver to be identified into the liver section automatic recognition model. Output the key anatomical structure detection results through the context and feature alignment network. The hierarchical anatomical reasoning engine infers the detection results based on clinical prior logic and determines the standard section category.

[0064] The context and feature alignment network includes a low-level feature extraction module, a mid-level semantic modeling module, and a high-level detection head. The low-level feature extraction module suppresses speckle noise in liver ultrasound images and preserves boundary information to obtain low-level features. The mid-level semantic modeling module extracts anisotropic features of tubular anatomical structures from the low-level features to obtain mid-level features. The high-level detection head alleviates the representational conflict between classification and localization tasks based on the mid-level features and outputs the detection results of key anatomical structures.

[0065] The hierarchical anatomical reasoning engine uses a hierarchical anatomical reasoning algorithm to transform clinical guidelines into a three-layer logical structure including a basic set of markers, an auxiliary feature set, and an exclusionary constraint set. Based on the three-layer logical structure, it performs hierarchical matching on the detection results of key anatomical structures to obtain the marker matching status corresponding to each candidate section. Finally, based on the marker matching status, it determines the standard section category to which the liver ultrasound image belongs.

[0066] Specifically, in the low-level feature extraction stage, the MedSutur module suppresses speckle noise through the collaborative design of feature enhancement and feature preservation paths. Its computational flow is as follows:

[0067] First, adaptive reweighting of the input features is performed in the enhancement path:

[0068]

[0069]

[0070] Among them, X in Indicates the input tensor. and These are the modulation functions for the channel and spatial dimensions, respectively. This represents element-wise multiplication; ; ;

[0071] Subsequently, the output features are obtained by fusion through the feature preservation path. :

[0072]

[0073] in, ⊕ represents a 1×1 convolution operation, and ⊕ represents element-wise addition fusion.

[0074] In the deep semantic modeling stage, the AxiPyS operator reconstructs the receptive field shape to adapt to the tubular anatomical structure. Its computational flow is as follows:

[0075] Feature aggregation is decoupled into a smooth support branch and an axial sensing branch. The axial sensing branch is represented as follows:

[0076]

[0077] in, For input features, and The kernels are orthogonal asymmetric convolution kernels. Use the Sigmoid activation function;

[0078] Finally, enhanced features P are generated through fusion. out :

[0079]

[0080] in, This represents the local context information obtained through cascaded 5×5 average pooling.

[0081] Specifically, the low-level feature extraction module suppresses speckle noise in liver ultrasound images while preserving boundary information to obtain low-level features. The calculation formula is as follows:

[0082] ;

[0083] ;

[0084] ;

[0085] ;

[0086] ;

[0087] ;

[0088] in, This represents a characteristic ultrasound image of the liver. The pixel value at position (i, j) in channel c of the liver ultrasound feature image; Represents the row index of the feature map; Column index representing the feature map; Indicates the height of the feature map; The width of the feature map is represented by z; the channel descriptor is represented by z. This represents the c-th element in the channel descriptor; This represents the first fully connected layer; This indicates the second fully connected layer; This represents a convolution transformation with a 7×7 kernel; Represents the Sigmoid function; Represents the channel attention map; Represents a spatial attention map; Represents the ReLU activation function; Represents the intermediate feature map; This indicates element-wise multiplication; and Indicates the middle feature map The channel dimension is used to apply average pooling and max pooling to obtain a two-dimensional spatial graph; Indicates a splicing operation; Indicates spatial alignment features; Represents underlying features; and All are learnable 1×1 convolutional projections.

[0089] Specifically, the mid-level semantic modeling module extracts the anisotropic features of the tubular anatomical structure from the low-level features to obtain the mid-level features. The calculation formula is as follows:

[0090] ;

[0091] ;

[0092] ;

[0093] in, Indicates a splicing operation; This represents the 5×5 average pooling operator; This represents a deep feature map generated based on low-level features; This represents a local context. Indicates directional characteristics; and This represents an asymmetric convolutional transform with both horizontal and vertical receptive fields. Represents the Sigmoid function; This represents a learnable 1×1 convolutional projection; This represents mid-level features.

[0094] Specifically, the high-level detection head mitigates the representational conflict between classification and localization tasks based on mid-level features, and outputs key anatomical structure detection results. The calculation formula is as follows:

[0095] ;

[0096] ;

[0097] ;

[0098] in, This represents the sequence after flattening the spatial feature map; Indicates linear projection; This represents the index of the linear transformation corresponding to the generated key vector k; Q* represents the linear transformation index corresponding to the generated value vector v; Q* represents the encoded global context representation; Q represents the learnable query vector; K and V represent the key and value generated based on mid-level features; Indicates the scaling factor; Represents the original spatial characteristics. Indicates a space broadcasting mechanism; This indicates the results of key anatomical structure examinations; represents the normalization exponential function, used to transform the calculated attention weights into a probability distribution; q represents the global descriptor; Represents a learnable linear transformation; Expand represents spatial broadcasting, used to match original spatial features. The resolution.

[0099] Specifically, hierarchical matching of key anatomical structure detection results is performed based on a three-layer logical structure, including:

[0100] The set D of credible anatomical structures in the current image is selected based on the detection confidence level.

[0101] Determine whether set D contains all the core structures in the basic label set; if no key structure is satisfied, then the aspect is deemed invalid.

[0102] Determine if set D intersects with the set of exclusionary constraints; if spatially mutually exclusive structures exist, exclude the corresponding candidate sections.

[0103] The number of matches between the statistical set D and the auxiliary feature set is counted; when both the basic label constraint and the exclusion constraint are satisfied, and the number of auxiliary feature matches reaches a preset threshold, the image is determined to belong to the corresponding standard section.

[0104] Specifically, in ultrasound imaging, speckle noise often overlaps with fine anatomical boundaries in the frequency domain. During convolutional downsampling, multiplicative noise and meaningful structural responses are often co-propagated, weakening the feature representation of anatomical boundaries. To mitigate this issue, we propose the MedSutur module, which combines dynamic feature recalibration with a lightweight projection branch to preserve structural continuity while suppressing noise responses. Deep within the feature pyramid, conventional isotropic convolution uniformly aggregates spatial information in all directions. When modeling elongated or anisotropic anatomical structures in ultrasound images, this aggregation can introduce irrelevant background responses and weaken directional continuity. To enhance orientation-sensitive representations while maintaining computational efficiency, inspired by lightweight orientation modeling strategies, we propose the AxiPyS spatial pooling operator.

[0105] Specifically, in the high-level feature alignment stage, QG-Head alleviates the representational conflict between classification and localization tasks through a lightweight global semantic modulation mechanism. First, it aggregates global contextual information through a cross-attention mechanism:

[0106]

[0107] Q is then mapped to a two-dimensional space for feature modulation via a spatial broadcasting mechanism:

[0108]

[0109] in, These are the aggregated query features. For the original spatial features, F() represents the broadcast enhancement operation.

[0110] Furthermore, the model employs the MPDIoU loss function in the bounding box regression stage, which improves the stability of localization by directly constraining the two-dimensional Euclidean distance between the diagonal key points of the predicted box and the ground truth box.

[0111] Specifically, such as Figure 5 The figure shown is a schematic diagram of the bounding box regression optimization principle in an embodiment of the present invention. The figure details the calculation logic of the MPDIoU loss function in the process of matching the predicted box with the ground truth label. By introducing constraints on the geometric characteristics of the bounding box, the gradient problem of conventional IoU in the state where the predicted box and the ground truth box intersect is effectively solved, thereby guiding the model to more accurately regress the boundary of the target three-dimensional anatomical structure.

[0112] Specifically, in this embodiment, the inference engine KG-HAR, based on prior clinical knowledge, can determine 13 types of standard liver ultrasound sections, namely: oblique longitudinal section at the 6th and 7th intercostal spaces (liver, gallbladder, main portal vein, inferior vena cava); high oblique transverse section below the costal margin at the second hepatic hilum (liver, hepatic vein, right portal vein); high transverse section at the level of the second hepatic hilum (liver, hepatic vein); longitudinal section of the left hepatic-abdominal aorta (liver, aorta, stomach, pancreas); and long-axis section of the portal vein at the first hepatic hilum (liver, portal vein). The KG-HAR reasoning logic transforms clinical rules into a three-layer logical structure: (Main trunk of the artery, common bile duct); longitudinal section of the liver and gallbladder (liver, gallbladder); longitudinal section of the liver and kidneys (liver, right kidney); low oblique transverse section at the level of the gallbladder and kidneys (liver, gallbladder, right kidney, inferior vena cava, aorta); longitudinal section of the left liver and stomach fundus (liver, stomach); low transverse section at the level of the liver and pancreas (liver, gallbladder, pancreas); mid-level oblique transverse section below the costal margin of the first hepatic hilum (liver, main trunk of the portal vein); mid-level transverse section at the level of the first hepatic hilum (liver, main trunk of the portal vein); longitudinal section of the hepatic-inferior vena cava (liver, inferior vena cava).

[0113] (1) Basic marker determination: If the core anatomical structure corresponding to the cut surface is missing, the cut surface is determined to be invalid and an active rejection mechanism is triggered;

[0114] (2) Auxiliary feature matching: By setting a minimum number of matches, the tolerance for individual differences is improved;

[0115] (3) Exclusion constraint: Use spatial mutual exclusion relation to exclude candidate sections that have semantic conflicts.

[0116] In summary, by constructing a Context and Feature Alignment Network (CFA-Net), accurate detection and feature extraction of key anatomical structures at the lower levels in ultrasound images were achieved. The MedSutur module was introduced in the lower-level feature extraction stage to effectively preserve the subtle boundary information of tissues while suppressing speckle noise interference. In the middle layer, the AxiPyS spatial operator was employed to accurately capture the geometric anisotropy features of tubular anatomical structures such as the hepatic vein by reconstructing the asymmetric receptive field. In the higher layer, a Query-Guided Detection Head (QG-Head) was used to significantly alleviate the representational conflict between classification and localization tasks through a global semantic modulation mechanism. This network effectively solves the cross-level "feature misalignment" dilemma in ultrasound imaging and comprehensively improves the detection accuracy and robustness of the model in complex acoustic environments.

[0117] Furthermore, a hierarchical anatomical reasoning engine (KG-HAR) was constructed based on authoritative clinical guidelines, realizing a paradigm shift from "black-box visual perception" to "white-box logical reasoning." This engine transforms descriptive clinical rules into multi-layered explicit logic containing basic markers, auxiliary features, and exclusionary constraints. This not only significantly improves the accuracy and clinical interpretability of standard section determination but also endows the system with an active rejection mechanism based on the absence of core markers, effectively avoiding the high-risk misdiagnosis risk caused by image quality degradation and greatly ensuring the clinical safety of the auxiliary diagnostic system. This invention effectively improves the detection stability of liver anatomical structures in complex acoustic environments through bottom-level noise reduction and fidelity preservation, mid-level geometric reconstruction, and top-level task decoupling. It also achieves safe clinical determination through a highly interpretable logical reasoning engine, establishing a new paradigm for high-risk ultrasound-assisted diagnosis that combines clinical logic with a safety baseline. This can assist doctors in improving diagnostic efficiency, reducing the risk of misdiagnosis and missed diagnosis, and providing effective support for the early detection and treatment of liver diseases.

[0118] Specifically, verification experiments were conducted on the embodiments of the present invention. The experimental environment included an Ubuntu 20.04 operating system and an NVIDIA RTX 3090 GPU. Experimental results showed that the average detection accuracy (mAP50) of the CFA-Net on the validation set reached 0.850. In independent blind human-machine testing, the system's cross-section determination accuracy reached 94.36%, and the Cohen's Kappa coefficient of consistency was 0.9390, demonstrating performance close to that of advanced experts. The system's single-frame inference latency was only 15.6ms (exceeding 60 FPS), meeting the requirements for real-time assisted diagnosis.

[0119] like Figure 6The figure shown is an example of a standard liver section and key anatomical structures labeled according to an embodiment of the present invention. The figure intuitively shows the characteristics of a typical liver imaging section and marks in detail the core anatomical structures (such as basic landmarks and auxiliary features) used for section identification, providing a basic morphological reference and judgment criterion for subsequent hierarchical anatomical reasoning.

[0120] Figure 7 This is a flowchart of the hierarchical anatomical reasoning process of KG-HAR in an embodiment of the present invention. The flowchart describes the logical progression from the output of candidate detection results by the model, through confidence screening, core structure matching of the basic label set, spatial mutual exclusion test of the exclusion constraint set, and statistical analysis of the number of matching auxiliary feature sets, to finally achieve rigorous reasoning on the standard section of the liver.

[0121] Figure 8 This is a schematic diagram of the organ detection output based on the liver according to an embodiment of the present invention. The figure shows the final output image with anatomical structure prediction boxes, classification confidence scores and section category determination results after the input original image passes through the CFA-Net network and inference process. It intuitively verifies the detection effect and localization accuracy of the method of the present invention in actual image processing.

[0122] like Figure 9 As shown, this embodiment also discloses a liver section recognition system based on feature alignment and prior reasoning, including:

[0123] The dataset construction module 91 is used to acquire liver ultrasound images and perform mask cropping preprocessing to construct the input dataset.

[0124] The model building module 92 is used to construct an automatic liver section recognition model, which includes a context and feature alignment network and a hierarchical anatomical inference engine; the automatic liver section recognition model is trained using the input dataset to obtain a trained automatic liver section recognition model;

[0125] The identification and judgment module 93 is used to input the ultrasound image of the liver to be identified into the liver section automatic identification model, output the key anatomical structure detection results through the context and feature alignment network, and infer the detection results based on clinical prior logic through the hierarchical anatomical reasoning engine to determine the standard section category.

[0126] The context and feature alignment network includes a low-level feature extraction module, a mid-level semantic modeling module, and a high-level detection head. The low-level feature extraction module suppresses speckle noise in liver ultrasound images while preserving boundary information to obtain low-level features. The mid-level semantic modeling module extracts anisotropic features of tubular anatomical structures from the low-level features and generates mid-level features based on these anisotropic features. The high-level detection head alleviates the representational conflict between classification and localization tasks based on the mid-level features and outputs the detection results of key anatomical structures.

[0127] The hierarchical anatomical reasoning engine uses a hierarchical anatomical reasoning algorithm to transform clinical guidelines into a three-layer logical structure including a basic set of markers, an auxiliary feature set, and an exclusionary constraint set. Based on the three-layer logical structure, it performs hierarchical matching on the detection results of key anatomical structures to obtain the marker matching status corresponding to each candidate section. Finally, based on the marker matching status, it determines the standard section category to which the liver ultrasound image belongs.

[0128] A specific implementation of a liver section recognition system based on feature alignment and prior reasoning is described in this embodiment, which is the same liver section recognition method based on feature alignment and prior reasoning.

[0129] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.

Claims

1. A liver slice recognition method based on feature alignment and priori reasoning, characterized in that, Includes the following steps: S1. Acquire liver ultrasound images and perform masking and cropping preprocessing to construct the input dataset; S2, Construct an automatic liver section recognition model, which includes a context and feature alignment network and a hierarchical anatomical inference engine; Train the automatic liver section recognition model using the input dataset to obtain a trained automatic liver section recognition model; S3. Input the ultrasound image of the liver to be identified into the liver section automatic recognition model. Output the key anatomical structure detection results through the context and feature alignment network. The hierarchical anatomical reasoning engine infers the detection results based on clinical prior logic and determines the standard section category. The context and feature alignment network includes a low-level feature extraction module, a mid-level semantic modeling module, and a high-level detection head; The low-level feature extraction module suppresses speckle noise in liver ultrasound images while preserving boundary information to obtain low-level features; the mid-level semantic modeling module extracts anisotropic features of tubular anatomical structures from the low-level features and generates mid-level features based on these anisotropic features. The high-level detection head mitigates the representational conflict between classification and localization tasks based on mid-level features, and outputs key anatomical structure detection results; The hierarchical anatomical reasoning engine uses a hierarchical anatomical reasoning algorithm to transform clinical guidelines into a three-layer logical structure including a basic set of markers, an auxiliary feature set, and an exclusionary constraint set. Based on the three-layer logical structure, it performs hierarchical matching on the detection results of key anatomical structures to obtain the marker matching status corresponding to each candidate section. Finally, based on the marker matching status, it determines the standard section category to which the liver ultrasound image belongs. 2.The liver slice identification method based on feature alignment and priori reasoning according to claim 1, characterized in that, In S1, the mask cropping preprocessing specifically includes: locating the effective imaging region of the ultrasound fan-shaped area corresponding to the liver ultrasound image, constructing the corresponding binary mask to remove the background area and additional information outside the fan-shaped area, retaining the effective acoustic area, and unifying the image into a square structure through zero padding. Finally, the liver ultrasound image is scaled to a specified resolution to construct the input dataset. 3.The liver slice identification method based on feature alignment and priori reasoning according to claim 1, characterized in that, In S3, the low-level feature extraction module suppresses speckle noise in the liver ultrasound image while preserving boundary information, thus obtaining the low-level features. The calculation formula is as follows: ; ; ; ; ; ; in, This represents a characteristic ultrasound image of the liver. The pixel value at position (i, j) in channel c of the liver ultrasound feature image; Represents the row index of the feature map; Column index representing the feature map; Indicates the height of the feature map; The width of the feature map is represented by z; the channel descriptor is represented by z. This represents the c-th element in the channel descriptor; This represents the first fully connected layer; This indicates the second fully connected layer; This represents a convolution transformation with a 7×7 kernel; Represents the Sigmoid function; Represents the channel attention map; Represents a spatial attention map; Represents the ReLU activation function; Represents the intermediate feature map; This indicates element-wise multiplication; and Indicates the middle feature map The channel dimension is used to apply average pooling and max pooling to obtain a two-dimensional spatial graph; Indicates a splicing operation; Indicates spatial alignment features; Represents underlying features; and All are learnable 1×1 convolutional projections.

4. The liver section recognition method based on feature alignment and prior reasoning according to claim 1, characterized in that, In S3, the mid-level semantic modeling module extracts anisotropic features of tubular anatomical structures from the low-level features to obtain mid-level features. The calculation formula is as follows: ; ; ; in, Indicates a splicing operation; This represents the 5×5 average pooling operator; This represents a deep feature map generated based on low-level features; This represents a local context. Indicates directional characteristics; and This represents an asymmetric convolutional transform with both horizontal and vertical receptive fields. Represents the Sigmoid function; This represents a learnable 1×1 convolutional projection; This represents mid-level features.

5. The liver section recognition method based on feature alignment and prior reasoning according to claim 1, characterized in that, In S3, the high-level detection head mitigates the representational conflict between classification and localization tasks based on mid-level features, and outputs the detection results of key anatomical structures. The calculation formula is as follows: ; ; ; in, This represents the sequence after flattening the spatial feature map; Indicates linear projection; This represents the index of the linear transformation corresponding to the generated key vector k; This represents the index of the linear transformation corresponding to the generated value vector v; The encoding represents the global context representation; Q represents the learnable query vector; K and V represent the keys and values ​​generated based on mid-level features; Indicates the scaling factor; Represents the original spatial characteristics. Indicates a space broadcasting mechanism; This indicates the results of the examination of key anatomical structures; represents the normalization exponential function, used to transform the calculated attention weights into a probability distribution; q represents the global descriptor; Represents a learnable linear transformation; Expand represents spatial broadcasting, used to match original spatial features. The resolution.

6. The liver section recognition method based on feature alignment and prior reasoning according to claim 1, characterized in that, In S3, hierarchical matching of key anatomical structure detection results is performed based on a three-layer logical structure, specifically including: The set D of credible anatomical structures in the current image is selected based on the detection confidence level. Determine whether set D contains all the core structures in the basic label set; if no key structure is satisfied, then the aspect is deemed invalid. Determine if set D intersects with the set of exclusionary constraints; if spatially mutually exclusive structures exist, exclude the corresponding candidate sections. The number of matches between the statistical set D and the auxiliary feature set is counted; when both the basic label constraint and the exclusion constraint are satisfied, and the number of auxiliary feature matches reaches a preset threshold, the image is determined to belong to the corresponding standard section.

7. A liver section recognition system based on feature alignment and prior reasoning, characterized in that, include: The dataset construction module is used to acquire liver ultrasound images and perform mask cropping preprocessing to construct the input dataset; The model building module is used to build an automatic liver section recognition model, which includes a context and feature alignment network and a hierarchical anatomical inference engine; the automatic liver section recognition model is trained using the input dataset to obtain a trained automatic liver section recognition model; The identification and judgment module is used to input the ultrasound image of the liver to be identified into the liver section automatic identification model, output the key anatomical structure detection results through the context and feature alignment network, and infer the detection results based on clinical prior logic through the hierarchical anatomical reasoning engine to determine the standard section category. The context and feature alignment network includes a low-level feature extraction module, a mid-level semantic modeling module, and a high-level detection head; The low-level feature extraction module suppresses speckle noise in liver ultrasound images while preserving boundary information to obtain low-level features; the mid-level semantic modeling module extracts anisotropic features of tubular anatomical structures from the low-level features and generates mid-level features based on these anisotropic features. The high-level detection head mitigates the representational conflict between classification and localization tasks based on mid-level features, and outputs key anatomical structure detection results; The hierarchical anatomical reasoning engine uses a hierarchical anatomical reasoning algorithm to transform clinical guidelines into a three-layer logical structure including a basic set of markers, an auxiliary feature set, and an exclusionary constraint set. Based on the three-layer logical structure, it performs hierarchical matching on the detection results of key anatomical structures to obtain the marker matching status corresponding to each candidate section. Finally, based on the marker matching status, it determines the standard section category to which the liver ultrasound image belongs.