Real-time detection method and system for multiple abnormal conditions of craniocerebral CT (computed tomography) guided by iconography priori knowledge
By introducing a symmetric information attention module and a context information aggregation module in the craniocerebral CT imaging abnormality detection model, the problem of complex alignment preprocessing and insufficient perception ability of contralateral regions in the prior art is solved, and more efficient and accurate craniocerebral image abnormality detection is achieved.
Patent Information
- Application Number
- CN202510021449.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The prior art has the problem of complex alignment preprocessing steps and insufficient perception of contralateral regions in the abnormal detection of craniocerebral CT images, making it difficult to effectively capture context information in the three-dimensional medical volume image of craniocerebral .
A real-time detection method for brain CT with a priori knowledge guided by imaging is designed. By introducing a symmetric information attention module into the backbone network, the contralateral difference graph attention network explicitly encodes brain symmetry to enhance the effectiveness of feature representation. At the same time, a context information aggregation module is used to combine 2D and 3D convolutions to effectively integrate the anatomical structure continuity between slices.
It improves the accuracy and reliability of cranial brain image abnormality detection, avoids complex alignment operations, enhances the ability to pay attention to information on the opposite side, and improves detection performance while ensuring computing efficiency.
Smart Images

Figure CN119941679A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method and system for real-time detection of multiple abnormalities in cranial CT scans guided by prior knowledge of imaging. Background Art
[0002] Many researchers have tried to incorporate brain imaging prior knowledge into the algorithm to improve the accuracy and reliability of brain abnormality detection models. This mainly includes the following two types of imaging prior knowledge:
[0003] The first is the imaging prior knowledge of hemispheric symmetry. Many studies have shown that the use of imaging prior knowledge of hemispheric symmetry can effectively guide model learning. However, the NCCT images (Non-Contrast CT, also known as plain scan CT) obtained in clinical practice often fail to show strict left-right symmetry due to differences in patient position during scanning. In response to this challenge, the following alignment strategies have emerged in the prior art: using a template to first align the slices to the standard space, and then inputting the anomaly detection model; or, obtaining the affine transformation matrix through the alignment network (or finding the brain midline through the brain midline positioning network, and then calculating the affine transformation matrix), using the affine transformation matrix to align to the standard space, and then inputting the anomaly detection model. The above alignment strategies are essentially to align the original image to the standard space, and the anomaly detection model itself does not have the ability to cope with the diverse spatial distribution in real scenes, and has the disadvantages of poor practicality and complex preprocessing process. In addition, there is also an alignment strategy in the prior art that embeds a network module that specifically utilizes symmetry knowledge in the deep structure of the model neural network, thereby avoiding complex preprocessing. However, even in the deep layer of the network, when the receptive field of the feature point becomes larger, due to individual differences, lesions and other factors in actual brain images, the feature map in the deep layer of the network cannot be strictly symmetrical, resulting in limited contribution of imaging prior knowledge of hemispheric symmetry to the accuracy and reliability of the anomaly detection model.
[0004] The second is the continuity of anatomical structures between slices of three-dimensional medical volumetric images of the brain. In the field of medical imaging, especially for NCCT sequence images of the brain, each slice maintains a close anatomical continuity. This feature makes it possible to aggregate rich contextual information from volumetric medical images. Traditional slice-based 2DCNN networks are difficult to achieve inductive bias in the cross-sectional direction because they cannot cross slice boundaries and are difficult to effectively capture and utilize contextual information in three-dimensional space. 3D convolutional neural networks mainly rely on 3D convolutions in the backbone network to extract features to utilize the continuity features between adjacent slices, but 3D convolutions also bring about a significant increase in model complexity. Compared with 2D convolutions, the number of parameters increases significantly and the inference speed slows down. In addition, 3D convolutions have the risk of exacerbating the attenuation of small abnormal features. Moreover, the NCCT images used in clinical practice are not obtained through thin-layer scanning, and their layer thickness is usually large (usually 5 mm). This feature poses a challenge to the performance of 3D CNN in processing NCCT images, because excessive aggregation of contextual information may cause small abnormal lesions confined to a single slice to be obscured or faded. As the downsampling process is performed multiple times, the features of these tiny lesions may gradually weaken or even be completely lost.
[0005] In the field of target detection, the YOLO series network architecture is generally composed of three core structures: backbone, neck, and detecthead. The workflow is as follows: First, the backbone is responsible for processing and extracting feature information from the preprocessed input image; then, in the neck part, the PANET (Path Aggregation Network) technology is used to fuse features from different scales to enhance the expressiveness of the features; finally, this information is passed to the three detection heads, which accurately predict the location coordinates, confidence scores, and categories of the target at different scales. Although there are subtle differences in the cranial and brain structures of different individuals, they generally present a relatively fixed morphology and anatomical structure. Compared with natural images, NCCT images have relatively low diversity and complexity in color, texture, and shape. Therefore, for the analysis task of NCCT images, a smaller network depth and width may be sufficient to provide sufficient feature representation. In this case, simply increasing the network depth and width may not be the most effective strategy. On the contrary, incorporating imaging prior knowledge into a smaller network architecture to fully explore and utilize the more unique feature information in cranial NCCT images may be a more efficient method.
[0006] In summary, in the deep learning anomaly detection model of cranial 3D medical volumetric images, many studies have been devoted to using imaging prior knowledge to guide the learning process of the model. However, there are still two aspects in this field that deserve further exploration and deepening: one is how to give the model a strong contralateral area perception ability while avoiding complex and tedious alignment preprocessing steps; the other is how to use the anatomical structure continuity between slices efficiently and effectively. On the other hand, the existing YOLO object detection network is designed and tuned for natural images. When used for the detection of various abnormalities of the brain, it is necessary to design it specifically for the unique imaging prior knowledge of cranial 3D medical volumetric images. Summary of the invention
[0007] The present invention aims to at least solve the technical problems existing in the prior art and provide a method and system for real-time detection of multiple abnormalities in cranial CT scans guided by prior knowledge of imaging.
[0008] In order to achieve the above-mentioned object of the present invention, according to the first aspect of the present invention, the present invention provides a method for real-time detection of multiple abnormalities in cranial CT guided by prior knowledge of imaging, comprising:
[0009] Real-time acquisition of brain CT image sequences;
[0010] Input the brain CT image sequence into a pre-trained anomaly detection model to obtain anomaly detection results; the anomaly detection model includes:
[0011] A backbone network is used to extract multi-scale features of a cranial CT image sequence, wherein the backbone network includes a plurality of cascaded feature extraction modules, and at least one feature extraction module is a symmetric information attention module;
[0012] The neck network fuses the features of different scales output by the backbone network to obtain multiple fused features;
[0013] The multi-detection head network processes multiple fusion features separately to obtain anomaly detection results;
[0014] Among them, the symmetric information attention module includes:
[0015] A contralateral difference map attention network or two or more cascaded contralateral difference map attention networks for extracting bilateral difference information of the input features of the symmetry information attention module;
[0016] The difference information acquisition unit obtains the difference between the bilateral difference information and the input features of the symmetric information attention module, which is recorded as the difference feature;
[0017] The feature enhancement unit enhances the input features of the symmetric information attention module based on the difference features to obtain bilateral difference enhanced features.
[0018] In order to achieve the above-mentioned purpose of the present invention, according to the second aspect of the present invention, the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the real-time detection method for multiple abnormalities of cranial CT guided by imaging prior knowledge as described in the first aspect of the present invention.
[0019] In order to achieve the above-mentioned object of the present invention, according to a third aspect of the present invention, the present invention provides an electronic equipment system, the electronic equipment system comprising:
[0020] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for real-time detection of multiple abnormalities in cranial CT guided by imaging prior knowledge as described in the first aspect of the present invention.
[0021] The above-mentioned technical solution provided by the present invention: The anomaly detection model of the present invention is an improvement on the existing YOL0 real-time target detection network. At least one symmetric information attention module is added to the backbone network. The symmetric information attention module includes one or two cascaded contralateral difference map attention networks. The contralateral difference map attention network explicitly encodes the imaging prior knowledge of brain symmetry into the model to effectively perceive the information difference of the same anatomical structure on the contralateral side to obtain bilateral difference information, and then uses the difference information acquisition unit to calculate the difference between the bilateral difference information and the input features of the symmetric information attention module to obtain the difference features, and the input features of the symmetric information attention module are calculated according to the difference features. The features are enhanced, the expression of important features is enhanced, while minor features are suppressed, and the effectiveness of feature representation is improved. It can be seen that the symmetry information attention module of the present invention explicitly encodes the imaging prior knowledge of brain symmetry into the backbone network, avoiding the complex alignment operation of the existing method, and also makes up for the problem that the existing method does not pay enough attention to the contralateral information, and increases the flexibility of matching bilaterally symmetrical anatomical structures. Even when the symmetry is destroyed to a certain extent, the abnormality detection model of the present invention can still capture the contralateral information to assist the detection of abnormal areas (such as lesions), thereby improving the accuracy and reliability of abnormality detection in cranial brain imaging. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flow chart of a method for real-time detection of multiple abnormalities in cranial CT guided by prior knowledge of imaging in a preferred embodiment of the present invention;
[0023] Figure 2 is a schematic diagram of the network structure of an anomaly detection model in a preferred embodiment of the present invention;
[0024] Figure 3 is a structural diagram of a symmetric information attention module in a preferred embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of the association between feature points and regions in a preferred embodiment of the present invention;
[0026] Figure 5 is a schematic diagram of the attention network of the contralateral difference map in a preferred embodiment of the present invention;
[0027] Figure 6 is a structural diagram of a context information aggregation module in a preferred embodiment of the present invention;
[0028] Figure 7 It is a comparison diagram of images generated by the traditional mosaic enhancement method and the 3D symmetrical mosaic enhancement method provided by the present invention;
[0029] Figure 8 It is a schematic diagram of a process of generating a slice by a 3D symmetric mosaic enhancement method in a preferred embodiment of the present invention;
[0030] Fig. 9 It is a structural schematic diagram of an electronic equipment system in a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0031] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0032] In the description of the present invention, it is necessary to understand that the terms "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0033] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal connection between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.
[0034] The execution subject of the method for real-time detection of multiple abnormalities in cranial CT guided by prior knowledge of imaging provided by the present invention includes but is not limited to at least one of the electronic devices such as the server and the terminal that can be configured to execute the method provided in the preferred embodiment of the present application. In other words, the method for real-time detection of multiple abnormalities in cranial CT guided by prior knowledge of imaging can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0035] The present invention provides a method for real-time detection of multiple abnormalities in cranial CT scans guided by prior knowledge of imaging. In a preferred embodiment, Figure 1 As shown, including:
[0036] Step S1, acquiring a brain CT image sequence in real time.
[0037] Preferably, an original cranial CT image sequence can be obtained from a CT device in real time, and the original cranial CT image sequence can be an NCCT sequence. The original cranial CT image sequence is preprocessed to generate multiple cranial CT image sequences, and at least one cranial CT image sequence is input into the abnormality detection model to obtain corresponding abnormality detection results. In this embodiment, the preprocessing includes:
[0038] Step S11, brain window processing and subdural window processing are performed on the original cranial CT image sequence. The image in the original cranial CT image sequence can be in DICOM (Digital Imaging and Communications in Medicine) format. For each slice in the original cranial CT image sequence, window processing is performed, and window processing is to select a specific sub-interval (i.e. window) within the wide grayscale range of the DICOM image, so as to highlight the information of interest. The present embodiment provides a brain window and a subdural window. In the brain window setting, by the brain window grayscale sub-interval set, the abnormal category to be detected can be intuitively observed, and the difference category includes at least two of the 8 abnormal categories of cerebral ischemia, cerebral hemorrhage, softening focus, leukoaraiosis, scalp hematoma, space-occupying, extracerebral effusion, and midline shift, so that the abnormal category area can be clearly displayed under the brain window. Studies have shown that the subdural window can also provide important medical details, particularly for subdural lesions, so the subdural grayscale sub-interval is set. This study applied two windowing processing strategies, brain window and subdural window, to the same DICOM image (i.e., slice). Specifically, the image data after brain window processing was assigned to the red (R) and green (G) channels of the RGB image, while the image data after subdural window processing was placed in the blue (B) channel, stacked into three channels of the RGB image. Windowing can adjust the grayscale value range of the image and highlight the grayscale value range of different tissues so that specific tissues can be observed more clearly. The grayscale range of brain windows and subdural windows is adjusted differently. Placing them on different channels can enrich the image information and help optimize the identification of anatomical structures and lesion areas. The slice after windowing processing is shown in the figure below. Figure 8 shown.
[0039] Step S12, sequentially dividing the original brain CT image sequence after brain window processing and subdural window processing into multiple subsequences with the same length, each subsequence is used as a brain CT image sequence to be input into the abnormality detection model, and there is at least one overlapping slice between adjacent subsequences.
[0040] Specifically, the image sequence can be divided into multiple subsequences of the same length according to the index order of the slices in the original cranial CT image sequence after the brain window processing and the subdural window processing. The length range of the subsequence can be 4 to 12, preferably 6. There is at least one overlapping slice between adjacent subsequences, which means that the tails of the two subsequences overlap, that is, the last slices of the previous subsequence are the same as the first slices of the next subsequence. In one example, the last two pictures of the previous subsequence are the same as the first two pictures of the next subsequence. If the last subsequence cannot completely contain pictures of the preset length, the number of overlapping pictures with the penultimate subsequence can be increased to ensure that the lengths of all subsequences reach the preset length.
[0041] Step S2, input the brain CT image sequence into the pre-trained anomaly detection model to obtain anomaly detection results. When the anomaly detection model detects an abnormal area, the anomaly detection result includes the position coordinates of each abnormal area (which can be the position coordinates of the abnormal area bounding box), the abnormal type of the abnormal area and the confidence score. At the same time, the abnormal area is marked with a bounding box in the brain CT image sequence, such as Figure 2 When the anomaly detection model does not detect an abnormal area, the output is NULL or normal.
[0042] In this implementation, the anomaly detection model is obtained based on the improvement of the Y0L0 series network architecture, such as the improvement based on the Y0L0V8 network structure.
[0043] like Figure 2 As shown, the anomaly detection model includes:
[0044] Backbone network( Figure 2 The backbone network is used to extract multi-scale features of the cranial CT image sequence. The backbone network includes multiple cascaded feature extraction modules, and at least one feature extraction module is a symmetric information attention module (Symmetric Information Attention, referred to as SIA module). Figure 2 The backbone network may include a C2f module, a SIA module, a CBS module, a C2f module, a CBS module, a C2f module and an SPPF module connected in sequence.
[0045] Neck network ( Figure 2 The neck network is a network that fuses features of different scales output by the backbone network to obtain multiple fusion features to enhance the expressiveness of the features. The neck network includes a first branch and a second branch. The first branch includes an upsample module Upsample, a connection layer Contact, a C2f module, an upsample module Upsample, a connection layer Contact, and a C2f module that are connected in sequence. The second branch includes a CBS module, a connection layer Contact, a C2f module, a CBS module, a connection layer Contact, and a C2f module that are connected in sequence.
[0046] Multi-sensor head network ( Figure 2 The multiple detection heads can accurately predict the location coordinates, confidence scores and anomaly categories of the target at different scales.
[0047] In this embodiment, the connection relationship between modules in the anomaly detection model can refer to Figure 2As shown, no further details are given here. The C2f module, CBS module, SPPF module and upsample module Upsample in the backbone network and the neck network can all refer to the Y0L0V8 network. The network structures in other prior arts can also be used for reference. For example, the CBS module can refer to the CBS module in the existing patent with publication number CN119090905A, the C2f module can refer to the C2f structure in the existing patent with publication number CN118982658A, and the SPPF module can refer to the SPPF structure in the existing patent with publication number CN117912062A, which will not be described here.
[0048] Among them, Figure 3 As shown in Figure 1, the SIA module is the symmetric information attention module, which includes:
[0049] A contralateral difference graph attention network or two or more cascaded contralateral difference graph attention networks are used to extract bilateral difference information of the input features of the symmetric information attention module. Bilateral Difference Graph attention network (BDGA network for short).
[0050] The difference information acquisition unit obtains the difference between the bilateral difference information and the input features of the symmetric information attention module, which is recorded as the difference feature. Figure 3 As shown, the difference information acquisition unit performs an element-level subtraction operation.
[0051] The feature enhancement unit enhances the input features of the symmetric information attention module based on the difference features to obtain bilateral difference enhanced features.
[0052] In this embodiment, only one contralateral difference map attention network (BDGA network) may be set in the symmetric information attention module (SIA module). In order to enable the model to focus on the contralateral feature area that has not been fully explored before and realize the effective compensation of the symmetric difference information, the symmetric information attention module (SIA module) may also set more than two serial contralateral difference map attention networks (BDGA networks), and at this time, the windows of the two or more contralateral difference map attention networks (BDGA networks) do not overlap, that is, the patch partition windows, such as Figure 4 As shown in , the window partition of the second contralateral difference graph attention network (BDGA network) will be shifted based on the first contralateral difference graph attention network (BDGA network). Through window adjustment, the contralateral difference graph attention network can cover nodes and directed edge combinations at different positions. Figure 4As shown, the first column of pictures shows the window partition of the first contralateral difference map attention network (BDGA network), and the second column of pictures shows the window partition of the second contralateral difference map attention network (BDGA network). The window partition of the second contralateral difference map attention network (BDGA network) is shifted compared to the window partition of the first contralateral difference map attention network (BDGA network).
[0053] When there are two cascaded contralateral difference map attention networks, the bilateral difference information X dif It is expressed as:
[0054] X dif =BDGA(BDGA(X));
[0055] Where X represents the input features of the symmetric information attention module and BDGA(·) represents the contralateral difference map attention network operation.
[0056] In a preferred embodiment, in order to extract different feature information and help improve the accuracy of anomaly detection, the contralateral difference map attention network BDGA uses K graph attention mechanisms to respectively extract graph attention features of the input features of the contralateral difference map attention network BDGA, splices the K graph attention features and uses the spliced features as the output features of the contralateral difference map attention network BDGA, where K is a positive integer.
[0057] In this embodiment, when K is 1, the graph attention features obtained by a single graph attention mechanism are used as the output features of the contralateral difference graph attention network BDGA: when K is greater than 1, the graph attention features of the input features of the contralateral difference graph attention network extracted by K graph attention mechanisms are concatenated, and the concatenated features are used as the output features of the contralateral difference graph attention network BDGA.
[0058] In this embodiment, the human brain anatomical structure is roughly bilaterally symmetrical, so a common diagnostic method used by clinicians is to identify abnormal areas by comparing the left and right brains. This actually establishes a connection between the left and right symmetrical anatomical structures of the brain. If this diagnostic process is simulated, the imaging knowledge of the bilateral symmetry of the cerebral hemispheres is incorporated into the deep learning model, and the key lies in building the relationship between the left and right symmetrical anatomical structures. Under ideal conditions, for the anatomical structure p, it is only necessary to compare it with the symmetrical anatomical structure. However, in actual clinical scenarios, due to differences in patient positioning and other reasons, the symmetry of brain images may be affected to a certain extent. If the original image is not aligned to the standard space, it is difficult to accurately find the It is more difficult. Even in the deep layer of the network, when the receptive field of the feature point becomes larger, it is difficult to ensure the perfect symmetry of the feature points on both sides using contralateral difference learning. This is due to individual differences, lesions and other factors in actual brain images. Based on this, the present invention provides a solution. Although the goal is only to establish p and But in order to increase the number of The present invention combines p with a possible Area To associate, such as Figure 4 As shown, the feature point p a Symmetrical with Area Similarly, p b and The advantage of this is that it increases the flexibility of matching bilaterally symmetrical anatomical structures, so that even when the symmetry is destroyed to a certain extent, the anomaly detection model of the present invention can still capture The information is used to assist the detection of the lesion area. In the process of constructing complex relationships within medical images, Graph Neural Networks (GNNs) have become an attractive choice due to their powerful representation learning ability and effective capture of potential structures and relationships in images. Therefore, the present invention proposes a contralateral difference graph attention network, which uses K graph attention mechanisms to perceive the contralateral difference respectively.
[0059] In this embodiment, refer to Figure 5 , the process of each graph attention mechanism extracting the graph attention features of the input features of the contralateral difference graph attention network includes:
[0060] Step A1, divide the input features of the contralateral difference map attention network into multiple patches without overlapping.
[0061] For the input feature F of the contralateral difference map attention network, a sliding window operation of size l×l is applied to divide it into non-overlapping patches, each patch is a collection of feature points i, j represent the horizontal and vertical coordinates of the feature point in the patch, d is the dimension of the feature point, and p ij Represents the features of the position point (i, j) in the patch.
[0062] In step A2, to achieve diversified features, each feature point of each patch is projected into the query embedding space, key embedding space and value embedding space to obtain query embedding, key embedding and value embedding.
[0063] Three separate linear projection layers are used to project these feature points in the patch into three different embedding spaces: query, key, and value. The process can be expressed as:
[0064] q ij =P ij W Q , k ij =P ij W K , v ij =P ij W v
[0065] Among them, W Q , W K , They are the mapping matrices of query embedding space, key embedding space and value embedding space respectively. ij , k ij and v ij Respectively represent p ij The query embedding, key embedding and value embedding of n are combined to form a graph node n. ij , which can be expressed as p ij Different representations of the anatomical structure in the receptive field. ij Used to evaluate the potential contribution or influence of the current node on other nodes, k ij The main goal is to explore and measure the similarity or association between the anatomical structures represented by other graph nodes and the current node, while v ij The role of is to provide information about the unique anatomical features or attributes represented by each graph node. k represents the index of the current graph attention mechanism, k∈[1,K].
[0066] In step A3, each feature point of each patch is taken as a node (i.e., graph node), and the query embedding, key embedding, and value embedding of the feature point are used as node attributes.
[0067] Step A4: construct a graph network for patch P to express the symmetric relationship of the anatomical structure through the graph network. The graph network includes a terminal node set, a starting node set, and a directed edge set. All nodes inside patch P constitute the terminal node set, and the patch that is symmetrical to patch P along the central axis of the input feature of the symmetry information attention module is All internal nodes form the starting node set, and the directed edge set includes all directed edges from the starting nodes to each terminal node. The attribute information of each directed edge is calculated through the attribute information of the starting node and the terminal node; P, Both represent the index of the patch and are positive integers.
[0068] In this embodiment, each patch is represented as a graph network, and the graph network of patch P is Among them, N is the set of all graph nodes in patch P, which is the terminal node set. and As the starting node set, is a set of directed edges. For node n in patch P ij , the present invention finds a patch that is symmetrical to the left and right of patch P along the central axis of the input feature F of the attention network of the contralateral difference map The set of all graph nodes in As node n ij Neighbor nodes of Indicates the starting node To the terminal node n ij Directed edge The attribute information of the graph network G. The degree of all nodes in the graph network G is fixed, which is l 2 , the graph network G constructs patch P and patch Calculate their scaled dot-product attention to quantify the terminal node n ij The key embedding k ij and the starting node Query embedding The similarity of the anatomical structure features represented by the two nodes is used to obtain the directed edge Attribute information Also called the attention coefficient:
[0069]
[0070] Among them, T represents the matrix transpose; attention coefficient Guidance Node How much information will be propagated to n ij ,Will As the starting node To the terminal node n ij Directed edge attribute information.
[0071] In step A5, the graph attention sub-feature of each node is obtained through the directed edge attribute information of each patch and the value embedding of each node.
[0072] Step A5 implements the contralateral difference perception mechanism. When the terminal node n ij When it contains lesion information, in its neighbor node set On the contrary, if the terminal node n ij If it does not contain lesion information, it can be Find similar features. This difference is directly reflected in the directed edge attribute information Therefore, for the starting node To the terminal node n ij The invention uses directed edge attribute information With terminal node n ij Specific anatomical representation of ij (i.e., value embedding) to aggregate neighbor information and obtain node n in the patch ij The graph attention sub-feature v ij ′:
[0073]
[0074] Among them, σ is the GELU activation function; W o is the output matrix, which is learnable.
[0075] In step A6, the graph attention sub-features of all nodes of all patches constitute the graph attention features output by the contralateral difference graph attention network under each graph attention mechanism.
[0076] In this embodiment, K graph attention mechanisms in the contralateral difference graph attention network BDGA respectively extract graph attention features of the input features of the contralateral difference graph attention network BDGA, concatenate the K graph attention features and use the concatenated features as the output features of the contralateral difference graph attention network BDGA. Specifically, the K graph attention sub-features obtained by each feature point under the K graph attention mechanisms are concatenated to obtain the concatenated sub-features. The concatenated sub-features of all nodes of all patches constitute the graph attention features output by the contralateral difference graph attention network. The node n ij The splicing sub-feature v ijK ' is expressed as:
[0077]
[0078] Among them, II represents splicing, is the attention coefficient of the k-th graph attention mechanism, is the weight matrix of the linear transformation of the output of the k-th graph attention mechanism.
[0079] In a preferred embodiment, referring to Figure 3 , the feature enhancement unit includes:
[0080] The global average pooling layer (GAP) processes the difference features to obtain pooled features. The difference features represent the difference or complementary information between the bilateral difference information and the original features (features input to the SIA module).
[0081] The Squeeze-and-Excitation (SE) processing unit obtains the channel weight map based on the pooled features. The Squeeze-and-Excitation (SE) processing unit consists of two fully connected layers. The SE module is constrained by the input feature X and the bilateral difference information. Its core role is to adaptively recalibrate the feature response of each channel by explicitly modeling the correlation between feature channels, thereby generating a weight map that can reflect the importance of each channel.
[0082] The multiplication unit multiplies the channel weight map with the input features of the symmetric information attention module to obtain weighted residual features. The multiplication process is element-by-element multiplication, which realizes weighted adjustment of the original features according to the difference between the contralateral difference information and the original features. This mechanism enhances the expression of important features, while suppressing minor features, and improves the effectiveness of feature representation.
[0083] The adding unit adds the weighted residual feature to the bilateral difference information to obtain the bilateral difference enhancement feature, and the addition is element-by-element addition.
[0084] In this embodiment, the structure of the symmetric information attention module (SIA module) is as follows: Figure 3 As shown in Figure 1, its working mechanism is as follows: First, the input feature X of the SIA module is directed to two contralateral difference map attention networks BGDA connected in series to perceive and extract the bilateral difference information X dif. Subsequently, this difference information is subtracted from the original input feature X to obtain the difference feature, which represents the difference or complementary information between the bilateral difference information and the original feature. The difference feature then undergoes a squeeze-and-excitation (SE) module that includes a global average pooling (GAP) operation and two fully connected layers. The SE module is constrained by the input feature X and the bilateral difference information here. Its core role is to adaptively recalibrate the feature response of each channel by explicitly modeling the correlation between feature channels, thereby generating a channel weight map that reflects the importance of each channel. Next, the channel weight map is element-wise multiplied with the original input feature X retained by the residual connection to obtain the weighted residual feature. This process realizes the weighted adjustment of the original feature according to the degree of difference between the contralateral difference information and the original feature. This mechanism enhances the expression of important features, while suppressing minor features and improving the effectiveness of feature representation. Finally, the weighted residual feature is combined with the bilateral difference information X. dif The final output feature of this module is obtained by adding the features, namely the bilateral difference enhancement feature. SIA can be expressed as:
[0085] SIA(x)=X dif +x×σ(MLP(σ(MLP(GAP(xX dif )))))
[0086] Among them, the above σ function represents the activation function of the fully connected layer.
[0087] 2D convolution and 3D convolution each have unique advantages. In some cases, it may be better to focus on a single slice feature, because the local features inside the slice are sufficient to provide sufficient information without introducing noise or interference that may be caused by other slices. Other features rely on contextual information and need to be captured and analyzed in three-dimensional space to reveal complex structural relationships. Although 3D CNN has the potential to adaptively adjust weights in the cross-sectional direction through learning, it has the ability of 2D convolution to a certain extent. However, strengthening the model's sensitivity and capture ability to abnormal features through explicit network structure design may make the model easier to learn and thus obtain better anomaly detection performance. Therefore, the present invention designs an effective and efficient context information aggregation CIA (Context Information Aggregation) module to replace the convolution downsampling module in the baseline model.
[0088] Therefore, in a preferred embodiment, Figure 2As shown in , at least one feature extraction module located in the shallow layer of the backbone network is a context information aggregation module. The shallow layer refers to multiple feature extraction modules in the backbone network close to the input end of the backbone network, such as Figure 2 As shown, the first, second and fourth feature extraction modules of the backbone network are set as context information aggregation modules.
[0089] In a preferred embodiment, referring to Figure 6 , the context information aggregation module (CIA module) includes:
[0090] The 2D convolution branch includes a 2D-CBS module and a first channel attention unit connected in sequence; the 2D-CBS module is the CBS module in the existing Y0L08 series network, which includes a two-dimensional convolution, a two-dimensional BatchNorm, and a SiLU activation function unit.
[0091] The 3D convolution branch includes a reshape and dimension reshaping unit, a 3D-CBS module, a dimension reshaping and reshape unit, and a second channel attention unit connected in sequence; the 3D-CBS module includes a 3D convolution.
[0092] Aggregation unit, used to aggregate the output features of the 2D convolution branch and the output features of the 3D convolution branch.
[0093] In this embodiment, refer to Figure 6 , the context information aggregation module (CIA module) consists of two branches: 2D convolution branch and 3D convolution branch. Suppose the input feature of CIA module is Where L = B × D, B is the batch size, D is the depth of the input sequence, that is, the number of slices, and H and W represent the height and width of the input feature map X respectively.
[0094] In the 2D convolution branch, the input feature X is first processed by the 2D-CBS module, where a 2D convolution with a 3×3 kernel and a stride of 2 downsamples the input feature X to obtain a 2D downsampled feature X2. Then, the first channel attention unit is used to dynamically adjust the weight of each channel to obtain the 2D channel weight CA2, which is used to rescale the channel features. This process enables the network to focus on more critical features, that is, those that are more discriminative for anomaly detection. Figure 6 ,The first channel attention unit includes a cascade of a GAP (global average pooling layer) and a fully connected layer.
[0095] X2=CBS2(X)
[0096] CA2=σ2(MLP2(GAP(X2))
[0097] Among them, CBS2(·) represents 2D-CBS module processing; MLP2(·) represents the fully connected layer of the first channel attention unit; σ2(·) represents the activation function of the fully connected layer of the first channel attention unit.
[0098] Compared with the 2D convolution branch, the 3D convolution branch focuses on capturing contextual information across slice dimensions. First, the input feature X is reshaped and reshaped by the reshape unit R&P to B×C×D×H×W to meet the input requirements of the 3D convolution. The 3D convolution is a 3D convolution module with a kernel size of 3×3×3 and a stride of 2, resulting in a scale of The convolutional features are then input into the dimension reshaping and reshaping unit P&R, and the dimension and shape are converted to a scale of The 3D downsampled feature X3 is then passed through the second channel attention unit to focus on the key features and obtain the 3D channel weight CA3. Figure 6 ,The second channel attention unit includes a cascaded GAP (global average pooling layer) and a fully connected layer.
[0099]
[0100] CA3=σ3(MLP3(GAP(X3)))
[0101] Among them, CBS represents convolution plus BN plus SiLu; CBS3(·) represents 3D-CBS module processing; MLP3 represents the fully connected layer of the second channel attention unit; R represents the reshape followed by permute operation; represents the permute followed by the reshape operation, MLP3(·) represents the fully connected layer of the second channel attention unit; σ3(·) represents the activation function of the fully connected layer of the second channel attention unit.
[0102] The aggregation unit fuses the output features X2×CA3 of the 2D convolution branch and the output features X3×CA3 of the 3D convolution branch by point-wise addition, realizing the weighted addition of the slice dimension features extracted by the 2D convolution and the cross-slice dimension features extracted by the 3D convolution. Therefore, the output feature CIA(X) of the context information aggregation module (CIA module) is:
[0103] CIA(X)=X2×CA3+X3×CA3
[0104] In this embodiment, a contextual information aggregation module (CIA module) is embedded in the anomaly detection model so that the model can adaptively select features extracted by two branches according to the input features, which not only fully utilizes the advantages of 3D convolution in contextual information aggregation, but also makes up for the limitations of 3D convolution in computational efficiency and local feature extraction through 2D convolution.
[0105] In this embodiment, the embedded context information aggregation module (CIA module) is set in the shallow layer of the backbone network in the anomaly detection model. This is mainly due to the following considerations:
[0106] 1. The shallow network is mainly responsible for extracting basic, low-level features. In the shallow network, the key information of the image sequence (such as texture, edges, and local spatial relationships) has not been over-abstracted or lost. It is hoped that when the information loss has not accumulated to a significant level, the advantages of the CIA module in aggregating contextual information and capturing detailed features can be fully utilized.
[0107] 2. As the network level deepens, the model gradually transitions from extracting basic visual features to extracting higher-level semantic information.
[0108] The CIA module incorporates a 3D convolution module, which significantly increases the computational complexity and number of parameters when used in the deep layers of the network, leading to an increased risk of overfitting. Therefore, in deep networks, we prefer to use simpler and more efficient 2D convolution operations to extract these semantic features in order to maintain the computational efficiency and generalization ability of the model.
[0109] In clinical practice, doctors often identify abnormal areas by comparing images on both sides of the brain. The present invention aims to simulate this process, which includes not only verification and testing, but also training. The premise of all this is to ensure that the left and right brains are roughly horizontally symmetrical in the input image, but this premise has encountered obstacles in the training process. Since YOLOv4 introduced the mosaic enhancement data enhancement strategy, the YOLO series models, including YOL0v8, have adopted this technology during the training process. Figure 7 The enhanced image in the figure is obtained by the traditional mosaic method. The traditional mosaic enhancement randomly uses multiple pictures, randomly scales them, and then randomly distributes them for splicing, which greatly enriches the detection data set. In particular, random scaling adds many small targets, making the network more robust. However, Figure 7As shown in the first image in the figure, after mosaic enhancement, the inventors found that the position distribution of brain regions became random, and the horizontal symmetry between the left and right brains could no longer be guaranteed. Paradoxically, mosaic enhancement can effectively improve model performance. This result undoubtedly highlights the importance of mosaic enhancement, but at the same time, the problem of symmetry destruction it brings must also be solved. In order to utilize the advantages of mosaic enhancement and ensure the symmetrical distribution of the left and right brains, the present invention improves traditional mosaic enhancement and provides a new mosaic enhancement method, which is named 3D symmetrical mosaic method. Figure 7 The second image is the enhanced image obtained by the 3D symmetric mosaic method. Figure 7 The third picture in the figure shows the distribution of the center points of all target bounding boxes in the dataset, where x and y are the relative coordinates to the original image size. It can be seen that the D-symmetric mosaic method ensures the symmetrical distribution of the left and right brains.
[0110] In a preferred embodiment, in order to improve the diversity of training data and increase the performance of the trained anomaly detection model, the anomaly detection model is trained by enhancing the image sequence sample set, and the method for obtaining the enhanced image sequence sample set includes:
[0111] Step B1, obtaining an image sequence sample set, each image sequence sample includes 6 consecutive slices.
[0112] Preferably, obtaining an image sequence sample set includes:
[0113] Step B11, obtaining an original cranial CT image sequence. Specifically, the original cranial CT image sequence can be obtained from a CT device, an MRI device, or the like.
[0114] Step B12, performing brain window processing and subdural window processing on the original cranial CT image sequence. Brain window processing and subdural window processing have been described in the above preferred embodiments and will not be repeated here.
[0115] Step B13, sequentially segmenting the original cranial CT image sequence after brain window processing and subdural window processing into multiple image sequence samples of the same length, there is at least one overlapping slice between adjacent image sequence samples, and the multiple image sequence samples constitute an image sequence sample set. The specific process of sequence segmentation has been described in the above preferred implementation mode and will not be repeated here.
[0116] Step B2: using a 3D symmetric mosaic method to obtain a plurality of enhanced image sequence samples, and adding the plurality of enhanced image sequence samples to an image sequence sample set to obtain an enhanced image sequence sample set.
[0117] in, Figure 8 This shows an example of generating a slice using the 3D symmetric mosaic method. Figure 8, the 3D symmetric mosaic method is: randomly select 4 image sequence samples from the image sequence sample set, and its sequence sample index set D: = {d i ∈N′|0≤i≤3}, where N′ represents the index set of the extracted sequence samples. represents the sth slice of sequence d, then the set of all slices in D can be expressed as Initialize the slice index s=0. When 0≤s≤5, then use the 3D symmetric mosaic technology to combine the pictures in the set X to generate a new sequence data of length 6, that is, an enhanced image sequence sample. Specifically, repeat steps a to c to generate the sth slice of the enhanced image sequence sample:
[0118] Step a: divide the blank image into 6 non-overlapping blank areas. Figure 8 As shown, the six blank areas are distributed in two rows, the first row is defined as the third blank area ③, the first blank area ① and the fourth blank area ④ from left to right, and the second row is defined as the fifth blank area ⑤, the second blank area ② and the sixth blank area ⑥ from left to right. The size of the blank image can be 1024*1024.
[0119] Step b, on the s-th slice of the first image sequence sample extracted, use the first random width and height box cropping that passes through the central axis and is symmetrical along the central axis, and place the cropped box image in the first blank area of the blank image; the central axis here is the central axis of the s-th slice of the first image sequence sample.
[0120] On the s-th slice of the second image sequence sample extracted, a second random width and height square frame passing through the central axis and symmetrical along the central axis is used for cropping, and the cropped square image is placed in the second blank area of the blank image; the central axis here is the central axis of the s-th slice of the second image sequence sample.
[0121] On the s-th slice of the third image sequence sample extracted, two third random width and height square frames that are symmetrical along the central axis are used for cropping, and the two square frame images obtained by cropping are placed in the third blank area and the fourth blank area of the blank image respectively; the central axis here is the central axis of the s-th slice of the third image sequence sample.
[0122] On the s-th slice of the fourth image sequence sample extracted, two fourth random width and height square frames that are symmetrical along the central axis are used for cropping, and the two square frame images obtained by cropping are placed in the fifth blank area and the sixth blank area of the blank image respectively; the central axis here is the central axis of the s-th slice of the fourth image sequence sample.
[0123] In step b, area ① is placed That is, the sth slice of sequence d0. Area ② is placed The cropping method is the same as that of area ①. Areas ③ and ④ come from It is obtained by cropping and sampling two boxes with random width and height that are symmetrical on the left and right. Similarly, areas ⑤ and ⑥ come from The cropping method is the same as ③ and ④. The process can be expressed as Here, Mθ(·) represents the 3D symmetric mosaic operation, and θ represents the size, coordinates and other parameters involved in the random cropping and placement of the sample image, including the coordinates of the lower left corner of area ①, the coordinates of the upper left corner of area ②, and four random widths and heights.
[0124] Step c: the blank image after placing the box images in the six blank areas is used as the s-th slice of the enhanced image sequence sample, and s=s+1.
[0125] In this embodiment, in order to ensure the continuity of the anatomical structure, the slice cropping areas of the above-mentioned 6 regions should be continuous in the cross-sectional direction (slice continuity direction), so the same random parameters are used for all slices, and the index order in the original slice is maintained. In addition, in order to maintain the continuity of the anatomical structure, the slice cropping of the above-mentioned six regions should be continuous in the cross-sectional direction. Therefore, the same random parameters are used for all slices, and their index order in the original slice is maintained unchanged. Finally, the enhanced image sequence sample can be expressed as: Finally, the 1024*1024 enhanced image sequence sample slice size is resized to 512*215.
[0126] The present invention provides a method for real-time detection of multiple abnormalities in cranial CT guided by prior knowledge of imaging, and its main improvements include:
[0127] (1) In order to simulate the diagnostic method commonly used by clinicians to identify abnormal areas by comparing the left and right brains, the present invention incorporates the imaging knowledge of the bilateral symmetry properties of the cerebral hemispheres into the deep learning model and designs a symmetry information attention module. This module avoids the complex alignment operations of existing methods and also makes up for the problem that existing methods do not pay enough attention to contralateral information. The imaging prior knowledge of brain symmetry is explicitly encoded into the model using a graph attention network to effectively perceive the information differences of the same anatomical structure on the contralateral side. The human brain anatomical structure is roughly bilaterally symmetrical, so identifying abnormal areas by comparing the left and right brains is a diagnostic method commonly used by clinicians. This actually establishes a connection between the left and right symmetrical anatomical structures of the brain. In order to simulate this diagnostic process, the present invention incorporates the imaging knowledge of the bilateral symmetry properties of the cerebral hemispheres into the deep learning model, and the key is to construct the relationship between the left and right symmetrical anatomical structures. Under ideal conditions, for the anatomical structure p, it only needs to be compared with the symmetrical anatomical structure. To increase the number of The present invention combines p with a possible an area To associate, such as Figure 4 This increases the flexibility of matching bilaterally symmetrical anatomical structures, allowing the model to capture the symmetry even when the symmetry is disrupted to a certain extent. information to assist in the detection of lesion areas.
[0128] (2) In order to make full use of the continuous relationship between the anatomical structures of adjacent slices and aggregate the contextual information of volumetric medical images, a simple and effective contextual information aggregation module CIA is designed. This module introduces the inductive bias in the cross-sectional direction of the volumetric medical image, achieves a balance in the feature aggregation within the slice and between adjacent slices, and realizes the effective integration of contextual information. The module is divided into two branches: 2D convolution and 3D convolution. The features of the two branches are fused by point-wise addition, realizing the weighted addition of the slice dimension features extracted by 2D convolution and the cross-slice dimension features extracted by 3D convolution. The model can adaptively select the features extracted by the two branches according to the input features, which not only fully utilizes the advantages of 3D convolution in contextual information aggregation, but also compensates for the limitation of 3D convolution that excessive aggregation of contextual information aggravates the attenuation of small abnormal features through 2D convolution. In addition, compared with the existing model that uniformly uses 2D or 3D convolution from shallow to deep layers, the present invention only uses this module in the shallow layer of the network and uses 2D convolution in the deep layer of the network. In this way, contextual information can be effectively aggregated while ensuring that the number of parameters and complexity of the model do not increase significantly.
[0129] (3) The present invention also proposes a 3D symmetric mosaic data enhancement strategy to ensure the left-right symmetry of the brain image during training. Different from natural images, the brain area in CT images is generally near the center of the image. Figure 7 As shown in the third figure, all the target bounding box center points in the data set are also distributed near the center of the image, and are basically symmetrical on the left and right, and rarely appear at the edge of the image. Comparing the traditional mosaic enhancement image and the 3D symmetrical mosaic enhancement image, it can be found that the 3D symmetrical mosaic enhanced image is closer to the above distribution, that is, the cranial brain area is basically concentrated in the central area, and is basically symmetrical on the left and right. The feature distribution of symmetrical mosaic enhancement is closer to the feature distribution of cranial brain NCCT images, so it can effectively improve the performance of the model in detecting cranial brain abnormalities.
[0130] It should be noted that the real-time detection method for multiple abnormalities of cranial CT guided by imaging prior knowledge provided by the present invention can also be applied to multiple abnormality detection of other types of cranial brain image sequences such as cranial brain MRI image sequences.
[0131] The present invention also discloses a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the real-time detection method for multiple abnormalities of cranial CT guided by the above-mentioned imaging prior knowledge provided by the present invention are implemented. The computer program product should be understood as a software product that mainly implements its solution through a computer program, such as a program product integrated in the cloud or a software library.
[0132] The present invention also discloses an electronic device system. In one embodiment, the electronic device system includes at least one processor; and a memory connected to the at least one processor in communication; wherein:
[0133] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the real-time detection method for multiple abnormalities of cranial CT guided by imaging prior knowledge provided by the present invention.
[0134] like Fig. 9 , which is a schematic diagram of the structure of an electronic device system for a method for real-time detection of multiple abnormalities in cranial CT guided by prior knowledge of imaging provided by an embodiment of the present invention. The electronic device system may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a program for real-time detection of multiple abnormalities in cranial CT guided by prior knowledge of imaging.
[0135] Among them, the processor 10 can be composed of an integrated circuit in some embodiments, for example, it can be composed of a single packaged integrated circuit, or it can be composed of multiple integrated circuits with the same function or different functions, including one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors and various control chips. The processor 10 is the control core (Control Unit) of the electronic device system, which uses various interfaces and lines to connect various components of the entire electronic device system, and executes or executes the programs or modules stored in the memory 11 (for example, executing a real-time detection method for multiple abnormalities of cranial CT guided by prior knowledge of imaging, etc.), and calls the data stored in the memory 11 to execute various functions of the electronic device system and process data.
[0136] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of an electronic device system, such as a mobile hard disk of the electronic device system. In other embodiments, the memory 11 may also be an external storage device of an electronic device system, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device system. Further, the memory 11 may also include both an internal storage unit of the electronic device system and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device system, such as the code of a method program for real-time detection of multiple abnormalities of cranial CT guided by prior knowledge of imaging, but can also be used to temporarily store data that has been output or is to be output.
[0137] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.
[0138] The communication interface 13 is used for communication between the above-mentioned electronic device system and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device system and other electronic device systems. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device system and to display a visual user interface.
[0139] Fig. 9 Only an electronic device system with components is shown, and those skilled in the art can understand that Fig. 9The structure shown does not constitute a limitation on the electronic device system, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0140] For example, although not shown, the electronic device system may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, and power status indicators. The electronic device system may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0141] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited by this structure.
[0142] Furthermore, if the module / unit integrated in the electronic device system 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0143] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", "an implementation", "a preferred implementation" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0144] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A real-time detection method for multiple abnormalities in cranial CT guided by prior knowledge of imaging, characterized in that: include: Real-time acquisition of brain CT image sequences; Input the brain CT image sequence into the pre-trained anomaly detection model to obtain anomaly detection results; The anomaly detection model includes: A backbone network is used to extract multi-scale features of a cranial CT image sequence, wherein the backbone network includes a plurality of cascaded feature extraction modules, and at least one feature extraction module is a symmetric information attention module; The neck network fuses the features of different scales output by the backbone network to obtain multiple fused features; The multi-detection head network processes multiple fusion features separately to obtain anomaly detection results; Among them, the symmetric information attention module includes: A contralateral difference map attention network or two or more cascaded contralateral difference map attention networks for extracting bilateral difference information of the input features of the symmetry information attention module; The difference information acquisition unit obtains the difference between the bilateral difference information and the input features of the symmetric information attention module, which is recorded as the difference feature; The feature enhancement unit enhances the input features of the symmetric information attention module based on the difference features to obtain bilateral difference enhanced features.
2. The real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge as claimed in claim 1, characterized in that: The contralateral difference map attention network adopts K map attention mechanisms to respectively extract the map attention features of the input features of the contralateral difference map attention network, concatenates the K map attention features and uses the concatenated features as the output features of the contralateral difference map attention network, where K is a positive integer.
3. The real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge as claimed in claim 2, characterized in that: The process of each graph attention mechanism extracting the graph attention features of the input features of the contralateral difference graph attention network includes: Divide the input features of the contralateral difference map attention network into multiple patches without overlapping; Project each feature point of each patch into the query embedding space, key embedding space and value embedding space respectively to obtain query embedding, key embedding and value embedding; Treat each feature point of each patch as a node, and use the query embedding, key embedding, and value embedding of the feature point as node attributes; Construct a graph network for patch P, which includes a terminal node set, a starting node set, and a directed edge set. All nodes inside patch P constitute the terminal node set, and the patch symmetric to the left and right of patch P along the central axis of the input feature of the symmetric information attention module All internal nodes form the starting node set, and the directed edge set includes all directed edges from the starting nodes to each terminal node. The attribute information of each directed edge is calculated through the attribute information of the starting node and the terminal node; P, Both represent the index of the patch, and are positive integers; The graph attention sub-feature of each node is obtained through the directed edge attribute information of each patch and the value embedding of each node; The graph attention sub-features of all nodes of all patches constitute the graph attention features output by the contralateral difference graph attention network under each graph attention mechanism.
4. The method for real-time detection of multiple abnormalities in cranial CT guided by imaging prior knowledge as claimed in any one of claims 1 to 3, characterized in that: The feature enhancement unit includes: The global average pooling layer processes the difference features to obtain pooling features; Compression and excitation processing unit, which obtains the channel weight map based on the pooling feature; The multiplication unit multiplies the channel weight map with the input features of the symmetric information attention module to obtain the weighted residual features; The adding unit adds the weighted residual feature and the bilateral difference information to obtain the bilateral difference enhancement feature.
5. The real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge as claimed in claim 4, characterized in that: At least one feature extraction module located in a shallow layer of the backbone network is a context information aggregation module.
6. The real-time detection method for multiple abnormalities in cranial CT scans guided by prior knowledge of imaging as claimed in claim 5, characterized in that: The context information aggregation module includes: 2D convolution branch, including 2D-CBS modules and first channel attention units connected in sequence; 3D convolution branch, including a reshape and dimension reshaping unit, a 3D-CBS module, a dimension reshaping and reshape unit, and a second channel attention unit connected in sequence; Aggregation unit, used to aggregate the output features of the 2D convolution branch and the output features of the 3D convolution branch.
7. The method for real-time detection of multiple abnormalities in cranial CT guided by imaging prior knowledge as claimed in claim 1, 2, 3, 5 or 6, characterized in that: The anomaly detection model is trained by an enhanced image sequence sample set, and the method for obtaining the enhanced image sequence sample set includes: Obtain an image sequence sample set, each image sequence sample includes 6 consecutive slices; A plurality of enhanced image sequence samples are obtained by using a 3D symmetric mosaic method, and the plurality of enhanced image sequence samples are added into an image sequence sample set to obtain an enhanced image sequence sample set; The 3D symmetric mosaic method is as follows: randomly select 4 image sequence samples from the image sequence sample set, initialize the slice index s=0; when 0≤s≤5, repeat steps a to c to generate the sth slice of the enhanced image sequence sample: Step a, dividing the blank image into six non-overlapping blank areas, the six blank areas are distributed in two rows, the first row is defined as the third blank area, the first blank area and the fourth blank area from left to right, and the second row is defined as the fifth blank area, the second blank area and the sixth blank area from left to right; Step b, on the s-th slice of the first image sequence sample extracted, cropping is performed using a first random width and height box passing through the central axis and symmetrical along the central axis, and placing the cropped box image in the first blank area of the blank image; On the s-th slice of the extracted second image sequence sample, a second random width and height box that passes through the central axis and is symmetrical along the central axis is used for cropping, and the cropped box image is placed in the second blank area of the blank image; On the s-th slice of the third image sequence sample extracted, two third random width and height boxes are used for cropping that are symmetrical along the central axis, and the two box images obtained by cropping are placed in the third blank area and the fourth blank area of the blank image respectively; On the s-th slice of the fourth image sequence sample extracted, two fourth random width and height boxes are used for cropping that are symmetrical along the central axis, and the two box images obtained by cropping are placed in the fifth blank area and the sixth blank area of the blank image respectively; Step c: the blank image after placing the box images in the six blank areas is used as the s-th slice of the enhanced image sequence sample, and s=s+1.
8. The real-time detection method for multiple abnormalities in cranial CT guided by prior knowledge of imaging as claimed in claim 7, characterized in that: The step of obtaining an image sequence sample set includes: Obtaining original brain CT image sequences; Perform brain window processing and subdural window processing on the original cranial CT image sequence; The original cranial CT image sequence after brain window processing and subdural window processing is sequentially divided into multiple image sequence samples with the same length, there is at least one overlapping slice between adjacent image sequence samples, and the multiple image sequence samples constitute an image sequence sample set.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for real-time detection of multiple abnormalities in cranial CT guided by prior knowledge of imaging described in any one of claims 1 to 8 are implemented.
10. An electronic equipment system, characterized in that: The electronic equipment system comprises: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for real-time detection of multiple abnormalities in cranial CT guided by imaging prior knowledge as described in any one of claims 1-8.
Citation Information
Patent Citations
Cow detection method based on ResMO-Dense-YOLO
CN117912062A
Infrared weak and small target detection method based on improved YOLO v8
CN118982658A
Target single edge detection method based on YOLO
CN119090905A
Cerebral stroke CT image segmentation method
CN113177943A
Medical image abnormal signal intensity detection method
CN115115576A
Cited By
Intelligent ultrasonic training system
CN120319090A
3D medical image anomaly detection and classification method
CN120635595A
Registration method and system for mouse brain tissue slice images
CN121837330A