Method and system for real-time detection of multiple abnormalities in cranial CT guided by imaging prior knowledge

By introducing a symmetric information attention module and a context information aggregation module in the craniocerebral CT image detection, combined with 3D symmetric mosaic enhancement, the problems of complex alignment and insufficient symmetry in the prior art are solved, and a more efficient abnormal detection effect is achieved.

CN119941679BActive Publication Date: 2025-07-22ARMY MEDICAL UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510021449.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-07-22
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The prior art has complex alignment preprocessing steps and insufficient symmetry in the detection of craniocerebral CT imaging abnormalities, and it is difficult to effectively use imaging prior knowledge to improve detection accuracy and reliability, especially under the influence of individual differences and lesion factors.

Method used

Using the imaging prior knowledge guidance method, by adding a symmetric information attention module to the backbone network, using the contralateral difference graph attention network and the difference information acquisition unit, enhance feature expression, and combine the context information aggregation module and the 3D symmetric mosaic enhancement strategy to improve the flexibility and accuracy of the detection model.

Benefits of technology

It improves the accuracy and reliability of abnormal detection of craniocerebral CT images, and can effectively capture abnormal areas when symmetry is damaged, simplifies alignment operations and improves model flexibility and detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941679B_ABST
    Figure CN119941679B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing, and provides a method for real-time detection of various abnormalities in cranial CT guided by prior knowledge of imaging, including: obtaining a real-time cranial CT image sequence; inputting the cranial CT image sequence into a pre-trained abnormality detection model to obtain an abnormality detection result; the abnormality detection model includes: a backbone network for extracting multi-scale features of the cranial CT image sequence, the backbone network includes a plurality of cascaded feature extraction modules, and at least one feature extraction module is a symmetric information attention module; a neck network and a multi-detection head network; the symmetric information attention module includes: a contralateral difference map attention network or two or more cascaded contralateral difference map attention networks, a difference information acquisition unit, and a feature enhancement unit. The present invention also discloses a computer program product and an electronic device system. The present invention increases the flexibility of matching bilateral symmetric anatomical structures and improves the accuracy and reliability of abnormality detection in cranial CT images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method and system for real-time detection of multiple abnormalities in cranial CT guided by prior knowledge of imaging. Background Art

[0002] Many researchers have tried to incorporate prior knowledge of cranial imaging into algorithms to improve the accuracy and reliability of cranial abnormality detection models. There are mainly the following two types of prior knowledge of imaging:

[0003] The first is the prior knowledge of imaging of cerebral hemisphere symmetry. Many studies have shown that using the prior knowledge of imaging of cerebral hemisphere symmetry can effectively guide model learning. However, in the NCCT images (Non-Contrast CT, that is, plain CT) obtained in clinical practice, due to the difference in the patient's scanning position, strict left-right symmetry is often not presented. In response to this challenge, the following alignment strategies have emerged in the prior art: using a template to register the slices to the standard space first, and then inputting them into the abnormality detection model; or, obtaining an affine transformation matrix through an alignment network (or finding the midline of the brain through a midline localization network of the brain, and then calculating the affine transformation matrix), using the affine transformation matrix to align to the standard space, and then inputting it into the abnormality detection model. The essence of the above alignment strategies is to align the original image to the standard space, and the abnormality detection model itself does not have the ability to handle the diverse spatial distributions in the real scene, and has the disadvantages of poor practicability and complex preprocessing process. In addition, there is another alignment strategy in the prior art, which is to embed a network module that specifically uses symmetry knowledge in the deep structure of the model neural network, avoiding the complex preprocessing process. However, even in the deep layer of the network, when the receptive field of the feature points becomes larger, due to factors such as individual differences and lesions in the actual brain images, the feature map in the deep layer of the network cannot be strictly symmetric, resulting in limited contribution of the prior knowledge of imaging of cerebral hemisphere symmetry to the accuracy and reliability of the abnormality detection model.

[0004] The second is the anatomical structure continuity between slices of three-dimensional medical volume images of the brain. In the field of medical imaging, especially for NCCT sequences of the brain, there is a tight anatomical structure continuity between slices. This characteristic makes it possible to aggregate rich contextual information from volumetric medical images. Traditional slice-based 2DCNN networks are difficult to achieve inductive bias in the cross-sectional direction because they cannot cross slice boundaries and are difficult to effectively capture and utilize contextual information in three-dimensional space. 3D convolutional neural networks mainly use 3D convolutions in the backbone network to extract features to utilize the continuity features between adjacent slices, but 3D convolutions also bring a significant increase in model complexity. Compared with 2D convolutions, the number of parameters increases significantly and the inference speed slows down. In addition, 3D convolutions carry the risk of exacerbating the attenuation of tiny abnormal features. Moreover, the NCCT images used clinically are not obtained by thin-layer scanning, and their slice thickness is usually large (usually 5mm). This characteristic poses a challenge to the performance of 3D CNNs when processing NCCT images because excessive aggregation of contextual information may cause tiny abnormal lesions confined to a single slice to be obscured or diluted. With multiple executions of the downsampling process, the features of these tiny lesions may gradually weaken or even be completely lost.

[0005] In the field of object detection, the YOLO series network architectures generally consist of three core structures: backbone, neck, and detect head. Its working process is as follows: First, the backbone is responsible for processing and extracting feature information from the preprocessed input image; subsequently, in the neck part, the PANET (Path Aggregation Network) technology is used to fuse features from different scales to enhance the expressiveness of features; finally, this information is passed to three detection heads, which accurately predict the position coordinates, confidence scores, and the categories to which the targets belong at different scales. Although there are slight differences in the cranial structures of different individuals, overall, they present relatively fixed shapes and anatomical structures. Compared with natural images, the diversity and complexity of NCCT images in terms of color, texture, and shape are relatively low. Therefore, for the analysis task of NCCT images, a relatively small network depth and width may be sufficient to provide sufficient feature representations. In this case, simply increasing the network depth and width may not be the most effective strategy. Instead, integrating imaging prior knowledge into a relatively small network architecture to fully explore and utilize the more unique feature information in cranial NCCT images may be a more efficient method.

[0006] In summary, in the anomaly detection models for deep learning of three-dimensional medical volume images of the brain, numerous studies have been dedicated to leveraging imaging prior knowledge to guide the learning process of the models. However, there are still two major aspects in this field that are worthy of further exploration and deepening: one is how to endow the model with a powerful contralateral region perception ability while avoiding complex and cumbersome alignment preprocessing steps; the other is how to efficiently and effectively utilize the anatomical structure continuity between slices. On the other hand, the existing YOLO object detection network is designed and optimized for natural images. When used for the detection of various anomalies in the brain, it needs to be specifically designed according to the unique imaging prior knowledge of three-dimensional medical volume images of the brain. Summary of the Invention

[0007] The present invention aims to at least solve the technical problems existing in the prior art, and provides a method and system for real-time detection of multiple anomalies in brain CT guided by imaging prior knowledge.

[0008] To achieve the above object of the present invention, according to the first aspect of the present invention, there is provided a method for real-time detection of multiple anomalies in brain CT guided by imaging prior knowledge, including:

[0009] Real-time acquisition of a brain CT image sequence;

[0010] Inputting the brain CT image sequence into a pre-trained anomaly detection model to obtain an anomaly detection result; the anomaly detection model includes:

[0011] A backbone network for extracting multi-scale features of the brain CT image sequence, the backbone network includes a plurality of cascaded feature extraction modules, and at least one feature extraction module is a symmetric information attention module;

[0012] A neck network for fusing features of different scales output by the backbone network to obtain a plurality of fused features;

[0013] A multi-detection head network for processing the plurality of fused features respectively to obtain an anomaly detection result;

[0014] Wherein, the symmetric information attention module includes:

[0015] One contralateral difference map attention network or two or more cascaded contralateral difference map attention networks for extracting bilateral difference information of the input features of the symmetric information attention module;

[0016] A difference information acquisition unit for acquiring the difference between the bilateral difference information and the input features of the symmetric information attention module, denoted as difference features;

[0017] A feature enhancement unit for enhancing the input features of the symmetric information attention module based on the difference features to obtain bilateral difference enhanced features.

[0018] To achieve the above object of the present invention, according to the second aspect of the present invention, there is provided a computer program product, including a computer program which, when executed by a processor, implements the steps of the method for real-time detection of multiple abnormalities in cranial CT guided by imaging prior knowledge described in the first aspect of the present invention.

[0019] To achieve the above object of the present invention, according to the third aspect of the present invention, there is provided an electronic device system, the electronic device system includes:

[0020] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the method for real-time detection of multiple abnormalities in cranial CT guided by imaging prior knowledge as described in the first aspect of the present invention.

[0021] The above technical solution provided by the present invention: The abnormality detection model of the present invention is an improvement on the existing YOLO real-time object detection network. At least one symmetric information attention module is added to the backbone network. The symmetric information attention module includes one or two cascaded contralateral difference map attention networks. The contralateral difference map attention network explicitly encodes the imaging prior knowledge of brain symmetry into the model to effectively perceive the information differences of the same anatomical structures on the contralateral side and obtain bilateral difference information. Then, the difference information acquisition unit calculates the difference between the bilateral difference information and the input features of the symmetric information attention module to obtain difference features, and enhances the input features of the symmetric information attention module according to the difference features, enhancing the expression of important features and suppressing secondary features at the same time, improving the effectiveness of feature representation. It can be seen that the symmetric information attention module of the present invention explicitly encodes the imaging prior knowledge of brain symmetry into the backbone network, avoiding the complex alignment operations of the existing methods, and also making up for the problem of insufficient attention to contralateral information in the existing methods, increasing the flexibility of matching bilateral symmetric anatomical structures, so that even when the symmetry is damaged to a certain extent, the abnormality detection model of the present invention can still capture contralateral information to assist in the detection of abnormal regions (such as lesions), thereby improving the accuracy and reliability of cranial image abnormality detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a schematic flow chart of the method for real-time detection of multiple abnormalities in cranial CT guided by imaging prior knowledge in a preferred embodiment of the present invention;

[0023] Figure 2 is a schematic network structure diagram of the abnormality detection model in a preferred embodiment of the present invention;

[0024] Figure 3 It is a schematic structural diagram of the symmetric information attention module in a preferred embodiment of the present invention;

[0025] Figure 4 It is a schematic diagram of the association between feature points and regions in a preferred embodiment of the present invention;

[0026] Figure 5 It is a schematic diagram of the contralateral difference map attention network in a preferred embodiment of the present invention;

[0027] Figure 6 It is a schematic structural diagram of the context information aggregation module in a preferred embodiment of the present invention;

[0028] Figure 7 It is a comparison diagram of the generated images between the traditional mosaic enhancement method and the 3D symmetric mosaic enhancement method provided by the present invention;

[0029] Figure 8 It is a schematic process diagram of generating a slice by the 3D symmetric mosaic enhancement method in a preferred embodiment of the present invention;

[0030] Figure 9 It is a schematic structural diagram of the electronic device system in a preferred embodiment of the present invention. Detailed implementation manners

[0031] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention.

[0032] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0033] In the description of the present invention, unless otherwise specified and defined, it should be noted that the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.

[0034] The execution subjects of the real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge provided by the present invention include, but are not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the preferred embodiment of the present application. In other words, the real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0035] The present invention provides a real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge. In a preferred embodiment, as Figure 1 shown, it includes:

[0036] Step S1, obtaining a cranial CT image sequence in real time.

[0037] Preferably, the original cranial CT image sequence can be obtained from a CT device in real time. The original cranial CT image sequence can be an NCCT sequence. The original cranial CT image sequence is preprocessed to generate multiple cranial CT image sequences, and at least one cranial CT image sequence is input into the abnormality detection model to obtain the corresponding abnormality detection result. In this embodiment, the preprocessing includes:

[0038] Step S11: Perform brain window processing and subdural window processing on the original cranial CT image sequence. The images in the original cranial CT image sequence can be in DICOM (Digital Imaging and Communications in Medicine) format. For each slice in the original cranial CT image sequence, windowing processing is performed. Windowing processing is to select a specific sub-interval (i.e., window) within the wide gray level range of the DICOM image to highlight the information of interest. In this embodiment, a brain window and a subdural window are set. In the brain window setting, through the set gray level sub-interval of the brain window, the abnormal categories to be detected can be intuitively observed. The different categories include at least two of the eight abnormal categories of cerebral ischemia, cerebral hemorrhage, softening focus, leukoaraiosis, scalp hematoma, mass, extra-cerebral fluid, and midline shift, so that the abnormal category areas can be clearly shown under the brain window. Research shows that the subdural window can also provide important medical details, especially for subdural lesions, so a subdural gray level sub-interval is set. In this study, two windowing processing strategies of the brain window and the subdural window are applied to the same DICOM image (i.e., slice). Specifically, the image data after brain window processing is assigned to the red (R) and green (G) channels of the RGB image, while the image data after subdural window processing is placed in the blue (B) channel, and stacked into the three channels of the RGB image. By windowing, the gray level value range of the image can be adjusted to highlight the gray level value ranges of different tissues, so as to more clearly observe specific tissues. The gray level ranges of the brain window and the subdural window are adjusted differently, and placing them on different channels can make the information of the image more abundant, which helps to optimize the recognition of anatomical structures and lesion areas. The slices after windowing processing are as Figure 8 shown.

[0039] Step S12: Sequentially divide the original cranial CT image sequence after brain window processing and subdural window processing into multiple subsequences with the same length. Each subsequence is used as a cranial CT image sequence to be input into the abnormal detection model, and there is at least one overlapping slice between adjacent subsequences.

[0040] Specifically, it can be divided into multiple subsequences with the same length according to the index order of the slices in the original cranial CT image sequence after brain window processing and subdural window processing. The length range of the subsequences can be from 4 to 12, preferably 6. There is at least one overlapping slice between adjacent subsequences, which means that the tails and heads of the two adjacent subsequences overlap, that is, the last slice or more of the previous subsequence is the same as the first slice or more of the next subsequence. In one example, the last two pictures of the previous subsequence are the same as the first two pictures of the next subsequence. If the last subsequence cannot completely contain the pictures with the preset length, the number of overlapping pictures with the penultimate subsequence can be increased to ensure that the length of all subsequences reaches the preset length.

[0041] Step S2: Input the sequence of cranial CT images into a pre-trained anomaly detection model to obtain the anomaly detection result. When the anomaly detection model detects an abnormal area, the anomaly detection result includes the position coordinates of each abnormal area (which can be the position coordinates of the bounding box of the abnormal area), the type of anomaly of the abnormal area, and the confidence score. At the same time, the abnormal area is marked with a bounding box in the sequence of cranial CT images, as Figure 2 shown. When the anomaly detection model does not detect an abnormal area, it outputs NULL or normal.

[0042] In this embodiment, the anomaly detection model is obtained by improving the Y0L0 series network architecture, such as improving based on the Y0L0V8 network structure.

[0043] As Figure 2 shown, the anomaly detection model includes:

[0044] The backbone network ( Figure 2 the backbone network in Figure 2 ) is used to extract multi-scale features of the sequence of cranial CT images. The backbone network includes a plurality of cascaded feature extraction modules, and at least one feature extraction module is a Symmetric Information Attention (SIA) module. As

[0045] shown, the backbone network may include a C2f module, an SIA module, a CBS module, a C2f module, a CBS module, a C2f module, and an SPPF module connected in sequence. Figure 2 The neck network (

[0046] the neck network in Figure 2 ) fuses the features of different scales output by the backbone network to obtain multiple fused features to enhance the expressiveness of the features. The neck network includes a first branch and a second branch. The first branch includes an Upsample module, a Contact layer, a C2f module, an Upsample module, a Contact layer, and a C2f module connected in sequence. The second branch includes a CBS module, a Contact layer, a C2f module, a CBS module, a Contact layer, and a C2f module connected in sequence.

[0047] In this embodiment, the connection relationship between the modules in the anomaly detection model can refer to Figure 2As shown, it will not be elaborated here. The C2f module, CBS module, SPPF module, and upsampling module Upsample in the backbone network and neck network can all refer to the Y0L0V8 network. The network structures in other existing technologies can also be referred to. For example, the CBS module can refer to the CBS module in the existing patent with the publication number CN119090905A. The C2f module can refer to the C2f structure in the existing patent with the publication number CN118982658A. The SPPF module can refer to the SPPF structure in the existing patent with the publication number CN117912062A. It will not be elaborated here.

[0048] Among them, as Figure 3 shown, the SIA module, i.e., the symmetric information attention module, includes:

[0049] One contralateral difference graph attention network or two or more cascaded contralateral difference graph attention networks, which are used to extract the bilateral difference information of the input features of the symmetric information attention module. The contralateral difference graph attention network (Bilateral Difference Graph attention network, abbreviated as BDGA network).

[0050] The difference information acquisition unit acquires the difference between the bilateral difference information and the input features of the symmetric information attention module, denoted as the difference feature. As Figure 3 shown, the difference information acquisition unit performs an element-wise subtraction operation.

[0051] The feature enhancement unit performs enhancement processing on the input features of the symmetric information attention module based on the difference feature to obtain bilateral difference enhanced features.

[0052] In this embodiment, only one contralateral difference graph attention network (BDGA network) can be set in the symmetric information attention module (SIA module). To enable the model to focus on the contralateral feature regions that have not been fully explored previously and achieve effective compensation of symmetric difference information, two or more serial contralateral difference graph attention networks (BDGA networks) can also be set in the symmetric information attention module (SIA module). At this time, the windows of the two or more contralateral difference graph attention networks (BDGA networks) do not overlap, that is, the patch division window. As Figure 4 shown, the window partition of the second contralateral difference graph attention network (BDGA network) will be shifted based on the first contralateral difference graph attention network (BDGA network). Through window adjustment, the contralateral difference graph attention network can cover different combinations of nodes and directed edges at different positions. As Figure 4As shown, the first column of pictures shows the window partition of the first contralateral difference graph attention network (BDGA network), and the second column of pictures shows the window partition of the second contralateral difference graph attention network (BDGA network). The window partition of the second contralateral difference graph attention network (BDGA network) is shifted compared to the window partition of the first contralateral difference graph attention network (BDGA network).

[0053] When there are two cascaded contralateral difference graph attention networks, the bilateral difference information X dif is expressed as:

[0054] X dif = BDGA(BDGA(X));

[0055] where X represents the input feature of the symmetric information attention module, and BDGA(·) represents the operation of the contralateral difference graph attention network.

[0056] In a preferred embodiment, in order to be able to extract different feature information and help improve the accuracy of anomaly detection, the contralateral difference graph attention network BDGA uses K graph attention mechanisms to respectively extract the graph attention features of the input features of the contralateral difference graph attention network BDGA, splices the K graph attention features, and uses the spliced features as the output features of the contralateral difference graph attention network BDGA, where K is a positive integer.

[0057] In this embodiment, when K is 1, the graph attention feature obtained by a single graph attention mechanism is used as the output feature of the contralateral difference graph attention network BDGA; when K is greater than 1, the graph attention features of the input features of the contralateral difference graph attention network extracted by K graph attention mechanisms are spliced, and the spliced features are used as the output features of the contralateral difference graph attention network BDGA.

[0058] In this embodiment, the cranial anatomical structure of the human body is roughly symmetrical left and right. Therefore, it is a commonly used diagnostic method for clinicians to identify abnormal regions by comparing the left and right brains, which actually establishes a connection between the left and right relatively symmetrical anatomical structures of the brain. If this diagnostic process is simulated and the imaging knowledge of the bilateral symmetry attribute of the cerebral hemisphere is incorporated into the deep learning model, the key lies in constructing the relationship of the left and right relatively symmetrical anatomical structures. Under ideal conditions, for the anatomical structure p, it is only necessary to associate it with the symmetric anatomical structure But in the actual clinical scenario, due to reasons such as patient positioning differences, the symmetry of the brain images may be affected to a certain extent. If the original images are not aligned to the standard space, then accurately finding is relatively difficult. Even when the receptive fields of feature points become larger in the deeper layers of the network, it is difficult to ensure the perfect symmetry of feature points on both sides using contralateral difference learning. This is due to factors such as individual differences and lesions existing in actual brain images. Based on this, the present invention provides a solution method. Although the goal is only to establish the association between p and but to increase the possibility of finding the present invention associates p with a region that may contain such as shown, associate the feature point p Figure 4 with the region where its left and right symmetric a is located. Similarly, associate p with the region where its left and right symmetric is located. The advantage of doing this is that it increases the flexibility of matching bilateral symmetric anatomical structures, so that even when the symmetry is damaged to a certain extent, the abnormal detection model of the present invention can still capture b and associate. The information of is used to assist in the detection of the lesion area. In the process of constructing the complex internal relationships of medical images, Graph Neural Networks (GNNs) have become a very attractive choice due to their powerful representation learning ability and effective capture of potential structures and relationships in images. Therefore, the present invention proposes a contralateral difference graph attention network, which uses K graph attention mechanisms to respectively perceive contralateral differences.

[0059] In this embodiment, referring to Figure 5 , the process of each graph attention mechanism extracting the graph attention features of the input features of the contralateral difference graph attention network includes:

[0060] Step A1, divide the input features of the contralateral difference graph attention network into multiple non-overlapping patches.

[0061] For the input feature F of the contralateral difference graph attention network, apply a sliding window operation with a size of l×l to divide it into non-overlapping patches, and each patch is a set of feature points i and j respectively represent the abscissa and ordinate of the feature point in the patch, d is the dimension of the feature point, and p ij represents the feature of the position point (i, j) in the patch.

[0062] Step A2, to achieve diverse features, project each feature point of each patch into the query embedding space, key embedding space, and value embedding space respectively to obtain query embeddings, key embeddings, and value embeddings.

[0063] These feature points in the patch are projected into three different embedding spaces of query, key, and value using three separate linear projection layers. This process can be expressed as:

[0064] q ij = P ij W Q ,k ij = P ij W K ,v ij = P ij W v

[0065] where W Q ,W K , are the mapping matrices of the query embedding space, key embedding space, and value embedding space respectively. q ij ,k ij and v ij represent the query embedding, key embedding, and value embedding of p ij respectively. Considering the combination of these three embeddings as a graph node n ij can represent different characterizations of the anatomical structure within the receptive field area where p ij is located. Among them, q ij is used to evaluate the potential contribution or influence of the current node on other nodes, k ij 's main goal is to explore and measure the similarity or correlation between the anatomical structures represented by other graph nodes and the current node, while v ij serves to provide the unique anatomical features or attributes represented by each graph node. k represents the index of the current graph attention mechanism, k ∈ [1, K].

[0066] Step A3: Take each feature point of each patch as a node (i.e., a graph node), and use the query embedding, key embedding, and value embedding of the feature point as node attributes.

[0067] Step A4: Construct a graph network for patch P to express the anatomical structure symmetry relationship through the graph network. The graph network includes a set of terminal nodes, a set of starting nodes, and a set of directed edges. Among them, all the nodes inside patch P form the set of terminal nodes, and the patch symmetric to patch P along the central axis of the input features of the symmetric information attention module and all the nodes inside it form the set of starting nodes. The set of directed edges includes directed edges from each starting node to each terminal node respectively. Calculate the attribute information of each directed edge through the attribute information of the starting node and the terminal node; P, Both represent the index of the patch and are positive integers.

[0068] In this embodiment, each patch is represented as a graph network. The graph network of patch P is where N is the set of all graph nodes within patch P, serving as the set of terminal nodes, while serves as the set of starting nodes, is the set of directed edges. For node n within patch P ij , the present invention finds the patch that is symmetric to patch P along the central axis of the input feature F of the contralateral difference graph attention network and all the graph nodes within serve as the neighbor nodes of node n ij ; represents the starting node to the directed edge ij from the starting node to the terminal node n of the attribute information. The degree of all nodes in graph network G is fixed at l 2 , and graph network G constructs the anatomical structure symmetry relationship between patch P and patch . Calculate their Scaled Dot-Product Attention to quantify the similarity degree of the anatomical structure features represented by the key embedding k ij of the terminal node n ij and the query embedding of the starting node , thereby obtaining the attribute information of the directed edge which is also called the attention coefficient:

[0069]

[0070] where T represents taking the matrix transpose; the attention coefficient guides the node to propagate how much information to n ij , and take as the directed edge attribute information from the starting node to the terminal node n ij .

[0071] Step A5, obtain the graph attention sub-feature of each node through the directed edge attribute information of each patch and the value embedding of each node.

[0072] Step A5 implements the contralateral difference perception mechanism. When the terminal node n ij contains lesion information, it may be difficult to find similar features in its neighbor node set . On the contrary, if the terminal node n ij does not contain lesion information, approximate features can be found in . This difference is directly reflected in the directed edge attribute information . Therefore, for the directed connection from the starting node to the terminal node n ij , the present invention uses the directed edge attribute information and the specific anatomical structure representation v ij of the terminal node n ij (i.e., value embedding) to aggregate neighbor information and obtain the graph attention sub-feature v ij ' of the node n in the patch: ij ':

[0073]

[0074] where σ is the GELU activation function; W o is the output matrix and is learnable.

[0075] Step A6. The graph attention sub-features of all nodes in all patches form the graph attention features output by the contralateral difference graph attention network under each graph attention mechanism.

[0076] In this embodiment, the K graph attention mechanisms in the contralateral difference graph attention network BDGA respectively extract the graph attention features of the input features of the contralateral difference graph attention network BDGA, splice the K graph attention features, and use the spliced features as the output features of the contralateral difference graph attention network BDGA. Specifically, the K graph attention sub-features obtained by each feature point under the K graph attention mechanisms are spliced to obtain the spliced sub-features. The spliced sub-features of all nodes in all patches form the graph attention features output by the contralateral difference graph attention network. The spliced sub-feature v ij ' of the node n ijK is expressed as:

[0077]

[0078] where II represents splicing, is the attention coefficient of the k-th graph attention mechanism, is the weight matrix of the output linear transformation of the k-th graph attention mechanism.

[0079] In a preferred embodiment, referring to Figure 3 , the feature enhancement unit includes:

[0080] The global average pooling layer (Global Average Pooling, GAP) processes the differential features to obtain pooled features. The differential features represent the differential or complementary information between the bilateral differential information and the original features (the features input to the SIA module).

[0081] The squeeze-and-excitation (SE) processing unit obtains a channel weight map based on the pooled features. The squeeze-and-excitation processing unit includes two fully connected layers. The SE module is jointly constrained by the input feature X and the bilateral differential information here. Its core role is to adaptively recalibrate the feature responses of each channel by explicitly modeling the correlations between feature channels, thereby generating a weight map that can reflect the importance of each channel.

[0082] The multiplication unit multiplies the channel weight map by the input features of the symmetric information attention module to obtain weighted residual features. The multiplication process is element-wise multiplication, which realizes the weighted adjustment of the original features according to the difference degree between the contralateral differential information and the original features. This mechanism enhances the expression of important features, suppresses secondary features at the same time, and improves the effectiveness of feature representation.

[0083] The addition unit adds the weighted residual features and the bilateral differential information to obtain bilateral differential enhanced features. The addition is element-wise addition.

[0084] In this embodiment, the structure of the symmetric information attention module (SIA module) is as shown in Figure 3 and its working mechanism is as follows: First, the input feature X of the SIA module is guided to two cascaded contralateral difference map attention networks BGDA to perceive and extract the bilateral differential information X difSubsequently, this difference information is subtracted from the original input feature X to obtain a difference feature, which represents the difference or complementary information between the bilateral difference information and the original feature. This difference feature then undergoes a processing by a Squeeze-and-Excitation (SE) module that includes a Global Average Pooling (GAP) operation and two fully-connected layers. The SE module is jointly constrained by the input feature X and the bilateral difference information here. Its core role is to adaptively recalibrate the feature responses of each channel by explicitly modeling the correlations between feature channels, thereby generating a channel weight map that can reflect the importance of each channel. Then, the channel weight map is element-wise multiplied with the original input feature X retained through the residual connection to obtain a weighted residual feature. This process realizes the weighted adjustment of the original feature according to the degree of difference between the contralateral difference information and the original feature. This mechanism enhances the expression of important features while suppressing secondary features, improving the effectiveness of feature representation. Finally, the weighted residual feature is added to the bilateral difference information X dif to perform an addition operation to obtain the final output feature of this module, that is, the bilateral difference enhanced feature. SIA can be expressed as:

[0085] SIA(x) = X dif + x × σ(MLP(σ(MLP(GAP(x - X dif )))))

[0086] where the above σ function represents the activation function of the fully-connected layer.

[0087] 2D convolution and 3D convolution each have unique advantages. In some cases, it may be better to focus on the features of a single slice because the local features within the slice are sufficient to provide enough information without introducing the noise or interference that other slices may bring. While some other features rely on context information and need to be captured and analyzed in three-dimensional space to reveal complex structural relationships. Although 3D CNN has the potential to adaptively adjust the weights in the cross-sectional direction through learning and thus has the ability of 2D convolution to a certain extent. However, by explicitly designing the network structure to enhance the sensitivity and capture ability of the model to abnormal features, the model may be easier to learn and thus obtain better abnormal detection performance. Therefore, the present invention designs an effective and efficient Context Information Aggregation (CIA) module to replace the convolutional downsampling module in the baseline model.

[0088] Therefore, in a preferred embodiment, as Figure 2As shown, at least one feature extraction module located in the shallow layer of the backbone network is a context information aggregation module. Shallow layer refers to multiple feature extraction modules in the backbone network close to the input end of the backbone network. For example, Figure 2 As shown, the first, second, and fourth feature extraction modules of the backbone network are set as context information aggregation modules.

[0089] In a preferred embodiment, referring to Figure 6 , the context information aggregation module (CIA module) includes:

[0090] The 2D convolution branch, including a 2D-CBS module and a first channel attention unit connected in sequence; the 2D-CBS module is the CBS module in the existing Y0L08 series network, which includes a two-dimensional convolution, a two-dimensional BatchNorm, and a SiLU activation function unit. Referring to

[0091] The 3D convolution branch, including a reshaping and dimension reshaping unit, a 3D-CBS module, a dimension reshaping and reshaping unit, and a second channel attention unit connected in sequence; the 3D-CBS module includes a 3D convolution.

[0092] An aggregation unit for aggregating the output features of the 2D convolution branch and the output features of the 3D convolution branch.

[0093] In this embodiment, referring to Figure 6 , the context information aggregation module (CIA module) consists of two branches: a 2D convolution branch and a 3D convolution branch. Let the feature input to the CIA module be where L = B×D, B is the batchsize, D is the depth of the input sequence, i.e., the number of slices, and H and W respectively represent the height and width of the input feature map X.

[0094] In the 2D convolution branch, the input feature X is first processed by the 2D-CBS module. The 2D convolution with a 3×3 convolution kernel and a 2-step stride in the 2D-CBS module downsamples the input feature X to obtain the 2D downsampled feature X2. Then, it passes through the first channel attention unit to dynamically adjust the weight of each channel, obtaining the 2D channel weight CA2 for rescaling the channel features. This process enables the network to focus on more critical features, that is, those features that are more discriminative for anomaly detection. Referring to Figure 6 , the first channel attention unit includes a cascaded GAP (global average pooling layer) and a fully connected layer.

[0095] X2 = CBS2(X)

[0096] CA2 = σ2(MLP2(GAP(X2))

[0097] Among them, CBS2(·) represents the processing of the 2D-CBS module; MLP2(·) represents the fully connected layer of the first channel attention unit; σ2(·) represents the activation function of the fully connected layer of the first channel attention unit.

[0098] Compared with the 2D convolution branch, the 3D convolution branch focuses on capturing context information across the slice dimension. First, the input feature X is reshaped and dimensionally reshaped by the reshaping and permuting unit R&P into B×C×D×H×W to meet the input requirements of 3D convolution. The 3D convolution is a 3D convolution module with a kernel size of 3×3×3 and a stride of 2, obtaining a convolutional feature of scale Then, the convolutional feature is input into the permuting and reshaping unit P&R to convert the dimension and shape into a 3D downsampled feature X3 of scale Then, it also passes through the second channel attention unit to focus on key features and obtain the 3D channel weight CA3. Referring to Figure 6 The second channel attention unit includes a cascaded GAP (global average pooling layer) and a fully connected layer.

[0099]

[0100] CA3 = σ3(MLP3(GAP(X3)))

[0101] Among them, CBS represents convolution plus BN plus SiLu; CBS3(·) represents the processing of the 3D-CBS module; MLP3 represents the fully connected layer of the second channel attention unit; R represents the operation of first reshaping and then permuting; represents the operation of first permuting and then reshaping, MLP3(·) represents the fully connected layer of the second channel attention unit; σ3(·) represents the activation function of the fully connected layer of the second channel attention unit.

[0102] The aggregation unit fuses the output feature X2×CA3 of the 2D convolution branch and the output feature X3×CA3 of the 3D convolution branch by point-wise addition, realizing the weighted addition of the slice dimension features extracted by 2D convolution and the cross-slice dimension features extracted by 3D convolution. Therefore, the output feature CIA(X) of the context information aggregation module (CIA module) is:

[0103] CIA(X) = X2×CA3 + X3×CA3

[0104] In this embodiment, embedding the context information aggregation module (CIA module) in the anomaly detection model enables the model to adaptively select the features extracted by the two branches according to the input features, making full use of the advantages of 3D convolution in context information aggregation and compensating for the limitations of 3D convolution in terms of computational efficiency and local feature extraction through 2D convolution.

[0105] In this embodiment, the embedded context information aggregation module (CIA module) is set in the shallow layer of the backbone network in the anomaly detection model, mainly for the following considerations:

[0106] 1. The shallow network is mainly responsible for extracting basic and low-level features. In the shallow layer of the network, the key information of the image sequence (such as texture, edges, and local spatial relationships, etc.) has not been overly abstracted or lost. It is hoped that the advantages of the CIA module in aggregating context information and capturing detailed features can be fully utilized before the information loss accumulates to a significant degree.

[0107] 2. As the network depth increases, the model gradually transitions from extracting basic visual features to extracting higher-level semantic information.

[0108] The CIA module incorporates a 3D convolution module. This characteristic makes its use in the deep layer of the network significantly increase the computational complexity and the number of parameters, thereby increasing the risk of overfitting. Therefore, in the deep network, we are more inclined to use more concise and efficient 2D convolution operations to extract these semantic features to maintain the computational efficiency and generalization ability of the model.

[0109] In clinical practice, doctors often identify abnormal regions by comparing the images of both sides of the brain. The present invention aims to simulate this process, which includes not only the verification and testing links but also the training link. And the premise of all this is to ensure that in the input images, the left brain and the right brain are roughly horizontally symmetric. However, this premise encounters obstacles in the training link. Since the introduction of the mosaic augmentation data augmentation strategy in YOLOv4, YOLO series models, including YOL0v8, have adopted this technology during the training process. As Figure 7 The enhanced image in [reference] is obtained through the traditional mosaic method. The traditional mosaic augmentation randomly uses multiple pictures, randomly scales them, and then randomly distributes and stitches them together, greatly enriching the detection dataset. In particular, the random scaling adds many small targets, making the network more robust. However, as Figure 7As shown in the first image, in the image after applying mosaic enhancement, the inventor found that the position distribution of the brain regions became random, and the horizontal symmetry between the left and right brains could no longer be guaranteed. Paradoxically, mosaic enhancement can effectively improve the model performance. This result undoubtedly highlights the importance of mosaic enhancement, but at the same time, the problem of symmetry disruption it brings must also be solved. In order to utilize the advantages of mosaic enhancement and ensure the symmetric distribution of the left and right brains, the present invention improves the traditional mosaic enhancement, provides a new mosaic enhancement method, and names it the 3D symmetric mosaic method. Figure 7 The second image is the enhanced image obtained by the 3D symmetric mosaic method. Figure 7 The third picture in shows the distribution of the center points of all target bounding boxes in the dataset, where x and y are the relative coordinates with respect to the original image size. It can be seen that the D symmetric mosaic method ensures the symmetric distribution of the left and right brains.

[0110] In a preferred embodiment, to improve the diversity of data in the training process and enhance the performance of the anomaly detection model after training, the anomaly detection model is trained through an enhanced image sequence sample set. The method for obtaining the enhanced image sequence sample set includes:

[0111] Step B1, obtain an image sequence sample set, where each image sequence sample includes 6 consecutive slices.

[0112] Preferably, obtaining the image sequence sample set includes:

[0113] Step B11, obtain the original cranial CT image sequence. Specifically, the original cranial CT image sequence can be obtained from devices such as CT devices and MRI devices.

[0114] Step B12, perform brain window processing and subdural window processing on the original cranial CT image sequence. The brain window processing and subdural window processing have been described in the previous preferred embodiment and will not be elaborated here.

[0115] Step B13, sequentially divide the original cranial CT image sequence after brain window processing and subdural window processing into multiple image sequence samples with the same length. There is at least one overlapping slice between adjacent image sequence samples, and the multiple image sequence samples form the image sequence sample set. The specific process of sequence division has been described in the previous preferred embodiment and will not be elaborated here.

[0116] Step B2, obtain multiple enhanced image sequence samples using the 3D symmetric mosaic method, and add the multiple enhanced image sequence samples to the image sequence sample set to obtain the enhanced image sequence sample set.

[0117] Among them, Figure 8 shows an example of generating a slice by the 3D symmetric mosaic method. Refer to Figure 8, the 3D symmetric mosaic method is as follows: randomly select 4 image sequence samples from the image sequence sample set, and its sequence sample index set D := {d i ∈ N′|0 ≤ i ≤ 3}, where N′ represents the index set of the selected sequence samples. Let represent the s-th slice of sequence d, then the set of all slices in D can be expressed as Initialize the slice index s = 0. When 0 ≤ s ≤ 5, next, use the 3D symmetric mosaic technique to combine the pictures in set X to generate a new sequence data with a length still of 6, that is, the enhanced image sequence sample. Specifically, repeat steps a to c to generate the s-th slice of the enhanced image sequence sample:

[0118] Step a, divide the blank image into 6 non-overlapping blank areas. As Figure 8 shown, the 6 blank areas are distributed in two rows up and down. The first row is defined as the third blank area ③, the first blank area ①, and the fourth blank area ④ from left to right in sequence. The second row is defined as the fifth blank area ⑤, the second blank area ②, and the sixth blank area ⑥ from left to right in sequence. The size of the blank image can be 1024 * 1024.

[0119] Step b, on the s-th slice of the first extracted image sequence sample, use a first randomly sized box that passes through the central axis and is symmetric about the central axis left and right to crop, and place the cropped box image in the first blank area of the blank image; the central axis here is the central axis of the s-th slice of the first extracted image sequence sample.

[0120] On the s-th slice of the second extracted image sequence sample, use a second randomly sized box that passes through the central axis and is symmetric about the central axis left and right to crop, and place the cropped box image in the second blank area of the blank image; the central axis here is the central axis of the s-th slice of the second extracted image sequence sample.

[0121] On the s-th slice of the third extracted image sequence sample, use two third randomly sized boxes that are symmetric about the central axis left and right to crop, and place the two cropped box images in the third blank area and the fourth blank area of the blank image respectively; the central axis here is the central axis of the s-th slice of the third extracted image sequence sample.

[0122] On the s-th slice of the fourth extracted image sequence sample, use two fourth randomly sized boxes that are symmetric about the central axis left and right to crop, and place the two cropped box images in the fifth blank area and the sixth blank area of the blank image respectively; the central axis here is the central axis of the s-th slice of the fourth extracted image sequence sample.

[0123] In step b, area ① is placed That is, the s-th slice of sequence d0. Region ② is placed The cropping method is the same as that of Region ①. Regions ③ and ④ are from Obtained by cropping and sampling two boxes with random width and height that are symmetric about the left and right. Similarly, Regions ⑤ and ⑥ are from The cropping method is the same as that of ③ and ④. This process can be expressed as where Mθ(·) represents the 3D symmetric mosaic operation, and θ represents the parameters such as size and coordinates involved in randomly cropping and placing the sample image, specifically including the coordinates of the lower left corner of Region ①, the coordinates of the upper left corner of Region ②, and 4 random widths and heights.

[0124] Step c, use the blank image after placing the square images in all 6 blank regions as the s-th slice of the enhanced image sequence sample, and let s = s + 1.

[0125] In this embodiment, in order to ensure the continuity of the anatomical structure, the slice cropping regions of the above-mentioned 6 regions should be continuous in the cross-sectional direction (the slice continuous direction). Therefore, the same random parameters are used for all slices, and the index order in the original slice is maintained. In addition, in order to maintain the continuity of the anatomical structure, the slice cropping of the above six regions should be continuous in the cross-sectional direction. Therefore, the same random parameters are used for all slices, and their index order in the original slice remains unchanged. Finally, the enhanced image sequence sample can be expressed as: Finally, resize the slice size of the 1024*1024 enhanced image sequence sample to 512*215.

[0126] The real-time detection method for multiple abnormalities in cranial CT guided by prior knowledge of medical imaging provided by the present invention has the following main improvements:

[0127] (1) To simulate the diagnostic method commonly used by clinicians to identify abnormal regions by comparing the left and right brains, the present invention incorporates the imaging knowledge of the bilateral symmetry attribute of the cerebral hemispheres into the deep learning model and designs a symmetric information attention module. This module avoids the complex alignment operations of existing methods and also makes up for the problem of insufficient attention to contralateral information in existing methods. It uses a graph attention network to explicitly encode the prior knowledge of brain symmetry in the imaging into the model to effectively perceive the information differences of the same anatomical structures on the contralateral side. The cranial anatomical structure of the human body is roughly symmetric left and right. Therefore, identifying abnormal regions by comparing the left and right brains is a commonly used diagnostic method by clinicians, which actually establishes a connection between the left and right relatively symmetric anatomical structures of the brain. To simulate this diagnostic process, the present invention incorporates the imaging knowledge of the bilateral symmetry attribute of the cerebral hemispheres into the deep learning model, and the key lies in constructing the relationship between the left and right relatively symmetric anatomical structures. In an ideal situation, for anatomical structure p, only need to compare it with the symmetric anatomical structure Just make the association. To increase the likelihood of finding , the present invention associates p with a region that may contain . For example, Figure 4 . This increases the flexibility of bilateral symmetric anatomical structure matching, enabling the model to still capture information to assist in the detection of the lesion area even when the symmetry is somewhat disrupted.

[0128] (2) To make full use of the continuous relationship of anatomical structures between adjacent slices and aggregate the context information of volumetric medical images, a simple and effective context information aggregation module CIA is designed. This module introduces an inductive bias in the cross-sectional direction of volumetric medical images, achieving a balance in feature aggregation within slices and between adjacent slices, and effectively integrating the context information. This module has two branches of 2D convolution and 3D convolution, and the features of the two branches are fused by point-wise addition, realizing the weighted addition of the slice-dimensional features extracted by 2D convolution and the cross-slice-dimensional features extracted by 3D convolution. This enables the model to adaptively select the features extracted by the two branches according to the input features, making full use of the advantages of 3D convolution in context information aggregation, and at the same time compensating for the limitation of 3D convolution over-aggregating context information and exacerbating the attenuation of tiny abnormal features through 2D convolution. In addition, compared with existing models that uniformly use 2D or 3D convolution from shallow to deep, the present invention only uses this module in the shallow layer of the network and uses 2D convolution in the deep layer of the network. In this way, not only can context information be effectively aggregated, but also the number of parameters and complexity of the model will not increase significantly.

[0129] (3) The present invention also proposes a data augmentation strategy for 3D symmetric mosaic to ensure the left-right symmetry of cranial images during the training process. Different from natural images, the cranial region in CT images is generally near the center of the image. As Figure 7 shown in the third figure, the center points of all target bounding boxes in the dataset are also distributed near the center of the image and are basically symmetric left and right, and rarely appear at the edges of the image. By comparing the traditional mosaic enhanced image and the 3D symmetric mosaic enhanced image, it can be found that the 3D symmetric mosaic enhanced image is closer to the above distribution, that is, the cranial region is basically concentrated in the central region and is basically symmetric left and right. The feature distribution of symmetric mosaic enhancement is closer to the feature distribution of cranial NCCT images, so it can effectively improve the performance of the model in detecting cranial abnormalities.

[0130] It should be noted that the method for real-time detection of multiple abnormalities in cranial CT guided by the imaging prior knowledge provided by the present invention can also be applied to the detection of multiple abnormalities in other types of cranial imaging sequences such as cranial MRI image sequences.

[0131] The present invention also discloses a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-mentioned real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge provided by the present invention. The computer program product should be understood as a software product that mainly realizes its solution through a computer program, such as a program product integrated in the cloud or a software library.

[0132] The present invention also discloses an electronic device system. In one embodiment, the electronic device system includes at least one processor; and a memory communicatively connected to the at least one processor; wherein,

[0133] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge provided by the present invention.

[0134] As Figure 9 shown, it is a schematic structural diagram of an electronic device system for the real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge provided by an embodiment of the present invention. The electronic device system may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a program for the real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge.

[0135] Among them, the processor 10 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device system, connecting various components of the entire electronic device system through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as executing the real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge, etc.), and calling data stored in the memory 11, to execute various functions of the electronic device system and process data.

[0136] The memory 11 at least includes one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device system, such as the mobile hard disk of the electronic device system. In some other embodiments, the memory 11 can also be an external storage device of the electronic device system, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device system. Further, the memory 11 can also include both the internal storage unit and the external storage device of the electronic device system. The memory 11 can be used not only to store application software installed in the electronic device system and various types of data, such as the code of the program for the real-time detection of multiple abnormalities in head CT guided by prior imaging knowledge, etc., but also to temporarily store the data that has been output or will be output.

[0137] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0138] The communication interface 13 is used for the communication between the above-mentioned electronic device system and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is usually used to establish a communication connection between this electronic device system and other electronic device systems. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device system and to display a visual user interface.

[0139] Figure 9 Only the electronic device system with components is shown, and those skilled in the art can understand that Figure 9The structures shown do not constitute a limitation on the electronic device system, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0140] For example, although not shown, the electronic device system may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device system may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0141] It should be understood that the embodiments are for illustrative purposes only and are not limited by this structure in the scope of the patent application.

[0142] Furthermore, if the modules / units integrated in the electronic device system 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0143] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one implementation manner", "one preferred implementation manner" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0144] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A real-time detection method for multiple abnormalities in cranial CT guided by imaging prior knowledge, characterized in that, Comprising: Obtaining a real-time sequence of cranial CT images; Inputting the cranial CT image sequence into a pre-trained anomaly detection model to obtain an anomaly detection result; The anomaly detection model includes: A backbone network for extracting multi-scale features of the cranial CT image sequence. The backbone network includes a plurality of cascaded feature extraction modules, and at least one feature extraction module is a symmetric information attention module; A neck network for fusing features of different scales output by the backbone network to obtain a plurality of fused features; A multi-detection head network for processing the plurality of fused features respectively to obtain an anomaly detection result; Among them, the symmetric information attention module includes: One contralateral difference map attention network or two or more cascaded contralateral difference map attention networks for extracting bilateral difference information of the input features of the symmetric information attention module; A difference information acquisition unit for acquiring the difference between the bilateral difference information and the input features of the symmetric information attention module, denoted as difference features; A feature enhancement unit for enhancing the input features of the symmetric information attention module based on the difference features to obtain bilateral difference enhanced features; Among them, the contralateral difference map attention network uses K graph attention mechanisms to extract graph attention features of the input features of the contralateral difference map attention network respectively, splices the K graph attention features, and uses the spliced features as the output features of the contralateral difference map attention network, where K is a positive integer; Among them, the process of each graph attention mechanism for extracting the graph attention features of the input features of the contralateral difference map attention network includes: Dividing the input features of the contralateral difference map attention network into multiple non-overlapping patches; Projecting each feature point of each patch into a query embedding space, a key embedding space, and a value embedding space respectively to obtain a query embedding, a key embedding, and a value embedding; Regarding each feature point of each patch as a node, and using the query embedding, key embedding, and value embedding of the feature point as node attributes; Construct a graph network for patch P. The graph network includes a set of terminal nodes, a set of starting nodes, and a set of directed edges. Among them, all the nodes inside patch P form the set of terminal nodes, and the patch symmetric to patch P along the central axis of the input features of the symmetry information attention module. All the nodes inside it form the set of starting nodes. The set of directed edges includes directed edges from each starting node to each terminal node. Calculate the attribute information of each directed edge through the attribute information of the starting node and the terminal node; P, both represent the index of the patch and are positive integers; Obtaining the graph attention sub-features of each node through the directed edge attribute information of each patch and the value embedding of each node; The graph attention sub-features of all nodes of all patches constitute the graph attention features output by the contralateral difference map attention network under each graph attention mechanism.

2. The real-time detection method for multiple abnormalities of cranial CT guided by prior imaging knowledge according to claim 1, wherein The feature enhancement unit includes: A global average pooling layer for processing the difference features to obtain pooled features; A squeeze-and-excitation processing unit for obtaining a channel weight map based on the pooled features; A multiplication unit for multiplying the channel weight map with the input features of the symmetric information attention module to obtain weighted residual features; An addition unit for adding the weighted residual features and the bilateral difference information to obtain bilateral difference enhanced features.

3. The real-time detection method for multiple abnormalities of head CT guided by prior imaging knowledge according to claim 2, wherein At least one feature extraction module located in the shallow layer of the backbone network is a context information aggregation module.

4. The real-time detection method for multiple abnormalities of head CT guided by prior imaging knowledge according to claim 3, characterized in that, The context information aggregation module includes: A 2D convolution branch including a 2D-CBS module and a first channel attention unit connected in sequence; A 3D convolution branch including a shape adjustment and dimension reshaping unit, a 3D-CBS module, a dimension reshaping and shape adjustment unit, and a second channel attention unit connected in sequence; An aggregation unit for aggregating the output features of the 2D convolutional branch and the output features of the 3D convolutional branch.

5. The real-time detection method for multiple abnormalities of head CT guided by prior imaging knowledge as claimed in claim 1 or 2 or 3 or 4, characterized in that The anomaly detection model is trained by an enhanced image sequence sample set, and the method for obtaining the enhanced image sequence sample set includes: Obtain an image sequence sample set, where each image sequence sample includes 6 consecutive slices; Use the 3D symmetric mosaic method to obtain multiple enhanced image sequence samples, and add the multiple enhanced image sequence samples to the image sequence sample set to obtain the enhanced image sequence sample set; Among them, the 3D symmetric mosaic method is: randomly select 4 image sequence samples from the image sequence sample set, and initialize the slice index s = 0; when 0 ≤ s ≤ 5, repeat steps a to c to generate the s-th slice of the enhanced image sequence sample: Step a, divide the blank image into 6 non-overlapping blank areas, and the 6 blank areas are distributed in two rows, with the first row defined as the third blank area, the first blank area, and the fourth blank area from left to right in sequence, and the second row defined as the fifth blank area, the second blank area, and the sixth blank area from left to right in sequence; Step b, on the s-th slice of the first image sequence sample extracted, use a first randomly sized box that passes through the central axis and is symmetric about the central axis for cropping, and place the cropped box image in the first blank area of the blank image; On the s-th slice of the second image sequence sample extracted, use a second randomly sized box that passes through the central axis and is symmetric about the central axis for cropping, and place the cropped box image in the second blank area of the blank image; On the s-th slice of the third image sequence sample extracted, use two third randomly sized boxes that are symmetric about the central axis for cropping, and place the two cropped box images in the third blank area and the fourth blank area of the blank image respectively; On the s-th slice of the fourth image sequence sample extracted, use two fourth randomly sized boxes that are symmetric about the central axis for cropping, and place the two cropped box images in the fifth blank area and the sixth blank area of the blank image respectively; Step c, use the blank image with box images placed in all 6 blank areas as the s-th slice of the enhanced image sequence sample, and let s = s + 1.

6. The real-time detection method for multiple abnormalities of cranial CT guided by prior imaging knowledge as described in claim 5, characterized in that, The obtaining of the image sequence sample set includes: Obtain the original cranial CT image sequence; Perform brain window processing and subdural window processing on the original cranial CT image sequence; Sequentially divide the original cranial CT image sequence after brain window processing and subdural window processing into multiple image sequence samples with the same length, and there is at least one overlapping slice between adjacent image sequence samples, and the multiple image sequence samples form the image sequence sample set.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for real-time detection of multiple anomalies in cranial CT guided by imaging prior knowledge according to one of claims 1-6.

8. An electronic device system, characterized in that, The electronic device system includes: At least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the method for real-time detection of multiple abnormalities in head CT guided by prior imaging knowledge as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Cow detection method based on ResMO-Dense-YOLO

    CN117912062A

  • Infrared weak and small target detection method based on improved YOLO v8

    CN118982658A

  • Target single edge detection method based on YOLO

    CN119090905A

  • Medical image abnormal signal intensity detection method

    CN115115576A

  • Classification method for intracranial great vessel blockage identification based on bilateral comparison difference information

    CN116403055A