Placenta implantation evaluation system based on multi-modal sign recognition and construction method and construction device thereof
Through a placental implantation evaluation system based on multimodal sign recognition, combined with deep neural network and biochemical detection data, automatic segmentation and evaluation of placental implantation symptoms are achieved, solving the problem of inaccurate evaluation in the prior art, and improving the accuracy and consistency of evaluation.
Patent Information
- Application Number
- CN202510607463.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing risk assessment methods for placenta implantation are inaccurate, relying on subjective segmentation and experience of ultrasound images, resulting in inconsistent evaluation results, especially in grassroots hospitals with insufficient accuracy due to individual differences.
A placental implantation evaluation system based on multimodal sign recognition is adopted. Through a deep neural network, a placental implantation biochemical detection objects and medical record data is combined with placental implantation biochemical detection objects and medical record data is used to automatically segment and locate the placental implantation signs to achieve accurate evaluation of placental implantation.
It improves the accuracy and consistency of placental implant risk assessment, reduces the experience dependence of doctors, provides more reliable placental implant severity assessment results, and helps medical staff develop reasonable surgical plans to avoid adverse outcomes.
Smart Images

Figure CN120147758A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of placenta accreta detection. Specifically, it relates to a placenta accreta evaluation system based on multi-modal sign recognition. Background Art
[0002] Placenta accreta refers to the pathological phenomenon that placental villi invade more than one-third of the myometrium, and even penetrate the myometrium or serosa. In severe cases, it may cause invasion of the rectum or bladder. The placenta accreta disease caused by it is an important cause of serious adverse outcomes such as hysterectomy, postpartum hemorrhage and death in women of childbearing age, and shows an increasing trend with the rising cesarean section rate in recent years. At present, the risk assessment of placenta accreta is the basis for the hierarchical management of placenta accreta disease and the determination of surgical decisions, and ultrasound examination is the most economical and commonly used preoperative assessment method for placenta accreta.
[0003] At present, retrospective studies have determined the evaluation criteria for placenta accreta for clinical ultrasound images. For example, the ultrasound scoring scale for placenta accreta types proposed in patent CN106308851B gives quantitative scoring criteria for multiple indicators in ultrasound images, and verifies that the use of this scoring criterion can effectively evaluate the type of placenta accreta, and accurately predict the risk of intraoperative bleeding and hysterectomy. However, the feasibility of this placenta accreta evaluation criterion is limited. For primary hospitals, the existing evaluation criteria are complex, and ultrasound examination more depends on the experience and skills of ultrasound instrument operators. Therefore, the scores of primary hospitals and general hospitals inevitably vary due to personnel subjectivity, resulting in inconsistent evaluation results, misdiagnosis in referrals, etc. In addition, in the existing ultrasound image evaluation system, the evaluation of some ultrasound signs still has the problem of strong subjectivity, and it is difficult to achieve stability and consistency in evaluation among different doctors, which makes the effect of placenta accreta risk assessment very dependent on the doctor's own ultrasound sign segmentation level.
[0004] The imaging quality of placenta ultrasound images is easily interfered by various factors, such as intestinal gas, body type, etc., and ultrasound images are prone to artifacts, resulting in relatively low resolution and signal-to-noise ratio. There are generally problems of insufficient accuracy caused by individual differences in the segmentation of placenta accreta signs using ultrasound images. Multi-modal data has complementarity and can better reflect individual differences from multiple angles. For example, biochemical marker detection data can show the depth of placental invasion into the myometrium, ultrasound blood flow imaging can evaluate the blood supply of the placenta, and CT images can also be used as a supplement to ultrasound images. Multi-modal data can help the neural network capture features more comprehensively and improve the positioning accuracy of the features of the placenta accreta area.
[0005] Therefore, there is a need for an evaluation system based on multi-modal placenta imaging for ultrasound sign recognition that can effectively and automatically segment placenta accreta signs to solve the problem of inaccurate existing placenta accreta risk assessment. Summary of the Invention
[0006] To solve the problems existing in the prior art, the purpose of this application is to provide a placenta accreta assessment system based on multi-modal sign recognition. By using a specific imaging method to obtain placenta ultrasound images, and utilizing a deep neural network combined with placenta accreta biochemical markers and medical record data to mine the pathological features in the placenta ultrasound images of the subjects, it realizes the precise automatic segmentation, localization and scoring of placenta accreta signs, so as to be suitable for the severity assessment of placenta accreta. Medical staff can quickly obtain effective assessment information such as the location, breadth, depth of placenta accreta and the invasion of surrounding organs by inputting the cross-sectional ultrasound images of the placenta of the subjects and biochemical test results into the system, which helps them quickly judge the situation of placenta accreta and estimate the intraoperative blood loss and the surgical plan for placenta accreta diseases, and avoid the occurrence of adverse maternal outcomes.
[0007] Specifically, this application relates to the following aspects:
[0008] According to one aspect of this application, there is provided a placenta accreta assessment system based on multi-modal sign recognition, including: a data collection module that collects abdominal ultrasound scan data of a subject to obtain a plurality of placenta ultrasound images; collects biochemical test and medical record data of the subject to obtain a plurality of test texts; a sign extraction module that inputs the plurality of placenta ultrasound images into a first feature extraction unit to obtain a first feature map, and inputs the plurality of test texts into a second feature extraction unit to obtain a second feature map; performs multiple upsamplings on the first feature map and the second feature map, and each time after upsampling, fuses the first feature map and the second feature map through a cross-attention unit, and obtains a third feature map after multiple upsamplings; uses a feature aggregation unit to skip connect and weighted-fuse the first feature map, the second feature map and the third feature map to obtain a fourth feature map; inputs the fourth feature map into a classification unit to obtain a plurality of segmentation masks of placenta accreta signs; an evaluation module that determines the severity of the placenta accreta signs corresponding to the plurality of segmentation masks, and obtains the placenta accreta assessment result of the subject based on the severity; wherein, collecting abdominal ultrasound scan data of the subject to obtain a plurality of placenta ultrasound images includes: taking the left and right sides and / or the upper and lower edges of the abdominal uterine contour of the subject in the supine position as boundaries, and scanning the sagittal placenta at intervals of a first distance and / or a second distance starting from one side and / or one edge until scanning to the other side and / or the other edge to obtain a plurality of placenta ultrasound images; the cross-attention unit includes a first cross-attention unit and a second cross-attention unit, which are used to calculate the similarity between the first feature map and the second feature map based on the first feature map and the second feature map respectively, so as to generate a first weight matrix and a second weight matrix respectively.
[0009] According to some embodiments of this application, the value range of the first distance is 1 cm to 10 cm; the value range of the second distance is 1 cm to 8 cm.
[0010] According to some embodiments of the present application, at least one of the first feature extraction unit and / or the second feature extraction unit includes a convolution layer with automatically adjustable shape; the convolution layer with automatically adjustable shape adjusts the shape of one or more of its convolution kernels according to the learned offsets of one or more of its convolution kernels at each sampling position.
[0011] According to some embodiments of the present application, the cross-attention unit is further configured to: extract placental implantation feature information of the first feature map and / or the second feature map based on the Top-k setting of the first weight matrix and the second weight matrix, so as to obtain a third feature map.
[0012] According to some embodiments of the present application, the first cross-attention unit and / or the second cross-attention unit includes a cross-attention decoder.
[0013] According to some embodiments of the present application, the feature aggregation unit establishes a skip connection among the output of the first feature extraction unit, the output of the second feature extraction unit, and the output of the cross-attention unit after the last upsampling among multiple upsamplings, and gates and weights the first feature map, the second feature map, and the third feature map to obtain a fourth feature map.
[0014] According to some embodiments of the present application, the feature aggregation unit includes a multi-scale feature aggregation block; the multi-scale feature aggregation block calculates the channel attention and spatial attention of the first feature map and the second feature map respectively before the skip connection between the first feature extraction unit and the second feature extraction unit, so as to optimize the channel weight distribution and spatial weight distribution of the first feature map and the second feature map.
[0015] According to some embodiments of the present application, determining the severity of placental implantation signs corresponding to multiple segmentation masks and obtaining the placental implantation evaluation result of the subject based on the severity includes: providing one of three classification labels to the placental implantation signs corresponding to each of the multiple segmentation masks through a classification unit to determine the severity of each placental implantation sign; and taking the sum of the severities of all placental implantation signs as the placental implantation evaluation result.
[0016] According to another aspect of the present application, there is provided a method for constructing a placenta accreta evaluation system based on multi-modal sign recognition, including: obtaining training data, annotating multiple placental ultrasound images to obtain a first training sample set, and annotating multiple detection texts to obtain a second training sample set; obtaining segmentation results, using a first feature extraction unit to process the first training sample set to obtain a first feature map, and using a second feature extraction unit to process the second training sample set to obtain a second feature map; performing multiple upsamplings on the first feature map and the second feature map, and fusing the first feature map and the second feature map through a cross-attention unit after each upsampling, to obtain a third feature map after multiple upsamplings; using a feature aggregation unit to perform skip connection and weighted fusion on the first feature map, the second feature map and the third feature map to obtain a fourth feature map; inputting the fourth feature map into a classification unit to obtain multiple segmentation masks of placenta accreta signs; iterating the segmentation model, and determining the loss values of the first feature extraction unit, the second feature extraction unit, the cross-attention unit and the feature aggregation unit according to the multiple segmentation masks, so as to iteratively update the weight parameters of the first feature extraction unit, the second feature extraction unit, the cross-attention unit and the feature aggregation unit by gradient backpropagation.
[0017] According to some embodiments of the present application, annotating multiple placental ultrasound images to obtain a first training sample set and annotating multiple detection texts to obtain a second training sample set includes: determining the labels of each of the multiple placental ultrasound images and the labels of each of the multiple detection texts based on evaluation parameters; wherein, the evaluation parameters include: placental position, placental thickness, posterior hypoechoic band of the placenta, bladder line, placental lacunae, basal blood flow of the placenta, cervical sinus, cervical morphology and / or history of cesarean section.
[0018] According to some embodiments of the present application, the method for constructing a placenta accreta evaluation system based on multi-modal sign recognition further includes: determining whether the iterated first feature extraction unit, second feature extraction unit, cross-attention unit and feature aggregation unit converge; in response to the convergence of the iterated first feature extraction unit, second feature extraction unit, cross-attention unit and feature aggregation unit, stopping the iteration of the first feature extraction unit, second feature extraction unit, cross-attention unit and feature aggregation unit.
[0019] According to some embodiments of the present application, the method for constructing a placenta accreta evaluation system based on multi-modal sign recognition further includes:
[0020] determining the maximum number of iterations of the first feature extraction unit, second feature extraction unit, cross-attention unit and feature aggregation unit; in response to the number of iterations of the first feature extraction unit, second feature extraction unit, cross-attention unit and feature aggregation unit reaching the maximum number of iterations, stopping the iteration of the first feature extraction unit, second feature extraction unit, cross-attention unit and feature aggregation unit.
[0021] According to another aspect of the present application, there is also provided a construction device for a placenta accreta evaluation system based on multi-modal sign recognition, including: a training data acquisition unit, which labels multiple placental ultrasound images to obtain a first training sample set and labels multiple detection texts to obtain a second training sample set; a segmentation result acquisition unit, which processes the first training sample set by using a first feature extraction unit to obtain a first feature map, and processes the second training sample set by using a second feature extraction unit to obtain a second feature map; performs multiple upsamplings on the first feature map and the second feature map, and fuses the first feature map and the second feature map through a cross-attention unit after each upsampling, and obtains a third feature map after multiple upsamplings; uses a feature aggregation unit to perform skip connection and weighted fusion on the first feature map, the second feature map and the third feature map to obtain a fourth feature map; inputs the fourth feature map into a classification unit to obtain multiple segmentation masks of placenta accreta signs; a segmentation model iteration unit, which determines the loss values of the first feature extraction unit, the second feature extraction unit, the cross-attention unit and the feature aggregation unit according to the multiple segmentation masks, so as to iteratively update the weight parameters of the first feature extraction unit, the second feature extraction unit, the cross-attention unit and the feature aggregation unit by using gradient backpropagation.
[0022] According to another aspect of the present application, there is also provided an electronic device, including: a processor; and a memory, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the foregoing construction method of the placenta accreta evaluation system based on multi-modal sign recognition.
[0023] According to another aspect of the present application, there is also provided a computer program product, including computer program instructions, and when the computer program instructions are run by the processor, the processor is caused to execute the foregoing construction method of the placenta accreta evaluation system based on multi-modal sign recognition.
[0024] According to another aspect of the present application, there is also provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the foregoing construction method of the placenta accreta evaluation system based on multi-modal sign recognition.
[0025] In this way, the placenta accreta assessment system based on multi-modal sign recognition provided by this application, as an intelligent assessment tool, can conveniently predict the risk of placenta accreta for subjects, especially those in areas with scarce medical resources, so as to achieve the following goals: 1. Based on the assessment results, guide the primary hospitals where the subjects are located to determine the dangerous level of placenta accreta and transfer them in time; 2. The superior hospitals formulate the preoperative preparation work required for the reasonable dangerous level according to the assessment results, including blood source preparation, intraoperative monitoring equipment and drugs, etc., to avoid insufficient material preparation or excessive waste; 3. According to the assessment results, formulate a detailed surgical plan before the operation, maximize the avoidance of hysterectomy for women of childbearing age, while reducing the bleeding volume and reducing the infusion of blood products, and improving the prognosis level of the subjects. Brief Description of the Drawings
[0026] Figure 1 The block diagram of the placenta accreta assessment system based on multi-modal sign recognition according to an embodiment of this application is illustrated.
[0027] Figure 2 The structural schematic diagram of the first feature extraction unit and the second feature extraction unit according to an embodiment of this application is illustrated.
[0028] Figure 3 The structural schematic diagram of the cross-attention unit according to an embodiment of this application is illustrated
[0029] Figure 4 The flowchart of the construction method of the placenta accreta assessment system based on multi-modal sign recognition according to an embodiment of this application is illustrated.
[0030] Figure 5 The block diagram of the construction device of the placenta accreta assessment system based on multi-modal sign recognition according to an embodiment of this application is illustrated.
[0031] Figure 6 The block diagram of the electronic device according to an embodiment of this application is illustrated.
[0032] Figure 7 The schematic diagram of the accuracy curve of the sign extraction module according to an embodiment of this application is illustrated. Detailed Description of the Embodiments
[0033] The following further illustrates this application in combination with embodiments. It should be understood that the embodiments are only used to further illustrate and explain this application, and are not used to limit this application.
[0034] Unless otherwise defined, the technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art. Although methods and materials similar or equivalent to those described herein can be used in experimental or practical applications, the materials and methods are described below. In case of conflict, the present specification, including its definitions, shall prevail. In addition, the materials, methods, and examples are for illustrative purposes only and not limiting. The present application will be further described below with reference to specific embodiments, but the scope of the present application is not limited thereby.
[0035] Overview of the Application
[0036] As described above, the existing ultrasound evaluation system usually relies on manual scoring of several items, during which the evaluation information is prone to distortion and loss. The ultrasound imaging evaluation of placenta accreta should be more based on the full and unified interpretation of the images collected from the subjects, and artificial intelligence has advantages in this regard. With the development of deep neural networks, computer vision technology has advanced by leaps and bounds, and the fields of application of image semantic segmentation and object detection are also becoming more and more extensive, providing technical support for computer recognition of medical image information. Specifically, Sun Yat-sen University in Guangdong has developed a device that uses machine learning algorithms to automatically extract the fetal nuchal translucency; at the present stage, through MRI / 3D imaging technology, it is also possible to image the placenta previa of patients and determine the selection of the surgical incision for placenta accreta through network analysis of MRI images. In addition, multiple prenatal features such as medical history and color Doppler ultrasound data can also be associated as observed state sequences and used to construct models such as hidden Markov algorithms to diagnose placenta accreta.
[0037] Although certain progress has been made in the prior art, most of them still rely on manual acquisition of placenta data to improve the sensitivity of the model to the sign textures in ultrasound images. These methods have not fully considered the global context features of the lesion areas in the images and have limitations in the accurate positioning of the lesion locations. That is, although a single-modal artificial intelligence model can effectively process the ultrasound image data of the placenta, due to the uneven image quality and the limitations of neural network sign recognition, there is still a lack of effective intelligent evaluation tools in the field of placenta accreta risk assessment.
[0038] The placenta accreta assessment system based on multi-modal sign recognition provided by this application can first perform ultrasonic scanning on the abdomen of the subject in a specific manner to obtain high-quality placental ultrasonic images. The placental ultrasonic images are input into the first feature extraction unit to obtain the first feature map, and the detection text is input into the second feature extraction unit to obtain the second feature map, completing the preliminary extraction of the information indicating the signs of placenta accreta in the images and text. The first feature map and the second feature map are upsampled multiple times to restore the resolution of the feature map. After each upsampling, the ultrasonic modality information of the first feature map and the text modality information of the second feature map are fused through a cross-attention unit to extract more effective feature information with high correlation from them, obtaining the third feature map. The feature aggregation unit is used to perform skip connection and weighted fusion on the first feature map, the second feature map, and the third feature map to further enhance the attention degree to the key sign regions in the feature information, obtaining the fourth feature map. The fourth feature map is input into the classification unit to obtain multiple segmentation masks of the signs of placenta accreta, so as to be used to evaluate the specific condition and severity of placenta accreta.
[0039] That is, this application introduces an attention mechanism into the convolutional layer of the neural network to capture more global context features of the placenta accreta lesion area in the feature map, reallocate the weights between the feature maps of the placental ultrasonic images extracted by the convolutional layer, specifically focus on the key features, and reduce the representation of non-important features, so as to more accurately identify the lesion area of placenta accreta and provide a solution for the segmentation of the signs of placenta accreta. In addition, a feature aggregation unit is introduced at the end of the decoder of the neural network for multi-scale feature fusion, fusing simple single-modal feature maps and complex multi-modal feature maps and optimizing the weight allocation of the feature maps to accurately locate the key sign regions and avoid noise interference or overfitting caused by simply splicing multi-modal features. In this way, the system obtains multiple ultrasonic signs, that is, the semantic segmentation results of multiple lesion areas, automatically calibrates the lesion positions instead of manual operation, and provides multiple key information for medical staff to perform the grading assessment of the severity of placenta accreta and the prediction of surgical risks.
[0040] Due to functions such as high-quality acquisition of ultrasonic image data, effective combination of the neural network and the attention mechanism, and optimization of the segmentation results using multi-modal placenta data, the system described in this application can more effectively capture the complex texture information in the image data under limited data, realize the in-depth extraction of the signs of placental lesions, and can provide comprehensive and accurate automatic support for corresponding diagnoses, medical decisions, etc., avoiding errors caused by subjective misjudgments of medical staff in the diagnosis situation, especially in primary medical institutions.
[0041] After introducing the basic principle of this application, various non-limiting embodiments of this application will be specifically introduced with reference to the accompanying drawings.
[0042] Exemplary System
[0043] Figure 1 The figure illustrates a placenta accreta assessment system based on multi-modal sign recognition according to an embodiment of the present application.
[0044] As Figure 1 shown, the placenta accreta assessment system based on multi-modal sign recognition according to an embodiment of the present application includes the following modules.
[0045] A data collection module, configured to collect abdominal ultrasound scan data of a subject to obtain a plurality of placental ultrasound images; and collect biochemical test and medical record data of the subject to obtain a plurality of test texts. This module supports ultrasound devices or endoscopic devices that conform to the existing DICOM standard. For example, existing ultrasound devices with video acquisition functions can cooperate with medical staff to collect static and / or dynamic ultrasound images of the subject's abdomen in real time. Data collection supports methods such as collection card, network port, or USB interface collection. This module can also save the data, for example, store it in various video data formats such as AVI and MP4 data, process the data through a collection card or any frame processing algorithm to obtain DICOM-format ultrasound images, and establish and store a set of placental ultrasound images of the subject.
[0046] In addition, the data collection module can be electrically connected to a computer storage device, which can be connected to a biochemical detection device and preprocess and store the detection results of the subject collected by the biochemical detection device. The data collection module described in the present application can obtain the detection results of the subject stored in the computer storage device. It can be understood that the preprocessing steps for the detection results of the subject can also be directly completed on the data collection module, such as denoising, enhancement, normalization, word segmentation, etc.
[0047] The system described in this application further includes a sign extraction module, which inputs the multiple placental ultrasound images into a first feature extraction unit to obtain a first feature map, and inputs the multiple detection texts into a second feature extraction unit to obtain a second feature map; performs multiple upsamplings on the first feature map and the second feature map, and fuses the first feature map and the second feature map through a cross-attention unit after each upsampling, and obtains a third feature map after the multiple upsamplings; uses a feature aggregation unit to perform skip connection and weighted fusion on the first feature map, the second feature map and the third feature map to obtain a fourth feature map; inputs the fourth feature map into a classification unit to obtain multiple segmentation masks of placental implantation signs. The sign extraction module can be exemplified as a device, an electronic device or a computer storage device with input / output interfaces, such as an MCU, a PLC, a computer, a server, etc. A first extraction unit for extracting image modality features, a second feature unit for extracting text modality features, a cross-attention unit, etc. can run on this module. The aforementioned multiple units are multi-parameter neural networks with complex structures, and the sign extraction module can run these units; the structures and functions of the aforementioned units will be described in detail below.
[0048] The system described in this application further includes an evaluation module, which is used to determine the severity of the placental implantation signs corresponding to the multiple segmentation masks, and obtain the placental implantation evaluation result of the subject based on the severity. In addition, the evaluation module described in this application can store the image segmentation masks output by the sign extraction module, and can be electrically connected to a display device, which is suitable for visualizing these image segmentation results on the display device, so as to facilitate medical staff to manually score the severity of the corresponding placental implantation lesion area according to the displayed results. The segmentation mask contains the semantic recognition results of the placental lesion signs of the subject, including multiple different mask values of the placental position, placental thickness, posterior hypoechoic band of the placenta, bladder line, placental lacuna, blood flow at the placental base, cervical blood sinus, etc.
[0049] In particular, abdominal ultrasound can clearly show the position, morphology of the placenta and its relationship with the uterine myometrium. Therefore, an optimal ultrasound scanning protocol is required to better capture characteristic signs of medical records such as rich blood flow behind the placenta, disappearance or abnormal dilation of the placental space, etc., and improve the quality of placental ultrasound images. Thus, according to an embodiment of the system described in the present application, the data collection module performing ultrasound scanning on the abdomen of the subject to obtain multiple placental ultrasound images can include the following two methods: The first scanning method is to use the ultrasound device or endoscope device of the data collection module to scan sagittal placental images at intervals of a first distance starting from one side, with the left and right sides of the uterine contour of the subject's supine abdomen as boundaries, until the other side of the abdominal uterine contour is scanned and the placental image disappears from the device's field of view, thus obtaining multiple placental ultrasound images; and the second scanning method in another embodiment is to use the ultrasound device or endoscope device of the data collection module to scan sagittal placental images at intervals of a second distance starting from one edge, with the upper and lower edges of the uterine contour of the subject's supine abdomen as boundaries, until the other side of the abdominal uterine contour is scanned and the placental image disappears from the device's field of view, so as to obtain multiple placental ultrasound images.
[0050] In this way, through the above implementation, the placental ultrasound images scanned and stored by the data collection module can maximally contain the placental integrity information of the subject, especially the information of the region where the signs are located. In addition, the data collection module can preset the overlapping regions scanned by the two scanning methods and remove duplicates from the images scanned in the overlapping regions to provide high-quality images to the sign extraction module.
[0051] Specifically, in the above scanning methods, the value range of the first distance can be selected from 1 to 10 cm. The smaller the first distance, the larger the data dimension of the set of placental ultrasound images obtained by this method. However, since the image size is overly subdivided, some texture features of placental implantation may not be prominent enough; the larger the first distance, the image may lose valid information. Therefore, preferably, the value range of the first distance is 4 cm to 5 cm, so that the information provided by the image is neither redundant nor missing, and the data volume is appropriate, which can improve the training efficiency of each unit in the sign extraction module. Similarly, the value range of the second distance can also be set to 1 cm to 8 cm; preferably, the value range of the second distance is 4 cm to 5 cm. Since there is usually no significant difference in the extraction of placental position features between horizontal or vertical scanning, therefore, more preferably, the first distance and the second distance are kept the same, avoiding asymmetry of placental implantation feature information in the two directions and avoiding affecting the generalization ability of the sign extraction module trained using placental ultrasound images.
[0052] Each of the multiple detection texts can be a concatenated text after preprocessing the detection data of multiple biochemical indicators and medical history query data of the subject, that is, a structured text; or it can be the original text of the detection data of multiple biochemical indicators and medical history query data of the subject respectively, that is, an unstructured text. Those skilled in the art can understand that the neural network can receive text data with a standard format after preprocessing (such as word embedding through Word2Vec or other structured processing), or directly process this data (such as selecting the second feature extraction unit as an existing natural language processing neural network). The detection text is used because it can be used as a supplement to the ultrasound image, reducing the problem of the variable lesion morphology caused by individual differences in the subject in the ultrasound image, enabling the entire sign extraction module to better adapt to the lesion morphology changes and suppressing the interference in the low signal-to-noise ratio area of the ultrasound image.
[0053] Preferably, the system of the present application selects the concentration of serum human chorionic gonadotropin (hCG), more preferably β-hCG, the concentration of placental growth factor (PIGF), the subject's cesarean section history, and the subject's uterine fibroids removal history as the detection text. It can be understood that after preprocessing, serializing, and vectorizing the above-mentioned structured / unstructured text, features can also be extracted using the convolutional layer. In the present application, the detection text mainly serves as a supplement to the ultrasound image, that is, it affects the feature selection of the feature map of the ultrasound image in the cross-attention unit, that is, the process of generating the third feature map.
[0054] Reference Figure 2 , the shape-adjustable convolutional layer can adjust the shape of one or more convolutional kernels according to the offset of one or more convolutional kernels at each sampling position in the convolutional layer, that is, adjust the size of the sampled original ultrasound image and / or feature map. Specifically, for a two-dimensional convolutional kernel of size K x K, an offset learning region with a size of 2 x K x K is set. The convolutional kernel learns the offsets in the x and y directions for each sampling point of the feature map through the offset learning region, so that each point on the feature map output by the two-dimensional convolutional kernel corresponds to all points within the K x K sampling range of the input feature map that have learned the offsets.
[0055] In this way, the convolution kernels of the first feature extraction unit and / or the second feature extraction unit for extracting sign features can adaptively adjust their sizes through offset learning in the previous stage, that is, adaptively adjust the receptive field during sampling, so as to automatically solve the problem of irregular infiltration / penetration boundaries between the placenta and the uterus in ultrasonic images. Compared with the standard convolutional layer, the convolutional layer of the first feature extraction unit or the second feature extraction unit is more sensitive to small lesions in ultrasonic images (such as the adhesion area of placental tissue), and can also better adapt to text instances with extreme aspect ratios (for example, non-standard format data existing in the input text data), thereby improving the quality of the first feature map and the second feature map.
[0056] That is, at least one of the first feature extraction unit and / or the second feature extraction unit included in the sign extraction module according to the embodiments of the present application includes a convolutional layer with automatic shape adjustment, and preferably both feature extraction units include convolutional layers with automatic shape adjustment.
[0057] Reference Figure 3 , the cross-attention unit is a bimodal cross-attention decoder, and its input is a sequence obtained by splicing the tokens obtained by mapping the first feature map or the second feature map to the input dimension respectively and the word tokens directly segmented by tools such as tokenizer, that is and ;
[0058] In this way, the cross-attention unit obtains the input features F U and F T , that is, the ultrasonic image features of the first feature map and the detection text features of the second feature map. Based on F U , calculate its cross-modal attention weights:
[0059] (Formula 1)
[0061] Where , , in this way, a feature weight matrix of the first feature map is established, and the detection text information can be fused into the ultrasonic sign information through the attention weights, and the contributions of different attention heads can be adjusted to select and retain high-contribution attention heads through the preset Top-k (that is, the top k features with the highest attention weights), so as to further extract the most effective ultrasonic sign features, thereby suppressing the interference of the mixed noise modality in the first feature map and improving the sign segmentation performance of the sign extraction module. Similarly, when the quantity level of the detection text is large, it can also be based on F T , and utilize F UFurther extract important features from the detection text; and, the features extracted from the two parts can also be weighted to retain more feature information. For the system described in this application, since usually four types of detection texts are selected and their confidence levels in clinical placenta accreta evaluation are relatively high, it is preferred to fuse the detection text information into the ultrasound sign information to obtain a feature weight matrix of one of the first feature map and the second feature map corresponding to the ultrasound image, so as to obtain a third feature map after fusing multi-modal feature information based on this feature map.
[0062] Continue to refer to Figure 3 , the dual-modal cross-attention decoding part can also calculate two attention weights in sequence to output important features, including channel cross-attention (CCA) and spatial cross-attention (SCA). CCA is mainly used to calculate the global channel dependence between the features of the feature maps corresponding to different modal data, and the channel weights can be calculated:
[0063] (Formula 2)
[0065] where σ is the activation function, W 1 , W 2 are the dimension reduction / restoration dimension weight matrices of the fully connected layers of CCA, and GAP(F) is the global average pooling result of each channel of the tokens corresponding to the selected feature map. In this way, a generated channel weight matrix is generated as the feature weight matrix, which can suppress irrelevant features in some channels and further enhance the weight of important features. Then, the generated features are passed through SCA to calculate the spatial mask in the feature map:
[0066] (Formula 3)
[0068] where f 7×7 is a 7x7 convolution operation, F avg is the average pooling result of the tokens corresponding to the selected feature map, and F max is the maximum pooling result of the tokens corresponding to the selected feature map. The weights corresponding to the spatial positions of the elements of each feature map calculated can help the dual-modal cross-attention decoding part focus on the key regions in the ultrasound image, the target ultrasound signs or the edge positions of the target ultrasound signs, etc., to improve the effective sign feature information amount of the extracted third feature map.
[0069] In particular, the attention calculation parts of CCA and SCA can also be set with residual connections, so that the cross-attention unit can better fuse the shallow input and the deep semantic features obtained through the attention head, and align the cross-modal of the features from the ultrasound image and the language embeddings from the detection text, improving the data utilization efficiency of the cross-attention unit.
[0070] Then, according to the present application, the cross-attention unit can reconstruct the feature map by mapping the output features, such as the attention map, to the original dimension of the input feature map, that is, output the fourth feature map, which is a technical solution that those skilled in the art can implement. In this way, the fourth feature map containing the deep multi-modal placenta accreta signs features is obtained, which can be skip-connected to the aforementioned first feature map and / or second feature map in the subsequent process to obtain a feature representation of the placenta accreta lesion area that highlights important signs and covers multi-modal lesion information, and can be used for effective automatic segmentation of these lesion areas.
[0071] That is, the cross-attention unit according to the embodiment of the present application includes a first cross-attention unit and a second cross-attention unit, which are used to calculate the similarity between the first feature map and the second feature map based on the first feature map and the second feature map respectively, so as to generate a first weight matrix and a second weight matrix respectively; and select the information of the first feature map and / or the second feature map based on the Top-k setting of the first weight matrix and the second weight matrix to obtain the third feature map.
[0072] And, the first cross-attention unit and / or the second cross-attention unit includes a cross-attention decoder.
[0073] In order to better fuse the low-level features of each modality and the multi-modal high-level features, skip connections can be established between the fourth feature map, the first feature map, and the second feature map. In the skip connections, a multi-scale feature aggregation block can also be introduced as described above, which also includes a CCA calculation part and an SCA calculation part. The fused first feature map, second feature map, and / or fourth feature map can be adjusted by channel attention weights and spatial attention weights to gate the weights of the weighted average of each feature map (such as using a pooling layer) to adjust the final fused feature, and then through operations such as bilinear interpolation and 1×1 convolution to compress the feature scale that those skilled in the art can understand, the fused feature is input into a classifier or a generator (such as a fully connected layer). In this way, the sign features extracted by the entire sign extraction module mainly contain multi-modal key information and do not miss the effective information of each modality input, so it can provide accurate segmentation of the placenta accreta signs.
[0074] That is, the feature aggregation unit according to the embodiment of the present application establishes a skip connection between the output of the first feature extraction unit, the output of the second feature extraction unit, and the output of the cross-attention unit after the last upsampling in the multiple upsamplings, and gates and weights the first feature map, the second feature map, and the third feature map to obtain the fourth feature map.
[0075] Moreover, the feature aggregation unit according to the embodiment of the present application includes a multi-scale feature aggregation block; before the skip connection between the first feature extraction unit and the second feature extraction unit, the multi-scale feature aggregation block calculates the channel attention and spatial attention of the first feature map and the second feature map respectively to optimize the channel weight distribution and spatial weight distribution of the first feature map and the second feature map.
[0076] Finally, the fourth feature map containing the fused multi-modal features is processed and input into the classification unit to obtain the classification labels of multiple placenta accreta ultrasound signs learned by the sign extraction module with the assistance of text. For example, the fourth feature map is respectively adjusted in channel number and size through a convolutional layer and a pooling layer with 1x1 convolution, and then appropriately input into the classification unit to complete multi-label classification. The classification unit can be exemplified as a combination of a fully connected layer and a softmax layer or a sigmoid layer, etc.; in this way, a classification label can be obtained for each placenta accreta ultrasound sign. The evaluation module can obtain these classification labels and use the value of each classification label as the severity of the corresponding sign to reflect the pathological characteristics of placenta accreta represented by the sign; the multi-labels can be 1, 2, 3, etc., and the value represented by the label is directly used as the severity of the sign. The larger the value, the more severe the placenta accreta lesion represented by the sign.
[0077] Table 1 Placenta Accreta Ultrasound Sign Evaluation Scale
[0078]
[0079] Table 1 shows the criteria for determining the placenta accreta evaluation results of the corresponding placenta ultrasound images according to the severity of the placenta accreta ultrasound signs. In an exemplary solution, the sum of the label values of all signs is used as the score to determine the evaluation result:
[0080] 1. If the score > 0 and ≤ 5, the evaluation result is that the patient corresponding to the placenta ultrasound image and the detection text has mild placenta accreta, and the mild type can be high-risk non-implantation / adhesion type;
[0081] 2. If the score > 5 and < 10, it is evaluated as a patient with moderate placenta accreta, and the moderate type can be high-risk implantation type;
[0082] 3. If the score ≥ 10, it is evaluated as a patient with severe placenta accreta, and the severe type can be high-risk penetration type;
[0083] 4. If the score is 0, it is evaluated as no placenta accreta. In this way, the placenta accreta evaluation results of the subjects from whom the placenta ultrasound image set and the detection text set are derived are automatically determined successively through the data collection module, the sign extraction module, and the evaluation module, which can be used as an effective basis for clinically assisting doctors in diagnosis.
[0084] That is, to determine the severity of the placenta accreta signs corresponding to each of the multiple segmentation masks, and obtain the placenta accreta evaluation result of the subject based on the severity, including: providing one of the three classification labels to the placenta accreta signs corresponding to each of the multiple segmentation masks through the classification unit to determine the severity of each placenta accreta sign; and taking the sum of the severities of all placenta accreta signs as the placenta accreta evaluation result.
[0085] Exemplary method
[0086] Figure 4 Illustrated is a method for constructing a placenta accreta evaluation system based on multi-modal sign recognition according to an embodiment of the present application.
[0087] As Figure 4 shown, the method for constructing a placenta accreta evaluation system based on multi-modal sign recognition according to an embodiment of the present application includes the following steps.
[0088] Step S110, obtain training data, label multiple placental ultrasound images in the placental ultrasound image set to obtain a first training sample set, and label multiple detection texts in the detection text set to obtain a second training sample set;
[0089] Step S120, obtain segmentation results, process the first training sample set by using a first feature extraction unit to obtain a first feature map, and process the second training sample set by using a second feature extraction unit to obtain a second feature map; perform multiple upsamplings on the first feature map and the second feature map, and provide a cross-attention unit to fuse the first feature map and the second feature map after each upsampling, and obtain a third feature map after the multiple upsamplings; use a feature aggregation unit to perform skip connection and weighted fusion on the first feature map, the second feature map, and the third feature map to obtain a fourth feature map; input the fourth feature map into a classification unit to obtain multiple segmentation masks of placenta accreta signs;
[0090] Step S130, iterate the segmentation model, determine the loss values of the first feature extraction unit, the second feature extraction unit, the cross-attention unit, and the feature aggregation unit according to the multiple segmentation masks, and use gradient backpropagation to iterate the weight parameters of the first feature extraction unit, the second feature extraction unit, the cross-attention unit, and the feature aggregation unit.
[0091] Among them, annotating multiple placental ultrasound images in the placental ultrasound image set to obtain a first training sample set, and annotating multiple detection texts in the detection text set to obtain a second training sample set includes: determining the labels of each of the multiple placental ultrasound images and the labels of each of the multiple detection texts based on evaluation parameters; wherein, the evaluation parameters include: placental position, placental thickness, posterior hypoechoic band of the placenta, bladder line, placental lacuna, blood flow at the placental basal part, cervical sinus, cervical morphology, and / or cesarean section history.
[0092] The specific way of iteratively segmenting the model is: determining whether the iterated first feature extraction unit, second feature extraction unit, cross-attention unit, and feature aggregation unit converge; in response to the convergence of the iterated first feature extraction unit, second feature extraction unit, cross-attention unit, and feature aggregation unit, stopping the iteration of the first feature extraction unit, second feature extraction unit, cross-attention unit, and feature aggregation unit.
[0093] Alternatively, an alternative way of iteratively segmenting the model can also be: determining the maximum number of iterations of the first feature extraction unit, second feature extraction unit, cross-attention unit, and feature aggregation unit; in response to the number of iterations of the first feature extraction unit, second feature extraction unit, cross-attention unit, and feature aggregation unit reaching the maximum number of iterations, stopping the iteration of the first feature extraction unit, second feature extraction unit, cross-attention unit, and feature aggregation unit.
[0094] It can be understood that the segmentation model mentioned above is a complete module including the first feature extraction unit, second feature extraction unit, cross-attention unit, feature aggregation unit, and classification unit, which can receive input data and output segmentation results. When iteratively updating the weight parameters of the first feature extraction unit, second feature extraction unit, cross-attention unit, and feature aggregation unit using gradient backpropagation, the cross-entropy loss function can be selected to measure the loss value of each unit.
[0095] In particular, the first feature extraction unit, second feature extraction unit, cross-attention unit, and / or feature aggregation unit in the placental implantation evaluation system based on multi-modal sign recognition according to the embodiments of the present application can be the first feature extraction unit, second feature extraction unit, cross-attention unit, and / or feature aggregation unit obtained by training the aforementioned segmentation model through the construction method of the placental implantation evaluation system based on multi-modal sign recognition according to the embodiments of the present application.
[0096] Embodiment
[0097] This application provides a general and / or specific description of the materials and experimental methods used in the experiments. Unless otherwise specified, all reagents and instruments used are conventional products that can be obtained commercially.
[0098] Example 1: Collection of training samples
[0099] Medical record data collected from the Third Hospital of Peking University was used: female patients diagnosed with placenta accreta between January 2021 and July 2024, excluding those with concurrent preeclampsia, HELLP syndrome, acute fatty liver of pregnancy, as well as those with heart disease above NYHA class III, end-stage renal disease, and those who had received uterine artery embolization within 3 months. Records of a total of 334 patients with placenta accreta (including penetrative placenta accreta) and 189 patients without placenta accreta (excluding adherent placenta accreta) were obtained; all patients signed informed consent forms agreeing to include ultrasound images in subsequent data calculations and modeling. Ultrasound imaging was obtained at intervals of the first distance and the second distance based on the scanning method described in this application, resulting in a total of 3156 placental ultrasound images, with the number of ultrasound images per person ranging from 7 to 12; professional doctors from the Third Hospital of Peking University manually segmented and labeled the placental ultrasound images, including: the area where placental tissue invaded the myometrium, the area where the serosa bulged or ruptured, the abnormal blood flow pool in the internal cavity, the serpentine vascular cluster, the area of bladder wall infiltration, and the local bulging of the uterine contour, corresponding to signs such as placental position, thickness, depression, hypoechoic band behind the placenta, blood flow at the placental base, bladder line, and cervical blood sinus. The ultrasound images were used as the first training sample set for input into the first feature extraction unit.
[0100] The electronic medical records of the enrolled patients were collected and the following were extracted: serum human chorionic gonadotropin (β-hCG) concentration, placental growth factor (PIGF) concentration, history of cesarean section, and history of myomectomy. The above text data was preprocessed, including: denoising, standardization, and processing with the BERT-WWM tokenizer to obtain the detection text; the placental accreta disease status of the source patients was labeled for all detection texts.
[0101] Example 2: Training and validation of the segmentation model
[0102] Serialization operations were performed on the labeled detection texts obtained in Example 1, including text truncation, word encoding, and position encoding, to align the sequences corresponding to each detection text; word vectors were constructed using Word2Vec to obtain the second training sample set for input into the second feature extraction unit.
[0103] The segmentation model includes an encoder part and a decoder part. The encoder part includes feature extraction units (i.e., the first feature extraction unit and the second feature extraction unit, the same below) and multiple downsamplings in cooperation. After each downsampling, a feature extraction unit is introduced; the decoder part includes cross-attention units and multiple upsamplings in cooperation. After each upsampling, a cross-attention unit is introduced. Skip connections are respectively made between each layer of the encoder part and the decoder part, and a feature aggregation unit is introduced on each skip connection. The obtained fourth feature map is input into a spatial pyramid pooling layer (with a 1x1 pooling kernel), and after compressing the number of channels using a 1x1 convolutional layer, it is input into a fully connected layer to output a segmentation mask.
[0104] The first training sample set and the second training sample set are respectively divided into a training set and a test set according to a ratio of 4:1. The loss function of the segmentation model is set as the cross-entropy loss function; the initial learning rate is set as 0.01, the Batch Size is set as 1, the optimizer is selected as Adam, and the number of training rounds is set as 150 epochs; the segmentation model automatically stops iterating after reaching the maximum number of iterations. The accuracy curve of the model is as Figure 7 shown. It can be seen that its accuracy has reached 0.95 after the number of training rounds reaches 10; in addition, the AUC of the ROC curve of the segmentation model is relatively high. This indicates that the segmentation model can effectively distinguish samples with different labeled regions in the test set through semantic segmentation; finally, the recall rate of the segmentation model is relatively high, enabling it to more effectively detect placenta accreta lesions and facilitating the screening of positive samples.
[0105] Exemplary device
[0106] Figure 5 The block diagram of the construction device of the placenta accreta evaluation system based on multi-modal sign recognition according to an embodiment of the present application is illustrated.
[0107] As Figure 5 shown, the construction device 200 of the placenta accreta evaluation system based on multi-modal sign recognition according to an embodiment of the present application includes:
[0108] A training data acquisition unit 210, which labels multiple placental ultrasound images to obtain a first training sample set and labels multiple detection texts to obtain a second training sample set;
[0109] The segmentation result acquisition unit 220 processes the first training sample set using the first feature extraction unit to obtain a first feature map, and processes the second training sample set using the second feature extraction unit to obtain a second feature map; performs multiple upsampling on the first feature map and the second feature map, and fuses the first feature map and the second feature map through a cross attention unit after each upsampling, and obtains a third feature map after the multiple upsampling; uses a feature aggregation unit to skip-connect and weightedly fuse the first feature map, the second feature map and the third feature map to obtain a fourth feature map; and inputs the fourth feature map into a classification unit to obtain multiple segmentation masks of placenta accreta signs;
[0110] The segmentation model iteration unit 230 determines the loss values of the first feature extraction unit, the second feature extraction unit, the cross attention unit and the feature aggregation unit according to the multiple segmentation masks, so as to iterate the weight parameters of the first feature extraction unit, the second feature extraction unit, the cross attention unit and the feature aggregation unit by using gradient back propagation.
[0111] Here, those skilled in the art will appreciate that the specific functions and operations of the various units and modules in the construction device 200 of the placenta accreta assessment system based on multimodal sign recognition have been described in the above reference. Figure 4 The present invention has been described in detail in the description of the construction method of the placenta accreta assessment system based on multimodal sign recognition, and therefore, its repeated description will be omitted.
[0112] As described above, the construction device 200 of the placenta accreta assessment system based on multimodal sign recognition according to the embodiment of the present application can be implemented in various terminal devices, such as a server for storing a training sample set, a plurality of segmentation masks, etc. In some examples, the construction device 200 of the placenta accreta assessment system based on multimodal sign recognition according to the embodiment of the present application can be integrated into the terminal device as a software module and / or a hardware module. For example, the construction device 200 of the placenta accreta assessment system based on multimodal sign recognition can be a software module in the operating system of the terminal device, or can be an application developed for the terminal device; of course, the construction device 200 of the placenta accreta assessment system based on multimodal sign recognition can also be one of the many hardware modules of the terminal device.
[0113] Alternatively, in other examples, the construction device 200 of the placenta accreta assessment system based on multimodal sign recognition and the terminal device may also be separate devices, and the construction device 200 of the placenta accreta assessment system based on multimodal sign recognition may be connected to the terminal device via a wired and / or wireless network, and transmit interactive information in accordance with an agreed data format.
[0114] Exemplary Electronic Device
[0115] Next, with reference to Figure 6 the electronic device according to an embodiment of the present application will be described.
[0116] Figure 6 The block diagram of the electronic device according to an embodiment of the present application is illustrated.
[0117] As Figure 6 shown, the electronic device 10 includes one or more processors 11 and a memory 12.
[0118] The processor 13 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 10 to perform desired functions.
[0119] The memory 12 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage medium, and the processor 11 can run the program instructions to implement the construction method of the placenta implantation assessment system based on multi-modal signs recognition in various embodiments of the present application as described above and / or other desired functions. Various contents such as placenta ultrasound images, detection texts, cross-attention units, etc. can also be stored in the computer-readable storage medium.
[0120] In one example, the electronic device 10 can further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0121] The input device 13 can include, for example, a keyboard, a mouse, etc.
[0122] The output device 14 can output various information to the outside, including calibration coefficients obtained through linear regression, etc. The output device 14 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0123] Of course, for simplicity, Figure 6 only some of the components related to the present application in the electronic device 10 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 10 can further include any other appropriate components.
[0124] Exemplary Computer Program Product and Computer Readable Storage Medium
[0125] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions that, when run on a processor, cause the processor to execute the steps in the method for constructing a placenta accreta assessment system based on multi-modal sign recognition according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.
[0126] The computer program product can be written in any combination of one or more programming languages for the program code to execute the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language, Python, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0127] Furthermore, an embodiment of the present application may also be a computer readable storage medium, on which computer program instructions are stored that, when run on a processor, cause the processor to execute the steps in the method for constructing a placenta accreta assessment system based on multi-modal sign recognition according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.
[0128] The computer readable storage medium may adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0129] The basic principles of the present application have been described in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. Additionally, the specific details disclosed above are only for illustrative and easy-to-understand purposes and not limitations. These details do not limit the present application to necessarily implement using the above specific details.
[0130] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0131] It should also be noted that in the devices, equipment, and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present application.
[0132] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0133] The above description has been given for purposes of illustration and description. In addition, this description does not intend to limit the embodiments of the present application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A placenta accreta assessment system based on multimodal sign recognition, characterized in that: include: A data collection module, collecting abdominal ultrasound scan data of the subject to obtain multiple placental ultrasound images; Collecting biochemical test and medical record data of subjects to obtain multiple test texts; A sign extraction module, inputting the multiple placental ultrasound images into a first feature extraction unit to obtain a first feature map, and inputting the multiple detection texts into a second feature extraction unit to obtain a second feature map; performing multiple upsampling on the first feature map and the second feature map, fusing the first feature map and the second feature map through a cross attention unit after each upsampling, and obtaining a third feature map after the multiple upsampling; Using a feature aggregation unit to skip-connect and weightedly fuse the first feature map, the second feature map, and the third feature map to obtain a fourth feature map; inputting the fourth feature map into a classification unit to obtain multiple segmentation masks of placenta accreta signs; an evaluation module, determining the severity of the placenta accreta signs corresponding to the multiple segmentation masks, and obtaining a placenta accreta evaluation result of the subject based on the severity; Wherein, collecting the abdominal ultrasound scan data of the subject to obtain a plurality of placental ultrasound images includes: Taking the left and right sides and / or the upper and lower edges of the uterine contour of the subject's abdomen in a supine position as boundaries, starting from one side and / or one edge, scanning the placenta in the sagittal position once every first distance and / or second distance until the other side and / or the other edge is scanned to obtain the multiple placental ultrasound images; The cross-attention unit includes a first cross-attention unit and a second cross-attention unit, which are used to calculate the similarity of the first feature map and the second feature map based on the first feature map and the second feature map, respectively, to generate a first weight matrix and a second weight matrix, respectively.
2. The placenta accreta assessment system based on multimodal sign recognition according to claim 1, characterized in that: The value range of the first distance is 1 cm to 10 cm; The second distance has a value range of 1 cm to 8 cm.
3. The placenta accreta assessment system based on multimodal sign recognition according to claim 1, characterized in that: At least one of the first feature extraction unit and / or the second feature extraction unit comprises a convolutional layer with automatic shape adjustment; The shape-automatically adjusted convolution layer adjusts the shapes of one or more convolution kernels according to the learned offset of the one or more convolution kernels at each sampling position.
4. The placenta accreta assessment system based on multimodal sign recognition according to claim 1, characterized in that: The cross attention unit is also used to extract placenta implantation feature information of the first feature map and / or the second feature map based on the Top-k setting of the first weight matrix and the second weight matrix to obtain a third feature map.
5. The placenta accreta assessment system based on multimodal sign recognition according to claim 1, characterized in that: The first cross-attention unit and / or the second cross-attention unit includes a cross-attention decoder.
6. The placenta accreta assessment system based on multimodal sign recognition according to claim 1, characterized in that: The feature aggregation unit establishes a jump connection between the output of the first feature extraction unit, the output of the second feature extraction unit, and the output of the cross-attention unit after the last upsampling in the multiple upsamplings, and obtains the fourth feature map by gated weighted fusion of the first feature map, the second feature map, and the third feature map.
7. The placenta accreta assessment system based on multimodal sign recognition according to claim 6, characterized in that: The feature aggregation unit includes a multi-scale feature aggregation block; The multi-scale feature aggregation block calculates the channel attention and spatial attention of the first feature map and the second feature map respectively before the jump connection between the first feature extraction unit and the second feature extraction unit to optimize the channel weight allocation and spatial weight allocation of the first feature map and the second feature map.
8. The placenta accreta assessment system based on multimodal sign recognition according to claim 1, characterized in that: Determining the severity of the placenta accreta signs corresponding to the multiple segmentation masks, and obtaining a placenta accreta assessment result of the subject based on the severity includes: providing, by the classification unit, one of three classification labels to the placenta accreta signs corresponding to each of the plurality of segmentation masks, so as to determine the severity of each placenta accreta sign; The sum of the severity of all signs of placenta accreta was used as the placenta accreta assessment outcome.
9. A method for constructing a placenta accreta assessment system based on multimodal sign recognition, characterized in that: include: Acquire training data, annotate multiple placental ultrasound images to obtain a first training sample set, and annotate multiple detection texts to obtain a second training sample set; Obtain a segmentation result, use a first feature extraction unit to process a first training sample set to obtain a first feature map, and use a second feature extraction unit to process a second training sample set to obtain a second feature map; perform multiple upsampling on the first feature map and the second feature map, and fuse the first feature map and the second feature map through a cross attention unit after each upsampling, and obtain a third feature map after the multiple upsampling; Using a feature aggregation unit to skip-connect and weightedly fuse the first feature map, the second feature map, and the third feature map to obtain a fourth feature map; inputting the fourth feature map into a classification unit to obtain multiple segmentation masks of placenta accreta signs; An iterative segmentation model is provided, wherein the loss values of the first feature extraction unit, the second feature extraction unit, the cross attention unit, and the feature aggregation unit are determined according to the multiple segmentation masks, so as to iterate the weight parameters of the first feature extraction unit, the second feature extraction unit, the cross attention unit, and the feature aggregation unit by using gradient back propagation.
10. The method for constructing a placenta accreta assessment system based on multimodal sign recognition according to claim 9, characterized in that: Labeling a plurality of placental ultrasound images to obtain a first training sample set, and labeling a plurality of detection texts to obtain a second training sample set includes: Determining a label of each of the plurality of placental ultrasound images and a label of each of the plurality of detection texts based on the evaluation parameters; The evaluation parameters include: placental position, placental thickness, retroplacental hypoechoic zone, bladder line, placental fossa, placental base blood flow, cervical sinusoids, cervical morphology and / or cesarean section history.
11. The method for constructing a placenta accreta assessment system based on multimodal sign recognition according to claim 9, characterized in that: Also includes: Determine whether the first feature extraction unit, the second feature extraction unit, the cross attention unit, and the feature aggregation unit converge after iteration; In response to the first feature extraction unit, the second feature extraction unit, the cross attention unit and the feature aggregation unit converging after iteration, the iteration of the first feature extraction unit, the second feature extraction unit, the cross attention unit and the feature aggregation unit is stopped.
12. The method for constructing a placenta accreta assessment system based on multimodal sign recognition according to claim 9, characterized in that: Also includes: Determining a maximum number of iterations of the first feature extraction unit, the second feature extraction unit, the cross attention unit, and the feature aggregation unit; In response to the number of iterations of the first feature extraction unit, the second feature extraction unit, the cross attention unit and the feature aggregation unit reaching the maximum number of iterations, the iterations of the first feature extraction unit, the second feature extraction unit, the cross attention unit and the feature aggregation unit are stopped.
13. A device for constructing a placenta accreta assessment system based on multimodal sign recognition, characterized in that: include: A training data acquisition unit, annotating a plurality of placental ultrasound images to obtain a first training sample set, and annotating a plurality of detection texts to obtain a second training sample set; a segmentation result acquisition unit, which processes the first training sample set using the first feature extraction unit to obtain a first feature map, and processes the second training sample set using the second feature extraction unit to obtain a second feature map; performs multiple upsampling on the first feature map and the second feature map, and fuses the first feature map and the second feature map through a cross attention unit after each upsampling, and obtains a third feature map after the multiple upsampling; Using a feature aggregation unit to skip-connect and weightedly fuse the first feature map, the second feature map, and the third feature map to obtain a fourth feature map; inputting the fourth feature map into a classification unit to obtain multiple segmentation masks of placenta accreta signs; A segmentation model iteration unit determines the loss values of the first feature extraction unit, the second feature extraction unit, the cross attention unit, and the feature aggregation unit according to the multiple segmentation masks, so as to iterate the weight parameters of the first feature extraction unit, the second feature extraction unit, the cross attention unit, and the feature aggregation unit by using gradient back propagation.
14. An electronic device, characterized in that: include: processor; as well as A memory, in which computer program instructions are stored, and when the computer program instructions are executed by the processor, the processor executes the method for constructing a placenta accreta assessment system based on multimodal sign recognition according to any one of claims 9 to 12.
15. A computer program product, characterized in that The method comprises computer program instructions, which, when executed by a processor, enable the processor to execute the method for constructing a placenta accreta assessment system based on multimodal sign recognition according to any one of claims 9 to 12.
16. A computer-readable storage medium, characterized in that: Computer program instructions are stored thereon, and when the computer program instructions are executed by a processor, the processor executes the method for constructing a placenta accreta assessment system based on multimodal sign recognition according to any one of claims 9 to 12.
Citation Information
Patent Citations
A method and apparatus for processing B-ultrasound images
CN106308851B
Processing method of B-scan ultrasonic image and device thereof
CN106308851A
Placenta implantation MRI sign detection and classification method and device based on deep neural network
CN116363081A
Lung CT image classification system based on domain knowledge and parallel separable convolution Swin Transform
CN117058448A
Mask type pre-training method and system for multi-modal medical data
CN117540339A
Cited By
Clinical test data query method and system
CN121092611A
Medical image processing method and system based on placenta implantation segmentation
CN121709164A