Fetal cleft lip and palate antenatal image auxiliary examination method, system and program product
Through the method of combining image recognition model and convolutional neural network with visual large language model, the fetal ultrasound images are automatically labeled and analyzed, which solves the problem of relying on doctors' experience and information in traditional methods, and achieves efficient and accurate diagnosis of fetal cleft lip and palate imaging.
Patent Information
- Application Number
- CN202510289118.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional fetal facial ultrasound analysis relies on doctor experience and lacks interpretability. The existing technology is difficult to provide efficient and accurate cleft lip and palate imaging diagnosis. Social media and search engines lack professionalism and targeting, resulting in insufficient information acquisition.
Pre-trained image recognition model is used to annotate and classify ultrasound image data, combine with convolutional neural network to extract global features, and compare the visual large language model with the standard image set of cleft lip and palate diagnostic knowledge to generate interpretable inference results.
It improves the accuracy and efficiency of early screening of fetal cleft lip and palate, provides scientific and reliable decision-making support, and achieves efficient and accurate image analysis.
Smart Images

Figure CN120356249A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of prenatal imaging diagnosis, and particularly to a prenatal imaging assisted examination method, system and program product for fetal cleft lip and palate. Background Art
[0002] Traditional ultrasonic analysis of the fetal face mainly relies on doctors' experience judgment and manual feature extraction. Although the application of deep learning technology in the field of medical image analysis is increasing day by day, in clinical practice, small models still face significant interpretability challenges, which limit their reliability and popularity. In addition, the popularization of knowledge about cleft lip and palate mostly relies on social media, search engine Q&A systems and knowledge graphs. Although these platforms can provide certain information support, they often lack a professional medical perspective and pertinence, and cannot provide personalized consultation and answers according to the specific situation of patients, resulting in insufficient comprehensiveness and effectiveness of information acquisition. Summary of the Invention
[0003] Embodiments of the present invention provide a prenatal imaging assisted examination method, system and program product for fetal cleft lip and palate to provide efficient, accurate and medically interpretable cleft lip and palate imaging analysis results.
[0004] To achieve the above object, on the one hand, a prenatal imaging assisted examination method for fetal cleft lip and palate is provided, and the method includes:
[0005] S1, obtaining prenatal facial ultrasonic image data of a fetus;
[0006] S2, annotating and classifying the ultrasonic image data through a pre-trained image recognition model to obtain an annotation result, including:
[0007] framing a section feature region in the ultrasonic image data through the image recognition model, performing type annotation on the section feature region according to a predetermined determination rule, and calculating the confidence of the section feature region, so as to obtain an annotation result;
[0008] performing global feature extraction on the ultrasonic image data through a pre-trained convolutional neural network model to generate a structured cleft lip and palate imaging report;
[0009] S3, inputting the ultrasonic image data, the annotation result and the cleft lip and palate imaging report into a pre-trained vision-language model for analysis to obtain an interpretable inference result, including:
[0010] S31. Compare and judge the ultrasonic image data with a predetermined standard image set of cleft lip and palate diagnosis knowledge through the visual large language model to obtain the section type and preliminary diagnosis result of the ultrasonic image data; wherein, the standard image set of cleft lip and palate diagnosis knowledge includes: a mid-sagittal section without abnormal features of cleft lip and palate, a mid-sagittal section with abnormal features of cleft lip and palate, a posterior nasal triangle section without abnormal features of cleft lip and palate, and a posterior nasal triangle section with abnormal features of cleft lip and palate; the section types include: mid-sagittal section and posterior nasal triangle section;
[0011] S32. Integrate the preliminary diagnosis result, the annotation result and the cleft lip and palate image report through the visual large language model to obtain the interpretable reasoning result.
[0012] Preferably, in the prenatal imaging auxiliary examination method for fetal cleft lip and palate, in the type annotation of the section feature area according to a predetermined determination rule in step S2, the determination rule includes:
[0013] Annotate the type of the section feature area with discontinuous palatal line in the mid-sagittal section as the abnormal type;
[0014] Annotate the type of the section feature area with a broken or discontinuous bottom edge in the posterior nasal triangle area in the posterior nasal triangle section as the abnormal type.
[0015] Preferably, in the prenatal imaging auxiliary examination method for fetal cleft lip and palate, in the global feature extraction of the ultrasonic image data through a pre-trained convolutional neural network model in step S2, the global features include: morphological features of the ultrasonic image, distribution of tissue structures, and predetermined landmark features of cleft lip and palate.
[0016] Preferably, in the prenatal imaging auxiliary examination method for fetal cleft lip and palate, the loss function of the image recognition model is:
[0017] Loss = λ class ·Loss class + λ bbox ·Loss bbox + λ conf ·Loss conf
[0018] Loss represents the loss function, Loss class represents the classification loss, λ class represents the weight parameter of the classification loss, Loss bbox represents the regression loss, λ bbox represents the weight parameter of the regression loss, Loss conf represents the confidence loss, λ confA weight parameter representing the confidence loss.
[0019] Preferably, in the prenatal imaging-assisted examination method for fetal cleft lip and palate, in step S2, a pre-trained convolutional neural network model is used to extract global features from the ultrasonic image data, and the generated structured cleft lip and palate imaging report includes:
[0020] During the decoding process of the pre-trained convolutional neural network, a relational memory module and a memory-driven conditional layer normalization are combined to generate a structured cleft lip and palate imaging report; wherein, the relational memory module is used to integrate the memory information of historical ultrasonic images during the generation of the cleft lip and palate imaging report.
[0021] Preferably, in the prenatal imaging-assisted examination method for fetal cleft lip and palate, in step S3, the visual large language model is used to compare and judge the ultrasonic image data with a predetermined standard image set of cleft lip and palate diagnosis knowledge, and the obtained cross-sectional type and preliminary diagnosis result of the ultrasonic image data include:
[0022] The visual large language model is analyzed in combination with a predetermined knowledge embedding mechanism, wherein the knowledge embedding mechanism is: embedding the ultrasonic diagnosis criteria for cleft lip and palate and the standard image set of cleft lip and palate diagnosis knowledge into the inference process of the visual large language model in a structured manner.
[0023] Preferably, in the prenatal imaging-assisted examination method for fetal cleft lip and palate, after the prenatal imaging-assisted examination method for fetal cleft lip and palate, it further includes: generating an answer according to the obtained text data regarding cleft lip and palate problems through a pre-trained second large language model, wherein the training data of the second large language model is obtained through the following steps:
[0024] Generating question-and-answer pairs recursively according to a pre-constructed corpus through a random strategy based on the length of independent paragraphs to obtain a first question-and-answer data set;
[0025] Using a publicly available medical data set to organize and expand the first question-and-answer data set through a pre-trained first large language model to obtain a second question-and-answer data set;
[0026] Training and optimizing the second large language model through the second question-and-answer data set.
[0027] Preferably, in the prenatal imaging-assisted examination method for fetal cleft lip and palate, in the recursive generation of question-and-answer pairs according to a pre-constructed corpus through a random strategy based on the length of independent paragraphs, the random strategy includes:
[0028] Determining the generated question-and-answer rounds by means of weighted random selection, wherein the generation probability of the question-and-answer rounds is:
[0029]
[0030] P(k) represents the generation probability of the Q&A round, where k represents the Q&A round, and length(S i ) represents the length of the independent paragraph, and w1(length(S i ), w2(length(S i ), and w3(length(S i )) represent the weights set according to the length of the independent paragraph.
[0031] On the other hand, another embodiment of the present invention provides a prenatal imaging-assisted examination system for fetal cleft lip and palate, including: a memory and a processor, where the memory stores at least one program, and the at least one program is executed by the processor to implement the prenatal imaging-assisted examination method for fetal cleft lip and palate as described in any one of the above.
[0032] On another aspect, another embodiment of the present invention provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the prenatal imaging-assisted examination method for fetal cleft lip and palate as described in any one of the above.
[0033] The above technical solutions have the following technical effects:
[0034] In the embodiments of the present invention, by acquiring prenatal ultrasound image data of the fetus, using a pre-trained image recognition model to automatically label and classify the section feature regions in these images, and calculating the confidence of each region to obtain accurate labeling results; using a convolutional neural network model to extract the global features of the ultrasound image data to generate a structured cleft lip and palate image report; inputting the ultrasound image data, labeling results, and generated image report into a pre-trained vision-language model, and by comparing with a predetermined standard image set of cleft lip and palate diagnosis knowledge, determining the section type and preliminary diagnosis result of the ultrasound image data, and conducting in-depth analysis in combination with the preliminary diagnosis result, labeling result, and image report, finally obtaining an interpretable reasoning result, improving the accuracy and efficiency of early screening for fetal cleft lip and palate, and also providing an intuitive and easy-to-understand analysis basis, providing scientific and reliable decision-making support for medical professionals and patients, thereby helping to achieve earlier intervention and treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flowchart of a prenatal imaging-assisted examination method for fetal cleft lip and palate according to an embodiment of the present invention;
[0036] Figure 2 is a schematic flowchart of a prenatal imaging-assisted examination method for fetal cleft lip and palate according to another embodiment of the present invention;
[0037] Figure 3 This is the framework diagram of the system and mini-program implemented by using the prenatal imaging assisted examination method for fetal cleft lip and palate in another embodiment of the present invention;
[0038] Figure 4 This is the schematic flowchart of the intelligent consultation and multi-dimensional popular science module in the system implemented by using the prenatal imaging assisted examination method for fetal cleft lip and palate in another embodiment of the present invention;
[0039] Figure 5 This is the visual page diagram of the mini-program implemented by using the prenatal imaging assisted examination method for fetal cleft lip and palate in another embodiment of the present invention;
[0040] Figure 6 This is the structural schematic diagram of the prenatal imaging assisted examination system for fetal cleft lip and palate in another embodiment of the present invention. Detailed implementation manners
[0041] To further illustrate each embodiment, the present invention provides accompanying drawings. These accompanying drawings are part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be used to explain the operating principle of the embodiments in conjunction with the relevant descriptions in the specification. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention. The components in the figures are not drawn to scale, and similar component symbols are usually used to represent similar components.
[0042] Now, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners.
[0043] Embodiment 1:
[0044] To provide efficient, accurate and medically interpretable cleft lip and palate imaging analysis results, an embodiment of the present invention provides a prenatal imaging assisted examination method for fetal cleft lip and palate. Figure 1 This is the flowchart of the prenatal imaging assisted examination method for fetal cleft lip and palate in an embodiment of the present invention. As Figure 1 shown, the method includes: S1, obtaining prenatal facial ultrasound image data of the fetus;
[0045] S2, annotating and classifying the ultrasound image data through a pre-trained image recognition model to obtain an annotation result, including:
[0046] Framing the section feature area in the ultrasound image data through the image recognition model, performing type annotation on the section feature area according to a predetermined determination rule, and calculating the confidence of the section feature area, so as to obtain the annotation result;
[0047] Performing global feature extraction on the ultrasound image data through a pre-trained convolutional neural network model to generate a structured cleft lip and palate imaging report;
[0048] S3. Input the ultrasound image data, annotation results, and cleft lip and palate imaging report into a pre-trained vision-language model for analysis to obtain interpretable inference results, including:
[0049] S31. Use the vision-language model to compare and judge the ultrasound image data with a pre-determined standard image set of cleft lip and palate diagnosis knowledge to obtain the section type and preliminary diagnosis result of the ultrasound image data; wherein, the standard image set of cleft lip and palate diagnosis knowledge includes: the median sagittal section without cleft lip and palate abnormal features, the median sagittal section with cleft lip and palate abnormal features, the posterior nasal triangle section without cleft lip and palate abnormal features, and the posterior nasal triangle section with cleft lip and palate abnormal features; the section types include: median sagittal section and posterior nasal triangle section;
[0050] S32. Integrate the preliminary diagnosis result, annotation result, and cleft lip and palate imaging report through the vision-language model to obtain interpretable inference results.
[0051] Embodiment 2:
[0052] Figure 2 This is a schematic flowchart of a method for prenatal imaging-assisted examination of fetal cleft lip and palate according to another embodiment of the present invention. As Figure 2 shown, the method includes:
[0053] S1. Obtain fetal prenatal facial ultrasound image data;
[0054] S2. Use a pre-trained image recognition model to annotate and classify the ultrasound image data to obtain the annotation result, that is, the probability of cleft lip and palate, including:
[0055] Use the image recognition model to frame the section feature area in the ultrasound image data, label the type of the section feature area according to the pre-determined judgment rule, and calculate the confidence of the section feature area, so as to obtain the annotation result;
[0056] Preferably, the image recognition model is the YOLO-V11 image recognition model, which accurately locates and classifies the cleft lip and palate features in the fetal prenatal facial ultrasound image data. This model can automatically identify and frame the abnormal areas and normal areas and perform corresponding annotations.
[0057] Preferably, the pre-determined judgment rule is: label the type of the section feature area with discontinuous palatal line in the median sagittal section as the abnormal type; label the type of the section feature area with a broken or discontinuous bottom edge in the posterior nasal triangle area in the posterior nasal triangle section as the abnormal type.
[0058] Preferably, based on the pre-trained weights of the YOLO-V11 image recognition model, it is adapted to the fetal prenatal facial ultrasound image dataset through fine-tuning. During the model training process, the bounding box B of each feature region is predicted through a regression task, that is, B = (x, y, w, h), where x and y represent the center coordinates of the bounding box, w represents the width of the bounding box, and h represents the height of the bounding box. At the same time, the model also performs classification prediction, annotates the cleft lip and palate status of the local feature region and calculates the confidence. This training process uses a multi-task loss function to optimize the object detection task, and the form of the loss function is as follows:
[0059] Loss = λ class ·Loss class + λ bbox ·Loss bbox + λ conf ·Loss conf
[0060] Loss represents the loss function, Loss class represents the classification loss, that is, the cross-entropy loss, which is used to optimize the class prediction; λ class represents the weight parameter of the classification loss, Loss bbox represents the regression loss, that is, the mean square error (MSE), which is used to optimize the prediction of the bounding box; λ bbox represents the weight parameter of the regression loss, Loss conf represents the confidence loss, which is used to evaluate the confidence of object detection, λ conf represents the weight parameter of the confidence loss.
[0061] The training process of the image recognition model can be briefly expressed as:
[0062] B, C, Confidence = f yolo (I, F region )
[0063] where I is the input fetal prenatal facial ultrasound image data, F region is the labeled local feature region, C is the classification result, C ∈ {abnormal, normal}, and confidence is the confidence.
[0064] In the actual prediction stage, the image recognition model directly locates and classifies the abnormal regions through the input fetal prenatal facial ultrasound image data without manually annotating the local feature regions. The prediction process is as follows:
[0065] B, C, confidence = f yolo (I).
[0066] In a specific embodiment, the image recognition model clearly calibrates the local features of cleft lip and palate in the ultrasound image data according to different types of ultrasound sections, where the types of ultrasound sections include:
[0067] Mid-sagittal section: In the sagittal section, the continuity of the palatal line is a key indicator for judging cleft lip and palate, and the interruption of the palatal line is usually an obvious abnormal manifestation of cleft lip and palate.
[0068] Post-nasal triangle section: This section mainly focuses on the continuity of the base of the post-nasal triangle area. If the base is broken or discontinuous, it is marked as abnormal.
[0069] The global features of the ultrasound image data are extracted through a pre-trained convolutional neural network model to generate a structured cleft lip and palate image report;
[0070] Preferably, the global features include: the morphological features of the ultrasound image, the distribution of tissue structures, and the predetermined landmark features of cleft lip and palate.
[0071] Preferably, during the decoding process of the pre-trained convolutional neural network, a structured cleft lip and palate image report is generated by combining a relational memory module and memory-driven conditional layer normalization; among them, the relational memory module is used to integrate the memory information of historical ultrasound images during the generation of the cleft lip and palate image report.
[0072] In a specific embodiment, the convolutional neural network model is the R2GenCMN model, which combines a relational memory network and memory-driven conditional layer normalization. The model extracts the global features of the ultrasound image and thus automatically generates a structured cleft lip and palate image report. The process includes three steps: a visual extractor, an encoder, and a decoder, where:
[0073] (1) Visual extractor
[0074] Input the prenatal fetal facial ultrasound image data, and use a pre-trained convolutional neural network (CNN), such as ResNet or VGG, to extract the visual features of the ultrasound image. The visual features extracted through convolutional operations include: the global morphology of the ultrasound image, the distribution of tissue structures, etc. These visual features can be expressed as:
[0075] F visual = CNN(I)
[0076] F visual represents the visual features, CNN represents the convolutional neural network, and I represents the prenatal fetal facial ultrasound image data.
[0077] (2) Encoder
[0078] The visual features obtained from the visual extractor are input into the Transformer encoder, which further processes these visual features to generate hidden states, and these hidden states will be used as the input for the subsequent decoder:
[0079] H encoder = Transformer Encoder(F visual )
[0080] H encoder represents the hidden state, and Transformer Encoder represents the encoder.
[0081] (3) Decoder
[0082] Based on the Transformer-based decoder structure, by combining the relational memory module and memory-driven conditional layer normalization (MCLN), a structured cleft lip and palate imaging report is generated:
[0083] Report = Transformer Decoder(H encoder , M memeory )
[0084] Report represents the cleft lip and palate imaging report, Transformer Decoder represents the decoder, and M memeory represents the memory module, which is used to integrate previous memory information during the generation of the cleft lip and palate imaging report to ensure that the generated report can more accurately reflect the clinical meaning of the image.
[0085] S3. Input the ultrasound image data, annotation results, and cleft lip and palate imaging report into a pre-trained vision-language model for analysis to obtain interpretable inference results, including:
[0086] S31. Compare and judge the ultrasound image data with a predefined set of standard images of cleft lip and palate diagnosis knowledge through the vision-language model to obtain the section type and preliminary diagnosis result of the ultrasound image data; among them, the set of standard images of cleft lip and palate diagnosis knowledge includes: the median sagittal section without cleft lip and palate abnormal features, the median sagittal section with cleft lip and palate abnormal features, the posterior nasal triangle section without cleft lip and palate abnormal features, and the posterior nasal triangle section with cleft lip and palate abnormal features;
[0087] The section types include: median sagittal section and posterior nasal triangle section;
[0088] Preferably, the ultrasonic image data is compared with a predetermined standard image set of cleft lip and palate diagnosis knowledge through a vision large language model to obtain the section type and preliminary diagnosis result of the ultrasonic image data, including: the vision large language model analyzes in combination with a predetermined knowledge embedding mechanism, wherein the knowledge embedding mechanism is: embedding the ultrasonic diagnosis criteria of cleft lip and palate and the standard image set of cleft lip and palate diagnosis knowledge into the reasoning process of the vision large language model in a structured manner.
[0089] In a specific embodiment, in order to make full use of the diagnostic experience in the clinical field, the ultrasonic diagnosis criteria of cleft lip and palate (characteristic patterns of the median sagittal section M1 and the posterior nasal triangle section M2) and related image representations are integrated into the reasoning process of the large language model in a structured manner. These knowledge are embedded through the context of the reasoning process and closely combined with the specific task of ultrasonic image analysis. The knowledge provided to the vision large language model includes not only standard diagnostic rules, but also images with detailed annotations of fetal facial ultrasound features, showing the key structural features of normal and abnormal.
[0090] S32. Integrate the preliminary diagnosis result, annotation result and cleft lip and palate imaging report through the vision large language model to obtain an interpretable reasoning result.
[0091] In a specific embodiment, the vision large language model integrates the preliminary diagnosis result, annotation result and cleft lip and palate imaging report through the following steps:
[0092] S321. Section identification: Section identification is the basis for cleft lip and palate feature analysis. Different standard sections of fetal facial ultrasound involve different cleft lip and palate features, so accurate section judgment is crucial. The vision large model compares the input ultrasonic image data with the injected standard image set of cleft lip and palate diagnosis knowledge to determine the type of the section.
[0093] S322. Key structure analysis: After completing the section identification, the next step is to conduct a detailed analysis of the key structures in the ultrasonic image data to ensure that the vision large language model can correctly identify abnormal features. Since different sections involve different structural features, the vision large language model must conduct targeted analysis based on the identified section type and combined with existing knowledge, and finally determine which image in the standard image set of cleft lip and palate diagnosis knowledge is most consistent with the input ultrasonic image data.
[0094] S323. Auxiliary information integration: In addition to the direct analysis of the ultrasonic image data, in order to make up for the limitations of the large language model in image analysis, external information of the annotation result and the cleft lip and palate imaging report is introduced to construct an information fusion mechanism.
[0095] S324, Comprehensive Judgment: The visual large language model is required to comprehensively consider during the reasoning process according to the steps of sectional plane recognition, key structure analysis, and auxiliary information integration, and give the final interpretable reasoning result. Among them, if the speculation results of each item are consistent, the visual large language model will draw a clear conclusion; if there are inconsistent situations, the visual large language model needs to give a reasonable explanation based on the differences between different pieces of information and give the final decision. At the same time, the visual large language model will provide the basis for each step and the confidence level of the final result, so as to ensure the accuracy and reliability of the imaging findings.
[0096] Preferably, the interpretable reasoning result includes: sectional plane type, label, confidence level, and analysis content.
[0097] Preferably, after the prenatal imaging auxiliary examination method for fetal cleft lip and palate, it further includes: generating an answer based on the obtained text data on cleft lip and palate problems through a pre-trained second large language model, where the training data of the second large language model is obtained through the following steps:
[0098] 1. Recursively generate question-and-answer pairs based on a randomly selected strategy based on the length of independent paragraphs from a pre-constructed corpus to obtain the first question-and-answer data set;
[0099] In a specific embodiment, recursively generating question-and-answer pairs based on a randomly selected strategy based on the length of independent paragraphs from a pre-constructed corpus to obtain the first question-and-answer data set includes:
[0100] 1.1 Construct a professional knowledge base covering six major disciplines, including stomatology, medical imaging, nursing, public health, basic medicine, and clinical medicine.
[0101] 1.2 Generate the first question-and-answer data set based on the cleft lip and palate factual literature in the professional knowledge base;
[0102] 1.2.1 To ensure clear data content and information unitization, divide the long documents in the professional knowledge base into independent paragraphs by chapter or theme, making the generation process more targeted; by removing irrelevant parts such as cited literature, the noise interference in the context is reduced in each generation process, ensuring that the generated result can focus more on the key information of the current paragraph.
[0103] 1.2.2 In the process of gradually generating the first question-and-answer data set, introduce a recursive generation strategy, and generate question-and-answer pairs for small units to ensure that each generation step covers the core facts and concepts of the current paragraph, thus ensuring the integrity and coherence of the data set; among them, the form of the recursive generation process is as follows:
[0104] (Q i,j , A i,j ) = f(Q i,j-1 , Ai,j-1 , S i )
[0105] Among them, Q i,j represents the question of the j-th question-and-answer pair in the i-th unit, and A i,j represents the answer of the j-th question-and-answer pair in the i-th unit. f(·) represents the generation function, and Q i,j-1 represents the question of the previous question-and-answer pair, and A i,j-1 represents the answer of the previous question-and-answer pair, and S i represents the current independent paragraph.
[0106] The core idea of recursive generation is to gradually expand and improve the question-and-answer pairs through each round of generation, ensuring that each question-and-answer pair is closely related to the core knowledge points in the paragraph.
[0107] 1.2.3 To enhance the diversity and flexibility of the generated content, a random strategy based on the paragraph length is introduced to control the number of generations of the question-and-answer pairs. By combining the length of the independent paragraph, the number of question-and-answer rounds for each paragraph is dynamically adjusted, thus avoiding the singularity and repetition of the generated content. The length of the independent paragraph affects the probability distribution of the number of generations. Therefore, a weighted random selection method is used to determine the number of generations of the question-and-answer pairs. Among them, the generation probability of the number of generations of the question-and-answer pairs is:
[0108]
[0109] P(k) represents the generation probability of the number of generations of the question-and-answer pairs, k represents the number of generations of the question-and-answer pairs, and length(S i ) represents the length of the independent paragraph, and w1(length(S i ), w2(length(S i ), and w3(length(S i )) represent the weights set according to the length of the independent paragraph. As the length of the independent paragraph increases, the probability of generating the number of generations of the question-and-answer pairs gradually increases (for example, k = 3). Conversely, shorter independent paragraphs tend to generate fewer generations of the question-and-answer pairs (for example, k = 1).
[0110] 1.2.4 Select the number of generations of the question-and-answer pairs with the highest generation probability as the final number of generations of the question-and-answer pairs, that is: m i represents the number of generations of the question-and-answer pairs generated for each paragraph.
[0111] 1.2.5 The first question-and-answer data set generated is:
[0112]
[0113] Q ref represents the first question-and-answer data set, and |D| represents the number of documents.
[0114] 2. Use the publicly available medical dataset to collate and augment the first Q&A dataset through a pre-trained first large language model to obtain a second Q&A dataset;
[0115] 2.1 To augment the first Q&A dataset, introduce the publicly available medical dataset and extract Q&A content related to cleft lip and palate.
[0116] 2.2 Perform polishing through the first large language model, and design prompt words to guide the first large language model to check the medical accuracy of each pair of Q&A. The specific process is as follows:
[0117]
[0118] Q refined represents the polished Q&A dataset, Q pub represents the medical dataset, P rompt represents the prompt words, LM represents the first large language model, (q′ j ,a′ j ) represents the Q&A pair obtained after being processed by the first large language model.
[0119] 2.3 Construct the second Q&A dataset through the publicly available medical dataset and the polished Q&A dataset, where:
[0120] Q = Q refined ∪Q pub
[0121] Q represents the second Q&A dataset.
[0122] 3. Train and optimize the second large language model through the second Q&A dataset.
[0123] Preferably, the second large language model is the Qwen2.5 - 7B - Instruct model, and the LoRA fine - tuning technology is used to enhance the second large language model's understanding of cleft lip and palate.
[0124] In a specific embodiment, Figure 3 is the framework diagram of the system and the mini - program implemented by applying the prenatal imaging assisted examination method for fetal cleft lip and palate in another embodiment of the present invention. As Figure 3 shown, a multi - modal large language model - driven cleft lip and palate knowledge popularization and teaching system implemented by applying the above - mentioned prenatal imaging assisted examination method for fetal cleft lip and palate, and a mini - program for doctor teaching and public science popularization constructed based on this teaching system. This teaching system includes:
[0125] An imaging discovery generation module, which is used to perform feature analysis on cleft lip and palate ultrasound images through fusing multiple traditional convolutional neural networks to generate interpretable inference results;
[0126] The intelligent consultation and multi-dimensional popular science module is used to generate answers according to the input cleft lip and palate consultation text through adaptive understanding and generation based on large language models. Figure 4 In the system implemented by the prenatal imaging-assisted examination method for fetal cleft lip and palate in another embodiment of the present invention, it is a schematic flowchart of the intelligent consultation and multi-dimensional popular science module. As Figure 4 shown, this intelligent consultation and multi-dimensional popular science module is used for: (1) recursively splitting the articles in the cleft lip and palate corpus, and generating recursive multi-round dialogues based on the content of the cleft lip and palate fact documents, and reviewing the generated content; (2) cleft lip and palate text perception, enhancing the model's understanding of cleft lip and palate knowledge by integrating LoRA fine-tuning technology and retrieval-augmented generation (RAG) technology;
[0127] The task scheduling center is used to obtain the input text and / or image data of the user, and parse the text and / or image data based on the pre-trained third large language model and natural language instructions, and input the text and / or image data into the imaging discovery generation module or the intelligent consultation and multi-dimensional popular science module according to the parsing results.
[0128] In addition, Figure 5 It is a visualization page diagram of a small program implemented by the prenatal imaging-assisted examination method for fetal cleft lip and palate in another embodiment of the present invention. Figure 5 Shows multiple page diagrams implemented by this small program. As Figure 5 shown, this small program includes: fetal cleft lip and palate imaging analysis, intelligent question-answering module, and popular science knowledge functional modules, among which:
[0129] The fetal cleft lip and palate imaging analysis module is constructed through a picture upload mechanism. This fetal cleft lip and palate imaging analysis module provides tools to assist doctors in identifying fetal facial features, and preliminarily judges the possibility of cleft lip and palate, providing support for early diagnosis and timely intervention;
[0130] The intelligent question-answering module is constructed based on a fine-tuned large language model, and includes: doctor intelligent assistant and user intelligent assistant, among which:
[0131] The doctor intelligent assistant supports doctors to ask professional questions and obtain multi-disciplinary intelligent answers, including providing medical advice and relevant knowledge queries;
[0132] The user intelligent assistant answers cleft lip and palate-related questions according to the needs of patients, and provides scientific and healthy advice;
[0133] The popular science knowledge module is constructed through a cleft lip and palate corpus, and includes: knowledge classification module and multimedia display module, among which:
[0134] The knowledge classification module is organized according to four categories of content: "Cleft Lip and Palate Classroom Knowledge", "Surgery", "Before Surgery", and "After Surgery", which is convenient for users to retrieve.
[0135] The multimedia display module presents popular science content in the form of a combination of pictures and texts, enhancing users' comprehension ability and learning interest.
[0136] Through the cleft lip and palate knowledge popularization and teaching system driven by the multi-modal large language model in the above specific embodiments and the small program constructed based on the teaching system, the diagnostic efficiency can be improved, medical resources can be saved, and the public's understanding and awareness of cleft lip and palate can be further enhanced.
[0137] Embodiment 3:
[0138] The present invention also provides a prenatal imaging-assisted examination system for fetal cleft lip and palate. As Figure 6 shown, the device includes a processor 601, a memory 602, a bus 603, and a computer program stored in the memory 602 and executable on the processor 601. The processor 601 includes one or more processing cores. The memory 602 is connected to the processor 601 through the bus 603. The memory 602 is used to store program instructions. When the processor 601 executes the computer program, it implements the steps in the above method embodiments of Embodiment 1 of the present invention.
[0139] Further, as an executable solution, the prenatal imaging-assisted examination system for fetal cleft lip and palate can be a computer unit, and this computer unit can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer unit may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above composition structure of the computer unit is only an example of the computer unit and does not constitute a limitation on the computer unit. It may include more or fewer components than the above, or combine some components, or different components. For example, the computer unit may further include input and output devices, network access devices, a bus, etc., and the embodiments of the present invention do not make any limitations in this regard.
[0140] Further, as an executable solution, the so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the computer unit and connects various parts of the entire computer unit using various interfaces and lines.
[0141] The memory can be used to store the computer program and / or modules. By running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory, the processor realizes various functions of the computer unit. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system and application programs required for at least one function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disks, memory, plug-in hard disks, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0142] Embodiment 4:
[0143] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it realizes the steps of the method as described above.
[0144] Although the present invention has been specifically shown and described in combination with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims, and all of them fall within the protection scope of the present invention.
Claims
1. A prenatal imaging-assisted examination method for fetal cleft lip and palate, characterized in that, Including: S1. Obtain prenatal facial ultrasound image data of the fetus; S2. Label and classify the ultrasound image data through a pre-trained image recognition model to obtain a labeling result, including: Use the image recognition model to frame the sectional feature area in the ultrasound image data, perform type labeling on the sectional feature area according to a predetermined determination rule, and calculate the confidence of the sectional feature area, so as to obtain the labeling result; Extract global features from the ultrasound image data through a pre-trained convolutional neural network model to generate a structured cleft lip and palate image report; S3. Input the ultrasound image data, the labeling result, and the cleft lip and palate image report into a pre-trained vision-language model for analysis to obtain an interpretable reasoning result, including: S31. Use the vision-language model to compare and judge the ultrasound image data with a predetermined standard image set of cleft lip and palate diagnosis knowledge to obtain the sectional type and preliminary diagnosis result of the ultrasound image data; wherein, the standard image set of cleft lip and palate diagnosis knowledge includes: a mid-sagittal sectional plane without cleft lip and palate abnormal features, a mid-sagittal sectional plane with cleft lip and palate abnormal features, a posterior nasal triangle sectional plane without cleft lip and palate abnormal features, and a posterior nasal triangle sectional plane with cleft lip and palate abnormal features; the sectional types include: mid-sagittal section and posterior nasal triangle section; S32. Use the vision-language model to integrate the preliminary diagnosis result, the labeling result, and the cleft lip and palate image report to obtain the interpretable reasoning result.
2. The prenatal imaging-assisted examination method for fetal cleft lip and palate according to claim 1, wherein In the step S2 of performing type labeling on the sectional feature area according to a predetermined determination rule, the determination rule includes: Label the type of the sectional feature area with discontinuous palatal line in the mid-sagittal section as an abnormal type; Label the type of the sectional feature area with a broken or discontinuous bottom edge in the posterior nasal triangle area in the posterior nasal triangle section as an abnormal type.
3. The prenatal imaging-assisted examination method for fetal cleft lip and palate according to claim 2, characterized in that, In the step S2 of extracting global features from the ultrasound image data through a pre-trained convolutional neural network model, the global features include: morphological features of the ultrasound image, distribution of tissue structures, and predetermined landmark features of cleft lip and palate.
4. The prenatal imaging-assisted examination method for fetal cleft lip and palate according to claim 1, wherein The loss function of the image recognition model is: Loss = λ class ·Loss class + λ bbox ·Loss bbox + λ conf ·Loss conf Loss represents the loss function, Loss class represents the classification loss, λ class represents the weight parameter of the classification loss, Loss bbox represents the regression loss, λ bbox represents the weight parameter of the regression loss, Loss conf represents the confidence loss, λ conf represents the weight parameter of the confidence loss.
5. The prenatal imaging-assisted examination method for fetal cleft lip and palate according to claim 1, wherein In step S2, extracting global features from the ultrasound image data through a pre-trained convolutional neural network model to generate a structured cleft lip and palate image report includes: During the decoding process of the pre-trained convolutional neural network, a structured cleft lip and palate image report is generated by combining a relational memory module and memory-driven conditional layer normalization; wherein, the relational memory module is used to integrate the memory information of historical ultrasound images during the generation of the cleft lip and palate image report.
6. The prenatal imaging-assisted examination method for fetal cleft lip and palate according to claim 1, wherein In step S3, using the vision-language model to compare and judge the ultrasound image data with a predetermined standard image set of cleft lip and palate diagnosis knowledge to obtain the sectional type and preliminary diagnosis result of the ultrasound image data includes: The visual large language model is combined with a predetermined knowledge embedding mechanism for analysis, wherein the knowledge embedding mechanism is: embedding the ultrasonic diagnosis criteria for cleft lip and palate and the standard image set of cleft lip and palate diagnosis knowledge into the inference process of the visual large language model in a structured manner.
7. The prenatal imaging-assisted examination method for fetal cleft lip and palate according to claim 1, wherein After the prenatal imaging-assisted examination method for fetal cleft lip and palate, it further includes: generating an answer based on the obtained text data regarding cleft lip and palate problems through a pre-trained second large language model, wherein the training data of the second large language model is obtained through the following steps: Recursively generating question-and-answer pairs according to a random strategy based on the independent paragraph length from a pre-constructed corpus to obtain a first question-and-answer data set; Using a publicly available medical data set to sort and expand the first question-and-answer data set through a pre-trained first large language model to obtain a second question-and-answer data set; Training and optimizing the second large language model through the second question-and-answer data set.
8. The prenatal imaging-assisted examination method for fetal cleft lip and palate according to claim 7, wherein In recursively generating question-and-answer pairs according to a random strategy based on the independent paragraph length from a pre-constructed corpus, the random strategy includes: Determining the generated question-and-answer rounds by means of weighted random selection, wherein the generation probability of the question-and-answer rounds is: P(k) represents the generation probability of the Q&A round, where k represents the Q&A round, and length(S i ) represents the independent paragraph length, and w1(length(S i ), w2(length(S i ), and w3(length(S i )) represent the weights set according to the independent paragraph length.
9. A prenatal imaging-assisted examination system for fetal cleft lip and palate, comprising: It includes a memory and a processor, and the memory stores at least one program, and the at least one program is executed by the processor to implement the prenatal imaging-assisted examination method for fetal cleft lip and palate as described in any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the prenatal imaging-assisted examination method for fetal cleft lip and palate as described in any one of claims 1 to 8.