A method for constructing a medical dataset based on AIGC image generation

By using multimodal semantic association and diffusion models to generate data in the construction of medical data sets, the problems of data bias and privacy restrictions in the construction of medical data sets are solved, and high-quality and diverse medical data set generation is achieved, which improves the diagnostic accuracy and generalization capabilities of the model.

CN119889596BActive Publication Date: 2025-06-13XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510381402.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-13
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The collection of medical data is subject to strict privacy regulations and ethical restrictions, and existing medical data sets have data bias problems, which affects the generalization ability of the model.

Method used

By establishing multimodal semantic associations between medical imaging features, pathological report text and anatomical markers, a structured propt template library is generated, and data is generated using diffusion models to verify anatomical structure rationality and cyclic correction, and finally, the representation space of the optimization model is extracted through dynamic weight mixing and shared feature.

Benefits of technology

The problems of data bias and privacy restrictions are solved, and the generated data sets have high medical rationality and accuracy, improving the model's diagnostic accuracy and generalization ability of various diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889596B_ABST
    Figure CN119889596B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a medical dataset based on AIGC image generation, belonging to the technical field of image generation models, and specifically including: constructing a semantic constraint matrix in the medical field, generating a structured prompt template library, using LORA fine-tuning to input real medical images and structured prompts into a diffusion model, establishing a mapping relationship to generate data, verifying the rationality of anatomical structures for the generated data, making the organ morphological parameters conform to medical prior knowledge through cyclic correction, generating different-scale variants according to the characteristics of the target lesion, forming a three-dimensional continuous parameter space, mixing real and generated data according to dynamic weights, using a hierarchical feature alignment algorithm to optimize the representation space of the diffusion model, extracting the shared feature base layer for pre-training, increasing the weight of real data in the later stage of training, monitoring the feature response pattern during the training process, and generating a test of the same type of evaluation dataset in real time when the indicators do not meet the standards until they reach the standard.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation models, and specifically relates to a method for constructing a medical data set based on AIGC image generation. Background Art

[0002] In the medical field, data is a key factor in driving medical research, disease diagnosis, and treatment plan formulation. High-quality, large-scale, and diverse medical data sets are crucial for developing accurate and reliable medical artificial intelligence models. However, the construction of current medical data sets faces many challenges, which limit the further development of medical artificial intelligence technology.

[0003] The collection of medical data is often subject to strict privacy regulations and ethical restrictions. Patients' personal information and medical records are highly sensitive data and require strict authorization and security measures to be used. This makes large-scale, multi-center data collection difficult, with high costs and long time consumption for data acquisition. Existing medical data sets often have the problem of data bias. For example, the sample size of some diseases is large, while that of some rare diseases is extremely small. This imbalance in data distribution will result in the model having a stronger ability to identify common diseases during training, but lower diagnostic accuracy for rare diseases. In addition, factors such as different races, genders, and ages will also cause differences in medical data. If the data set cannot fully cover these factors, the generalization ability of the model will be limited and it cannot be accurately applied to different patient groups.

[0004] In recent years, significant progress has been made in image generation technology based on artificial intelligence-generated content (AIGC). In view of the problems in data set construction in the medical field, the present invention proposes a method for constructing a medical data set based on AIGC image generation. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for constructing a medical data set based on AIGC image generation, and solve the following technical problems:

[0006] The collection of medical data is often subject to strict privacy regulations and ethical restrictions, and medical data sets often have the problem of data bias.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] A method for constructing a medical data set based on AIGC image generation includes the following steps:

[0009] S1. Establish a multi-modal semantic association among medical image features, pathological report texts, and anatomical markers, construct a semantic constraint matrix in the medical field, and generate a structured prompt template library based on the semantic constraint matrix;

[0010] S2. The controllable fine-tuning module based on LORA inputs the real medical images and the structured prompt into the diffusion model, establishes the mapping relationship between medical features and the latent space, and the diffusion model generates data.

[0011] S3. Verify the rationality of the anatomical structure of the generated data, and judge whether the anatomical features conform to the biological features. If they do not conform to the biological features, the organ morphological parameters are made to conform to the medical prior knowledge through the loop correction mechanism.

[0012] S4. Generate variants of different scales according to the target lesion features. The scales include anatomical position offset, tissue density gradient, and lesion stage evolution, and a three-dimensional continuous parameter space is formed according to the scale change.

[0013] S5. Mix the real data and the generated data according to the dynamic weights, and optimize the representation space of the diffusion model through the hierarchical feature alignment algorithm.

[0014] S6. Extract the shared feature base layer of the real data and the generated data, and apply the shared feature base layer to the pre-training stage of the diffusion model; in the later stage of the diffusion model training, gradually increase the training weight of the real data to the preset ratio, and freeze the feature channels corresponding to the generated data.

[0015] S7. Monitor the feature response mode during the training process of the diffusion model. When it is recognized that the set index of the model does not meet the preset standard, an evaluation data set of the same type as the current training sample is generated in real time to test the model until the set index of the model reaches the preset standard.

[0016] As a further solution of the present invention: in the step S1, the construction of the structured prompt template library includes:

[0017] Extract the anatomical marker information from the DICOM-format medical data, and establish the basic anatomical structure framework according to the anatomical marker information; based on the anatomical structure framework, extract the key description statements in the pathology report, and construct the pathology feature dictionary according to the key description statements; combine the pathology feature dictionary, analyze different medical image features, and establish the multimodal feature mapping relationship; based on the multimodal feature mapping relationship, construct the attention weight matrix of feature fusion; according to the attention weight matrix, generate the prompt grammar structure with constraint conditions; finally, verify and optimize the prompt grammar structure through the semantic verification module of the medical knowledge graph.

[0018] As a further solution of the present invention: in the step S2, the controllable fine-tuning module includes:

[0019] Construct a medical feature embedding space, map the structured prompt into the latent space, in the embedding space, set an anatomical structure attention guidance mechanism to strengthen the features of the set key regions; based on the attention guidance mechanism, generate a pathological feature diffusion path constraint algorithm to control the generation path of lesion features and optimize the tissue texture of the generated data, set a multi-scale feature consistency loss function to calculate the difference degree of key region features at different resolutions, and adjust the model parameters until the loss function converges; verify whether the generated data is consistent with the structured prompt through a semantic alignment validator, otherwise repeat the above process.

[0020] As a further solution of the present invention: in the step S4, generating variants of different scales according to the target lesion features specifically includes:

[0021] Perform basic geometric transformations on the generated data through an anatomical space transformation engine, including translation, rotation, and scaling; during the geometric transformation process, adjust the tissue density in the generated data;

[0022] Set a lesion time axis according to medical knowledge, divide the lesion process into several stages, and in each stage, adjust the lesion features in the image according to the feature changes in the current stage, including the size, shape, and density of the lesion features, where the size change and shape change of the lesion features are simulated by performing geometric transformations on the generated data, and the density of the lesion features is simulated by adjusting the tissue density in the generated image;

[0023] Obtain the image parameters at known time points in the lesion time axis, perform interpolation calculations between the parameters at adjacent time points to obtain the parameter values at intermediate time points, generate a smooth lesion evolution sequence with a preset time point density, and add artifact features that conform to the physiological movement law to the evolution sequence.

[0024] As a further solution of the present invention: in the step S5, the dynamic weight mixing includes:

[0025] Real-time calculate the feature distribution distance between the generated data and the real data, evaluate the data domain difference between the generated data and the real data according to the feature distribution distance; according to the feature distribution distance, adaptively adjust the mixing ratio of the generated data and the real data using KL divergence; according to the mixing ratio, design a hierarchical feature space alignment loss function for optimizing the feature representation of the data; evaluate the confidence of the generated data and remove the generated data with a confidence lower than the threshold;

[0026] Calculate the similarity between the medical image features and the pathological report features through the loss function, and adjust the model parameters by minimizing the loss function to align the features of the two modal data in a common feature space.

[0027] As a further solution of the present invention: In the step S6, the underlying layer for extracting the shared feature basis of the real data and the generated data includes:

[0028] Construct a shared feature encoder to extract the shared features of the generated data and the real data; based on the shared features, distinguish the respective unique features of the generated data and the real data; introduce a cross-domain feature alignment loss function to align the spatial distribution of the same features between the two types of images; according to the aligned feature distribution, shuffle the features and integrate the feature information at different levels and scales.

[0029] As a further solution of the present invention: In the step S6, gradually increasing the training weight of the real data to a preset ratio includes:

[0030] Evaluate the importance of each feature channel in the model, identify the key feature channels, freeze the parameters of the key feature channels during the model training process, adaptively adjust the learning rate of the model training according to the changes in the training stage and the model performance parameters, during the model training process, gradually increase the proportion of the real data in each training batch, and perform feature reconstruction and classification on the real data.

[0031] As a further solution of the present invention: In the step S7, generating an evaluation data set of the same type as the current training sample in real time to test the model until the set index of the model reaches the preset standard includes:

[0032] Perform transformation operations on the current training sample, including rotation, flipping, scaling, adding noise, and sample from the original training data set to form a new evaluation data set with the transformed training sample. Input the generated evaluation data set into the diffusion model being currently trained, record the prediction results of the diffusion model, and calculate the performance of the diffusion model on the evaluation data set according to the set index. According to the test results, adjust the hyperparameters of the model, and modify the number of neurons and the number of layers of the model. After adjusting the model, generate an evaluation data set of the same type as the current training sample again to test the model; repeat the above process until the set index of the model reaches the preset standard.

[0033] The beneficial effects of the present invention:

[0034] (1) The present invention simultaneously establishes multi-modal semantic associations among medical images, pathological report texts, and anatomical markers, solving the integration and sharing obstacles caused by inconsistent data formats and standards in different medical institutions. The generated data comes with semantic information and feature annotations, reducing the need for manual annotation by a large number of professional medical personnel, lowering the annotation workload and cost, and performing semantic verification based on structured prompt templates and medical knowledge graphs to improve the annotation consistency and quality. Simulating different scale changes according to the characteristics of the target lesion, generating diverse data that conforms to medical prior knowledge, covering rare diseases and special situations, solving the problem that traditional data augmentation methods cannot simulate real physiological and pathological changes, and improving the diagnostic accuracy and generalization ability of the model for various diseases.

[0035] (2) The present invention verifies and circularly corrects the anatomical structure rationality of the generated data to ensure its high medical rationality and accuracy, and optimizes image features and textures through various mechanisms in the controllable fine-tuning module. The progressive fine-tuning mechanism enables the model to adaptively adjust according to the training situation, improving the training efficiency and stability. Real-time monitoring of the model training process, generating an evaluation data set for testing when the indicators do not meet the preset standards, and ensuring that the model meets the standards by adjusting parameters and structures. The dynamic test set generation engine identifies the vulnerable features of the model, constructs challenging samples, and expands the test coverage, realizing the continuous evolution of the test set and enhancing the reliability and stability of the model in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The present invention will be further described below with reference to the accompanying drawings.

[0037] Figure 1 It is a schematic flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] Please refer to Figure 1 shown. The present invention is a method for constructing a medical data set based on AIGC image generation, including the following steps:

[0040] S1. Establish multi-modal semantic associations among medical image features, pathological report texts, and anatomical markers, construct a semantic constraint matrix in the medical field, and generate a structured prompt template library based on the semantic constraint matrix;

[0041] S2. The LORA-based controllable fine-tuning module inputs the real medical images and the structured prompt into the diffusion model to establish the mapping relationship between medical features and the latent space, and the diffusion model generates data;

[0042] S3. Verify the anatomical structure rationality of the generated data, and judge whether the anatomical features conform to the biological features. If they do not conform to the biological features, the organ morphological parameters are made to conform to the medical prior knowledge through the loop correction mechanism;

[0043] S4. Generate variants of different scales according to the target lesion features. The scales include anatomical position offset, tissue density gradient, and lesion stage evolution, and form a three-dimensional continuous parameter space according to the scale change;

[0044] S5. Mix the real data and the generated data according to the dynamic weight, and optimize the representation space of the diffusion model through the hierarchical feature alignment algorithm;

[0045] S6. Extract the shared feature base layer of the real data and the generated data, and apply the shared feature base layer to the pre-training stage of the diffusion model; in the later stage of the diffusion model training, gradually increase the training weight of the real data to the preset ratio, and freeze the feature channels corresponding to the generated data;

[0046] S7. Monitor the feature response mode during the training process of the diffusion model. When it is recognized that the set index of the model does not meet the preset standard, an evaluation data set of the same type as the current training sample is generated in real time to test the model until the set index of the model reaches the preset standard.

[0047] In another preferred embodiment of the present invention, the construction of the structured prompt template library in step S1 is a rigorous and complex process, which is an important basis for subsequent generation of high-quality medical images based on AIGC. The following are the detailed steps:

[0048] First of all, the present invention focuses on medical data in DICOM (Digital Imaging and Communications in Medicine) format. As a widely used standard data format in the medical field, DICOM contains a large amount of accurate information about the human anatomical structure. The present invention uses advanced data parsing technology to accurately extract anatomical marker information from these DICOM-format medical data. These marker information cover key contents such as the position, size, shape of each organ in the human body and their relative relationships.

[0049] Based on the extracted anatomical marker information, the present invention begins to establish a basic anatomical structure framework. This is like building a solid skeleton for subsequent work, which presents the anatomical structure of the human body in a systematic and orderly manner. Through this framework, the present invention can clearly understand the hierarchical relationships and spatial layouts of various anatomical parts, providing an accurate reference basis for integrating other medical information in the future.

[0050] After having the basic anatomical structure framework, the present invention shifts its focus to the pathology report. The pathology report is a written description obtained by doctors after a detailed examination of diseased tissues, which contains many key information about disease characteristics. The present invention uses natural language processing technology to deeply analyze the pathology report and extract the key descriptive statements. These statements may describe important information such as the type, degree, and distribution of the lesion.

[0051] According to the extracted key descriptive statements, the present invention constructs a pathology feature dictionary. This dictionary is like a knowledge base that classifies and organizes various pathology features, providing a rich vocabulary resource for establishing multimodal feature mapping relationships in the future. Through the pathology feature dictionary, the present invention can more accurately describe and express the characteristics of different diseases, providing strong support for generating prompts that conform to medical practice.

[0052] Combined with the constructed pathology feature dictionary, the present invention carefully analyzes different medical image features. Medical image features can be obtained through various imaging techniques (such as X-ray, CT, MRI, etc.), which reflect the visual information of the internal structure and lesions of the human body. The present invention uses machine learning and data analysis methods to find the internal connections between medical image features and pathology features and establish multimodal feature mapping relationships.

[0053] This mapping relationship can associate and integrate data of different modalities, enabling the present invention to comprehensively understand the characteristics of diseases from multiple perspectives. For example, through the mapping relationship, the present invention can know that a specific feature presented in a certain medical image corresponds to a certain type of lesion described in the pathology report, thus providing a basis for generating more accurate prompts in the future.

[0054] Based on the established multimodal feature mapping relationship, the present invention further constructs an attention weight matrix for feature fusion. In the process of multimodal data fusion, different features have different importance for generating accurate medical images. The role of the attention weight matrix is to assign a reasonable weight to each feature to highlight important features and suppress irrelevant features.

[0055] The present invention utilizes the attention mechanism in deep learning. Through a large number of experiments and trainings, the weight values of each feature are calculated. These weight values reflect the importance of the feature in the process of generating medical images. Through the attention weight matrix, the present invention can more effectively fuse multi-modal features and improve the quality and accuracy of the generated images.

[0056] According to the attention weight matrix, the present invention generates a prompt grammar structure with constraint conditions. This grammar structure stipulates the format and rules of the prompt, and at the same time incorporates various constraint conditions and mapping relationships established in the previous steps. The prompt with constraint conditions can more precisely guide the AIGC model to generate medical images that conform to medical practice.

[0057] For example, the prompt grammar structure may stipulate that when describing a certain disease, it must include specific anatomical locations, pathological features, and the relationships between them and other information. In this way, the present invention can ensure that the generated medical images have high medical rationality and accuracy.

[0058] Finally, the present invention verifies and optimizes the generated prompt grammar structure through the semantic verification module of the medical knowledge graph. The medical knowledge graph is a database containing rich medical knowledge and semantic relationships, which can comprehensively check and evaluate the prompt grammar structure.

[0059] The semantic verification module will check whether the prompt grammar structure conforms to medical logic and semantic rules, and whether there are contradictions or unreasonable places. If problems are found, they will be corrected and optimized in a timely manner. Through this verification and optimization process, the present invention can ensure that the generated prompt grammar structure is accurate and provides a reliable guarantee for the subsequent generation of high-quality medical images.

[0060] In another preferred embodiment of the present invention, the controllable fine-tuning module in step S2 is a key link to realize the generation of accurate medical data based on AIGC. It can finely regulate the diffusion model to ensure that the generated data meets medical requirements. The following are the specific steps:

[0061] First, the present invention constructs a medical feature embedding space. This space is a high-dimensional mathematical space that can represent various medical features in a compact and effective manner. Medical features include anatomical features, pathological features, imaging features, etc. Through the embedding space, the present invention can uniformly encode and process these features.

[0062] Map the structured prompt into the latent space. The latent space is an abstract space inside the diffusion model, which contains the latent representation of the model for data. By mapping the structured prompt into the latent space, the present invention can guide the diffusion model to generate medical data related to the prompt. This mapping process is like giving the model a clear instruction, telling it what kind of medical data the present invention expects to generate.

[0063] In the medical feature embedding space, the present invention sets up an anatomical structure attention guidance mechanism. The anatomical structure of the human body is very complex, and different anatomical parts have different importance for the diagnosis and treatment of diseases. The role of the attention guidance mechanism is to make the model pay more attention to the set key regional features.

[0064] The present invention makes the model focus more attention on the key anatomical structures during the data generation process by defining attention weights and attention regions. For example, when generating medical images of the lungs, the present invention can set up an attention guidance mechanism to make the model pay more attention to key features such as nodules and inflammations in the lungs, thereby improving the accuracy and pertinence of the generated data.

[0065] Based on the anatomical structure attention guidance mechanism, the present invention generates a pathological feature diffusion path constraint algorithm. The development of diseases usually has certain rules and paths, and the role of the pathological feature diffusion path constraint algorithm is to control the generation path of lesion features to make it conform to the medical disease development law. Constrain and guide the generation process of lesion features. For example, when generating medical data on the development of tumors, the algorithm will control features such as the growth direction and size change of the tumor to make it conform to the actual development process of the tumor inside the human body. At the same time, the algorithm will also optimize the tissue texture of the generated data to make it more realistically simulate real pathological tissues.

[0066] To ensure that the generated data has high quality and consistency at different resolutions, the present invention sets up a multi-scale feature consistency loss function. In medical images, different resolutions may present different feature information. The role of the multi-scale feature consistency loss function is to calculate the difference degree of key regional features at different resolutions.

[0067] By minimizing this loss function, the present invention can adjust the parameters of the model so that the generated data can maintain feature consistency at different resolutions. For example, when generating high-resolution and low-resolution lung images, the loss function will ensure that key features such as nodules and textures in the lungs can be accurately presented at both resolutions.

[0068] Finally, the present invention verifies whether the generated data is consistent with the structured prompt through a semantic alignment validator. The semantic alignment validator analyzes and compares the semantic information of the generated data to check whether it meets the requirements specified in the structured prompt.

[0069] If it is found that the generated data is inconsistent with the prompt, the present invention will repeat the above process to further adjust and optimize the model until the generated data is consistent with the structured prompt. Through this process of verification and iterative optimization, the present invention can ensure that the generated medical data accurately reflects the medical information described in the structured prompt, improving the quality and reliability of the generated data.

[0070] In another preferred embodiment of the present invention, generating variants of different scales according to the target lesion characteristics in step S4 is specifically implemented as follows:

[0071] First, the basic geometric transformations are performed on the generated data with the help of an anatomical space transformation engine. These transformations mainly include translation, rotation, and scaling operations. The translation operation can simulate the movement of the lesion in the human anatomical space, such as the displacement of a tumor in the body. The rotation operation can be used to present the morphological characteristics of the lesion at different angles, just like observing the lesion tissue from different perspectives. The scaling operation can simulate the change in the size of the lesion, which is crucial for studying the growth or atrophy process of the lesion. During the process of these geometric transformations, the present invention also synchronously adjusts the tissue density in the generated data. The change in tissue density is often closely related to the development degree of the lesion. For example, the internal tissue density of a tumor may be different at different stages. By reasonably adjusting the tissue density, the generated data can be made closer to the real medical scenario.

[0072] Secondly, a lesion time axis is set based on rich medical knowledge. A large number of research results on the development processes of various lesions have been accumulated in the medical field. The present invention uses this knowledge to divide the lesion process into several stages in detail. Each stage has unique characteristic changes, which reflect the state of the lesion at different time points. In each stage, the present invention adjusts the lesion characteristics in the image according to the characteristic changes of the current stage. The lesion characteristics mainly involve aspects such as size, shape, and density. For the size change and shape change of the lesion characteristics, the present invention simulates them by performing geometric transformations on the generated data. For example, in the early stage of lesion development, the tumor may be small and regular in shape. As time goes by, the tumor gradually grows and its shape may become irregular. The present invention can simulate this change through geometric transformations such as scaling and translation. For the density change of the lesion characteristics, the present invention realizes it by adjusting the tissue density in the generated image. For example, during the deterioration of the tumor, the tissue density inside it may gradually increase, and the present invention correspondingly increases the tissue density in the generated data.

[0073] Finally, the present invention obtains the image parameters of the known time points in the lesion time axis. These image parameters of the known time points are obtained based on actual medical research or clinical data and have high accuracy and representativeness. Interpolation calculations are performed between the parameters of adjacent time points, and the parameter values of the intermediate time points are estimated through mathematical methods. In this way, a smooth lesion evolution sequence with the density at the preset time points can be generated. This sequence can clearly show the continuous change process of the lesion from one stage to another and is of great significance for medical research and model training. Moreover, in this evolution sequence, the present invention also adds artifact features that conform to the physiological movement law. The organs and tissues inside the human body are in continuous motion, and these motions will produce some artifacts in medical images. Adding artifact features that conform to the physiological movement law can make the generated data more realistic and further improve the adaptability of the model to the real medical scenario.

[0074] In another preferred embodiment of the present invention, the dynamic weight mixing in step S5 is a key step to optimize the representation space of the diffusion model and improve the model performance, and the specific operation is as follows:

[0075] First, the present invention calculates the feature distribution distance between the generated data and the real data in real time. The feature distribution distance is an important index to measure the difference degree of the feature distributions of two groups of data. Through advanced data analysis methods, the present invention can accurately calculate the feature distribution distance between the two. According to this feature distribution distance, the present invention can evaluate the data domain difference between the generated data and the real data. The data domain difference reflects the distribution difference situation of the generated data and the real data in the feature space. If the difference is too large, it may cause the model to deviate when processing data.

[0076] Next, according to the calculated characteristic distribution distance, the present invention uses the Kullback-Leibler divergence to adaptively adjust the mixing ratio of the generated data and the real data. The Kullback-Leibler divergence is a method for measuring the difference between two probability distributions, which can help the present invention dynamically adjust the mixing ratio according to the actual situation of the data characteristic distribution. When the characteristic distribution difference between the generated data and the real data is large, the present invention will appropriately reduce the proportion of the generated data and increase the proportion of the real data to ensure that the model can learn more characteristics of the real data. Conversely, the proportion of the generated data will be appropriately increased to make full use of the diversity of the generated data.

[0077] Then, according to the adjusted mixing ratio, the present invention designs a hierarchical feature space alignment loss function. The role of this loss function is to optimize the feature representation of the data so that the generated data and the real data can be better aligned in the feature space. Hierarchical feature space alignment takes into account the characteristics and importance of different-level features. By minimizing the loss function, the model can more effectively learn the essential features of the data.

[0078] After that, the present invention will evaluate the confidence of the generated data. The confidence reflects the reliability and accuracy of the generated data. The present invention will set a confidence threshold, and for the generated data with a confidence lower than this threshold, the present invention will remove it. This can avoid the negative impact of low-quality generated data on model training and improve the efficiency and accuracy of model training.

[0079] Finally, the present invention calculates the similarity between the medical image features and the pathological report features through the loss function. The medical image features and the pathological report features are two different modalities of data, which reflect the information of the lesion from different angles. The similarity between them is measured through the loss function, and then the model parameters are adjusted by minimizing this loss function. The purpose of this is to align the features of the two modalities of data in a common feature space, so that the model can better fuse different modalities of data and improve the model's comprehensive analysis ability and diagnostic accuracy for the lesion.

[0080] In another preferred embodiment of the present invention, the process of extracting the shared feature base layer of the real data and the generated data in step S6 is a complex and crucial process, which is of great significance for improving the model performance and data fusion effect. The specific steps are as follows:

[0081] First, the present invention constructs a shared feature encoder. This encoder is designed based on deep learning technology, and its core purpose is to accurately extract shared features from generated data and real data. Shared features refer to the information that exists in both generated data and real data and can reflect the essential features of the data. By carefully designing the structure and parameters of the encoder, it can deeply mine and analyze data from different sources to find out the common features among them. For example, in medical image data, shared features may include some basic anatomical structure features, common lesion texture features, etc.

[0082] Based on the extracted shared features, the present invention further distinguishes the unique features of the generated data and the real data respectively. Unique features are the feature information that is unique to each data set and different from other data sets. By comparing the shared features with the original data, the present invention can use feature analysis algorithms and pattern recognition techniques to find out the features that only exist in the generated data or the real data. This step helps the present invention deeply understand the characteristics and differences between the two types of data, providing more targeted information for subsequent feature processing and model training.

[0083] Next, the present invention introduces a cross-domain feature alignment loss function. Since the generated data and the real data may come from different data sources or have different distribution characteristics, the spatial distribution of the same features may be different between the two types of images. The role of the cross-domain feature alignment loss function is to measure this difference and minimize this loss function through an optimization algorithm, so as to achieve the spatial distribution alignment of the same features between the two types of images. This is like calibrating the positions of the same object in two different coordinate systems so that they have a consistent representation in a unified space. Through feature alignment, the model can better integrate the information of the generated data and the real data, improving the processing ability of different types of data.

[0084] Finally, according to the aligned feature distribution, the present invention performs a shuffling operation on the features. The purpose of shuffling the features is to break the original order and association between the features, preventing the model from relying too much on a specific feature arrangement. Then, the present invention integrates the feature information at different levels and different scales. Features at different levels reflect the information of the data at different levels of abstraction, while features at different scales cover the features of the data at different sizes and ranges. By integrating this feature information, the present invention can provide a more comprehensive and rich input for the model, thereby improving the performance and generalization ability of the model.

[0085] In another preferred embodiment of the present invention, gradually increasing the training weight of the real data to a preset ratio in step S6 is a key link in dynamically adjusting the model training process, and the specific operation is as follows:

[0086] First, the present invention evaluates the importance of each feature channel in the evaluation model. A feature channel is a channel in the model used to process different feature information, and different feature channels may have different impacts on the output results of the model. The present invention can use feature importance evaluation algorithms, such as gradient-based feature importance analysis, information gain-based feature selection methods, etc., to determine the importance scores of each feature channel. Based on these scores, the present invention can identify the key feature channels. Key feature channels are those channels that have a greater impact on the model performance, and they contain the most critical information in the data.

[0087] During the model training process, the present invention freezes the parameters of the key feature channels. Freezing the parameters means not updating these parameters during the training process. The purpose of doing this is to keep the information learned by the key feature channels stable and avoid the decline of the model performance caused by parameter updates in the later stage of training. At the same time, according to the changes in the training stage and the model performance parameters, the present invention adaptively adjusts the learning rate of the model training. The learning rate is an important hyperparameter that controls the update step size of the model parameters. In the initial stage of training, the present invention can set a larger learning rate to enable the model to converge quickly; as the training progresses, when the model approaches the optimal solution, the present invention gradually reduces the learning rate to avoid the model skipping the optimal solution. By adaptively adjusting the learning rate, the present invention can improve the training efficiency and stability of the model.

[0088] During the model training process, the present invention gradually increases the proportion of real data in each training batch. Real data has high credibility and representativeness. Increasing its proportion in the training batch can enable the model to learn more real-world data features and improve the generalization ability of the model. At the same time, the present invention performs feature reconstruction and classification on the real data. Feature reconstruction refers to recombining and representing the features of real data to extract more valuable feature information; feature classification is to divide the real data into different categories so that the model can better learn the feature differences of different category data.

[0089] In another preferred embodiment of the present invention, in step S7, generating an evaluation data set of the same type as the current training sample in real time and testing the model until the set indicators of the model reach the preset standard is an important step to ensure the model performance and reliability. The specific process is as follows:

[0090] The present invention performs a series of transformation operations on the current training samples. These transformation operations include rotation, flipping, scaling, and adding noise. The rotation operation can simulate the presentation of an object at different angles, the flipping operation can increase the symmetry variation of the data, the scaling operation can change the size ratio of the data, and adding noise can simulate the interference factors that may occur in actual applications. Through these transformation operations, the present invention can generate diverse sample data and expand the diversity of the data. At the same time, the present invention samples from the original training dataset. The sampling process follows certain rules, such as random sampling or stratified sampling, to ensure that the sampled data can represent the characteristic distribution of the original dataset. The sampled data and the transformed training samples are combined to form a new evaluation dataset.

[0091] Next, the present invention inputs the generated evaluation dataset into the currently trained diffusion model. The model processes each sample in the evaluation dataset and outputs prediction results. The present invention records the prediction results of the diffusion model in detail and calculates the performance of the model on the evaluation dataset according to the set metrics. The set metrics can be accuracy, recall, F1 value, etc., and these metrics can comprehensively reflect the performance of the model.

[0092] According to the test results, the present invention adjusts the hyperparameters of the model. Hyperparameters are parameters that need to be preset before model training, such as learning rate, batch size, regularization coefficient, etc. By adjusting the hyperparameters, the present invention can optimize the training process of the model and improve the performance of the model. At the same time, the present invention also modifies the number of neurons and the number of layers in the model. Adjusting the number of neurons and the number of layers will affect the complexity and expressive ability of the model. According to the test results, the present invention can increase or decrease the number of neurons and the number of layers to find the most suitable model structure for the current task.

[0093] After adjusting the model, the present invention generates an evaluation dataset of the same type as the current training samples again to test the model. The present invention continuously repeats the above process, just like constantly polishing a work of art, and gradually optimizes the performance of the model. Until the set metrics of the model reach the preset standard, at this time the model has high accuracy and reliability and can meet the requirements of actual applications.

[0094] The above has described in detail an embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A method for constructing a medical dataset based on AIGC image generation, characterized in that: The following steps are involved: S1. Establish multimodal semantic associations between medical image features, pathology report texts and anatomical markers, construct a semantic constraint matrix in the medical field, and generate a structured prompt template library based on the semantic constraint matrix; S2, a controllable fine-tuning module based on LORA, inputs the real medical image and the structured prompt into the diffusion model, establishes a mapping relationship between medical features and latent space, and the diffusion model generates data; S3. Verify the rationality of the anatomical structure of the generated data to determine whether the anatomical features are consistent with the biological features. If not, make the organ morphological parameters consistent with medical prior knowledge through a cyclic correction mechanism; S4, generating variants of different scales according to the characteristics of the target lesion, wherein the scales include anatomical position offset, tissue density gradient, and lesion stage evolution, and forming a three-dimensional continuous parameter space according to the scale changes; S5, mixing the real data and the generated data according to the dynamic weights, and optimizing the representation space of the diffusion model by a hierarchical feature alignment algorithm; S6, extracting a shared feature base layer of the real data and the generated data, and applying the shared feature base layer to the pre-training stage of the diffusion model; in the later stage of the diffusion model training, gradually increasing the training weight of the real data to a preset ratio, and freezing the feature channel corresponding to the generated data; S7. Monitor the characteristic response patterns of the diffusion model training process. When it is identified that the set indicators of the model do not meet the preset standards, generate an evaluation data set of the same type as the current training sample in real time and test the model until the set indicators of the model meet the preset standards.

2. The method for constructing a medical data set based on AIGC image generation according to claim 1, characterized in that: In step S1, the construction of the structured prompt template library includes: Extract anatomical marking information from medical data in DICOM format, and establish a basic anatomical structure framework based on the anatomical marking information; based on the anatomical structure framework, extract key description sentences in the pathology report, and build a pathology feature dictionary based on the key description sentences; combine the pathology feature dictionary to analyze different medical imaging features and establish a multimodal feature mapping relationship; based on the multimodal feature mapping relationship, build a feature fusion attention weight matrix; based on the attention weight matrix, generate a prompt grammatical structure with constraints; finally, verify and optimize the prompt grammatical structure through the semantic verification module of the medical knowledge graph.

3. The method for constructing a medical data set based on AIGC image generation according to claim 1, characterized in that: In step S2, the controllable fine-tuning module includes: Construct a medical feature embedding space, map the structured prompt into the latent space, set an anatomical structure attention guidance mechanism in the embedding space, and strengthen the set key area features; based on the attention guidance mechanism, generate a pathological feature diffusion path constraint algorithm to control the generation path of lesion features, optimize the tissue texture of the generated data, set a multi-scale feature consistency loss function, calculate the difference of key area features at different resolutions, and adjust the model parameters until the loss function converges; verify whether the generated data is consistent with the structured prompt through a semantic alignment verifier, otherwise repeat the above process.

4. The method for constructing a medical data set based on AIGC image generation according to claim 1, characterized in that: In step S4, generating variants of different scales according to the target lesion characteristics specifically includes: Performing basic geometric transformations on the generated data through the anatomical space transformation engine, including translation, rotation, and scaling; during the geometric transformation process, adjusting the tissue density in the generated data; A lesion timeline is set according to medical knowledge, and the lesion process is divided into several stages. At each stage, the lesion features in the image are adjusted according to the feature changes of the current stage, including the size, shape and density of the lesion features, wherein the size change and shape change of the lesion features are simulated by geometrically transforming the generated data, and the density of the lesion features is simulated by adjusting the tissue density in the generated image; The image parameters of known time points in the lesion timeline are obtained, and interpolation calculations are performed between the parameters of adjacent time points to obtain the parameter values ​​of the intermediate time points, thereby generating a smooth lesion evolution sequence with a preset time point density, and adding artifact features that conform to the laws of physiological movement to the evolution sequence.

5. The method for constructing a medical data set based on AIGC image generation according to claim 1, characterized in that: In step S5, the dynamic weight mixing includes: Calculate the feature distribution distance between the generated data and the real data in real time, and evaluate the data domain difference between the generated data and the real data according to the feature distribution distance; use KL divergence to adaptively adjust the mixing ratio of the generated data and the real data according to the feature distribution distance; design a hierarchical feature space alignment loss function according to the mixing ratio to optimize the feature representation of the data; evaluate the confidence of the generated data, and remove the generated data with a confidence lower than a threshold; The similarity between medical image features and pathology report features is calculated through the loss function. By minimizing the loss function, the model parameters are adjusted to align the features of the two modal data in a common feature space.

6. The method for constructing a medical data set based on AIGC image generation according to claim 1, characterized in that: In step S6, extracting the shared feature base layer of the real data and the generated data includes: A shared feature encoder is constructed to extract the shared features of generated data and real data. Based on the shared features, the unique features of the generated data and the real data are distinguished. A cross-domain feature alignment loss function is introduced to align the spatial distribution of the same features between the two types of images. According to the aligned feature distribution, the features are shuffled to integrate feature information at different levels and scales.

7. The method for constructing a medical data set based on AIGC image generation according to claim 1, characterized in that: In step S6, gradually increasing the training weight of the real data to a preset ratio includes: Evaluate the importance of each feature channel in the model, identify key feature channels, freeze the parameters of key feature channels during model training, adaptively adjust the learning rate of model training according to the changes in the training stage and model performance parameters, and gradually increase the proportion of real data in each training batch during model training, and reconstruct and classify the features of real data.

8. The method for constructing a medical data set based on AIGC image generation according to claim 1, characterized in that: In step S7, an evaluation data set of the same type as the current training sample is generated in real time, and the model is tested until the set index of the model reaches the preset standard, including: Perform transformation operations on the current training sample, including rotation, flipping, scaling, and adding noise, and sample from the original training data set to form a new evaluation data set with the changed training sample. Input the generated evaluation data set into the currently trained diffusion model, record the prediction results of the diffusion model, and calculate the performance of the diffusion model on the evaluation data set according to the set indicators. According to the test results, adjust the hyperparameters of the model, and modify the number of neurons and layers of the model. After adjusting the model, generate an evaluation data set of the same type as the current training sample again to test the model. Repeat the above process until the set indicators of the model reach the preset standards.

Citation Information

Patent Citations

  • Histopathology image generation method and device based on hidden space diffusion model

    CN117994374A

  • Medical image generation method based on adversarial probability diffusion model

    CN118247374A