Method and apparatus for generating medical three-dimensional model, and electronic device and storage medium
Through the generation network of deep learning models, from medical text to medical three-dimensional models, the problems of insufficient authenticity and editing efficiency of existing medical three-dimensional models are solved, and a high-reality and rapid generation of medical three-dimensional models are achieved.
Patent Information
- Application Number
- PCT/CN2023/140413
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
The existing three-dimensional medical models have shortcomings in terms of authenticity and editing efficiency, and cannot truly reflect the changes of patients in different situations and states.
A generative network using deep learning models is automatically generated by combining the first generative network from medical text to medical two-dimensional images and the second generative network from medical two-dimensional images to medical three-dimensional models.
It realizes the rapid and automatic generation of medical three-dimensional models with high reality, and can dynamically adjust according to different scenarios and states, solving the problem of poor authenticity of the three-dimensional models of traditional Chinese medicine in the prior art.
Smart Images

Figure CN2023140413_26062025_PF_FP_ABST
Abstract
Description
Method, device, electronic device and storage medium for generating medical three-dimensional model Technical Field
[0001] The present application relates to the field of computer technology. Specifically, the present application relates to a method, device, electronic device and storage medium for generating a medical three-dimensional model. Background Art
[0002] At present, a large number of medical 3D models are needed in both medical training and medical teaching. These medical 3D models are usually manually created by 3D model content companies. Not only is the realism of the virtual content limited, but content editing also takes a lot of time, and it cannot truly reflect the changes of patients in different scenarios and conditions.
[0003] From the above, we can see that how to improve the realism of medical three-dimensional models remains to be solved.
[0004] Summary of the Invention
[0005] The embodiments of the present application provide a method, device, electronic device, and storage medium for generating a medical three-dimensional model, which can solve the problem of poor realism of medical three-dimensional models in related technologies. The technical solution is as follows:
[0006] According to one aspect of the present application, a method for generating a medical three-dimensional model includes: obtaining medical text in response to an input operation in an interactive interface; calling a first generation network, and under the guidance of the medical text, learning a first tensor input to the first generation network to obtain a medical two-dimensional image that conforms to the description of the medical text; the first generation network is a trained deep learning model with the ability to generate medical two-dimensional images from medical text; the medical two-dimensional image is input into a second generation network to generate a medical three-dimensional model; the second generation network is a trained deep learning model with the ability to generate medical two-dimensional images from medical three-dimensional models; and the medical three-dimensional model generated by the second generation network is displayed in the interactive interface.
[0007] According to one aspect of the present application, a device for generating a medical three-dimensional model includes: a text acquisition module for acquiring medical text in response to an input operation in an interactive interface; an image generation module for calling a first generation network, and under the guidance of the medical text, learning a first tensor input into the first generation network to obtain a medical two-dimensional image that conforms to the description of the medical text; the first generation network is a trained deep learning model with the ability to generate medical two-dimensional images from medical text; a model generation module for inputting the medical two-dimensional image into a second generation network to generate a medical three-dimensional model; the second generation network is a trained deep learning model with the ability to generate medical two-dimensional images from medical three-dimensional models; and a model display module for displaying the medical three-dimensional model generated by the second generation network in the interactive interface.
[0008] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the method for generating a medical three-dimensional model as described above.
[0009] According to one aspect of the present application, a storage medium stores computer-readable instructions thereon, and the computer-readable instructions are executed by one or more processors to implement the method for generating a medical three-dimensional model as described above.
[0010] According to one aspect of the present application, a computer program product includes computer-readable instructions, which are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the method for generating a medical three-dimensional model as described above.
[0011] The beneficial effects of the technical solution provided by this application are:
[0012] In the above technical solution, after obtaining the medical text, a first generative network capable of generating a medical two-dimensional image from the medical text can be called to learn the first tensor input into the first generative network under the guidance of the medical text to obtain a medical two-dimensional image that conforms to the description of the medical text. Then, based on a second generative network capable of generating a medical three-dimensional model from the medical two-dimensional image, the medical two-dimensional image is input into the second generative network to generate a medical three-dimensional model, and finally the medical three-dimensional model is generated and displayed in the interactive interface. It can be seen that, on the one hand, the use of two generative networks can automatically and quickly generate the medical three-dimensional model required for medical training or medical teaching. On the other hand, according to the actual needs of different medical training or medical teaching, simple medical text can be input into the interactive interface to guide the generated medical three-dimensional model to truly reflect the changes of the patient in different scenarios and states, thereby effectively solving the problem of poor authenticity of medical three-dimensional models existing in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts.
[0014] FIG1 is a schematic diagram of an implementation environment according to the present application;
[0015] FIG2 is a flow chart showing a method for generating a medical three-dimensional model according to an exemplary embodiment;
[0016] FIG3 is a flowchart showing a generation process of a first generation network from medical text to a medical two-dimensional image according to an exemplary embodiment;
[0017] FIG4 is a schematic diagram showing conversion of medical text into a medical three-dimensional model according to an exemplary embodiment;
[0018] FIG5 is a flowchart showing a process of constructing a first training set and a second training set according to an exemplary embodiment;
[0019] FIG6 is a schematic diagram of the first training set and the second training set in the embodiment corresponding to FIG5 ;
[0020] FIG7 is a schematic diagram of a specific interaction of a method for generating a medical three-dimensional model in an application scenario;
[0021] FIG8 is a structural block diagram of a device for generating a medical three-dimensional model according to an exemplary embodiment;
[0022] FIG9 is a hardware structure diagram of an electronic device according to an exemplary embodiment;
[0023] Fig. 10 is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0024] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.
[0025] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0026] The following is an introduction and explanation of several terms involved in this application:
[0027] AIGC, the full English name is AI-Generated Content, and its Chinese meaning is generative artificial intelligence.
[0028] Stable Diffusion, which means stable diffusion in Chinese.
[0029] VR, the full English name is Virtual Reality, and its Chinese meaning is virtual reality.
[0030] MR, the full English name is Mixed Reality, and its Chinese meaning is mixed reality.
[0031] As mentioned above, current medical 3D models are mainly created manually by 3D model content companies. They are usually produced using 3D software or generated by medical image segmentation, reconstruction, and rendering. They still have the defect of low realism. Not only are the operation steps complicated and the production cost high, but there is no way to freely change and edit the content style controlled by the suggested text description. Moreover, it is difficult to adapt to the needs of different content styles in various scenarios such as medical training or medical teaching.
[0032] With the rapid development of deep learning models, generative artificial intelligence (AIGC) has begun to emerge. The core concept of AIGC technology is to use artificial intelligence algorithms to generate content with a certain degree of creativity and quality. By training models and learning from large amounts of data, AIGC can generate relevant content based on input conditions or guidance. For example, by inputting keywords, descriptions, or samples, AIGC can generate matching articles, images, audio, and more.
[0033] The current AIGC works well in natural image models and natural scenes. However, due to the limited number of medical content datasets, it cannot be trained on large-scale medical content datasets, which means it still cannot achieve satisfactory results in generating medical-related content.
[0034] From the above, we can see that the relevant technologies still have the defect of poor realism of medical three-dimensional models, which makes it difficult to conduct accurate quantitative evaluation in application scenarios such as medical training or medical teaching.
[0035] To this end, the medical three-dimensional model generation method provided in this application can effectively improve the realism of the medical three-dimensional model. Accordingly, the medical three-dimensional model generation method is suitable for a medical three-dimensional model generation device, which can be deployed on an electronic device. The electronic device can be a computer device that deploys the von Neumann architecture, for example, the computer device includes a desktop computer, a laptop computer, a server, etc.
[0036] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0037] Figure 1 is a schematic diagram of an implementation environment involved in a method for generating a medical three-dimensional model. It should be noted that this implementation environment is only an example adapted for this application and cannot be considered as providing any limitation on the scope of use of this application.
[0038] The implementation environment includes a client 110 and a server 130 .
[0039] Specifically, the user terminal 110 may be an electronic device that provides a display function. For example, the electronic device may be a desktop computer, a laptop computer, a server, etc. equipped with a display screen, or a smart phone, a tablet computer, etc. equipped with a touch screen.
[0040] Server 130 can be an electronic device such as a desktop computer, laptop computer, or server. It can also be a computer cluster consisting of multiple servers, or even a cloud computing center consisting of multiple servers. Server 130 is used to provide background services, such as, but not limited to, the generation of medical 3D models.
[0041] A communication connection is established in advance between the server 130 and the client 110 via a wired or wireless method, and data is transmitted between the server 130 and the client 110 via the communication connection. The transmitted data includes but is not limited to: the first generation network, the second generation network, etc.
[0042] In one application scenario, for the server 130, the medical three-dimensional model generation service is called to deploy the first generation network and the second generation network used to generate the medical three-dimensional model to the user terminal 110. Specifically, the first generation network is obtained by training based on a pre-constructed first training set, and the second generation network is obtained by training based on a pre-constructed second training set, and the first generation network and the second generation network are sent to the user terminal 110.
[0043] As the user terminal 110 interacts with the server terminal 130, the user terminal 110 receives the first generation network and the second generation network, and then completes the deployment of the first generation network and the second generation network. Then, after the user inputs the medical text through the interactive interface of the user terminal 110, the first generation network and the second generation network can be called to generate a medical three-dimensional model under the guidance of the medical text, so that the medical three-dimensional model can truly reflect the changes of the patient in different scenarios and states, and finally the medical three-dimensional model is displayed in the interactive interface, thereby effectively solving the problem of poor authenticity of medical three-dimensional models existing in related technologies.
[0044] Please refer to FIG2 . An embodiment of the present application provides a method for generating a medical three-dimensional model. The method is applicable to an electronic device, which may be the user terminal 110 in the implementation environment shown in FIG1 .
[0045] In the following method embodiments, for ease of description, the execution subject of each step of the method is taken as an electronic device as an example for illustration, but this does not constitute a specific limitation.
[0046] As shown in FIG2 , the method may include the following steps:
[0047] Step 310: Responding to an input operation in the interactive interface, obtaining medical text.
[0048] First, it should be noted that the interactive interface essentially refers to the page that interacts with the user during the medical 3D model generation process. In some embodiments, the interactive interface is provided by a client running on an electronic device. It is understood that as the client runs on the electronic device, the interactive interface can be displayed on the screen configured for the electronic device. The client can be in the form of an application or a web page, and accordingly, the page can be in the form of a program window or a web page, without limitation herein.
[0049] Secondly, medical text is used to describe the content style, especially the desired style of medical 2D images or 3D models. In other words, the medical text should truly reflect the changes in the patient's condition under different scenarios and conditions. For example, a medical text can be described as "lung with tumor" or more precisely as "tumor in the right lung with a diameter of 20mm."
[0050] In order to understand the content style of the medical two-dimensional image or medical three-dimensional model expected by the user, in some embodiments, an input portal is provided in the interactive interface so that the user can input medical text through the input portal.
[0051] For example, an input box is displayed in an interactive interface, and a user can enter medical text in the input box. The electronic device can then detect the input operation in the input box and understand the user's desired content style for the medical 2D image or 3D model. The input box is considered an input entry provided in the interactive interface, and the input operation is considered an input operation in the interactive interface.
[0052] Of course, in other embodiments, the input operation may vary depending on the input components configured in the electronic device. For example, if the electronic device is a desktop computer with a keyboard, the input operation may refer to a mechanical operation such as a single click on the keyboard; if the electronic device is a tablet computer with a touch screen, the input operation may refer to a gesture operation such as a click or slide on the touch screen. This is not intended to constitute a specific limitation.
[0053] Step 330 : Call the first generative network, and under the guidance of the medical text, learn the first tensor input into the first generative network to obtain a medical two-dimensional image that conforms to the description of the medical text.
[0054] The first generative network is a trained deep learning model that has the ability to generate medical text into two-dimensional medical images. That is, by training the deep learning model with the first training set, a first generative network capable of generating medical text into two-dimensional medical images can be obtained, wherein the first training set is constructed from a large number of two-dimensional medical images and corresponding medical texts. Based on this, the first generative network essentially reflects the mathematical mapping relationship between medical text and two-dimensional medical images. Therefore, by calling the first generative network, the corresponding two-dimensional medical image can be mapped from the medical text based on the mathematical mapping relationship between the medical text and the two-dimensional medical image reflected by the first generative network.
[0055] In some embodiments, the first generative network includes a text encoder, a diffusion model, and an image decoder, wherein the initial diffusion model may be a pre-trained stable diffusion model.
[0056] Figure 3 shows a schematic diagram of the generation process of the first generative network from medical text to medical 2D images. In Figure 3, the text encoder first converts the medical text into a second tensor and obtains a randomly generated first tensor. Then, the second tensor is controlled to guide the diffusion model to perform diffusion learning on the first tensor according to the content style described in the medical text, resulting in a third tensor. Finally, the image decoder is used to convert the third tensor into a medical 2D image.
[0057] The diffusion learning process specifically refers to: as shown in FIG3 , the first tensor is input into the diffusion model and denoised through the inverse process of the diffusion model; the second tensor is used as a guiding condition, and the guiding condition is introduced into the denoising process of the first tensor to obtain the third tensor.
[0058] In the above process, since the medical text essentially describes the user's desired content style of the medical 2D image, the second tensor converted from the medical text can also reflect the user's desired content style. Introducing this second tensor during the denoising process of the first tensor guides the diffusion model to diffuse learning of the first tensor toward the user's desired content style, ultimately resulting in a medical 2D image that conforms to the description of the medical text. In other words, the medical 2D image conforms to the user's desired content style. For example, if the medical text is "lung with tumor," the resulting medical 2D image will not only contain the lungs but also display the tumor within them. Alternatively, if the medical text is "20mm diameter tumor in the right lung," the resulting medical 2D image will not only contain both the left and right lungs but also display a 20mm diameter tumor within the right lung. This allows the medical text to guide the 2D image and truly reflect the patient's changes in different scenarios and states.
[0059] It's worth noting that the first tensor is an image tensor of a set size. Similarly, the third tensor is also of a set size. This means that the size of the first tensor controls the size of the third tensor, and thus the size of the medical 2D image. In other words, the set size can be flexibly set based on the actual image size requirements of the medical 2D image in the application scenario and is not limited here.
[0060] Step 350: Input the medical two-dimensional image into the second generation network to generate a medical three-dimensional model.
[0061] The second generation network is a trained deep learning model capable of generating a medical three-dimensional model from a medical two-dimensional image. In some embodiments, the deep learning model can be a pre-trained end-to-end deep learning model Pixel2Mesh.
[0062] In other words, by training the deep learning model with a second training set, a second generative network capable of generating medical 2D images into medical 3D models can be obtained, wherein the second training set is constructed from a large number of medical 2D images and corresponding medical 3D models, both from the same or different perspectives. It should be noted that the correspondence between the medical 2D images and the medical 3D models refers to the fact that the medical 2D images and the medical 3D models, both from the same or different perspectives, have matching medical texts. Alternatively, it can be understood that the medical 2D images and the medical 3D models, both from the same or different perspectives, conform to the content style described by the matching medical texts.
[0063] Based on this, the first generation network actually reflects the mathematical mapping relationship between the medical two-dimensional image and the medical three-dimensional model. Therefore, by inputting the medical two-dimensional image into the second generation network, the corresponding medical three-dimensional model can be obtained by mapping the medical two-dimensional image based on the mathematical mapping relationship between the medical two-dimensional image and the medical three-dimensional model reflected by the second generation network.
[0064] Figure 4 shows a schematic diagram of the conversion from medical text to a medical three-dimensional model. In Figure 4, the medical text is "a lung with a tumor". Through the first generation network, a picture of the lung with a tumor is obtained, that is, a medical two-dimensional image. Based on this medical two-dimensional image, through the second generation network, a three-dimensional mesh model of the lung with a tumor is obtained, that is, a medical three-dimensional model.
[0065] Step 370: Display the medical three-dimensional model generated by the second generation network in the interactive interface.
[0066] After obtaining the medical 3D model, it can be displayed to the user in an interactive interface. It should be noted that due to differences in display capabilities provided by different electronic devices, such as different display screen resolutions, the generated medical 3D model must be encoded in a data format compatible with the electronic device before it can be displayed in the interactive interface. This ensures that the model can be output in a format suitable for the electronic device.
[0067] Through the above process, on the one hand, the two generation networks can be used to automatically and quickly generate the medical three-dimensional models required for medical training or medical teaching. On the other hand, according to the actual needs of different medical training or medical teaching, simple medical text can be input in the interactive interface to guide the generated medical three-dimensional model to truly reflect the changes of patients in different scenarios and states, thereby effectively solving the problem of poor authenticity of medical three-dimensional models existing in related technologies.
[0068] Referring to FIG. 5 , in an exemplary embodiment, the above method may further include the following steps:
[0069] Step 410 : Acquire an original medical image, and annotate the original medical image with medical text to obtain a medically annotated image.
[0070] Among them, medically annotated images refer to original medical images that carry medical text.
[0071] Regarding the acquisition of medical original images, they can be obtained from medical images publicly available on the Internet, from private medical imaging data of major hospitals / medical schools and other organizations, or from various public medical competition data; further, based on the medical original images obtained above, they can also be segmented according to different organs, tissues, lesion sites, etc. For example, a medical original image containing the left and right lungs can be segmented into two medical original images according to the organs, thereby forming a large-scale medical original image dataset with various organs, tissues, and lesion sites.
[0072] Secondly, it should be noted that annotation refers to adding medical text to the original medical image. In some embodiments, the medical text can be added to the original medical image as a text label or by naming it as a file name, which is not limited here.
[0073] Step 430 : Perform three-dimensional image reconstruction calculation on the medically annotated image to obtain a three-dimensional annotated model.
[0074] The three-dimensional annotation model carries medical text corresponding to the medical annotation image.
[0075] Specifically, based on the medically annotated image, the original medical image and the medical text it carries are determined; three-dimensional image reconstruction technology is used to perform three-dimensional image reconstruction calculation on the original medical image to obtain a medical original three-dimensional model; the original medical three-dimensional model is annotated using the medical text carried by the original medical image to obtain a three-dimensional annotated model.
[0076] That is to say, by using three-dimensional image reconstruction technology, each medically annotated image can obtain a corresponding three-dimensional annotated model. It should be noted that corresponding means that the medically annotated image and the three-dimensional annotated model have matching medical texts. It can also be understood that the original medical image and the original medical three-dimensional model both conform to the content style described by the matching medical text.
[0077] Step 450 : Decompose the three-dimensional annotation model into multiple two-dimensional annotation images according to different viewing angles.
[0078] Each two-dimensional annotated image corresponds to a different perspective and carries medical text corresponding to the three-dimensional annotated model.
[0079] Specifically, based on the three-dimensional annotation model, the original medical three-dimensional model and the medical text it carries are determined; the original medical three-dimensional model is decomposed into multiple original medical two-dimensional images under different perspectives; and the multiple original medical two-dimensional images under different perspectives are annotated using the medical text carried by the original medical three-dimensional model to obtain multiple two-dimensional annotated images under different perspectives.
[0080] It is explained here that for each three-dimensional annotated model, there are corresponding multiple two-dimensional annotated images under different perspectives. The three-dimensional annotated model and its corresponding multiple two-dimensional annotated images have matching medical texts. It can also be understood that the original medical three-dimensional model and its corresponding multiple original medical two-dimensional images all conform to the content style described by the matching medical text.
[0081] Step 470: construct a first training set based on each two-dimensional annotated image and the medical text it carries, and construct a second training set based on the three-dimensional annotated model and each two-dimensional annotated image carrying the medical text corresponding to the three-dimensional annotated model.
[0082] Among them, the two-dimensional annotated image refers to the original medical two-dimensional image carrying medical text; the three-dimensional annotated model refers to the original medical three-dimensional model carrying medical text.
[0083] As shown in Figure 6, the first training set is composed of a large number of original medical two-dimensional images and the medical texts they carry, that is, the first training set is a data set from medical text to medical two-dimensional images; and the second training set is composed of a large number of original medical three-dimensional models, corresponding multiple original medical two-dimensional images from different perspectives and the medical texts they carry, that is, the second training set is a data set from medical two-dimensional images to medical three-dimensional models.
[0084] After obtaining the first training set, the first generative network can be trained based on the first training set, which can specifically include the following steps: obtaining an initial diffusion model; the initial diffusion model is a pre-trained stable diffusion model; based on the first training set, the initial diffusion model is parameter tuned and trained, and if the parameter tuning training of the initial diffusion model meets the set conditions, a diffusion model that has been trained is obtained; based on the text encoder, the trained diffusion model, and the image decoder, a first generative network is constructed.
[0085] After obtaining the second training set, the second generative network can be obtained based on the second training set training. Specifically, the following steps may be included: obtaining a deep learning model pre-trained using a natural image training set; performing parameter tuning training on the deep learning model based on the second training set; if the parameter tuning training of the deep learning model meets the set conditions, the second generative network is obtained.
[0086] The set conditions can be flexibly set based on the actual needs of the application scenario. For example, in one application scenario, the set condition may mean that the parameters are optimized to improve the training accuracy of the model; in another application scenario, the set condition may mean that the number of iterations reaches a threshold to improve the training efficiency of the model. These conditions are not limited here. The parameters can be optimized through loss functions, etc., which is also not limited here.
[0087] Under the influence of the above-mentioned embodiments, by constructing a medical text-medical three-dimensional model dataset and applying it to a pre-trained model for training, an end-to-end medical text to medical three-dimensional model network is realized, which can be applied to multiple clinical scenarios and multiple categories in medicine, and improve the performance of the original three-dimensional AIGC generation model based on text prompts in medical content generation, so as to provide new technologies and tools for subsequent VR / MR clinical practice training and assessment.
[0088] The current traditional medical education system, which primarily relies on animal and human specimens and teaching aids, faces challenges in training medical professionals, including insufficient resources and a high risk of harm to patients. Virtual reality (VR) technology uses computers to generate virtual three-dimensional scenes, providing users with a sense of immersion through visual, auditory, and tactile sensations. With the recent development of virtual reality and mixed reality (MR) technologies, their application in medical education, training, and assessment has become a new trend. Compared to traditional teaching models, VR and MR technologies offer significant advantages. By collecting multi-dimensional data from the real human body and constructing digital models of the human body or target tissue through simulation modeling, they enable low-cost, repeatable, and quantifiable digital teaching. This allows students to learn and grow in a repeatable practice environment, provides rich case studies, and provides scientifically standardized simulation materials, effectively alleviating the challenges of insufficient teaching resources and the difficulty of quantitative assessment. Furthermore, the use of mixed reality technology can achieve better results in clinical teaching and preoperative planning.
[0089] Current mixed reality medical training and teaching tasks require a large number of medical 3D models. On the one hand, the mainstream method for existing medical 3D virtual models is to directly produce them using geometric modeling software. Some solutions also use medical images to obtain 3D models through segmentation, reconstruction, rendering and other steps. However, the operation steps are complicated and the production cost is high. There is no way to achieve simple text descriptions to control the free change and editing of content style; on the other hand, the current AIGC 3D generation large model works well in natural image models and natural scenes, but it cannot achieve satisfactory results in generating medical-related content, and it is difficult to adapt to the needs of different content styles in various different scenarios of medical training.
[0090] FIG7 shows a schematic diagram of a medical three-dimensional model in an application scenario. As shown in FIG7 , in this application scenario, the service end can be a server, etc., and the first user end and the second user end can both be electronic devices capable of interacting with the user. For example, the first user end can be a desktop computer, and the second user end can be a laptop computer, etc. Then, the first user end can generate a medical three-dimensional model from the medical text input by the user through user interaction, and the second user end can conduct medical training or medical teaching assessments for the user through user interaction. It is worth mentioning that the first user end and the second user end can also be deployed on the same electronic device, which does not constitute a specific limitation here.
[0091] Specifically, the server constructs a first training set and a second training set, so as to obtain a first generated network based on the first training set and a second generated network based on the second training set, and deploys both to the first user terminal.
[0092] After the first user terminal completes the deployment of the first generation network and the second generation network, the user can use the first user terminal to input corresponding medical text according to the content style of the desired medical three-dimensional model, so that the first user terminal generates a medical two-dimensional image from the medical text by calling the first generation network, and generates a medical three-dimensional model from the medical two-dimensional image by calling the second generation network. The medical three-dimensional model thus obtained meets the content style expected by the user, and then the medical three-dimensional model is transmitted to the second user terminal.
[0093] As the client in the second user terminal runs, the second user terminal will show the user a virtual scene with an imported medical three-dimensional model. The virtual scene is a digital scene constructed using computer technology for medical training or medical teaching assessments to simulate the environment required for medical training or medical teaching (such as a medical operating laboratory), and then conduct medical training or medical teaching assessments on the user by capturing the user's simulated operations in the simulated environment.
[0094] For example, when a user wishes to participate in a medical training assessment, they can launch the client and enter a corresponding virtual scene. For example, the virtual scene could be a simulated medical surgical laboratory, and the imported medical 3D model could be a patient's lung with a tumor. The user's simulated operations on the patient's lung with a tumor include, but are not limited to, simulated manipulation of surgical instruments and responses to assessment questions. As the user's simulated operations progress, the second user terminal can also capture corresponding operation videos using image acquisition devices and corresponding sensor data using mixed reality devices to assess the user's medical training.
[0095] In this application scenario, by applying AIGC technology to the medical field, AIGC technology is applied to the generation of medical mixed reality models, filling the application gap in this area; it can also make up for the shortcomings of AIGC in the generation of medical content models due to the lack of medical content learning. At the same time, the use of this generation solution can solve the three-dimensional model requirements in different scenarios of mixed reality medical simulation training and teaching; in addition, through this text generation method, medical three-dimensional content under various conditions can be quickly obtained, reducing the generation cost of three-dimensional models for mixed reality medical development.
[0096] The following are embodiments of the device of the present application, which can be used to execute the method for generating a medical three-dimensional model involved in the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method for generating a medical three-dimensional model involved in the present application.
[0097] Please refer to Figure 8. An embodiment of the present application provides a medical three-dimensional model generation device 900, including but not limited to: a text acquisition module 910, an image generation module 930, a model generation module 950, and a model display module 970.
[0098] The text acquisition module 910 is used to acquire medical text in response to input operations in the interactive interface.
[0099] Image generation module 930 is configured to invoke a first generative network and, guided by medical text, learn the first tensor input to the first generative network to generate a two-dimensional medical image that conforms to the medical text description. The first generative network is a trained deep learning model capable of generating medical two-dimensional images from medical text.
[0100] The model generation module 950 is configured to input the medical 2D image into a second generation network to generate a medical 3D model. The second generation network is a trained deep learning model capable of generating a medical 3D model from the medical 2D image.
[0101] The model display module 970 is used to display the medical three-dimensional model generated by the second generation network in the interactive interface.
[0102] In an exemplary embodiment, the first generative network includes a text encoder, a diffusion model, and an image decoder.
[0103] Among them, the image generation module 930 is also used to use the text encoder to convert the medical text into a second tensor and obtain the randomly generated first tensor; control the second tensor to guide the diffusion model to perform diffusion learning on the first tensor according to the content style described in the medical text to obtain a third tensor; use the image decoder to convert the third tensor into the medical two-dimensional image.
[0104] In an exemplary embodiment, the image generation module 930 is further used to input the first tensor into the diffusion model and perform denoising through the inverse process of the diffusion model; use the second tensor as a guiding condition, and introduce the guiding condition into the denoising process of the first tensor to obtain the third tensor.
[0105] In an exemplary embodiment, the apparatus 900 further includes: a training set construction module.
[0106] Among them, the training set construction module is used to obtain a medical original image and annotate the medical original image with medical text to obtain a medical annotated image; the medical annotated image refers to a medical original image carrying medical text; a three-dimensional image reconstruction calculation is performed on the medical annotated image to obtain a three-dimensional annotated model; the three-dimensional annotated model carries medical text corresponding to the medical annotated image; the three-dimensional annotated model is decomposed into multiple two-dimensional annotated images according to different perspectives; each of the two-dimensional annotated images corresponds to a different perspective and carries medical text corresponding to the three-dimensional annotated model; a first training set is constructed based on each of the two-dimensional annotated images and the medical text they carry, and a second training set is constructed based on the three-dimensional annotated model and each of the two-dimensional annotated images carrying medical text corresponding to the three-dimensional annotated model; wherein the first training set is used to train the first generation network, and the second training set is used to train the second generation network.
[0107] In an exemplary embodiment, the apparatus 900 further includes: a first training module.
[0108] Among them, the first training module is used to obtain an initial diffusion model; the initial diffusion model is a pre-trained stable diffusion model; based on the first training set, the initial diffusion model is parameter tuned and trained to obtain the trained diffusion model; based on the text encoder, the trained diffusion model, and the image decoder, the first generation network is constructed.
[0109] In an exemplary embodiment, the apparatus 900 further includes: a second training module.
[0110] Among them, the second training module is used to obtain a deep learning model obtained by pre-training using a natural image training set; based on the second training set, the deep learning model is subjected to parameter tuning training; if the parameter tuning training of the deep learning model meets the set conditions, the second generation network is obtained.
[0111] In an exemplary embodiment, the apparatus 900 further includes an assessment module.
[0112] Among them, the assessment module is used to import the medical three-dimensional model into a constructed virtual scene; the virtual scene is constructed for medical training or medical teaching; based on the target object's simulated operation on the medical three-dimensional model in the virtual scene, the medical training or medical teaching of the target object is assessed.
[0113] It should be noted that the medical three-dimensional model generation device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when generating a medical three-dimensional model. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the medical three-dimensional model generation device will be divided into different functional modules to complete all or part of the functions described above.
[0114] In addition, the medical three-dimensional model generation device and the medical three-dimensional model generation method provided in the above embodiments belong to the same concept, and the specific manner in which each module performs operations has been described in detail in the method embodiment and will not be repeated here.
[0115] Fig. 9 shows a schematic diagram of the structure of an electronic device according to an exemplary embodiment. The electronic device is applicable to the user terminal 110 in the implementation environment shown in Fig. 1 .
[0116] It should be noted that the electronic device is only an example adapted for the present application and should not be considered to provide any limitation on the scope of use of the present application. The electronic device should not be interpreted as needing to rely on or necessarily having one or more components in the exemplary electronic device 2000 shown in FIG9 .
[0117] The hardware structure of the electronic device 2000 may vary greatly due to different configurations or performances. As shown in FIG9 , the electronic device 2000 includes a power supply 210 , an interface 230 , at least one memory 250 , and at least one central processing unit (CPU) 270 .
[0118] Specifically, the power supply 210 is used to provide operating voltage for various hardware devices on the electronic device 2000 .
[0119] The interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. Of course, in other examples adapted by this application, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, and at least one USB interface 237, as shown in FIG9 , and this is not a specific limitation.
[0120] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon include an operating system 251, application 253 and data 255, etc. The storage method can be temporary storage or permanent storage.
[0121] Among them, the operating system 251 is used to manage and control the various hardware devices and application programs 253 on the electronic device 2000 to enable the central processing unit 270 to calculate and process the massive data 255 in the memory 250. It can be Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0122] Application 253 is a computer-readable instruction that performs at least one specific task based on operating system 251. It may include at least one module (not shown in FIG9 ), each of which may contain computer-readable instructions for electronic device 2000. For example, a device for generating a three-dimensional medical model may be considered application 253 deployed on electronic device 2000.
[0123] The data 255 may be photos, pictures, etc. stored in a disk, or may be a first generated network, a second generated network, etc. stored in the memory 250 .
[0124] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer-readable instructions stored in the memory 250, thereby performing operations and processing on the massive amount of data 255 in the memory 250. For example, the method for generating a medical three-dimensional model may be completed by the central processing unit 270 reading a series of computer-readable instructions stored in the memory 250.
[0125] In addition, the present application can also be implemented through hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present application is not limited to any specific hardware circuits, software, or a combination of the two.
[0126] Please refer to FIG. 10 . An electronic device 4000 is provided in an embodiment of the present application. The electronic device 4000 may include a desktop computer, a laptop computer, a server, etc.
[0127] In FIG. 10 , the electronic device 4000 includes at least one processor 4001 and at least one memory 4003 .
[0128] Data exchange between the processor 4001 and the memory 4003 can be achieved via at least one communication bus 4002. The communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. The communication bus 4002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG10 shows only one thick line, but this does not mean that there is only one bus or one type of bus.
[0129] Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
[0130] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0131] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program instructions or codes in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited to these.
[0132] Computer-readable instructions are stored in the memory 4003 , and the processor 4001 can read the computer-readable instructions stored in the memory 4003 through the communication bus 4002 .
[0133] The computer-readable instructions are executed by one or more processors 4001 to implement the method for generating a medical three-dimensional model in the above-mentioned embodiments.
[0134] In addition, an embodiment of the present application provides a storage medium having computer-readable instructions stored thereon, and the computer-readable instructions are executed by one or more processors to implement the method for generating a medical three-dimensional model as described above.
[0135] In an embodiment of the present application, a computer program product is provided. The computer program product includes computer-readable instructions, which are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the method for generating a medical three-dimensional model as described above.
[0136] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0137] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for generating a medical three-dimensional model, characterized in that, The method includes: In response to an input operation in the interaction interface, obtain a medical text; Invoke a first generation network, and under the guidance of the medical text, learn a first tensor input to the first generation network to obtain a medical two-dimensional image that conforms to the description of the medical text; the first generation network is a deep learning model that has been trained and has the ability to generate from a medical text to a medical two-dimensional image; Input the medical two-dimensional image into a second generation network for generating a medical three-dimensional model; the second generation network is a deep learning model that has been trained and has the ability to generate from a medical two-dimensional image to a medical three-dimensional model; Display the medical three-dimensional model generated by the second generation network in the interaction interface.
2. The method according to claim 1, characterized in that, The first generation network includes a text encoder, a diffusion model, and an image decoder; The invoking the first generation network, and under the guidance of the medical text, learning a first tensor input to the first generation network to obtain a medical two-dimensional image that conforms to the description of the medical text includes: Use the text encoder to convert the medical text into a second tensor, and obtain the randomly generated first tensor; Control the second tensor to guide the diffusion model to perform diffusion learning on the first tensor according to the content style described by the medical text to obtain a third tensor; Use the image decoder to convert the third tensor into the medical two-dimensional image.
3. The method according to claim 2, wherein The controlling the second tensor to guide the diffusion model to perform diffusion learning on the first tensor according to the content style described by the medical text to obtain a third tensor includes: Input the first tensor into the diffusion model and perform denoising through the reverse process of the diffusion model; Use the second tensor as a guiding condition and introduce the guiding condition into the denoising process of the first tensor to obtain the third tensor.
4. The method according to claim 1, wherein The method further includes: Obtain a medical original image, and perform annotation on the medical original image with respect to a medical text to obtain a medical annotated image; the medical annotated image refers to a medical original image carrying a medical text; Perform three-dimensional image reconstruction calculation on the medical annotated image to obtain a three-dimensional annotated model; the three-dimensional annotated model carries a medical text corresponding to the medical annotated image; Decompose the three-dimensional annotated model into a plurality of two-dimensional annotated images according to different perspectives; each of the two-dimensional annotated images corresponds to a different perspective and carries a medical text corresponding to the three-dimensional annotated model; Construct a first training set based on each of the two-dimensional annotated images and the medical text carried by them, and construct a second training set based on the three-dimensional annotated model and each of the two-dimensional annotated images carrying a medical text corresponding to the three-dimensional annotated model; Among them, the first training set is used to train the first generation network, and the second training set is used to train the second generation network.
5. The method according to claim 4, characterized in that, Before the invoking the first generation network, and under the guidance of the medical text, learning a first tensor input to the first generation network to obtain a medical two-dimensional image that conforms to the description of the medical text, the method further includes: Obtain an initial diffusion model; the initial diffusion model is a pre-trained Stable Diffusion model. Based on the first training set, perform parameter tuning training on the initial diffusion model to obtain the trained diffusion model. Based on the text encoder, the trained diffusion model, and the image decoder, construct the first generation network.
6. The method according to claim 4, wherein Before inputting the medical two-dimensional image into the second generation network for generating a medical three-dimensional model, the method further includes: Obtain a deep learning model pre-trained using a natural image training set. Based on the second training set, perform parameter tuning training on the deep learning model. If the parameter tuning training of the deep learning model meets the set conditions, obtain the second generation network.
7. The method according to any one of claims 1 to 6, characterized in that, After displaying the medical three-dimensional model generated by the second generation network in the interaction interface, the method includes: Import the medical three-dimensional model into the constructed virtual scene; the virtual scene is constructed for medical training or medical teaching. Based on the simulation operation of the target object on the medical three-dimensional model in the virtual scene, Conduct an assessment of the medical training or medical teaching of the target object.
8. A generating device for a medical three-dimensional model, characterized in that The device includes: A text acquisition module for acquiring medical text in response to an input operation in the interaction interface. An image generation module for calling the first generation network to learn the first tensor input into the first generation network under the guidance of the medical text to obtain a medical two-dimensional image that conforms to the description of the medical text; the first generation network is a deep learning model that has been trained and has the ability to generate from medical text to medical two-dimensional images. A model generation module for inputting the medical two-dimensional image into the second generation network to generate a medical three-dimensional model; the second generation network is a deep learning model that has been trained and has the ability to generate from medical two-dimensional images to medical three-dimensional models. A model display module for displaying the medical three-dimensional model generated by the second generation network in the interaction interface.
9. An electronic device, characterized in that, Includes: At least one processor and at least one memory, where The memory stores computer-readable instructions. The computer-readable instructions are executed by one or more of the processors, enabling the electronic device to implement the method for generating a medical three-dimensional model according to any one of claims 1 to 7.
10. A storage medium having computer-readable instructions stored thereon, characterized in that, The computer-readable instructions are executed by one or more processors to implement the method for generating a medical three-dimensional model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for webpage displaying of three-dimensional medical model and system thereof
CN102880454A
Key structure reconstruction method and device based on three-dimensional image and computer equipment
CN115115772A
Object generation method, device and system
CN116228959A
Style digital human generation method, device and equipment and readable storage medium
CN116310113A
Text generation 3D printing model method based on big data deep learning
CN116580156A