Breast cancer segmentation architecture generation method based on text processing
By generating a vector database and matching the knowledge base with a large language model, the breast cancer segmentation neural network architecture is automatically designed, which solves the problems of high resource consumption and high computing costs in the existing technology, and improves the accuracy and reliability of architecture generation.
Patent Information
- Application Number
- CN202510126908.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-27
AI Technical Summary
The prior art is difficult to automatically design efficient neural network architectures in breast cancer tumor segmentation, resulting in large resource consumption, high computing costs, high time costs and difficult to discover the optimal structure.
By collecting medical image articles, a vector database is generated, and a large language model is used to match the knowledge base to generate the latest results, and finally screen the optimal solution through the neural network to generate the breast cancer segmentation architecture.
Reduces the requirements for professional knowledge, improves the accuracy and reliability of architecture generation, reduces resource consumption and time costs, and enables more comprehensive utilization of modules in the knowledge base.
Smart Images

Figure CN119988523A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and in particular relates to a method for generating a breast cancer segmentation framework based on text processing. Background Art
[0002] Breast cancer is one of the most common malignant tumors in women. Regular breast screening is very important for developing diagnosis and treatment plans and improving survival rates. At present, ultrasound imaging is widely used in the clinical diagnosis of breast lesions due to its advantages of high flexibility, non-invasiveness and low cost. However, complex ultrasound patterns, similar intensity distribution, variable tumor morphology and fuzzy boundaries pose great challenges to the segmentation of breast tumors. To solve the above problems, CNN (convolutional neural network) with strong nonlinear learning ability is widely used in breast cancer tumor segmentation. In order to adapt to the characteristics of fuzzy boundaries of breast ultrasound images and uncertain size and shape of target areas, many researchers have designed and proposed various modules by manually designing neural networks, and have made certain improvements on the original neural network architecture. However, there are the following problems with manually designing neural network architecture: 1. Deep professional knowledge is required. Effective design of neural network architecture requires in-depth understanding of deep learning theory, different types of network structures (such as convolutional neural networks, recurrent neural networks, graph neural networks, etc.) and various optimization techniques. Designers who lack relevant experience and knowledge may find it difficult to design a network with excellent performance. 2. It is difficult to find the optimal structure. Human intuition and experience may not be sufficient to find the optimal or nearly optimal network structure in some cases. Automated methods (such as neural architecture search) can explore a wider design space and discover innovative architectures that human designers may overlook. 3. High resource consumption: Manually designing network architectures requires a lot of manpower and computing resources, especially when trying multiple different architectures to find the best solution. This is particularly prominent when resources are limited, limiting the scope of application of manual design methods.
[0003] In order to solve the difficulty of manually designing neural network architectures, automated neural architecture search (NAS) was proposed. Although NAS reduces manual participation, it still has the following problems: 1. High computational cost: huge resource consumption. NAS usually requires a lot of computing resources, especially when the search space is large. The time and computing power required to train and evaluate a large number of candidate architectures are very high. This makes NAS difficult to promote in resource-limited environments. 2. High time cost: The complete NAS process may take days or even weeks, which is unrealistic for rapid iteration and real-time applications. 3. Complex search space design: The difficulty of defining a reasonable search space is still difficult for a medical worker who does not understand deep learning. An overly large search space will make the search process more complicated and time-consuming, while an overly small search space may limit the possibility of finding the optimal architecture. Designing a balanced search space that covers sufficient diversity still requires deep expertise and experience.
[0004] In recent years, in the field of artificial intelligence, large language models have made significant breakthroughs. Large language models have powerful knowledge encoding and storage capabilities, text and code understanding and generation capabilities, and complex task reasoning capabilities. The large language model's powerful knowledge encoding and storage capabilities and complex task reasoning capabilities make it potential for neural network architecture design. However, general large language models lack relevant professional knowledge accumulation. Without corresponding prompts and knowledge, they can only generate existing simple architectures based on basic knowledge during training, such as U-net. When more complex functions are required, LLMs find it difficult to break out of the original framework to meet the requirements. Summary of the invention
[0005] In order to solve the problem that LLM is difficult to jump out of the original framework to meet the requirements when text processing needs to realize more complex functions, the present invention proposes a breast cancer segmentation architecture generation method based on text processing.
[0006] The technical solution of the present invention is: a method for generating a breast cancer segmentation architecture based on text processing comprises the following steps:
[0007] S1. Collect several medical imaging articles, process each medical imaging article, and generate a vector database;
[0008] S2, using each knowledge base of the vector database to process the initial query statement and generate several latest results;
[0009] S3. Use a neural network to select the optimal solution from the latest results, obtain the final neural network architecture, and use the final neural network architecture to process breast cancer data.
[0010] Furthermore, S1 includes the following sub-steps:
[0011] S11, collect several medical imaging articles and generate a set of knowledge items;
[0012] S12, cleaning the knowledge item set, and performing vectorization operation on the cleaned knowledge item set using the embedding model;
[0013] S13. Store the knowledge item set after the vectorization operation into the vector database.
[0014] Furthermore, in S11, the expression of the knowledge item set P is:
[0015] P = {p1, p2, p3, ..., p i ,…,p n};
[0016] Connect(S i , M i )→p i ;
[0017] In the formula, p1, p2, p3, p i and p n represents the knowledge items of the 1st, 2nd, 3rd, i and nth medical imaging articles respectively, S i represents the summary content of the i-th medical imaging article, M i It represents the method content of the i-th medical imaging article, and Connect(·) means connecting the summary content with the method content.
[0018] Furthermore, in S13, the expression of the vector database D is:
[0019] D = {d1, d2, d3, ..., d k};
[0020]
[0021] In the formula, d1, d2, d3, d k and d j Respectively represent the 1st, 2nd, 3rd, kth and jth knowledge bases of the vector database, v j1 、v j2 、v j3 and Represent the 1st, 2nd, 3rd and Nth in the jth knowledge base respectively j knowledge items, k represents the total number of knowledge bases, N j Indicates the total number of knowledge items contained in the knowledge base.
[0022] Furthermore, S2 includes the following sub-steps:
[0023] S21, using each knowledge base of the vector database to perform semantic matching on the initial query statement to generate a semantic matching result;
[0024] S22, generating a number of initial results as candidate sets according to the semantic matching results;
[0025] S23. Optimize each initial result of the candidate set to generate several latest results.
[0026] Furthermore, in S21, the semantic matching result R i The expression of (Q′) is:
[0027] R i (Q′) = {r1, r2, ..., r n};
[0028] Q′=OptimizePrompt(Q);
[0029] In the formula, r1 represents the first result of semantic matching in the knowledge base, r2 represents the second result of semantic matching in the knowledge base, and r n represents the nth result of semantic matching in the knowledge base, Q represents the initial query statement, OptimizePrompt(·) represents the combination of the initial query statement and the pre-prepared prompt, and Q′ represents the result obtained after feature engineering optimization.
[0030] Furthermore, in S22, the jth initial result S j The expression is:
[0031]
[0032] In the formula, LLM j represents the j-th large language model, R j represents the matching result between the j-th large language model and the knowledge base, m represents the number of groups of matching results between the large language model and the knowledge base, and GenerateInitialScheme(·) represents the initial scheme generated by the j-th large language model according to the matching result of the corresponding knowledge base.
[0033] Furthermore, in S23, the kth latest result S′ k The expression is:
[0034]
[0035] In the formula, SelectAndBuildScheme(·) indicates that the k-th large language model selects a suitable module from the M set and rebuilds a new scheme, M represents the candidate set composed of the initial scheme set, m represents the number of groups of matching results between the large language model and the knowledge base, and LLMk Represents the k-th large language model.
[0036] Furthermore, the optimal solution S * The expression is:
[0037]
[0038] In the formula, S′ k represents the kth latest result, Score(·) represents the score of the kth latest result, argmax represents the corresponding choice of the solution with the highest score, and S represents the specific solution with the highest score.
[0039] The beneficial effects of the present invention are:
[0040] (1) The present invention builds a neural network architecture suitable for breast cancer tumor segmentation, which greatly reduces the requirements for user professional knowledge, improves the accessibility and application scope of the technology, and enables the large language model to have sufficient prior knowledge, which is conducive to the large language model to perform better in the vertical field of neural network architecture generation. The establishment of a decentralized sub-knowledge base can ensure that the modules in the knowledge base can be fully used;
[0041] (2) The present invention can improve the accuracy and reliability of architecture generation, and can more comprehensively retrieve and utilize modules in the knowledge base, avoiding the problem of modules being ignored due to the large knowledge base or token restrictions in traditional methods. The present invention can also effectively reduce the "hallucination" phenomenon that may be caused by a single model, thereby improving the accuracy and reliability of the generated neural network architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Flowchart of the method for text-based breast cancer segmentation architecture generation. DETAILED DESCRIPTION
[0043] The embodiments of the present invention will be further described below in conjunction with the accompanying drawings.
[0044] like Figure 1 As shown, the present invention provides a method for generating a breast cancer segmentation architecture based on text processing, comprising the following steps:
[0045] S1. Collect several medical imaging articles, process each medical imaging article, and generate a vector database;
[0046] S2, using each knowledge base of the vector database to process the initial query statement and generate several latest results;
[0047] S3. Use a neural network to select the optimal solution from the latest results, obtain the final neural network architecture, and use the final neural network architecture to process breast cancer data.
[0048] In this embodiment of the present invention, S1 includes the following sub-steps:
[0049] S11, collect several medical imaging articles and generate a set of knowledge items;
[0050] S12, cleaning the knowledge item set, and performing vectorization operation on the cleaned knowledge item set using the embedding model;
[0051] S13. Store the knowledge item set after the vectorization operation into the vector database.
[0052] In an embodiment of the present invention, relevant papers on a specified medical imaging data set are collected and downloaded, and stored locally in the form of processable files such as PDF. Data preprocessing is performed on the collected relevant papers. Specifically, a large language model (neural network) is first used to perform a rough extraction on each paper, and the title, abstract, method, and conclusion of the paper are summarized to obtain Summary (S), and the method part of the original paper is extracted separately to obtain Methods (M), which will include the specific implementation steps of the module in the paper. Methods (M) is combined with Summary (S) as auxiliary data to obtain a processed paper collection.
[0053] In the embodiment of the present invention, in S11, the expression of the knowledge item set P is:
[0054] P = {p1, p2, p3, ..., p i ,…,p n};
[0055] Connect(S i , M i )→p i ;
[0056] In the formula, p1, p2, p3, p i and p n represents the knowledge items of the 1st, 2nd, 3rd, i and nth medical imaging articles respectively, S i represents the summary content of the i-th medical imaging article, M i It represents the method content of the i-th medical imaging article, and Connect(·) means connecting the summary content with the method content.
[0057] In the embodiment of the present invention, in S13, the expression of the vector database D is:
[0058] D = {d1, d2, d3, ..., d k};
[0059]
[0060] In the formula, d1, d2, d3, d k and d j Respectively represent the 1st, 2nd, 3rd, kth and jth knowledge bases of the vector database, v j1 、v j2 、v j3 and Represent the 1st, 2nd, 3rd and Nth in the jth knowledge base respectively j knowledge items, k represents the total number of knowledge bases, N j Indicates the total number of knowledge items contained in the knowledge base.
[0061] In this embodiment of the present invention, S2 includes the following sub-steps:
[0062] S21, using each knowledge base of the vector database to perform semantic matching on the initial query statement to generate a semantic matching result;
[0063] S22, generating a number of initial results as candidate sets according to the semantic matching results;
[0064] S23. Optimize each initial result of the candidate set to generate several latest results.
[0065] In the embodiment of the present invention, in S21, the semantic matching result R i The expression of (Q′) is:
[0066] R i (Q′) = {r1, r2, ..., r n};
[0067] Q′=OptimizePrompt(Q);
[0068] In the formula, r1 represents the first result of semantic matching in the knowledge base, r2 represents the second result of semantic matching in the knowledge base, and r n represents the nth result of semantic matching in the knowledge base, Q represents the initial query statement, OptimizePrompt(·) represents the combination of the initial query statement and the pre-prepared prompt, and Q′ represents the result obtained after feature engineering optimization.
[0069] In the embodiment of the present invention, in S22, the jth initial result S j The expression is:
[0070]
[0071] In the formula, LLM j represents the j-th large language model, R jrepresents the matching result between the j-th large language model and the knowledge base, m represents the number of groups of matching results between the large language model and the knowledge base, and GenerateInitialScheme(·) represents the initial scheme generated by the j-th large language model according to the matching result of the corresponding knowledge base.
[0072] In the embodiment of the present invention, in S23, the kth latest result S′ k The expression is:
[0073]
[0074] In the formula, SelectAndBuildScheme(·) indicates that the k-th large language model selects a suitable module from the M set and rebuilds a new scheme, M represents the candidate set composed of the initial scheme set, m represents the number of groups of matching results between the large language model and the knowledge base, and LLM k Represents the k-th large language model.
[0075] In the embodiment of the present invention, the optimal solution S * The expression is:
[0076]
[0077] In the formula, S′ k represents the kth latest result, Score(·) represents the score of the kth latest result, argmax represents the corresponding choice of the solution with the highest score, and S represents the specific solution with the highest score.
[0078] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A method for generating a breast cancer segmentation architecture based on text processing, characterized in that: The following steps are involved: S1. Collect several medical imaging articles, process each medical imaging article, and generate a vector database; S2, using each knowledge base of the vector database to process the initial query statement and generate several latest results; S3. Use a neural network to select the optimal solution from the latest results, obtain the final neural network architecture, and use the final neural network architecture to process breast cancer data.
2. The method for generating a breast cancer segmentation architecture based on text processing according to claim 1, characterized in that: The S1 comprises the following sub-steps: S11, collect several medical imaging articles and generate a set of knowledge items; S12, cleaning the knowledge item set, and performing vectorization operation on the cleaned knowledge item set using the embedding model; S13. Store the knowledge item set after the vectorization operation into the vector database.
3. The method for generating a breast cancer segmentation architecture based on text processing according to claim 2, characterized in that: In S11, the expression of the knowledge item set P is: P={p1,p2,p3,…,p i ,...,p n }; Connect(S i ,M i )→p i ; In the formula, p1, p2, p3, p i and p n represents the knowledge items of the 1st, 2nd, 3rd, i and nth medical imaging articles respectively, S i represents the summary content of the i-th medical imaging article, M i It represents the method content of the i-th medical imaging article, and Connect(·) means connecting the summary content with the method content.
4. The method for generating a breast cancer segmentation architecture based on text processing according to claim 2, characterized in that: In S13, the expression of the vector database D is: <h2 style=";text-align:left;direction:ltr">D = {d1, d2, d3,…, d<h2 style=";text-align:left;direction:ltr"> k <h2 style=";text-align:left;direction:ltr">}; In the formula, d1, d2, d3, d k and d j Respectively represent the 1st, 2nd, 3rd, kth and jth knowledge bases of the vector database, v j1 、v j2 、v j3 and Represent the 1st, 2nd, 3rd and Nth in the jth knowledge base respectively j knowledge items, k represents the total number of knowledge bases, N j Indicates the total number of knowledge items contained in the knowledge base.
5. The method for generating a breast cancer segmentation architecture based on text processing according to claim 1, characterized in that: The S2 comprises the following sub-steps: S21, using each knowledge base of the vector database to perform semantic matching on the initial query statement to generate a semantic matching result; S22. Generate several initial results as candidate sets according to the semantic matching results; S23. Optimize each initial result of the candidate set to generate several latest results.
6. The method for generating a breast cancer segmentation architecture based on text processing according to claim 5, characterized in that: In S21, the semantic matching result R i The expression of (Q′) is: R i (Q′)={r1,r2,…,r n }; Q′=OptimizePrompt(Q); In the formula, r1 represents the first result of semantic matching in the knowledge base, r2 represents the second result of semantic matching in the knowledge base, and r n represents the nth result of semantic matching in the knowledge base, Q represents the initial query statement, OptimizePrompt(·) represents the combination of the initial query statement and the pre-prepared prompt, and Q′ represents the result obtained after feature engineering optimization.
7. The method for generating a breast cancer segmentation architecture based on text processing according to claim 5, characterized in that: In S22, the jth initial result S j The expression is: In the formula, LLM j represents the j-th large language model, R j represents the matching result between the j-th large language model and the knowledge base, m represents the number of groups of matching results between the large language model and the knowledge base, and GenerateInitialScheme(·) represents the initial scheme generated by the j-th large language model according to the matching result of the corresponding knowledge base.
8. The method for generating a breast cancer segmentation architecture based on text processing according to claim 5, characterized in that: In S23, the kth latest result S′ k The expression is: In the formula, SelectAndBuildScheme(·) indicates that the k-th large language model selects a suitable module from the M set and rebuilds a new scheme, M represents the candidate set composed of the initial scheme set, m represents the number of groups of matching results between the large language model and the knowledge base, and LLM k Represents the kth large language model.
9. The method for generating a breast cancer segmentation architecture based on text processing according to claim 1, characterized in that: The expression of the optimal solution is: In the formula, S * represents the optimal solution; S′ k represents the kth latest result, Score(·) represents the score of the kth latest result, argmax represents the corresponding choice of the solution with the highest score, and S represents the specific solution with the highest score.
Citation Information
Patent Citations
Tumor image auxiliary interpretation system based on radiomics knowledge base
CN117594200A
Label calculation method based on hybrid expert LLM
CN118760736A
Knowledge discovery system training method and device, electronic equipment and storage medium
CN118820430A
Network architecture searching method and device, equipment and storage medium
CN118839720A
Method and system for multi-level artificial intelligence supercomputer design
US12001462B1