Oral medical image report generation method based on large model
By constructing an image report instruction dataset and optimizing a large model using a multi-dimensional reward function, the problem of report inaccuracy caused by template fixation in existing technologies is solved, achieving efficient and accurate oral image report generation and improving the model's generation capability and the accuracy of diagnostic information.
Patent Information
- Application Number
- CN202510997888.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-20
- Publication Date
- 2025-10-31
AI Technical Summary
When faced with complex oral imaging conditions or the diagnosis of rare diseases, the fixed templates of existing technologies may affect the accuracy and completeness of reports. They are difficult to handle comprehensive cases with multiple coexisting diseases, lack flexibility, and cannot meet the needs of the oral medical field for efficient, accurate, and intelligent image report generation.
A dataset of oral medical image report instructions was constructed. Model parameters were calculated using a scoring function based on contribution, importance, and influence on the final layer. Some layers were frozen and fine-tuned. The large model was optimized by combining the reinforcement learning algorithm GRPO. A multidimensional reward function was designed to evaluate the accuracy and structure of the generated reports.
It significantly improves the model's ability to generate structured and professional reports, enhances the accuracy of key diagnostic information and the clinical logic of personalized reports, reduces reliance on manually designed rules, and lowers maintenance and usage costs.
Smart Images

Figure CN120878024A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oral medicine technology, specifically to a method for generating oral medical image reports based on a large model. Background Technology
[0002] Automatic generation of oral medical imaging reports aims to utilize computer technology to create text reports based on medical images, thereby reducing the workload of doctors and improving diagnostic efficiency and accuracy. Existing solutions are template-based, with the core being the pre-design of imaging report templates for various oral diseases, marking relevant fields, and filling in corresponding information in the appropriate fields according to specific circumstances. These templates contain fixed text frames and placeholders for filling in imaging examination results and diagnostic information, while also supporting dynamic adjustment of the template structure to adapt to different clinical scenarios. When generating a report, the system first analyzes oral images using image recognition technology to extract key information such as the number, location, and shape of teeth, periodontal condition, and jawbone status, and cross-validates this information with patient history data and clinical guidelines. Subsequently, based on preset rules and adaptive logic, the extracted information is filled into the corresponding placeholders. For example, if a missing tooth is detected, the system will fill in the missing tooth number and other specific information in the placeholder corresponding to the tooth location in the template, while simultaneously triggering an association analysis module to assess secondary changes such as tilting of adjacent teeth or elongation of opposing teeth. The advantage of this method is that it ensures the standardization and uniformity of the report structure, and can quickly generate reports in batches with a uniform format, thus saving doctors' time and effort. Its disadvantages are that when faced with complex oral imaging situations or the diagnosis of rare diseases, the fixed nature of the template may affect the accuracy and completeness of the report, and it is difficult to handle comprehensive cases with multiple coexisting diseases, lacking flexibility and thus failing to meet the needs of the oral healthcare field for efficient, accurate, and intelligent image report generation. Summary of the Invention
[0003] In view of the above-mentioned shortcomings in the prior art, the oral medical image report generation method based on large models provided by the present invention solves the problem that the fixed nature of the template may affect the accuracy and completeness of the report when dealing with complex oral imaging conditions or the diagnosis of rare diseases.
[0004] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A method for generating oral medical image reports based on a large model is provided, which includes the following steps: S1. Construct an oral medical image report instruction dataset. The input data in the instruction dataset consists of the patient's medical records and oral image keywords, and the output data consists of the image report written by the doctor. S2. Based on the oral medical image report instruction dataset, calculate the comprehensive score of each layer of the pre-trained model parameters using a scoring function based on contribution, importance, and the degree of influence on the final layer. S3. Select the model parameters of the top K% layers with the highest comprehensive scores, freeze the model parameters of the remaining layers, and fine-tune the pre-trained model with frozen model parameters using the oral medical image report instruction dataset. S4. Input the acquired patient medical records and oral imaging keywords into the fine-tuned large model to generate the patient's imaging report.
[0005] Furthermore, the expression for the scoring function based on contribution, importance, and degree of influence on the final layer is as follows: in, This is a comprehensive score for the parameters of the l-th layer model; All are hyperparameters; This is a min-max normalization operation; for traces; For the first Layer model parameters The degree of contribution; For the first Layer model parameters The importance of; For the first Layer model parameters The degree of influence on the final layer; Contribution The expression is: in, This is a dataset of oral medical image report instructions, where x represents the input data and y represents the output data. The set of all model parameters for the pre-trained model; For the first Layer model parameters; For the first Gradients of layer model parameters; For pre-trained models, in model parameters Given the input data Generate output data The probability; T is the transpose; This represents the expectation of all input-output pairs in the instruction dataset; importance The expression is: in, For the first Layer weight matrix; Let L be the L2 norm of the matrix; degree of impact The expression: in, The hidden layer representation is obtained by replacing the parameters of the l-th layer of the pre-trained model with Gaussian noise and performing forward propagation. The hidden layer representation is obtained by forward propagation when the parameters of the l-th layer of the pre-trained model have not been replaced.
[0006] Furthermore, when fine-tuning the large model, the goal is to minimize the loss function for the specialized task, by fine-tuning the selected set of trainable parameters; the expression for minimizing the loss function for the specialized task is: in, These are trainable parameters, i.e., the model parameters of the first K% layers selected in step S3; To minimize the loss function for specialized tasks; This is the length of the output sequence; For the first Output words at each position; For the first All output tokens before the given position.
[0007] Furthermore, between steps S3 and S4, the large model is fine-tuned using the GRPO reinforcement learning algorithm based on population relative policy optimization. The fine-tuned large model generates multiple responses based on the current input data, and the quality of these responses is evaluated by a reward function to obtain the corresponding reward for the current batch of responses; By normalizing the rewards, we can measure the quality of the generated results relative to the average level. in, This is an estimate of the merits of the i-th response; and These are the mean and standard deviation of the current batch of rewards, respectively; By optimizing the objective function of GRPO to increase the probability of generating high-reward responses and decrease the probability of generating low-reward responses, the fine-tuned large model is further optimized to obtain the final fine-tuned large model.
[0008] Furthermore, the expression for the objective function of GRPO is: in, For multiple responses generated; This is the length of the output sequence; For the first The first reply Each word element; This is the current input data; For the first All output words preceding the given word; This is the model that is currently being optimized; This is the large model after fine-tuning; For pre-trained models; For the first The first reply The advantage of the step; This is the KL divergence term.
[0009] Furthermore, the expression for the reward function is: in, A reward for replying; Rewards are given for the accuracy of the facts; Rewards are given for semantic consistency of content. Awarded for the simplicity of the report structure; y represents the set of correct clinical concepts; x represents the input data, and y represents the output data. Final fact accuracy reward The expression is: Where k represents all keywords in y that match the correct clinical concept; The positive basic reward for k; As a penalty amplification factor; This is the penalty coefficient for omissions; The i-th correct clinical concept in the set of correct clinical concepts; Content semantic consistency reward The expression is: in, In order to be in The set of clinical concepts correctly identified in the text; For the concept Importance weights; Cosine similarity; For text embedding models; Describe the concept in y Sentences; Standard descriptive sentences written for experts; Award for concise report structure The expression is: in, for The number of lexical units in the text; for The number of redundant words used; These are adjustable hyperparameters.
[0010] Furthermore, the image report in step S1 includes image description and image diagnosis results, wherein the oral image is an X-ray and / or CBCT image.
[0011] The beneficial effects of this invention are as follows: In order to align the language style of the large model with the professional requirements of oral medical imaging reports, this solution quantifies the correlation between each layer of the large model and the generation of oral medical imaging reports through multiple aspects such as contribution, importance and influence on the final layer, and performs targeted fine-tuning on the most relevant layers. While retaining the general language capability, it significantly improves the model's ability to generate structured and professional reports.
[0012] When fine-tuning the pre-trained model, this solution only requires previously written reports by doctors and a small number of diagnostic reward rules to optimize the large model, without relying on manually designing a large number of rules and templates, thus greatly reducing long-term maintenance and usage costs. By introducing a large model with massive knowledge, this solution can realize the complex correlation between image features and diseases, and generate more personalized reports that conform to clinical logic.
[0013] To enhance the accuracy of key diagnostic information, this solution employs a multi-dimensional reward reinforcement learning strategy. By defining a positive / negative keyword system and reward function (including assessments of image conceptual accuracy, semantic appropriateness of content, and report structural simplicity), combined with the GRPO algorithm, the model-generated reports maintain fluency while significantly improving the ability to accurately describe positive / negative findings. Attached Figure Description
[0014] Figure 1 This is a flowchart of a method for generating oral medical image reports based on a large model. Detailed Implementation
[0015] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0016] refer to Figure 1 , Figure 1A flowchart illustrating a method for generating oral medical image reports based on a large model is shown, such as... Figure 1 As shown, the method S includes steps S1 to S4.
[0017] In step S1, an oral medical image report instruction dataset is constructed. The input data in the instruction dataset consists of the patient's medical records and oral image keywords, and the output data consists of an image report written by the doctor. The image report includes an image description and an image diagnosis result. The oral images are X-ray films and / or CBCT images.
[0018] In step S2, based on the oral medical image report instruction dataset, a comprehensive score for the parameters of each layer of the pre-trained model is calculated using a scoring function based on contribution, importance, and the degree of influence on the final layer. In one embodiment of the present invention, the expression for the scoring function based on contribution, importance, and degree of influence on the final layer is: in, This is a comprehensive score for the parameters of the l-th layer model; All are hyperparameters; This is a min-max normalization operation; for traces; For the first Layer model parameters The degree of contribution; For the first Layer model parameters The importance of; For the first Layer model parameters The degree of influence on the final layer; Contribution The expression is: in, This is a dataset of oral medical image report instructions, where x represents the input data and y represents the output data. The set of all model parameters for the pre-trained model; For the first Layer model parameters; For the first Gradient of layer parameters; For pre-trained models, in model parameters Given the input data Generate output data The probability; T is the transpose; This solution is calculated. traces This is used to obtain a scalar value that represents the average contribution of the parameters of this layer to the task loss function. The larger the value, the more drastic the change in the loss function can be caused by a small change in the parameters of that layer, and therefore the more important that layer is.
[0019] importance The expression is: in, For the first Layer weight matrix; The L2 norm of the matrix; importance As a static prior, it can help identify those layers that occupy a core position in pre-trained large models.
[0020] degree of impact The expression: in, The hidden layer representation is obtained by replacing the parameters of the l-th layer of the pre-trained model with Gaussian noise and performing forward propagation. The hidden layer representation obtained by forward propagation when the parameters of the l-th layer of the pre-trained model have not been replaced; This represents the expected value (average value) of all input-output pairs in the instruction dataset.
[0021] In step S3, the model parameters of the top K% layers with the highest comprehensive scores are selected, the model parameters of the remaining layers are frozen, and the pre-trained model with frozen model parameters is fine-tuned using the oral medical image report instruction dataset.
[0022] In implementation, this scheme preferably minimizes the loss function for specialized tasks when fine-tuning large models, by selecting a set of trainable parameters for fine-tuning; the expression for minimizing the loss function for specialized tasks is: in, These are trainable parameters, i.e., the model parameters of the first K% layers selected in step S3; To minimize the loss function for specialized tasks; This is the length of the output sequence; For the first Output words at each position; For the first All output tokens before the given position.
[0023] By optimizing this objective, this targeted fine-tuning strategy can significantly improve the basic ability of large models to generate structured and professional oral imaging reports. At the same time, since most of the layers carrying general knowledge are not changed, the generalization performance of the model can be preserved to the greatest extent.
[0024] In step S4, the acquired patient medical records and oral imaging keywords are input into the finely tuned large model to generate the patient's imaging report.
[0025] This approach uses a dataset consisting of input and output data forming "instruction-output" pairs to fine-tune the pre-trained model. This improves the model's ability to understand specific instructions in oral imaging descriptions and generate imaging report texts that conform to professional standards, have accurate language style, and contain necessary diagnostic information, while maintaining its original general language understanding and generation capabilities.
[0026] While the fine-tuned large model shows improvements in language style and basic content generation, it may still exhibit instability or omissions in the accuracy of key diagnostic information, particularly in the precise description of positive / negative findings for specific diseases. To enhance the stability and comprehensiveness of the key diagnostic information output by the large model, this scheme includes optimizing the fine-tuned large model using the Group Relative Policy Optimization (GRPO) reinforcement learning algorithm between steps S3 and S4. GRPO improves the quality of the large model's output by evaluating the reward of a set of outputs relative to other outputs and encouraging the large model to generate high-reward outputs.
[0027] In one embodiment of the present invention, the detailed implementation steps of the large model after optimization and fine-tuning using the reinforcement learning algorithm GRPO based on population relative policy optimization include: The fine-tuned large model generates multiple responses based on the current input data, and the quality of these responses is evaluated by a reward function to obtain the corresponding reward for the current batch of responses; By normalizing the rewards, we can measure the quality of the generated results relative to the average level. in, This is an estimate of the merits of the i-th response; and These are the mean and standard deviation of the current batch of rewards, respectively; By optimizing the objective function of GRPO to increase the probability of generating high-reward responses and decrease the probability of generating low-reward responses, the fine-tuned large model is further optimized to obtain the final fine-tuned large model.
[0028] In implementation, the preferred expression for the objective function of GRPO in this scheme is: in, For multiple responses generated; This is the length of the output sequence; For the first The first reply Each word element; This is the current input data; For the first All output words preceding the given word; This is the model that is currently being optimized; This is the large model after fine-tuning; For pre-trained models; For the first The first reply The advantage of the step; This is the KL divergence term.
[0029] To effectively leverage reinforcement learning to optimize the oral medical image report generation capabilities of large models, this scheme designs a novel and targeted reward function. This function aims to evaluate the generated output book from three dimensions: the accuracy of image concepts, the semantic consistency of content, and the simplicity of the report structure. (Based on input data) A comprehensive assessment will be conducted. The assessment criteria will incorporate structured real-world annotations from oral imaging. That is, the set of correct clinical concepts. Pre-built by senior dental radiology experts, it contains a set of all clinical concepts that need to be reported in the images. Each of the concepts Each of these is associated with a set of positive diagnostic keywords and corresponding expert standard descriptive sentences.
[0030] In one embodiment of the present invention, the expression of the reward function is: in, A reward for replying; Rewards are given for the accuracy of the facts; Rewards are given for semantic consistency of content. Awarded for the simplicity of the report structure; y represents the set of correct clinical concepts; x represents the input data, and y represents the output data. Final fact accuracy reward The expression is: Where k represents all keywords in y that match the correct clinical concept; The positive basic reward for k; As a penalty amplification factor; This is the penalty coefficient for omissions; It is the i-th correct clinical concept in the set of correct clinical concepts.
[0031] Ultimately, the reward for accuracy of fact is central to the reward function, designed to evaluate a report's description of key facts in a clear and rewarding manner. It first identifies... All of the correct clinical concepts Matching keywords For each match, a positive basic reward is given based on the diagnostic importance. Next, identification. It is mentioned but does not exist in the correct set of clinical concepts. The pathological keywords in the text. Each occurrence of such a hallucination is accompanied by a significant negative punishment. (in It is a penalty amplification factor to strongly suppress the model from fabricating facts. Furthermore, it examines the set of true concepts. What are the necessary concepts in The middle part was completely omitted. For each omitted concept... Each of them is subject to a negative penalty based on the weight of its most important keyword. ( (For omission penalty coefficients), to encourage the model to generate comprehensive reports.
[0032] Content semantic consistency reward The expression is: in, In order to be in The set of clinical concepts correctly identified in the text; For the concept Importance weights; Cosine similarity; For text embedding models; Describe the concept in y Sentences; Standard descriptive sentences written for experts; Award for concise report structure The expression is: in, for The number of lexical units in the text; for The number of redundant words used; These are adjustable hyperparameters.
[0033] This approach utilizes the GRPO reinforcement learning method based on a multi-dimensional reward function. The large model learns to generate report texts that are not only fluent in language but also contain key diagnostic information and have accurate semantics, thereby significantly improving its accuracy, reliability, and clinical applicability in oral medical image report generation tasks.
Claims
1. A method for generating oral medical image reports based on a large model, characterized in that, Including the following steps: S1. Construct an oral medical image report instruction dataset. The input data in the instruction dataset consists of the patient's medical records and oral image keywords, and the output data consists of the image report written by the doctor. S2. Based on the oral medical image report instruction dataset, calculate the comprehensive score of each layer of the pre-trained model parameters using a scoring function based on contribution, importance, and the degree of influence on the final layer. S3. Select the model parameters of the top K% layers with the highest comprehensive scores, freeze the model parameters of the remaining layers, and fine-tune the pre-trained model with frozen model parameters using the oral medical image report instruction dataset. S4. Input the acquired patient medical records and oral imaging keywords into the fine-tuned large model to generate the patient's imaging report.
2. The method for generating oral medical image reports based on a large model according to claim 1, characterized in that, The expression for the scoring function based on contribution, importance, and influence on the final layer is as follows: in, This is a comprehensive score for the parameters of the l-th layer model; All are hyperparameters; This is a min-max normalization operation; for traces; For the first Layer model parameters The degree of contribution; For the first Layer model parameters The importance of; For the first Layer model parameters The degree of influence on the final layer; Contribution The expression is: in, This is a dataset of oral medical image report instructions, where x represents the input data and y represents the output data. The set of all model parameters for the pre-trained model; For the first Layer model parameters; For the first Gradients of layer model parameters; For pre-trained models in model parameters Given the input data Generate output data The probability; T is the transpose; This represents the expectation of all input-output pairs in the instruction dataset; importance The expression is: in, For the first Layer weight matrix; Let L be the L2 norm of the matrix; degree of impact The expression: in, The hidden layer representation is obtained by replacing the parameters of the l-th layer of the pre-trained model with Gaussian noise and performing forward propagation. The hidden layer representation is obtained by forward propagation when the parameters of the l-th layer of the pre-trained model have not been replaced.
3. The method for generating oral medical image reports based on a large model according to claim 1, characterized in that, When fine-tuning a large model, the goal is to minimize the loss function for a specific task, by selecting a set of trainable parameters for fine-tuning. The expression for minimizing the loss function for a specific task is: in, These are trainable parameters, i.e., the model parameters of the first K% layers selected in step S3; To minimize the loss function for specialized tasks; This is the length of the output sequence; For the first Output words at each position; For the first All output tokens before the given position.
4. The method for generating oral medical image reports based on a large model according to any one of claims 1-3, characterized in that, Between steps S3 and S4, the large model is further optimized and fine-tuned using the GRPO reinforcement learning algorithm based on population relative policy optimization. The fine-tuned large model generates multiple responses based on the current input data, and the quality of these responses is evaluated by a reward function to obtain the corresponding reward for the current batch of responses; By normalizing the rewards, we can measure the quality of the generated results relative to the average level. in, This is an estimate of the merits of the i-th response; and These are the mean and standard deviation of the current batch of rewards, respectively; By optimizing the objective function of GRPO to increase the probability of generating high-reward responses and decrease the probability of generating low-reward responses, the fine-tuned large model is further optimized to obtain the final fine-tuned large model.
5. The method for generating oral medical image reports based on a large model according to claim 4, characterized in that, The objective function of GRPO is expressed as follows: in, For multiple responses generated; This is the length of the output sequence; For the first The first reply Each word element; This is the current input data; For the first All output words preceding the given word; This is the model that is currently being optimized; This is the large model after fine-tuning; For pre-trained models; For the first The first reply Estimating the advantage of each step; This is the KL divergence term.
6. The method for generating oral medical image reports based on a large model according to claim 4, characterized in that, The expression for the reward function is: in, A reward for replying; Rewards are given for the accuracy of the facts; Rewards are given for semantic consistency of content. Awarded for the simplicity of the report structure; y represents the set of correct clinical concepts; x represents the input data, and y represents the output data. Final fact accuracy reward The expression is: Where k represents all keywords in y that match the correct clinical concept; The positive basic reward for keyword k; As a penalty amplification factor; This is the penalty coefficient for omissions; The i-th correct clinical concept in the set of correct clinical concepts; Content semantic consistency reward The expression is: in, In order to be in The set of clinical concepts correctly identified in the text; For the concept Importance weights; Cosine similarity; For text embedding models; Describe the concept in y Sentences; Standard descriptive sentences written for experts; Award for concise report structure The expression is: in, for The number of lexical units in the text; for The number of redundant words used; These are adjustable hyperparameters.
7. The method for generating oral medical image reports based on a large model according to any one of claims 1-3 and 5-6, characterized in that, The image report in step S1 includes image description and image diagnosis results, wherein the oral images are X-ray films and / or CBCT images.