Clinical research report automatic writing system and method based on artificial intelligence
Through data acquisition and relationship mapping, the large language model is optimized, and the problem of large-scale computing resource consumption and equipment data disconnection in the generation of clinical research reports is solved, and efficient and accurate automatic report writing is achieved.
Patent Information
- Application Number
- CN202510962042.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-14
AI Technical Summary
The existing clinical research report generation model consumes a lot of computing resources during the generation process, destroying the integrity of pre-training knowledge, lacking explicit constraints on adverse reaction characteristics, resulting in inaccurate reporting, and disconnection of equipment data affects diagnostic accuracy.
The data acquisition module extracts the reporting elements, establishes the relationship mapping module to obtain mapping relationships, uses a multi-condition encoder to fine-tune the large language model, builds a medical knowledge graph, strengthens the mapping of the lesion area and the internal structure of the equipment, and optimizes the large language model.
It realizes automated mapping from text to structured dimensions, improves the accuracy and efficiency of generating reports, significantly improves the sensitivity to adverse reaction characteristics and controllability of equipment status, and ensures the integrity of model knowledge.
Smart Images

Figure CN120473070A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to an artificial intelligence-based automatic writing system and method for clinical research reports. Background Art
[0002] A clinical study report (CSR) is a comprehensive regulatory report that describes the data and results observed in a clinical study. It is generally required to be written at the end of the study, but may also be generated at other times during the study. Automatic report generation is currently a widely used technology. However, in the field of clinical medicine, most clinical study report generation models initially consider only medical images or medical reports. This leads to a poor match between the feature information of medical images and medical reports, resulting in insufficient extraction of key semantics and reduced model robustness. In recent years, with the rapid development of artificial intelligence technology, especially the successful application of large language models in various natural language processing tasks, automatic report writing has also gained new development opportunities. Large language models can automatically generate corresponding reports based on natural language expressions. In the process of generating clinical research reports, large language models are usually trained and cross-modal attention mechanisms are proposed to fine-tune all parameters. This method consumes a lot of computing resources and easily destroys the knowledge integrity of the pre-trained large language model. It lacks an explicit constraint mechanism for adverse reaction characteristics, has a high false positive rate, and ultimately leads to inaccurate reports. In addition, the generation of modern clinical medical reports is highly dependent on the test data output by medical equipment, including imaging equipment (CT, MRI), in vitro diagnostic equipment (biochemical analyzers, hematology analyzers) and vital signs monitoring equipment (ECG monitors, ventilators), etc. These devices should theoretically provide objective and real-time data support for the generation of clinical research reports; however, the existing report generation process has a significant problem of device data disconnection, which leads to impaired diagnostic accuracy, thereby affecting the accuracy of clinical research reports and causing medical errors to a certain extent. Summary of the Invention
[0003] (1) Technical problems solved In response to the deficiencies of the prior art, the present invention provides an artificial intelligence-based automatic writing system and method for clinical research reports, which solves the problems raised in the background technology.
[0004] (2) Technical solution To achieve the above objectives, the present invention is implemented through the following technical solutions: In the first aspect, the present application provides an artificial intelligence-based automatic writing system for clinical research reports, comprising: The data collection module collects clinical research data and device profile data uploaded by users and extracts report elements; The relationship mapping module, based on the report elements, obtains the entities, entity types, and corresponding parameter points that have a mapping relationship between the clinical research process and the report elements, obtains the adverse reaction types and lesion areas that have a mapping relationship between the clinical research process and the report elements, and obtains the internal structure of the research equipment that has a mapping relationship with the report elements from the clinical research process; The training fine-tuning module sets a multi-condition encoder based on the mapping relationship, constructs corresponding fine-tuning instructions, performs multi-condition fine-tuning on the large language model, and trains and updates it to obtain an optimized large language model.
[0005] Furthermore, the steps to extract report elements include: Extract research requirements from user-uploaded clinical research data, including research literature, research topics, research plans, research progress, target diseases, and target conditions; Extract equipment requirements from equipment profile data, including equipment aging requirements and equipment performance requirements; Combine study requirements and equipment requirements to generate report elements.
[0006] Furthermore, obtaining the mapping relationship also includes: Analyze report elements, including at least data cleaning and data conversion; Identify the types of studies to be uploaded, use the BioBERT model to extract disease conditions, symptoms, and drug entities, and obtain corresponding parameter points. Use the rule engine to match dimensions, execute a quantitative evaluation strategy, obtain the corresponding evaluation feature set and training sample size, and use a preset large language model for pre-training. The evaluation dataset includes matching efficiency, matching accuracy, and fusion coverage. Taking each entity as a node, retrieve each entity type and construct a medical knowledge graph that includes target disease, target symptom, target drug and research type. Based on the medical knowledge graph, obtain the parameter point corresponding to any entity, standardize the corresponding value of the parameter point into a relative position index, draw a parameter change curve and superimpose the dynamic standard range. When at least two parameter points are detected to be outside the dynamic standard range, a warning signal is generated and injected into the medical knowledge graph as a dynamic attribute. Obtaining the weight of the mapping relationship between the lesion area and the report element; Obtain the over-limit thresholds that have a mapping relationship between the internal structure of the research equipment and the report elements.
[0007] Furthermore, a quantitative evaluation strategy is implemented based on the following formula: ; In the formula, N represents the number of training samples, D typerepresents the research type, including parallel design, factorial design, crossover design and mixed design, φ(·) represents the nonlinear weighted result of the characteristic index of the corresponding research type, and the characteristic index includes at least one of the balance characteristic, interaction intensity characteristic and individual difference characteristic, h represents the standardization factor, β type represents the weight coefficient corresponding to the characteristic index, Es represents the basic sample size, and λ represents the adjustment factor.
[0008] Furthermore, the steps of constructing fine-tuning instructions include: First fine-tuning: Calculate the value index based on the evaluation feature set, compare and analyze the value index with the preset standard value range [jz1, jz2], automatically match the template paragraph, and call the NLP explanatory paragraph generator under the conditions of triggering the paragraph basic template and paragraph standard template, and automatically annotate; Secondary fine-tuning: Establish a joint embedding space, including a visual encoder and a text encoder. Input the weights of the mapping relationship between the lesion area and the report elements into the joint embedding space, set the scene feature matrix for the lesion area and retain the alignment loss. Establish a conditional encoder, input the weights and scene feature matrix into the conditional encoder, and enforce constraints when the adverse reaction type is detected. Three fine-tuning steps: Obtain the timestamp that exceeds the threshold, randomly extract several moments before and after the corresponding time stamp to form a time series, and the time series is a dynamic change quantity; extract the average and fluctuation values of the corresponding values of the associated parameter points in the time series, and take a weighted sum to calculate the risk level; then retrieve the medical knowledge graph to determine whether the numerical results corresponding to the parameter points are correct; Content is identified and categorized based on different risk levels, and the risk levels are divided into three levels. The numerical value corresponding to each risk level is data-bound to the generated report content label, and Internet big data is provided for deep learning of risk labels. For parameter points below level two, the corresponding calibration value of the parameter point is provided. Otherwise, the value corresponding to the parameter point is marked as invalid, and an equipment maintenance instruction is generated.
[0009] Furthermore, the value index is compared with the preset standard value range [jz1, jz2]: Mark the value index as value; When value < jz1, match the paragraph basic template; When jz1≤value<jz2, match the paragraph standard template; When value ≥ jz2, matches the paragraph top-level template.
[0010] Furthermore, the second fine-tuning step includes: Establish a joint embedding space, use the visual encoder to extract the image features of the lesion area, use the text encoder to extract the text features, and map the features corresponding to the lesion area and report elements, as well as the weights of the mapping relationship between the lesion area and the report elements, into a unified space; The image features and text features of each lesion area are concatenated into a joint feature vector, a scene feature matrix is set for the lesion area, and the alignment loss is retained; When an adverse reaction type is detected, the constraint is enforced: LOA = max(LOA, 1.5); Where LOA represents the alignment loss.
[0011] Furthermore, the step of extracting the associated parameter points in the time sequence includes: when establishing the association relationship between the time sequence and the parameter points, first providing a judgment criterion for the numerical value corresponding to each parameter point, and then extracting the corresponding trigger relationship between each judgment criterion and the report element.
[0012] In a second aspect, the present application provides an artificial intelligence-based method for automatically writing clinical research reports, comprising the following steps: receiving an instruction to generate a report; Determine the device profile data and clinical research data that generate report instructions, and extract report elements from them; Based on the report elements, entities, entity types, and corresponding parameter points that have a mapping relationship between the clinical research process and the report elements are obtained, as well as adverse reaction types and lesion areas that have a mapping relationship between the clinical research process and the report elements. The internal structure of the research equipment that has a mapping relationship with the report elements is also obtained from the clinical research process; Based on the mapping relationship, a multi-condition encoder is set up, corresponding fine-tuning instructions are constructed, multi-condition fine-tuning is performed on the large language model, and training and updating are performed to obtain an optimized large language model.
[0013] (3) Beneficial effects The present invention provides an artificial intelligence-based automatic writing system and method for clinical research reports, which has the following beneficial effects: 1. The present invention obtains entities, entity types, and corresponding parameter points that have a mapping relationship between the clinical research process and report elements, and obtains adverse reaction types and lesion areas that have a mapping relationship between the clinical research process and report elements. It also obtains the internal structure of the research equipment that has a mapping relationship with the report elements from the clinical research process. This not only realizes the automated mapping from text to structured dimensions, but also continuously improves the matching accuracy through dynamic learning and multimodal fusion, forming an intelligent analysis engine with clinical decision-making value. By establishing a medical knowledge graph, extracting parameter points corresponding to entities, drawing parameter change curves and superimposing dynamic standard intervals, and dynamically updating the medical knowledge graph, it can realize intelligent and standardized monitoring of clinical research quality and significantly improve the efficiency and accuracy of abnormality detection. 2. The present invention first pre-trains a large language model, and based on the mapping relationship, sets a multi-condition encoder, constructs corresponding fine-tuning instructions, fine-tunes specific parameters, and optimizes the large language model, which not only ensures the knowledge integrity of the large language model but also improves computational efficiency; in the first fine-tuning, an evaluation feature set is constructed based on matching efficiency, matching accuracy, and fusion coverage, and a value index is calculated to automatically match template paragraphs; in the second fine-tuning, a joint embedding space is established, a scene feature matrix is set for the lesion area, and the alignment loss is retained to strengthen the lesion, facilitate the later extraction of local features that are crucial for diagnosis, significantly improve the sensitivity to adverse reaction characteristics, and have a deeper understanding of the semantics of generated reports; in the third fine-tuning, the internal structure of the device and the over-limit threshold are integrated into the large language model, solving the long-standing problem of invisible and uncontrollable impact of device status in clinical research. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a module diagram of a system for automatically writing clinical research reports according to an exemplary embodiment; Figure 2 The figure is a schematic diagram of the steps of a method for automatically writing a clinical research report according to an exemplary embodiment. DETAILED DESCRIPTION
[0015] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0016] Example 1 The embodiment of the present invention provides an artificial intelligence-based automatic writing system for clinical research reports; Figure 1is a module diagram of a system for automatically writing clinical research reports according to an exemplary embodiment; Figure 1 The system includes: a data acquisition module, a relationship mapping module and a training and fine-tuning module, and the data acquisition module, the relationship mapping module and the training and fine-tuning module are communicatively connected; Description of the functional modules of this system: Data collection module: collects clinical research data and device profile data uploaded by users and extracts report elements from them; The steps to extract report elements include: Extract research requirements from user-uploaded clinical research data, including research literature, research topics, research plans, research progress, diseases, conditions, and drugs; Extract equipment requirements from equipment portrait data, including equipment aging requirements and equipment performance requirements; Device profile data, which is a summary of device status labels, can also be replaced by other labeling technologies. For example, using big data analysis to generate device profiles ensures that device selection is more closely aligned with the actual needs of clinical research, enabling the automated generation of clinical research reports throughout the entire process, significantly improving research efficiency and report quality while meeting medical data compliance requirements. Combine study requirements with equipment requirements to obtain report elements; The following are explanations of the relevant nouns involved: Research literature: refers to academic resources directly related to clinical research, including published clinical trial reports, case studies, diagnosis and treatment guidelines, meta-analysis, etc.; Research topics: research topics proposed around clinical problems (such as disease diagnosis and treatment, disease management, etc.), which must clearly define intervention methods and clinical endpoints; Research protocols: operational documents that standardize the implementation of clinical research and must comply with ethical review and regulatory requirements (such as ICH-GCP standards); Research progress: a summary of interim achievements and challenges during the implementation of clinical research, used to adjust the trial protocol or meet regulatory reporting requirements; Diseases: the specific disease entity targeted in the study, which must clearly define its clinical classification (for example, in breast molecular, Luminal A type, Luminal B type, etc.), stage (such as: stage 0, stage I, stage II, stage III and stage IV) and treatment status; Symptoms: the clinical manifestations targeted in the study, usually used as primary / secondary endpoints or safety indicators; Drugs: the treatment indicators configured in the study based on the clinical manifestations; Equipment aging requirements: The management standards set for the performance degradation of equipment due to long-term use, environmental exposure, or natural wear and tear in clinical studies. Equipment performance requirements: The technical specifications and functional standards that equipment must meet in clinical studies to ensure that it can accurately and stably perform the operations or measurements required for the study. The above data can also be combined with crawler technology to obtain more relevant data from public network platforms. After screening and cleaning, the data dimensions are enriched, and the breadth and depth of data collection are further improved. At the same time, combined with the user's device usage habits, the proficiency and correctness of the device usage are judged. Through mutual combination, the statistical device status labels will be more accurate, thereby providing users with more professional report automatic writing solutions. This system improves resource utilization and the efficiency of automatic report writing by optimizing the user's acquisition method of report elements and the accuracy of demand; The relationship mapping module obtains entities, entity types, and corresponding parameter points that have a mapping relationship between the clinical research process and the report elements based on the report elements, and obtains adverse reaction types and lesion areas that have a mapping relationship between the clinical research process and the report elements. It also obtains the internal structure of the research equipment that has a mapping relationship with the report elements from the clinical research process; It should be noted that a mapping relationship refers to establishing an association between the two. Once one of them is referenced, the other one will also change accordingly, or be referenced accordingly. By obtaining entities and entity types, a mapping relationship is established between indicator points and report elements in the clinical research process. By obtaining adverse reaction types, a mapping relationship is established between lesion areas and report elements in the clinical research process. A mapping relationship is also established between the internal structure of the research equipment and report elements. The acquisition of the mapping relationship also includes: Analyze report elements, including at least data cleaning and data conversion; Data cleaning: missing data processing, outlier detection, duplicate data processing, data correction, used to ensure the accuracy and reliability of data. It eliminates noise, errors and inconsistencies in the data, so that subsequent analysis can draw reliable conclusions. Data transformation: variable conversion, data aggregation and grouping, deriving new variables, time series processing, used to convert raw data into a form more suitable for analysis, or generate new variables for further analysis; Identify the types of studies to be uploaded, use the BioBERT model to extract disease conditions, symptoms, and drug entities, and obtain corresponding parameter points. Use the rule engine to match dimensions, execute a quantitative evaluation strategy, obtain the corresponding evaluation feature set and training sample size, and use a preset large language model for pre-training. The evaluation dataset includes matching efficiency, matching accuracy, and fusion coverage. The steps for implementing the quantitative evaluation strategy include: Setting formula: ; In the formula, N represents the number of training samples, Indicates the study type, including parallel design, factorial design, crossover design, and mixed design. It represents the nonlinear weighted result of the characteristic index of the corresponding research type, which is used to quantify the characteristic index of the research type. The characteristic index includes at least one characteristic index obtained by quantifying the balance characteristic, the interaction intensity characteristic, and the individual difference characteristic. h represents the standardization factor, and its value is greater than 0, which is used to control the quantification of different characteristic indices. represents the weight coefficient corresponding to the characteristic index, and Greater than 0, e represents the natural base, represents the basic sample size, It represents the adjustment factor, and its value range is [0.01, 0.1], which is used to control the rationality of the research type combination; Specifically, if the parallel design is marked as A, the factorial design is marked as B, and the crossover design is marked as C, there are 7 research type combinations: [A], [B], [C], [AB], [AC], [BC], and [ABC]; When there is only a single research type: [A] corresponds to the balance feature, [B] corresponds to the interaction intensity feature, and [C] corresponds to the individual difference feature; when there are two research types, [AB] corresponds to the balance feature and the interaction intensity feature, [AC] corresponds to the balance feature and the individual difference feature, and [BC] corresponds to the interaction intensity feature and the individual difference feature; when there are three research types [ABC], they correspond to the balance feature, the interaction intensity feature, and the individual difference feature; For example: If it is a mixed design, the corresponding It can be decomposed into h1 (balanced feature) * h2 (interaction strength feature) * h3 (individual difference feature), and h1, h2 and h3 represent normalization factors; Assuming the calculated φ is 1.8 and Es is 50, is 0.05; Then the sample size = 1.8*e 0.05*50 =22, the sample size is adjusted to 22; By inputting the characteristics of each study type corresponding to each group into the above formula model as monitoring vectors, the system automatically determines the sample size through the combination of balance characteristics, interaction strength characteristics, and individual difference characteristics, rather than relying solely on qualitative experience. This significantly improves the interpretability and practicality of large language models in clinical research. In addition, the parallel design: random grouping improves the balance between groups, and the corresponding training sample size is large; the factorial design: analyzes the interaction of multiple factors, and the corresponding training sample size is moderate; the crossover design: reduces the differences between individuals, and the corresponding training sample size is small; the balance feature is measured by calculating the standardized mean difference of the covariates and the entropy balance index, the interaction intensity feature uses the hierarchical causal forest model to estimate the interaction effect value between variables, and the individual difference feature is quantified by the random intercept-slope mixed model. The specific process is not described in detail here; Taking each entity as a node, retrieve each entity type and construct a medical knowledge graph including target disease - target symptom - target drug - research type; Based on the medical knowledge graph, the parameter point q corresponding to any entity is obtained, and the value corresponding to any parameter point q is marked as Lq. Based on the dynamic standard interval QZ[QZ1, QZ2] of the parameter point q, the value corresponding to the parameter point q is standardized as a relative position indicator. The parameter change curve is drawn and the dynamic standard interval is superimposed and displayed. When at least two parameter points are detected to be outside the dynamic standard interval, an early warning signal is generated and injected into the medical knowledge graph as a dynamic attribute; The dynamic standard interval QZ corresponds to the parameter point q one by one, and the dynamic standard interval is generated by at least one of the following methods: Linear adjustment based on research progress time: ; Where, represents the dynamic standard interval at a certain research time t, represents the upper and lower limits of the initial standard interval, k represents the time slope coefficient, which is expressed as a percentage, and t represents the research progress time; Statistical intervals based on sample size adaptation: ; Where, 、 Indicates the upper and lower limits corresponding to the current time point t. Represents the moving average corresponding to the historical parameter points, represents the standard deviation corresponding to the parameter point, Indicates the current cumulative sample size; Fuzzy logic intervals based on expert rules: ; Where QZ represents the dynamic standard interval, It represents the union symbol, which is used to control multiple rules of fuzzy logic and means merging the results of multiple rules. n1 represents the number of rules. represents the mth expert rule, represents the rule weight, Represents the implication relationship in fuzzy logic, that is, if the condition is met , then use weight ; Specifically, because report elements are usually multimodal data, the CRF-BERT hybrid model is used for medical entity recognition and relationship extraction, and weakly supervised learning is used to probabilistically fill in missing label parameters. Different subgroups have different standard intervals for corresponding parameter points, and each parameter point is associated with a corresponding numerical value and report position. By drawing parameter change curves, the severity corresponding to the parameter points can be analyzed, and the medical knowledge graph can be updated, which can realize intelligent and standardized monitoring of clinical research quality and significantly improve the efficiency and accuracy of anomaly detection. The system not only realizes automatic mapping from text to structured dimensions, but also continuously improves matching accuracy through dynamic learning and multimodal fusion, forming an intelligent analysis engine with clinical decision-making value. In addition, in actual implementation, it is necessary to configure the parameter rule base in combination with the specific research type and deeply integrate it with existing clinical data platforms (such as Medidata Rave). Obtaining the weight of the mapping relationship between the lesion area and the report element; Obtaining the exceeding threshold value with mapping relationship between the internal structure of the research equipment and the report elements; The training fine-tuning module sets a multi-condition encoder based on the mapping relationship, constructs corresponding fine-tuning instructions, performs multi-condition fine-tuning on the large language model, and trains and updates it to obtain an optimized large language model; The steps to build fine-tuning instructions include: One-time fine-tuning: Calculate the value index based on the evaluation feature set: ; Where, represents the value index, Indicates the weight correction coefficient of the preset value index, e represents the natural base number, 、 、 They are matching efficiency, matching accuracy and fusion coverage, 、 、 They are matching efficiency threshold, matching accuracy threshold and fusion coverage threshold respectively. 、 、 are weight ratio coefficients, 、 、 are all greater than 0, and α1+α2+α3=1, represents the weight correction coefficient of the preset number of training samples, and greater than 0; Formula Explanation: The larger the values corresponding to matching efficiency, matching accuracy, and fusion coverage, the larger the calculated value index, indicating higher quality of the generated report output by the system and better training of the large language model; The value index Comparative analysis with the preset standard value range [jz1, jz2]: when <jz1, matches the paragraph basic template; When jz1≤ <jz2, matches the paragraph standard template; when When ≥jz2, it matches the paragraph top-level template; Specifically, the basic paragraph template is suitable for primary content generation, has low term density, and cannot control paragraph length, resulting in low relevance, which cannot meet the basic requirements of most clinical research reports. The standard paragraph template is suitable for structured report writing and has fixed segmentation, but it cannot provide detailed descriptions for more complex clinical studies (such as research images), resulting in weak relevance. The top-level paragraph template has medium to high term density, mandatory inclusion of professional terms, and adaptive adjustment of the number of evidence points, resulting in strong relevance. Under the conditions of triggering the basic paragraph template and the standard paragraph template, since the corresponding templates show a general degree of correlation, NLP is called for intelligent interpretation and multi-dimensional automatic annotation (for example: data traceability, confidence, guideline basis, data calibration) to assist in the fully automated generation of clinical reports; the automatic annotation content includes annotation type, implementation method and example value, for example:
[0017] Through interval judgment and dynamic template matching, a context-aware progressive architecture is implemented, and content of different depths is dynamically generated to improve the correlation strength between semantics in report writing. Secondary fine-tuning: Establish a joint embedding space, including a visual encoder and a text encoder, input the weights of the mapping relationship between the lesion area and the report elements into the joint embedding space, set the scene feature matrix for the lesion area and retain the alignment loss; establish a conditional encoder, input the weights and scene feature matrix into the conditional encoder, enforce constraints when the adverse reaction type is detected, and fine-tune the weights with low rank; Establish a joint embedding space, use the visual encoder to extract the image features of the lesion area, use the text encoder to extract the text features, and map the features corresponding to the lesion area and report elements, as well as the weights of the mapping relationship between the lesion area and the report elements, into a unified space; Specifically, in the visual encoder, Mask R-CNN is used to enhance the lesion area; in the text encoder, a bidirectional long short-term memory (LSTM) network is used to train short text interpretation. When the text length is 10-50 words, the number of memory units is set to 128-256; when the text length is 50-100 words, the number of memory units is set to 256-512. When the text involves medical diagnosis, the number of memory units is set to 512, and the time window sliding step is 15 words. The (Transformer) network is used to train long text interpretation, and the number of memory units is set to 512, and the time window sliding step is 50 words. The specific steps are not explained here. The image features and text features of each lesion area are spliced into a joint feature vector, and a scene feature matrix is set for the lesion area. The representation form is: Where, represents vector concatenation, n3 represents the number of lesion regions, and d represents the dimension of the embedding space; The calculation formula for alignment loss is: ; Where, represents the alignment loss, which is an indicator to measure the quality of alignment between image and text features. sim(·) represents the similarity function. Represents the image features corresponding to the positive sample, Represents the text features corresponding to the positive sample, Represents the image features corresponding to all samples, including unmatched negative samples, The weight representing the mapping relationship between lesion area i and report element j, represents the weight between the lesion region i and the unmatched negative sample j, and 、 are greater than 0, represents the adjustment coefficient, and Greater than 0, used to control the overall distribution shape of the sample to prevent dynamic weights (such as: ) introduces numerical instabilities; When an adverse reaction type is detected, constraints are imposed: ; Formula Explanation: By adjusting the LOA, the alignment of key features is strengthened. Regardless of whether an adverse reaction is detected, the minimum LOA value is forced to 1.5, strengthening the alignment of features related to adverse reactions (such as bleeding, necrosis, and edema). Among them, the adverse reaction type is represented by a performance label, y = {1, 2, ..., P}, where P is a positive integer greater than 0. Adverse reaction-aware gating is constructed based on adverse reaction type and alignment loss, and is introduced into the scene feature matrix to perform low-rank fine-tuning of weights. Specifically, the establishment of a scene feature matrix is used to maintain the semantic relevance between lesion region vectors. The alignment loss is used to allow the model to learn to distinguish the correct cross-modal associations by comparing positive samples with a large number of negative samples. That is, through the synergy of the scene feature matrix and low-rank conditional constraints, while ensuring cross-modal semantic alignment, the complex associations of clinical scenarios are effectively captured. While maintaining the overall stability of the model, it can not only prevent important lesion features and important text features from being over-compressed in the embedding space, but also significantly improve the sensitivity to adverse reaction features, and have a deeper understanding of the semantics of the generated report. Three fine-tunings: Obtain a timestamp that exceeds the threshold, randomly extract several moments before and after the corresponding time stamp to form a time series, and the time series is a dynamically changing quantity; extract the average and fluctuation values of the corresponding values of the associated parameter points in the time series, and sum them up weightedly to calculate the risk level, and call the medical knowledge graph to determine whether the numerical results corresponding to the parameter points are correct; the fluctuation value is half of the difference between the maximum and minimum values in the time series; The step of extracting the associated parameter points in the time series includes: obtaining parameter points with a mapping relationship between the clinical research process and the report elements. When establishing the association relationship, first, providing a judgment standard for the value corresponding to each parameter point, including a value range, a change rate, and a synergistic constraint condition of the associated parameters; then extracting the corresponding trigger relationship between each judgment standard and the report element, for example: establishing a trigger rule mapping table, dynamically associating the judgment standard conditions of the parameter with the report element (for example, warning level, recommended measures, chart type), when the value corresponding to the parameter point exceeds the preset value range (the value corresponding to the parameter point is not within the preset dynamic standard interval) or the change rate of the parameter point exceeds the preset change rate, it is determined that the numerical result corresponding to the parameter point is incorrect, and the parameter point needs to be calibrated to obtain a calibrated value. The calculation model is: calibrated value = g (original value, error model), wherein the error model includes zero offset, gain error, nonlinear error, and random noise. The zero offset correction value, gain correction coefficient, and nonlinear compensation coefficient are determined using least squares method or weighted regression. g(·) can be expressed as: ; In the formula, jz represents the calibration value of the parameter point, ys represents the original value of the parameter point, represents random noise, a represents the gain correction coefficient, b represents the zero offset correction value, and c represents the nonlinear compensation coefficient. The specific calculation steps are not described in detail here. Content is identified and categorized based on different risk levels. Risk levels are divided into three levels: 1, 2, and 3, with increasing risk severity. The numerical value corresponding to each risk level is data-bound to the generated report content label, and Internet big data is used to conduct in-depth learning of risk labels. For parameter points below two levels (below level 1 or below level 2), the corresponding calibration value of the parameter point is provided. Otherwise, the corresponding value of the parameter point is marked as invalid, and an equipment maintenance instruction is generated. When receiving equipment maintenance instructions, immediately repair the equipment; Finally, based on the pre-trained large language model, several blank new connection layers are added after the pre-training layer of the pre-trained initial model. The blank connection layers are trained based on the model's adjusted input features. After three rounds of fine-tuning and iterative training, the evaluation feature set and training sample size are continuously updated. Based on the model, a new evaluation feature set and the corresponding training sample size are input, and the fully connected layers are adjusted forward from the new connection layer to achieve large language model optimization. Specifically, large language models can comprehensively process image and text information to generate more comprehensive and accurate medical reports. These models are rich in medical domain knowledge and language representation capabilities, helping to better understand and express medical texts. Automatic report writing is often highly dependent on the quality and quantity of training data, so equipment is crucial for collecting clinical data. For example, in a continuous blood glucose monitor, sensor detachment due to equipment problems can lead to false hypoglycemia, ultimately resulting in inaccurate reports. Furthermore, the difficulty of understanding natural semantics is increased during the reporting process. By integrating the internal structure of the device, the accuracy of semantic report writing can be further improved. In addition, the pre-trained large language model can be set as a capsule network to extract key semantic information through the capsule network to ensure that the model can effectively learn the core semantics in the text; attention pooling technology is used to focus on document-level information in the text to enhance the recognition and understanding of medical professional terms and concepts.
[0018] Example 2 The embodiment of the present invention provides an artificial intelligence-based method for automatically writing clinical research reports; Figure 2 is a schematic diagram of the steps of a method for automatically writing a clinical research report according to an exemplary embodiment; Figure 2 , the method comprises the following steps: S1. Receive a report generation instruction; S2. Determine the device profile data and clinical research data that generate report instructions and extract report elements from them; S3. Based on the report elements, obtain the entities, entity types, and corresponding parameter points that have a mapping relationship between the clinical research process and the report elements, obtain the adverse reaction types and lesion areas that have a mapping relationship between the clinical research process and the report elements, and obtain the internal structure of the research equipment that has a mapping relationship with the report elements from the clinical research process; S4. Based on the mapping relationship, set up a multi-condition encoder, construct corresponding fine-tuning instructions, perform multi-condition fine-tuning on the large language model, and train and update it to obtain an optimized large language model.
[0019] The weight coefficient is determined using the coefficient of variation method, which assigns weights to each indicator based on the degree of variation between its current value and the target value. If the numerical difference between an indicator is large, it can clearly distinguish between the evaluated objects, indicating that the indicator has rich information for distinguishing between them, and therefore it should be given a larger weight. Conversely, if the numerical difference between the evaluated objects on a certain indicator is small, then the ability of this indicator to distinguish between the evaluated objects is weak, and therefore it should be given a smaller weight. This method directly utilizes the information contained in each indicator to calculate the indicator weight, and therefore is objective. In the application, the several formulas involved are all calculated by taking their numerical values after removing the dimensions, and the formula is a formula of the most recent real situation obtained by collecting a large amount of data and performing software simulation. The formula is set by technical personnel in this field according to actual conditions.
[0020] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.
[0021] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.
[0022] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. An artificial intelligence-based automatic writing system for clinical research reports, characterized by: include: The data collection module collects clinical research data and device profile data uploaded by users and extracts report elements; The relationship mapping module, based on the report elements, obtains the entities, entity types, and corresponding parameter points that have a mapping relationship between the clinical research process and the report elements, obtains the adverse reaction types and lesion areas that have a mapping relationship between the clinical research process and the report elements, and obtains the internal structure of the research equipment that has a mapping relationship with the report elements from the clinical research process; The training fine-tuning module sets a multi-condition encoder based on the mapping relationship, constructs corresponding fine-tuning instructions, performs multi-condition fine-tuning on the large language model, and trains and updates it to obtain an optimized large language model.
2. The artificial intelligence-based automatic writing system for clinical research reports according to claim 1, characterized in that: The steps to extract report elements include: Extract research requirements from user-uploaded clinical research data, including research literature, research topics, research plans, research progress, target diseases, and target conditions; Extract equipment requirements from equipment profile data, including equipment aging requirements and equipment performance requirements; Combine study requirements and equipment requirements to generate report elements.
3. The artificial intelligence-based automatic writing system for clinical research reports according to claim 1, characterized in that: Obtaining the mapping relationship also includes: Analyze report elements, including at least data cleaning and data conversion; Identify the types of studies to be uploaded, use the BioBERT model to extract disease conditions, symptoms, and drug entities, and obtain corresponding parameter points. Use the rule engine to match dimensions, execute a quantitative evaluation strategy, obtain the corresponding evaluation feature set and training sample size, and use a preset large language model for pre-training. The evaluation dataset includes matching efficiency, matching accuracy, and fusion coverage. Using each entity as a node, retrieve each entity type and construct a medical knowledge graph that includes target disease, target symptom, target drug, and research type. Based on the medical knowledge graph, obtain the parameter points corresponding to any entity, standardize the corresponding values of the parameter points into relative position indicators, draw parameter change curves, and overlay and display the dynamic standard range. When at least two parameter points are detected to be outside the dynamic standard range, a warning signal is generated and injected into the medical knowledge graph as a dynamic attribute. Obtaining the weight of the mapping relationship between the lesion area and the report element; Obtain the over-limit thresholds that have a mapping relationship between the internal structure of the research equipment and the report elements.
4. The artificial intelligence-based automatic writing system for clinical research reports according to claim 3 is characterized in that: The quantitative evaluation strategy is implemented based on the formula: ; In the formula, N represents the number of training samples, D type represents the research type, including parallel design, factorial design, crossover design and mixed design, φ(·) represents the nonlinear weighted result of the characteristic index of the corresponding research type, and the characteristic index includes at least one of the balance characteristic, interaction intensity characteristic and individual difference characteristic, h represents the standardization factor, β type represents the weight coefficient corresponding to the characteristic index, Es represents the basic sample size, and λ represents the adjustment factor.
5. The artificial intelligence-based automatic writing system for clinical research reports according to claim 3 is characterized in that: The steps to build fine-tuning instructions include: First fine-tuning: Calculate the value index based on the evaluation feature set, compare and analyze the value index with the preset standard value range [jz1, jz2], automatically match the template paragraph, and call the NLP explanatory paragraph generator under the conditions of triggering the paragraph basic template and paragraph standard template, and automatically annotate; Secondary fine-tuning: Establish a joint embedding space, including a visual encoder and a text encoder. Input the weights of the mapping relationship between the lesion area and the report elements into the joint embedding space, set the scene feature matrix for the lesion area and retain the alignment loss. Establish a conditional encoder, input the weights and scene feature matrix into the conditional encoder, and enforce constraints when the adverse reaction type is detected. Three fine-tuning steps: Obtain the timestamp that exceeds the threshold, randomly extract several moments before and after the corresponding time stamp to form a time series, and the time series is a dynamic change quantity; extract the average and fluctuation values of the corresponding values of the associated parameter points in the time series, and take a weighted sum to calculate the risk level; then retrieve the medical knowledge graph to determine whether the numerical results corresponding to the parameter points are correct; Content is identified and categorized based on different risk levels, and the risk levels are divided into three levels. The numerical value corresponding to each risk level is data-bound to the generated report content label, and Internet big data is provided for deep learning of risk labels. For parameter points below level two, the corresponding calibration value of the parameter point is provided. Otherwise, the value corresponding to the parameter point is marked as invalid, and an equipment maintenance instruction is generated.
6. The artificial intelligence-based automatic writing system for clinical research reports according to claim 5, characterized in that: Compare and analyze the value index with the preset standard value range [jz1, jz2]: Mark the value index as value; When value < jz1, match the paragraph basic template; When jz1≤value<jz2, match the paragraph standard template; When value ≥ jz2, matches the paragraph top-level template.
7. The artificial intelligence-based automatic writing system for clinical research reports according to claim 1, characterized in that: The steps of secondary fine-tuning include: Establish a joint embedding space, use the visual encoder to extract the image features of the lesion area, use the text encoder to extract the text features, and map the features corresponding to the lesion area and report elements, as well as the weights of the mapping relationship between the lesion area and the report elements, into a unified space; The image features and text features of each lesion area are concatenated into a joint feature vector, a scene feature matrix is set for the lesion area, and the alignment loss is retained; When an adverse reaction type is detected, the constraint is enforced: LOA = max(LOA, 1.5); Where LOA represents the alignment loss.
8. The artificial intelligence-based automatic writing system for clinical research reports according to claim 5, characterized in that: The step of extracting the associated parameter points in the time sequence includes: when establishing the association relationship between the time sequence and the parameter points, first providing a judgment standard for the numerical value corresponding to each parameter point, and then extracting the corresponding trigger relationship between each judgment standard and the report element.
9. An artificial intelligence-based automatic writing method for clinical research reports, characterized in that: The steps include: receiving an instruction to generate a report; Determine the device profile data and clinical research data that generate report instructions, and extract report elements from them; Based on the report elements, entities, entity types, and corresponding parameter points that have a mapping relationship between the clinical research process and the report elements are obtained, as well as adverse reaction types and lesion areas that have a mapping relationship between the clinical research process and the report elements. The internal structure of the research equipment that has a mapping relationship with the report elements is also obtained from the clinical research process; Based on the mapping relationship, a multi-condition encoder is set up, corresponding fine-tuning instructions are constructed, multi-condition fine-tuning is performed on the large language model, and training and updating are performed to obtain an optimized large language model.
Citation Information
Patent Citations
Large language model system for clinical test data analysis and interpretation
CN118116536A
Artificial intelligence automatic report evaluation method and system
CN119517274A
Radiology report auxiliary writing method and system based on intelligent follow-up visit
CN119943249A
Medical image report generation method based on multi-modal large model preference alignment technology
CN120032790A
Domain-adaptive pre-training of instruction-tuned llms for radiology report impression generation
US20240387014A1
Cited By
Multi-modal data driven general report generation method and system based on large model
CN121031543A