Case data simulation generation method and system based on large model and medium

Through a large-model-based case data simulation generation method, and using data management relationship analysis and clustering algorithms to evaluate data quality, the problems of low data quality and difficulty in capturing nonlinear interactive relationships were solved, and high-quality medical record data simulation was achieved.

CN120767007APending Publication Date: 2025-10-10NORTH CHINA DIGITAL HEALTH TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510628546.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In the existing technology, the data quality generated by case data simulation is not high, and manual verification is required to see whether the fitted data deviates from the real data. It is also difficult to capture the nonlinear interactions in complex biological systems.

Method used

By obtaining the actual medical record data after desensitization, using the data management relationship analysis program to extract the distribution characteristics and correlation relationships, combining the large model to generate fitting medical record data, and evaluating the data quality through clustering algorithm and distance calculation, it is automatically determined whether the fitting data is qualified.

Benefits of technology

It improves the quality of simulation data, reduces the need for manual verification, can more accurately capture nonlinear interactions in complex biological systems, and the generated data is more consistent with actual data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120767007A_ABST
    Figure CN120767007A_ABST
Patent Text Reader

Abstract

The invention discloses a case data simulation generation method and system based on a large model, and a medium, mainly relates to the technical field of case data simulation, and is used for solving the problems that the quality of data generated by an existing simulation normal form is not high; whether fitting data deviates from real data or not needs to be manually verified, and a simulation normal form is difficult to capture a nonlinear interaction relation in a complex biological system in the fitting process. Comprising the following steps: extracting an association relationship between distribution characteristics of actual medical record data and preset variables; obtaining a preset clustering number of first clustering centers; cue words of the simulated case data are obtained, and a preset number of fitting case data are generated through large model simulation; adding the fitting medical record data to the actual medical record data to obtain total medical record data, and obtaining a preset clustering number of second clustering centers; the distance between the first clustering center and the corresponding second clustering center is calculated, and when the preset clustering number of distances is smaller than the preset maximum distance, it is determined that the fitting medical record data is qualified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of case data simulation, and in particular to a method, system and medium for generating case data simulation based on a large model. Background Art

[0002] In medical research, data simulation analysis, a key pillar of the quantitative research paradigm, has deeply penetrated key areas such as clinical trial design, disease prediction modeling, and precision medicine decision-making. Leveraging breakthroughs in distributed computing frameworks (such as Hadoop / Spark) and deep learning algorithms, medical data processing has expanded from traditional structured data to multimodal biomedical data (including radiomics, genomics, and time-series data from wearable devices). However, the sensitivity of medical data is constrained by data privacy regulations such as HIPAA and GDPR, resulting in significant technical barriers to cross-institutional data sharing and the creation of "data silos."

[0003] The current medical data simulation paradigm is primarily based on classic statistical modeling frameworks (such as GLM and Cox regression), and its technical implementation relies heavily on the manual feature engineering and programming skills of statistical analysts (R / Python). This manually driven modeling approach faces two challenges: first, the data quality generated by the simulation paradigm is low, requiring manual verification to verify whether the fitted data deviates from the real data; second, the simulation paradigm has difficulty capturing nonlinear interactions (such as gene-environment interactions) in complex biological systems during the fitting process. Therefore, a method, system, and medium for generating case data simulations based on large models are urgently needed to address the problems of low data quality generated by the simulation paradigm, the need for manual verification to verify whether the fitted data deviates from the real data, and the difficulty of capturing nonlinear interactions in complex biological systems during the fitting process. Summary of the Invention

[0004] In response to the above-mentioned deficiencies in the prior art, the present application provides a method, system and medium for simulating and generating case data based on a large model to solve the problems existing in the existing solutions: the data quality generated by the simulation paradigm is not high, manual verification is required to determine whether the fitted data deviates from the real data, and the simulation paradigm is difficult to capture the nonlinear interactions in complex biological systems during the fitting process.

[0005] In a first aspect, the application provides a case data simulation generation method based on a large model, the method comprising: obtaining a plurality of actual medical record data after desensitization processing; obtaining a data management relationship analysis program, and then using the data management relationship analysis program to extract the correlation between the distribution characteristics of the actual medical record data and the preset variables; obtaining a first clustering center of a preset clustering number by using the actual medical record data through a preset clustering algorithm; obtaining a prompt word of simulated case data, and using a large model to simulate and generate a preset number of fitted medical record data according to the prompt word, the correlation between the distribution characteristics of the actual medical record data and the preset variables; adding the fitted medical record data to the actual medical record data to obtain total medical record data, taking the total medical record data as the input of the preset clustering algorithm, and obtaining a second clustering center of a preset clustering number; calculating the distance between the first clustering center and the corresponding second clustering center, and determining that the fitted medical record data is qualified when the preset clustering number of distances are all less than a preset maximum distance; generating an abnormal alarm prompt when any distance is greater than or equal to the preset maximum distance, and uploading the abnormal alarm prompt to a preset maintenance terminal; wherein the actual medical record data of the corresponding cluster is obtained according to the abnormal alarm prompt.

[0006] In an implementation manner of the application, the plurality of actual medical record data after desensitization processing is obtained, specifically comprising: extracting an actual medical record data from the plurality of actual medical record data, and obtaining an actual data type involved in the current actual medical record data; obtaining a preset desensitization data type and a desensitization program corresponding to each preset desensitization data type; wherein the desensitization program at least includes a data replacement program, a data masking program and a data shielding program; based on the preset desensitization data type, the corresponding desensitization program is called to perform desensitization processing on the data corresponding to the preset desensitization data type in all actual medical record data.

[0007] In an implementation manner of the application, before obtaining the data management relationship analysis program and then using the data management relationship analysis program to extract the correlation between the distribution characteristics of the actual medical record data and the preset variables, the method further comprises: creating a data management relationship analysis program corresponding to a Bayesian non-parametric model to automatically identify the distribution characteristics corresponding to the actual medical record data; taking the actual medical record data as the input of a preset prediction model to complete the training of the preset prediction model; using a DeepSHAP interpreter to calculate the interaction effect diagram of any two distribution characteristics on the prediction result of the preset prediction model; identifying whether the scatter points in the interaction effect diagram have an upward or downward trend, and determining that the current two distribution characteristics have a correlation relationship when there is an upward or downward trend; determining that the current two distribution characteristics do not have a correlation relationship when there is no upward or downward trend; inputting the specific data corresponding to the two distribution characteristics having a correlation relationship in the medical record data into a preset language large model algorithm to obtain the specific correlation relationship.

[0008] In one implementation of the present application, a preset number of cluster centers are obtained using actual medical record data through a preset clustering algorithm, specifically including: obtaining a K value as the preset number of clusters; using a preset clustering algorithm, randomly selecting K data points from the actual medical record data as initial cluster centers; and obtaining K first cluster centers after the algorithm converges.

[0009] In one implementation of the present application, the distance between the first cluster center and the corresponding second cluster center is calculated. When the preset number of cluster distances are all less than the preset maximum distance, it is determined that the fitted medical record data is qualified, specifically including: The vectors corresponding to the first cluster center and the corresponding second cluster center are input into the Euclidean distance calculation formula to obtain the distance between the first cluster center and the corresponding second cluster center; when all K distances are less than the preset maximum distance, the fitted medical record data is determined to be qualified.

[0010] In one implementation of the present application, when any distance is greater than or equal to a preset maximum distance, after generating an abnormal alarm prompt and uploading it to a preset maintenance terminal, the method further includes: When no processing information returned by the preset maintenance terminal is received within a preset time period, the actual medical record data corresponding to the first cluster center with a distance greater than or equal to the preset maximum distance is extracted; the specific numerical range corresponding to the preset keyword in the actual medical record data corresponding to the first cluster center is extracted; the specific numerical range is used to replace the specific numerical range corresponding to the preset keyword in the prompt word; according to the updated prompt word, the distribution characteristics of the actual medical record data and the correlation between the preset variables, a large model is used to simulate and generate a number of fitting medical record data corresponding to the first cluster center; the Euclidean distance calculation formula is used to calculate the distance between the several fitting medical record data corresponding to the first cluster center and any two data in the actual medical record data to obtain the maximum distance; when the maximum distance is less than the preset maximum distance, an abnormal alarm prompt cancellation message is sent to the preset maintenance terminal; when the maximum distance is greater than or equal to the preset maximum distance, an abnormal alarm prompt is sent to the preset maintenance terminal again.

[0011] In a second aspect, the present application provides a case data simulation generation system based on a large model, the system comprising: The first obtaining module is configured to obtain a plurality of actual medical record data after desensitization processing; a data management relationship analysis program is acquired, and then the data management relationship analysis program is used to extract a correlation between distribution features of the actual medical record data and preset variables; a preset clustering algorithm is used to obtain a preset clustering number of first clustering centers by using the actual medical record data; the fitting generation module is configured to acquire prompt words of simulated case data, acquire a preset number of fitting medical record data by using a large model according to the prompt words, the correlation between the distribution features of the actual medical record data and the preset variables; the second obtaining module is configured to add the fitting medical record data to the actual medical record data to obtain total medical record data, and use the total medical record data as an input of the preset clustering algorithm to obtain a preset clustering number of second clustering centers; the qualified determination module is configured to calculate distances between the first clustering centers and the corresponding second clustering centers, and determine that the fitting medical record data is qualified when the preset clustering number of distances are all less than a preset maximum distance; the abnormality prompt module is configured to generate an abnormality alarm prompt when any distance is greater than or equal to the preset maximum distance, and upload the abnormality alarm prompt to a preset maintenance terminal; wherein the abnormality alarm prompt acquires actual medical record data of a corresponding cluster.

[0012] In an implementation manner of the present application, the first obtaining module comprises a desensitization unit, The first obtaining module is configured to extract an actual medical record data from a plurality of actual medical record data, acquire an actual data type related to the current actual medical record data; acquire a preset desensitization data type and a desensitization program corresponding to each preset desensitization data type; wherein the desensitization program at least includes a data replacement program, a data masking program and a data shielding program; based on the preset desensitization data type, the corresponding desensitization program is called to perform desensitization processing on the data corresponding to the preset desensitization data type in all actual medical record data.

[0013] In an implementation manner of the present application, the first obtaining module comprises a data obtaining unit, The first obtaining module is configured to create a data management relationship analysis program corresponding to a Bayesian non-parametric model to automatically identify distribution features corresponding to the actual medical record data; use the actual medical record data as an input of a preset prediction model to complete training of the preset prediction model; use a DeepSHAP interpreter to calculate an interaction effect diagram of any two distribution features on a prediction result of the preset prediction model; identify whether there is an upward or downward trend in the scatter points in the interaction effect diagram, and when there is an upward or downward trend, it is determined that the current two distribution features have a correlation; when there is no upward or downward trend, it is determined that the current two distribution features have no correlation; input specific data corresponding to the two distribution features having a correlation in the medical record data into a preset language large model algorithm to obtain a specific correlation.

[0014] In a third aspect, the present application provides a non-volatile computer storage medium having computer instructions stored thereon, the computer instructions, when executed, implementing a case data simulation generation method based on a large model according to any one of the above.

[0015] Those skilled in the art can understand that the present application has at least the following beneficial effects: 1. Improved quality of simulation data: By obtaining the actual medical record data after desensitization processing, and using the data management relationship analysis program to extract the distribution characteristics of the data and the correlation between variables, a real and reliable basis is provided for simulation generation. Further, through the association between the prompt words, the distribution characteristics of the actual medical record data and the preset variables, the large model is guided to more accurately capture the internal laws and characteristics of the data, thereby improving the quality of the generated data.

[0016] In addition, through clustering algorithm and distance calculation, the quality of the fitted medical record data is evaluated. When the clustering center distance between the fitted data and the actual data is within the preset range, the data is determined to be qualified, which ensures the consistency of the generated data and the actual data in terms of group characteristics, and further improves the data quality.

[0017] 2. Reduced need for manual verification: The present application reduces the need for manual participation through an automated process, including data extraction, simulation generation, quality evaluation and other steps. In particular, in the quality evaluation stage, the clustering center distance is calculated to automatically determine whether the fitted data is qualified, avoiding the problem of requiring a large amount of manual verification of whether the fitted data deviates from the actual data in the traditional method.

[0018] 3. Ability to capture non-linear interaction in complex biological systems: The application of the large model enables the present method to more effectively capture non-linear interactions in complex biological systems. In addition, through the prompt words, the distribution characteristics of the actual medical record data and the correlation between variables, the large model can simulate the generation of fitted medical record data that is closer to the actual situation, thereby capturing non-linear interaction effects that are difficult to capture by traditional methods. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0020] Figure 1 is a flowchart of a case data simulation generation method based on a large model provided by an embodiment of the present application.

[0021] Figure 2 FIG. 1 is a schematic diagram of an internal structure of a case data simulation generation system based on a large model according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] It should be understood by those skilled in the art that the embodiments described below are only preferred embodiments of the present disclosure, and do not represent that the present disclosure can only be implemented by the preferred embodiments. The preferred embodiments are only used to explain the technical principles of the present disclosure, and are not used to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by those skilled in the art without creative labor should still fall within the protection scope of the present disclosure.

[0023] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or apparatus that includes a list of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent in such process, method, article, or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0024] The technical solutions of the embodiments of the present application will be described in detail below with reference to the drawings.

[0025] The embodiments provide a case data simulation generation method based on a large model, as shown in FIG. 1, the method provided by the embodiments of the present application mainly includes the following steps: Figure 1 As shown in FIG. 1, the method provided by the embodiments of the present application mainly includes the following steps: Step 110, obtaining a plurality of actual medical record data after desensitization processing; obtaining a data management relationship analysis program, and then using the data management relationship analysis program to extract the correlation between the distribution characteristics of the actual medical record data and the preset variables; obtaining a preset clustering number of first clustering centers by using the actual medical record data through a preset clustering algorithm.

[0026] In some embodiments, the desensitization processing scheme can be specifically as follows: Extracting an actual medical record data from a plurality of actual medical record data, obtaining an actual data type involved in the current actual medical record data; obtaining a preset desensitization data type and a desensitization program corresponding to each preset desensitization data type; wherein the desensitization program at least includes: a data replacement program, a data masking program, and a data shielding program; based on the preset desensitization data type, calling the corresponding desensitization program to perform desensitization processing on the data corresponding to the preset desensitization data type in all actual medical record data.

[0027] It should be noted that the data extraction process can be implemented by existing technologies, and this application does not impose too many restrictions on this.

[0028] The data management relationship analysis program in the step may be a data management relationship analysis program corresponding to the Bayesian non-parametric model, so as to automatically identify the distribution characteristics corresponding to the actual medical record data.

[0029] In addition, the relationship between the distribution characteristics of the actual medical record data and the preset variables can be specifically extracted as follows: Actual medical record data is used as the input of the preset prediction model to complete the training of the preset prediction model; the DeepSHAP (SHapley Additive exPlanations for Deep Learning) interpreter is used to calculate the interaction effect diagram of any two distribution features for the prediction results of the preset prediction model; identify whether the scatter points in the interaction effect diagram have an upward or downward trend; when there is an upward or downward trend, it is determined that the two current distribution features have a correlation relationship; when there is no upward or downward trend, it is determined that the two current distribution features have no correlation relationship; the specific data corresponding to the two distribution features with a correlation relationship in the medical record data is input into the preset language large model algorithm to obtain the specific correlation relationship.

[0030] It should be noted that the solution for identifying whether the scattered points in the interaction effect diagram have an upward or downward trend can be implemented by existing image recognition programs.

[0031] The preset number of clustering centers is obtained by using the actual medical record data through a preset clustering algorithm, which can be specifically: Obtain a K value as the preset number of clusters; use a preset clustering algorithm to randomly select K data points from the actual medical record data as initial cluster centers; after the algorithm converges, obtain K first cluster centers.

[0032] It should be noted that the process of obtaining the K value can be directly obtained through a preset acquisition interface, or there is a default K value. When the K value cannot be obtained from the outside, the default K value is used. In addition, the K value can be any feasible value, such as 7.

[0033] Step 120: Obtain prompt words of the simulated case data, and use the large model to simulate and generate a preset number of fitting medical record data based on the prompt words, the distribution characteristics of the actual medical record data, and the correlation between the preset variables.

[0034] It should be noted that the prompt words can be: “You are an expert in medical data simulation generation. Please design and generate a comprehensive simulated medical dataset that may cover the following key aspects: Patient Basic Information: Create multiple virtual patient profiles, each containing a fictional name, age, gender, contact information (ensure privacy protection), occupation, residence, and medical history. Ensure patient information is diverse to reflect different demographic characteristics.

[0035] Physiological and Health Data: Generate detailed physiological data for each patient, including but not limited to heart rate, blood pressure, blood glucose, cholesterol levels, weight, height, BMI, etc. These data should reflect normal fluctuation ranges and contain some abnormal values to simulate real-world health condition changes.

[0036] Medical Images and Diagnoses: Generate simulated medical image data such as X-rays, CT scans, MRI images, etc., along with detailed medical diagnosis reports. Ensure that the image data and descriptions in the diagnosis report are consistent and accurately reflect various disease characteristics.

[0037] Clinical Records and Medical History: Construct a complete electronic medical record system data, including admission records, daily ward rounds records, laboratory examination results, imaging examination reports, medical order records, treatment plans, operation records, discharge summaries, etc. Ensure that the medical history content is coherent, the timeline is clear, and it conforms to the actual medical process.

[0038] Medications and Treatment Plans: Generate detailed medication use records and treatment plans for each patient, including drug names, dosages, administration routes, treatment cycles, etc. At the same time, simulate patient responses to different drugs, including efficacy and adverse reactions.

[0039] Public Health and Epidemiological Data: Generate simulated public health data sets, including vaccine coverage rates, disease incidence rates, mortality rates, medical resource allocation in different regions and age groups. These data should reflect the impact of specific public health events and contain preliminary analysis reports based on these data.

[0040] Data Compliance and Privacy Protection: When generating all data, ensure compliance with relevant laws, regulations and ethical requirements, especially those related to data privacy and security. Ensure that all personal information is presented in an anonymous form to avoid leaking sensitive information.

[0041] Data Quality and Verification: The generated data should undergo strict quality control and verification processes to ensure its accuracy, completeness and consistency. This includes checking the logical self-consistency of the data, the continuity of the timeline, and the compliance with medical common sense.

[0042] Data Format and Accessibility: The provided data should be presented in an easily understandable and accessible format, such as spreadsheets, databases or specialized medical data platforms. At the same time, provide necessary data dictionaries and documents to explain the meaning and structure of the data.

[0043] Extensibility and flexibility: The designed dataset should be extensible and flexible so that new data types or patient information can be added when needed. In addition, the dataset should support a variety of analysis methods and tools to meet the needs of different researchers.

[0044] It should be noted that the process of obtaining the prompt words of the simulated case data can be directly obtained through a preset acquisition interface.

[0045] Step 130: Add the fitted medical record data to the actual medical record data to obtain the total medical record data, use the total medical record data as the input of the preset clustering algorithm, and obtain a preset number of clustering second cluster centers.

[0046] Step 140: Calculate the distance between the first cluster center and the corresponding second cluster center. When the preset number of cluster distances are all less than the preset maximum distance, determine that the fitted medical record data is qualified.

[0047] It should be supplemented that the preset maximum distance can be any feasible value, which can be obtained by those skilled in the art through multiple experiments.

[0048] In some embodiments, this step may specifically include: The vectors corresponding to the first cluster center and the corresponding second cluster center are input into the Euclidean distance calculation formula to obtain the distance between the first cluster center and the corresponding second cluster center; when all K distances are less than the preset maximum distance, the fitted medical record data is determined to be qualified.

[0049] Step 150: When any distance is greater than or equal to the preset maximum distance, an abnormal alarm is generated and uploaded to the preset maintenance terminal.

[0050] It should be noted that the abnormal alarm prompts the acquisition of the actual medical record data of the corresponding cluster.

[0051] In addition, to avoid the preset maintenance terminal failing to handle the problem in a timely manner, this application includes a backup solution, specifically: When any distance is greater than or equal to the preset maximum distance, an abnormal alarm prompt is generated and uploaded to the preset maintenance terminal. When no processing information returned by the preset maintenance terminal is received within the preset time period, the actual medical record data corresponding to the first cluster center with a distance greater than or equal to the preset maximum distance is extracted; the specific numerical range corresponding to the preset keyword in the actual medical record data corresponding to the first cluster center is extracted; the specific numerical range is used to replace the specific numerical range corresponding to the preset keyword in the prompt word; according to the updated prompt word, the distribution characteristics of the actual medical record data and the correlation between the preset variables, a large model is used to simulate and generate several fitting medical record data corresponding to the first cluster center; the Euclidean distance calculation formula is used to calculate the distance between the several fitting medical record data corresponding to the first cluster center and any two data in the actual medical record data to obtain the maximum distance; when the maximum distance is less than the preset maximum distance, the abnormal alarm prompt cancellation information is sent to the preset maintenance terminal; when the maximum distance is greater than or equal to the preset maximum distance, the abnormal alarm prompt is sent to the preset maintenance terminal again.

[0052] In addition, this application Figure 2 The embodiment of the present application provides a case data simulation generation system based on a large model. Figure 2 As shown, the system provided in the embodiment of the present application mainly includes: The first acquisition module 210 is used to obtain a number of actual medical record data after desensitization processing; obtain a data management relationship analysis program, and then use the data management relationship analysis program to extract the distribution characteristics of the actual medical record data and the correlation between preset variables; through a preset clustering algorithm, use the actual medical record data to obtain a preset number of first cluster centers.

[0053] The first acquisition module 210 includes a desensitizing unit, which is used to extract a piece of actual medical record data from a number of actual medical record data to obtain the actual data type involved in the current actual medical record data; obtain the preset desensitized data type, and the desensitizing program corresponding to each preset desensitized data type; wherein the desensitizing program includes at least: a data replacement program, a data masking program, and a data shielding program; based on the preset desensitized data type, call the corresponding desensitizing program to desensitize the data corresponding to the preset desensitized data type in all actual medical record data.

[0054] The first acquisition module 210 includes a data acquisition unit, which is used to create a data management relationship analysis program corresponding to the Bayesian non-parametric model to automatically identify the distribution characteristics corresponding to the actual medical record data; use the actual medical record data as the input of the preset prediction model to complete the training of the preset prediction model; use the DeepSHAP interpreter to calculate the interaction effect diagram of any two distribution characteristics for the prediction results of the preset prediction model; identify whether there is an upward or downward trend in the scatter points in the interaction effect diagram, and when there is an upward or downward trend, determine that the two current distribution characteristics have a correlation relationship; when there is no upward or downward trend, determine that the two current distribution characteristics do not have a correlation relationship; input the specific data corresponding to the two distribution characteristics with a correlation relationship in the medical record data into the preset language large model algorithm to obtain a specific correlation relationship.

[0055] The fitting generation module 220 is used to obtain prompt words of the simulated case data, and use the large model to simulate and generate a preset number of fitting medical record data based on the prompt words, the distribution characteristics of the actual medical record data and the correlation between the preset variables.

[0056] The second obtaining module 230 is used to add the fitted medical record data to the actual medical record data to obtain the total medical record data, and use the total medical record data as the input of the preset clustering algorithm to obtain a preset number of second cluster centers.

[0057] The qualified determination module 240 is used to calculate the distance between the first cluster center and the corresponding second cluster center. When the preset number of cluster distances are all less than the preset maximum distance, it is determined that the fitted medical record data is qualified.

[0058] The abnormality prompt module 250 is used to generate an abnormality alarm prompt when any distance is greater than or equal to the preset maximum distance, and upload it to the preset maintenance terminal; wherein the abnormality alarm prompt obtains the actual medical record data of the corresponding cluster.

[0059] In addition, an embodiment of the present application further provides a non-volatile computer storage medium on which executable instructions are stored. When the executable instructions are executed, a method for simulating and generating case data based on a large model as described above is implemented.

[0060] Thus far, the technical solutions of the present disclosure have been described in conjunction with the foregoing multiple embodiments. However, it is easy for those skilled in the art to understand that the scope of protection of the present disclosure is not limited to these specific embodiments. Without departing from the technical principles of the present disclosure, those skilled in the art may split and combine the technical solutions in the above-mentioned various embodiments, and may also make equivalent changes or replacements to the relevant technical features. Any changes, equivalent replacements, improvements, etc. made within the technical concepts and / or technical principles of the present disclosure will fall within the scope of protection of the present disclosure.

Claims

1. A method for simulating and generating case data based on a large model, characterized in that: The method comprises: Obtaining a number of actual medical record data after desensitization processing; obtaining a data management relationship analysis program, and then using the data management relationship analysis program to extract the distribution characteristics of the actual medical record data and the correlation between preset variables; using a preset clustering algorithm, using the actual medical record data, obtaining a preset number of cluster first cluster centers; Obtain prompt words from simulated case data, and use a large model to simulate and generate a preset number of fitted medical record data based on the prompt words, the distribution characteristics of the actual medical record data, and the correlation between the preset variables; Add the fitted medical record data to the actual medical record data to obtain the total medical record data, use the total medical record data as the input of the preset clustering algorithm, and obtain the preset number of cluster second cluster centers; Calculate the distance between the first cluster center and the corresponding second cluster center. When the preset number of cluster distances are all less than the preset maximum distance, determine that the fitted medical record data is qualified. When any distance is greater than or equal to the preset maximum distance, an abnormal alarm prompt is generated and uploaded to the preset maintenance terminal; wherein, the abnormal alarm prompt obtains the actual medical record data of the corresponding cluster.

2. The method for simulating and generating case data based on a large model according to claim 1, characterized in that: Obtained some actual medical records after desensitization, including: Extracting a piece of actual medical record data from a number of actual medical record data to obtain the actual data type involved in the current actual medical record data; Obtaining preset desensitizing data types and desensitizing programs corresponding to each preset desensitizing data type; wherein the desensitizing program includes at least: a data replacement program, a data masking program, and a data shielding program; Based on the preset desensitized data type, the corresponding desensitization program is called to desensitize the data corresponding to the preset desensitized data type in all actual medical record data.

3. The method for simulating and generating case data based on a large model according to claim 1, characterized in that: Before obtaining the data management relationship analysis program and then using the data management relationship analysis program to extract the correlation between the distribution characteristics of the actual medical record data and the preset variables, the method further includes: Create a data management relationship analysis program corresponding to the Bayesian non-parametric model to automatically identify the distribution characteristics corresponding to the actual medical record data; Using actual medical record data as input to the preset prediction model to complete the training of the preset prediction model; Use the DeepSHAP interpreter to calculate the interaction effect plot of any two distribution features on the prediction results of the preset prediction model; Identify whether there is an upward or downward trend in the scatter points in the interaction effect diagram. If there is an upward or downward trend, determine that there is a correlation between the two current distribution characteristics; When there is no upward or downward trend, it is determined that there is no correlation between the two current distribution characteristics; The specific data corresponding to the two distribution features with an associated relationship in the medical record data are input into the preset language large model algorithm to obtain the specific associated relationship.

4. The method for simulating and generating case data based on a large model according to claim 1, characterized in that: By using the preset clustering algorithm and actual medical record data, the first cluster centers of the preset number of clusters are obtained, specifically including: Get the K value as the preset number of clusters; Using the preset clustering algorithm, K data points are randomly selected from the actual medical record data as the initial cluster centers; After the algorithm converges, K first cluster centers are obtained.

5. The method for simulating and generating case data based on a large model according to claim 1, characterized in that: Calculate the distance between the first cluster center and the corresponding second cluster center. When the preset number of cluster distances are all less than the preset maximum distance, determine that the fitted medical record data is qualified, specifically including: Input the vectors corresponding to the first cluster center and the corresponding second cluster center into the Euclidean distance calculation formula to obtain the distance between the first cluster center and the corresponding second cluster center; When all K distances are smaller than the preset maximum distance, the fitted medical record data is determined to be qualified.

6. The method for simulating and generating case data based on a large model according to claim 1, characterized in that: When any distance is greater than or equal to a preset maximum distance, an abnormal alarm is generated and uploaded to a preset maintenance terminal, the method further includes: When no processing information is received from the preset maintenance terminal within the preset time period, Extracting actual medical record data corresponding to the first cluster center whose distance is greater than or equal to a preset maximum distance; Extracting a specific numerical range corresponding to a preset keyword in the actual medical record data corresponding to the first cluster center; Use a specific numerical range to replace the specific numerical range corresponding to the preset keyword in the prompt word; Based on the updated prompt words, the distribution characteristics of the actual medical record data and the correlation between the preset variables, a large model is used to simulate and generate a number of fitted medical record data corresponding to the first cluster center; Using the Euclidean distance calculation formula, calculate the distance between any two data in the fitted medical record data and the actual medical record data corresponding to the first cluster center to obtain the maximum distance; When the maximum distance is less than the preset maximum distance, an abnormal alarm prompt cancellation message is sent to the preset maintenance terminal; When the maximum distance is greater than or equal to the preset maximum distance, an abnormal alarm prompt is sent to the preset maintenance terminal again.

7. A case data simulation generation system based on a large model, characterized in that: The system comprises: The first acquisition module is used to obtain a number of actual medical record data after desensitization processing; obtain a data management relationship analysis program, and then use the data management relationship analysis program to extract the distribution characteristics of the actual medical record data and the correlation between preset variables; through a preset clustering algorithm, using the actual medical record data, obtain a preset number of first cluster centers; The fitting generation module is used to obtain prompt words of simulated case data, and use the large model to simulate and generate a preset number of fitting medical record data based on the correlation between the prompt words, the distribution characteristics of the actual medical record data and the preset variables; A second obtaining module is used to add the fitted medical record data to the actual medical record data to obtain the total medical record data, and use the total medical record data as the input of a preset clustering algorithm to obtain a preset number of second cluster centers; A qualified determination module is used to calculate the distance between the first cluster center and the corresponding second cluster center, and when the preset number of cluster distances are all less than the preset maximum distance, it is determined that the fitted medical record data is qualified; The abnormality prompt module is used to generate an abnormality alarm prompt when any distance is greater than or equal to the preset maximum distance, and upload it to the preset maintenance terminal; wherein the abnormality alarm prompt obtains the actual medical record data of the corresponding cluster.

8. The large model-based case data simulation generation system according to claim 7, characterized in that: The first acquisition module includes a desensitization unit, Used to extract a piece of actual medical record data from a number of actual medical record data; Obtain the actual data type involved in the current actual medical record data; Obtaining preset desensitizing data types and desensitizing programs corresponding to each preset desensitizing data type; wherein the desensitizing program includes at least: a data replacement program, a data masking program, and a data shielding program; Based on the preset desensitized data type, the corresponding desensitization program is called to desensitize the data corresponding to the preset desensitized data type in all actual medical record data.

9. The large model-based case data simulation generation system according to claim 7, characterized in that: The first acquisition module includes a data acquisition unit, A data management relationship analysis program for creating a Bayesian nonparametric model to automatically identify the distribution characteristics corresponding to actual medical record data; Using actual medical record data as input to the preset prediction model to complete the training of the preset prediction model; Use the DeepSHAP interpreter to calculate the interaction effect plot of any two distribution features on the prediction results of the preset prediction model; Identify whether there is an upward or downward trend in the scatter points in the interaction effect diagram. If there is an upward or downward trend, determine that there is a correlation between the two current distribution characteristics; When there is no upward or downward trend, it is determined that there is no correlation between the two current distribution characteristics; The specific data corresponding to the two distribution features with an associated relationship in the medical record data are input into the preset language large model algorithm to obtain the specific associated relationship.

10. A non-volatile computer storage medium, characterized in that Computer instructions are stored thereon, and when the computer instructions are executed, they implement a method for simulating and generating case data based on a large model as described in any one of claims 1 to 6.