Patient report outcome management method and system

By optimizing patient report content through permission allocation and biomedical pre-training models, the problem of real-time updating in existing technologies is solved, real-time optimization and medical compliance of content are achieved, and the credibility and clinical application value of data are improved.

CN120708794APending Publication Date: 2025-09-26TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510852076.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing patient-reported outcome management methods cannot be updated in real time according to medical standards and contain outdated or erroneous statements, which affects the credibility and clinical value of the data.

Method used

Through permission allocation, behavior coding, random forest algorithm and biomedical pre-training model, content degradation index and medical compliance index are generated to optimize patient report content and ensure that it complies with medical standards.

Benefits of technology

It achieves real-time optimization of patient report content, improves the credibility and clinical application value of data, reduces the terminology calibration costs of medical staff, and ensures content quality and compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708794A_ABST
    Figure CN120708794A_ABST
Patent Text Reader

Abstract

The invention discloses a patient report outcome management method and system, and relates to the technical field of medical information. The method comprises the following steps: firstly, dividing role permissions, and defining an information viewing and operating range; key operation behaviors are collected and coded according to a unified format to generate a role behavior chain; extracting a content degradation factor from the behavior chain, determining a weight through a random forest algorithm, and carrying out weighted calculation on a content degradation index; a term vector is generated by using a biomedical model, and a medical compliance index is calculated through a cosine distance; content optimization indexes are calculated by fusing the two indexes, and fields needing to be optimized are marked by comparing threshold values; when the field content is the content optimization item, optimizing the original field content; and comparing data of the experimental group and the control group through A / B test, calculating to obtain an optimization effect index, and automatically replacing the optimization effect index with optimization content if the optimization effect index reaches the standard to form a final patient report. According to the optimization effect index, whether the optimized field content is replaced with the original field content or not is determined, and a final patient report is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical information technology, and in particular to a patient-reported outcome management method and system. Background Art

[0002] Patient-reported outcomes are self-reported health status assessment tools used extensively in chronic disease management, drug clinical trials, and efficacy evaluation. The quality of patient-reported outcomes directly determines the credibility and clinical value of the data. Reports with ambiguous terminology, confusing logic, or missing information can lead doctors to misjudge disease progression, develop inaccurate treatment plans, and even introduce medication risks. Furthermore, obscure language can reduce patient understanding, weakening their trust in treatment plans and willingness to cooperate, ultimately impacting recovery outcomes.

[0003] However, existing patient-reported outcome management methods usually rely on expert experience to optimize report content and lack a data-driven effect verification mechanism, which results in the inability to update patient-reported outcomes in real time according to medical standards and the presence of outdated or incorrect content. Summary of the Invention

[0004] In response to the deficiencies of the existing technology, the present invention provides a patient-reported outcome management method and system, which solves the problem that patient-reported outcomes cannot be updated in real time according to medical standards and contain outdated or erroneous content.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A patient-reported outcome management method comprises the following steps: Step S1: assign permissions to each role in the patient report to obtain the information visibility and operation boundaries of each role; Step S2: Based on the information visibility range and operation boundary of each role, collect the key operation behaviors of each role to obtain key operation behavior data, and obtain the role behavior chain by encoding the key operation behavior data in a unified format; Step S3: extracting indicators from the character behavior chain to obtain a content degradation factor, performing weight analysis on the content degradation factor using a random forest algorithm to obtain a content degradation factor weight, and calculating a content degradation index by combining the original value of the content degradation factor and the content degradation factor weight; Step S4: Collecting the field contents in the medical standard terminology library and the patient report, inputting the field contents in the medical standard terminology library and the patient report into the biomedical pre-training model, respectively obtaining corresponding medical standard term vectors and patient report term vectors, and calculating the medical standard term vectors and patient report term vectors using the cosine distance formula to obtain a medical compliance index; Step S5: Calculate a content optimization index by combining the content degradation index and the medical compliance index, compare the content optimization index with a preset threshold to obtain a comparison result, and classify the field content according to the comparison result to obtain a field content label; Step S6: When the field content is a content optimization item, the original field content is optimized to obtain the optimized field content, which is included in the optimization candidate pool. The optimized field content is tested by a comparative test method to obtain the experimental group indicators and the control group indicators. Based on the experimental group indicators and the control group indicators, the optimization effect index is calculated. Based on the optimization effect index, it is decided whether to extract the optimized field content from the optimization candidate pool and replace it with the original field content to obtain the final patient report.

[0006] Preferably, the step S1 includes: Assign permissions to each role in the patient report to obtain the information visibility and operation boundaries of each role, including patients, doctors, and managers; Patient role permissions: Information visibility: Subjective rating items: Patients can view their subjective ratings given in various questionnaires; Simple trend chart: Generate a simple trend chart for the patient to show the changing trend of certain key indicators of the patient over a period of time; Operational boundaries: Within the scope of their authority, patients can only confirm their subjective scores or re-score within the specified time. In addition, patients can view their own trend charts, but cannot modify the trend chart generation method or data; Doctor role permissions: Information visibility: Subjective scoring items and simple trend charts: Doctors can see all the subjective scoring items and simple trend charts that patients can see, which helps them understand the patient's initial assessment of their condition from the patient's perspective; Risk prediction score: Based on the patient's various questionnaire data, historical diagnosis and treatment records, and medical statistical models, the doctor generates a risk prediction score for the patient; Historical deviation range: The doctor is shown the deviation range of each patient's indicators from the normal range or the patient's own historical average value; Operational Boundaries: Doctors can modify labels on patient reports and annotate key data in patient reports. Doctors can view changes in patient data over different time periods and conduct comparative analysis. In addition, doctors can adjust the risk prediction model parameters in the system based on the patient's latest condition to improve the accuracy of risk prediction. Data manager role permissions: Information visibility: Data managers can view complete report data for all patients, including patients' subjective scores, trend charts, doctors' risk prediction scores, historical deviation intervals, and doctors' label modifications and annotations. They can access all data records stored in the system, including original questionnaire data, intermediate results during data processing, and various reports and analytical data generated in the end; Operational Boundaries: Data integrity review: Data management personnel are responsible for checking the integrity of patient report data and ensuring that all required fields have data entered; Problem field backtracking: When data anomalies or suspicious situations are found, data management personnel can backtrack the relevant fields.

[0007] Preferably, the step of encoding the key operation behavior data in a unified format to obtain a role behavior chain includes: The system standardizes the coding of each click, question skipping, modification, and delayed submission behavior of the patient while filling out the questionnaire; at the same time, the frequency of doctors reviewing patient reports, label modifications, and annotation behaviors are also included in the same behavior record chain; the field correction or suspicious marking behavior of the data management personnel is incorporated into the behavior chain as an independent node; all behaviors are indexed by timestamps and stored in a chain structure, and combined to form a role behavior chain. Each chain can trace back to core elements such as the source role, operation intention, affected field, and trigger time.

[0008] The encoding format is standardized as <operation type, operation object, timestamp>.

[0009] Preferably, the content degradation factor obtained based on the index extraction of the character behavior chain includes: Extract indicators from the character behavior chain to obtain content degradation factors: behavior density, redundant operation frequency, topic skipping frequency, early termination rate, and field repeated revision rate; Definition of behavior density: Without considering the time factor, behavior density can be defined as the ratio of the actual number of operations to the theoretical minimum number of operations, for patients, doctors, and data managers;

[0010] in, is the behavior density, It refers to the actual number of operations performed by patients, doctors, and data managers during the operation process; The theoretical minimum number of operations is the theoretical minimum number of operations required based on the structure of the operation. It represents the number of operations required to complete the task under the most ideal and simplified circumstances. Definition of redundant operation frequency: The ratio of meaningless repeated operations to the total number of operations, quantifying invalid interactions, for patients, doctors, and data managers; formula:

[0011] in, is the frequency of redundant operations, is the number of times you click the same option repeatedly. is the number of times the operation is repeated after undoing, It is the total number of operations; repeated confirmation of key operations in special scenarios is not counted as redundancy; Definition of field repetition revision rate: refers to the weighted average of the number of revisions of each field and its corresponding revision weight in a group of fields. This indicator is used to measure the frequency and importance of field revisions, targeting the modification behavior of patients, doctors, and data managers; formula:

[0012] in, is the field repetition correction rate, It is Number of field corrections, It is Field correction weights; is the total number of fields; modify the weight More weight is given to recent revisions; Definition of question skipping frequency: It refers to the ratio of the number of questions that a character actively skips in a set of questions to the number of questions that can actually be skipped. It is used to measure the frequency of the patient's question skipping behavior; formula:

[0013] in, is the frequency of skipping questions, is the number of active question skipping, is the number of logical jump questions, is the total number of questions; Definition of early termination rate: the proportion of incomplete questionnaires, combined with the completion progress to calculate the severity, targeting patient operation behavior; formula:

[0014] in, is the early termination rate, It ends at Number of patients with the problem; is the total number of questions in the questionnaire; is the number of ending positions, is the total number of patients; Preferably, the calculation of the content degradation index by combining the original value of the content degradation factor and the weight of the content degradation factor includes: Random forest algorithm training and weight determination: Feature Importance Matrix

[0015] in, is the number of random forest trees, is the characteristic dimension; Content degradation factor weight:

[0016] in, The final weight vector, each element of which corresponds to the weight of a content degradation factor; is the domain expert correction coefficient, which adjusts the weight to match clinical needs. , formal representation; is the number of trees in the random forest It is The first tree The importance value of a feature is expressed in the formula , , formal representation; Content degradation index calculation:

[0017] in, is the content degradation index, It is Content degradation factor values: ; is the weight corresponding to the i-th content degradation factor, is the time decay coefficient; is the number of days since the last optimization, The value range is ,The higher the value is, the more serious the quality of the content has deteriorated, that is, the more profound the role has ,modified the content of this patient report.

[0018] Preferably, the step of inputting the field content in the patient report and the medical standard terminology library into the biomedical pre-training model to obtain corresponding medical standard term vectors and patient report term vectors comprises: Construct the input feature vector: Word vector embedding: assign a unique word vector to each unit after word segmentation; Segment embeddings: Distinguish the source of input text: patient reports or standard medical terminology. For example, if both "dizziness" and "R42 dizziness and vertigo" are input, different segment embeddings are used to identify them, making it easier for the model to distinguish semantic context. Position vector embedding: Considering the position information of words in the text, such as "dizzy" at the beginning or in the middle of a sentence, its position encoding is different, helping the model understand the context order; Finally, the three embeddings are added together to form the model's input vector = word vector embedding + segment vector embedding + position vector embedding, which is then input into the BERT biomedical pre-training model. Model structure: BERT is based on a multi-layer bidirectional Transformer architecture. The biomedical pre-trained model is based on the general BERT and is further trained on a large amount of biomedical literature, which can more accurately capture the semantics of medical terminology. Feature extraction: The input vector passes through multiple layers of Transformer blocks in sequence, and each layer calculates the association between words through the self-attention mechanism; Output high-dimensional semantic vector: Take the hidden state of the last layer of the model as the semantic vector of the term; assuming the model output dimension is 768, then the terms such as "dizziness" and "fatigue" are finally converted into 768-dimensional vector patient report term vectors and Standard Medical Terminology .

[0019] Preferably, the medical compliance index obtained by calculating the medical standard term vector and the patient report term vector using the cosine distance formula includes: The average distance between the semantic vector of the term in the field and the standard term is calculated according to the cosine distance formula:

[0020] in, The semantic vector representing the terms in the field semantic vector field With the standard term vector The average semantic distance of is the number of standard terms, i.e., the same as the field Corresponding medical standard terminology set The number of terms in is the semantic vector representing the i-th term in the field; It is Semantic vectors of standard terms, Used to calculate the semantic vector of the i-th term in the field With the standard term vector The cosine similarity of the two words is used to measure the semantic similarity between them; Medical Compliance Index Calculation: Introducing the standard term vector space radius : )

[0021] in, Medical Compliance Index, The value range is ,The higher the value, the more the content complies with medical standards. Represents field semantic vector With the standard term vector The average semantic distance of is the radius of the introduced standard term vector space; for example, if a reported term completely matches a standard term, ,but , indicating that the field content is completely consistent with medical standards; if the deviation is extremely large, then d=r, , indicating that the field content does not comply with medical standards.

[0022] Preferably, the content optimization index calculated by combining the content degradation index and the medical compliance index includes: Combining the content degradation index and the medical compliance index, the content optimization index is calculated:

[0023] Among them, COI is the content optimization index; CDI is the content degradation index; MCI is the medical compliance index; is the weight coefficient, which is determined by the analytic hierarchy process.

[0024] Preferably, the optimization effect index calculated based on the experimental group index and the control group index includes: Optimization effect index calculation:

[0025] in To optimize the effect index, It is The weight of each indicator; The experimental group indicator values; : Control group indicator values; : No. The relative rate of change of an indicator.

[0026] Preferably, the step of determining whether to extract the optimized field content from the optimization candidate pool and replace it with the original field content based on the optimization effect index to obtain the final patient report includes: Version update decision: Replace: If , will optimize the field content Extract the original field content from the optimization candidate pool Replace and update patient report versions, records, time, strategies, and optimization effect indices; Fallback: If ,reserve , adjust the strategy and retest until it meets the target or is marked as an invalid strategy.

[0027] A patient-reported outcome management system, characterized in that the system comprises: The authority allocation module is used to allocate authority to each role in the patient report and obtain the information visibility and operation boundaries of each role; A role behavior chain generation module collects key operation behaviors of each role based on the information visibility range and operation boundaries of each role, obtains key operation behavior data, and obtains a role behavior chain by encoding the key operation behavior data in a unified format; a content degradation index generation module, configured to extract indicators from the character behavior chain to obtain a content degradation factor, perform weight analysis on the content degradation factor using a random forest algorithm to obtain a content degradation factor weight, and calculate a content degradation index by combining the original value of the content degradation factor and the content degradation factor weight; A medical compliance index generation module is configured to collect field contents from a medical standard terminology library and patient reports, input the field contents from the medical standard terminology library and patient reports into a biomedical pre-training model, obtain corresponding medical standard term vectors and patient report term vectors, respectively, and calculate the medical standard term vectors and patient report term vectors using a cosine distance formula to obtain a medical compliance index; a field content label generation module, configured to calculate a content optimization index by combining the content degradation index and the medical compliance index, compare the content optimization index with a preset threshold to obtain a comparison result, and classify the field content according to the comparison result to obtain a field content label; The decision-making module is used to optimize the original field content when the field content is a content optimization item, obtain the optimized field content, and include it in the optimization candidate pool. The optimized field content is tested by the comparative test method to obtain the experimental group indicators and the control group indicators. The optimization effect index is calculated based on the experimental group indicators and the control group indicators. Based on the optimization effect index, it is decided whether to extract the optimized field content from the optimization candidate pool and replace it with the original field content to obtain the final patient report.

[0028] Beneficial effects The present invention provides a patient-reported outcome management method involving machine learning and deep learning technologies, which has the following beneficial effects: (1) This patient-reported outcome management method clarifies the information visibility and operation boundaries through hierarchical role permissions, prevents data leakage and unauthorized operations, complies with medical data security standards, and improves system credibility.

[0029] (2) This patient-reported outcome management method uses a biomedical pre-trained model to vectorize patient-reported terms and medical standard terms, combines cosine distance to calculate semantic matching, and generates a medical compliance index, thereby automatically verifying whether the report content meets clinical standards. Compared with traditional manual verification, this process is efficient and complete, ensuring that patient-reported data can be directly used for clinical diagnosis and research, and reducing the cost of terminology calibration for medical staff.

[0030] (3) The patient report outcome management method uses the content optimization index to integrate the content degradation index and the medical compliance index, and then uses the threshold judgment to form a field content label. This mechanism breaks the limitations of single-dimensional evaluation and achieves a comprehensive quantitative analysis of content quality, ensuring that the optimization direction takes into account both role experience and medical compliance, avoiding the contradiction of improved experience but deviation from standard terminology or compliance but complex operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a flow chart of a patient-reported outcome management system proposed by the present invention.

[0032] Figure 2 This is another flow chart of a patient-reported outcome management system proposed by the present invention.

[0033] Figure 3 This is another flow chart of the patient-reported outcome management system proposed by the present invention. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0035] See also Figure 1 The present invention provides a technical solution: a patient-reported outcome management method. Specifically, the following patient-reported outcome management method is provided, please refer to Figure 1 , the method comprises the following steps: Step S1: Assign permissions to each role in the patient report to obtain the information visibility and operation boundaries of each role.

[0036] Assign permissions to each role in the patient report to obtain the information visibility and operation boundaries of each role, including patients, doctors, and data managers; Patient role permissions: Information visibility Subjective scoring items: Patients can view their subjective scores given in various questionnaires, such as the score of their pain level in the pain assessment questionnaire (0-10 points), and the subjective score of their sleep quality, emotional state, etc. in the quality of life assessment questionnaire.

[0037] Simple trend charts: Generate simple trend charts for patients to show the changing trends of certain key indicators over a period of time (such as the past month). For example, a line chart can be used to show the changes in the patient's average pain score per week, or a bar chart can be used to show the patient's monthly mood state (such as the number of days with positive emotions compared to negative emotions). Operational Boundaries: Within their authority, patients can only confirm their subjective scores or re-score within a specified timeframe (e.g., 24 hours). Additionally, patients can view their own trend graphs but cannot modify the graph generation method or data.

[0038] Doctor role permissions: Information visibility: Subjective scoring items and simple trend charts (inheriting the patient's perspective); doctors can see all the subjective scoring items and simple trend charts that patients can see, which helps doctors understand the patient's preliminary assessment of their own condition from the patient's perspective.

[0039] Risk Prediction Score: Based on various questionnaire data, historical medical records, and medical statistical models, a risk prediction score is generated for doctors. For example, in a cardiovascular disease risk assessment, the patient's age, blood pressure, blood lipids, lifestyle, and other data are used to calculate the patient's risk of cardiovascular disease within the next year (e.g., a 30% risk of developing the disease). This score is then presented to the doctor.

[0040] Historical Deviation Range: This displays the deviation range of each patient's indicators from the normal range or the patient's historical average. For example, for a patient's blood sugar indicator, the doctor can see the deviation of the patient's current blood sugar value from the normal blood sugar range (such as a fasting blood sugar of 3.96.1mmol / L) and the deviation from the patient's average blood sugar value over the past three months (such as a current blood sugar value that is 1.5mmol / L higher than the historical average).

[0041] Operational Boundaries: Doctors can modify labels in patient reports, such as changing a patient's condition from "stable" to "mild fluctuations," to better track changes in their condition. Doctors can also annotate key data in patient reports, such as noting "further investigation required" next to an abnormal test result or adding possible diagnostic options next to a patient's symptom description.

[0042] Doctors can view changes in patient data over different time periods and conduct comparative analysis. Furthermore, doctors can adjust the risk prediction model parameters in the system (within a reasonable range authorized by the system) based on the patient's latest condition to improve the accuracy of risk prediction.

[0043] Data manager role permissions Visibility: Data managers can view complete report data for all patients, including their subjective scores, trend charts, physician risk prediction scores, historical deviation intervals, and physician label modifications and annotations. They also have access to all data records stored in the system, including raw questionnaire data, intermediate results during data processing, and ultimately generated reports and analytical data.

[0044] Operational Boundaries: Data integrity review: Data management personnel are responsible for verifying the completeness of patient-reported data, ensuring that all required fields are entered. For example, in a 10-question patient questionnaire, they will verify that each question has been answered by the patient or has a reasonable default. If missing data are found, they will take appropriate action (such as contacting the patient or physician to supplement the data).

[0045] Problem Field Backtracking: When data anomalies or suspicious situations are discovered (e.g., an unreasonable extreme value for a patient's indicator), data managers can backtrack the relevant fields. They can review the entire data generation process for the field, including the patient's original input, the system's data processing algorithms, and any manual modifications, to determine the root cause of the problem.

[0046] Step S2: Based on the information visibility range and operation boundary of each role, collect the key operation behaviors of each role to obtain key operation behavior data, and obtain the role behavior chain by encoding the key operation behavior data in a unified format.

[0047] The system standardizes the coding of each click, question skipping, modification, delayed submission, and incomplete questionnaire behavior of the patient during the questionnaire filling process; at the same time, the frequency of doctors reviewing patient reports, label modifications, and annotation behaviors are also included in the same behavioral record chain; the field correction or suspicious marking behavior of the data management personnel is incorporated into the behavioral chain as an independent node.

[0048] Key operational behavior data: patient operational behavior code, doctor operational behavior code, and data management personnel operational behavior code.

[0049] All key operational behavior data are timestamp-indexed and stored in a chain structure, which is combined to form a role behavior chain. Each chain can trace back to core elements such as the source role, operational intention, impact field, and trigger time.

[0050] The patient's operation behavior is standardized and coded in a unified and standardized format as <operation type, operation object, timestamp>: Click behavior: Code each click performed by the patient on the questionnaire. For example, clicking on a questionnaire option or clicking the button to view a trend chart will be recorded. The coding format can be <action type: click, action object: questionnaire option A, timestamp: 2025052310:00:00>, where the action type clearly indicates a click and the action object indicates a specific option in the questionnaire. The timestamp records the exact time the action occurred.

[0051] Question skipping: This behavior is recorded when a patient skips a question in the questionnaire. Codes such as <Action Type: Skip, Action Target: Question 3, Timestamp: 2025052310:05:00> clearly indicate which question the patient skipped and when. Question skipping is categorized into active and logical jumps. Active skipping is recorded when a character changes the order of questions by clicking navigation buttons (such as "Previous Question," "Next Question," or "Go to Specific Question"). Specific event listeners can be added to these navigation buttons. Each click records the direction (forward or backward) and the number of questions skipped. Logical jumps are recorded when a character enters a different question based on logical conditions (e.g., the answer to the previous question determines whether the next question appears). During questionnaire design, flags can be set for each question with logical jump conditions. When the character enters or skips these questions, the corresponding jumps are recorded.

[0052] Modification behavior: If the patient modifies their own answers, record the content and time of the modification. For example, <Operation type: Modify, Operation object: Headache score (original answer: 5, new answer: 7), Timestamp: 2025052310:10:00>. This facilitates subsequent analysis of changes in the patient's self-assessment.

[0053] Delayed submission: When a patient submits a questionnaire beyond the specified time (e.g., a questionnaire requires completion within 30 minutes), this delay is recorded. Coded as <Action type: Delayed submission, Action object: Questionnaire (delay time: 15 minutes), Timestamp: 2025052310:35:00>. This helps understand the patient's efficiency in completing the questionnaire and any difficulties they may have.

[0054] Incomplete questionnaire behavior: When the patient encounters difficulties or gives up the questionnaire, the incomplete questionnaire behavior is recorded and coded as <operation type: incomplete questionnaire, operation object: questionnaire title, timestamp: 2025052310:35:00>, which helps to understand the patient's completion rate of the questionnaire.

[0055] The doctor's operation behavior is standardized and coded in a unified and standardized format of <operation type, operation object, timestamp>: Frequency of patient report review: This records the time and frequency of a doctor's review of patient reports. For example, <Operation Type: Review Report, Operation Target: Patient A's Report (Review Frequency: 3rd), Timestamp: 2025052314:00:00>. Analyzing the review frequency can help you understand the doctor's level of attention to the patient's condition.

[0056] Label modification behavior: When a doctor modifies the condition label in a patient report, the record format is such as <operation type: label modification, operation object: patient A's condition label (original label: stable; new label: fluctuating), timestamp: 2025052314:10:00>, which helps track changes in the doctor's judgment of the patient's condition.

[0057] Annotation behavior: Doctors' comments on patient reports are also recorded. For example, <Operation type: Annotation, Operation object: Patient A's test results (Annotation content: Recommended renal function reexamination), Timestamp: 2025052314:15:00>, making it easier to understand the doctor's diagnosis and treatment ideas and recommendations for further examinations.

[0058] The data management personnel's operation behavior is standardized and coded in a unified and standardized format of <operation type, operation object, timestamp>: Data modification behavior: When the data management personnel modify the data fields in the patient report, the record format is such as <Operation type: Field modification, Operation object: Patient A's blood glucose value (original data: 8.0mmol / L, new data: 7.5mmol / L), Timestamp: 2025052315:00:00>, which helps track data changes and ensure data accuracy.

[0059] Suspicious marking behavior: If the data management personnel find suspicious situations in patient data and mark them, the record format is such as <Operation type: Suspicious marking, Operation object: Patient B's blood pressure value (Suspicious reason: excessive fluctuation of the value), Timestamp: 2025052315:30:00>, so that the suspicious data can be further investigated and processed.

[0060] In addition, the character behavior chain not only contains standardized coding of character operation behaviors, but also records other related data: The actual number of operations performed by patients, doctors, and data managers during the operation ; Theoretically minimum number of operations required based on the structure of the operation ; Topic skipping behavior data , the number of logical jump problems ; Field modification data: Number of field modifications (The range of roles includes patients, doctors, and data managers); Redundant operation data (data derived from the operation behavior of patients, doctors, and data managers): records the number of times patients, doctors, and data managers repeatedly click on the same option and the number of redo operations after undo , and the total number of operations; Unfinished questionnaire behavior data: Record ends at Number of patient roles in the question , the total number of questionnaire questions n, The number of ending positions m.

[0061] Step S3: extract indicators from the role behavior chain to obtain a content degradation factor, perform weight analysis on the content degradation factor using a random forest algorithm to obtain a content degradation factor weight, and calculate a content degradation index by combining the original value of the content degradation factor and the content degradation factor weight.

[0062] Indicators were extracted from the character behavior chain to obtain content degradation factors: behavior density, frequency of redundant operations, frequency of question skipping, early termination rate, and rate of repeated field revisions. Essentially, through structured analysis of character interaction behaviors, these abstract factors—behavior complexity, ineffective interactions, frequency of field revisions, frequency of question skipping, and incomplete questionnaires—were converted into concrete, calculable, and actionable indicators. Behavior Density: Definition: Without considering time, behavior density can be defined as the ratio of the actual number of operations to the theoretical minimum number of operations, for patients, doctors, and data managers. Behavior density quantifies the complexity of different roles in the operation process. Higher behavior density values ​​indicate more complex or redundant operations, while lower behavior density values ​​indicate simpler and more efficient operations.

[0063]

[0064] in, is the behavior density, It refers to the actual number of operations performed by the roles (patients, doctors, data managers) during the operation process; Theoretical minimum number of operations, which is the theoretical minimum number of operations required based on the structure of the operation. It represents the number of operations required to complete the task under the most ideal and simplified circumstances.

[0065] Redundant Operation Frequency (RedundantOperationFrequency): Definition: The ratio of meaningless repeated operations to the total number of operations, quantifying invalid interactions, for patients, doctors, and data managers.

[0066] formula:

[0067] in, is the frequency of redundant operations, is the number of times you click the same option repeatedly. is the number of times the operation is repeated after undoing, It is the total number of operations; repeated confirmation of key operations (such as submission) in special scenarios is not counted as redundancy.

[0068] Field Revision Rate (FieldRevisionRate) Definition: This is the weighted average of the number of field revisions and their corresponding revision weights within a set of fields. This metric measures the frequency and importance of field revisions, targeting the modification behavior of patients, physicians, and data managers.

[0069] formula:

[0070] in, is the field repetition correction rate, It is Number of field corrections, It is The weight of each field is modified (the weight of the most recent modification is higher); is the total number of fields; modify the weight The most recent revision is given a higher weight (e.g. the last revision has a weight of 3, the previous one has a weight of 2, and the first one has a weight of 1).

[0071] Skipping Frequency Definition: It refers to the ratio of the number of questions a character actively skips in a set of questions to the number of questions actually available for skipping (i.e., the total number of questions minus the number of logical jump questions). It is used to measure the frequency of the patient's question-skipping behavior.

[0072] formula:

[0073] in, is the frequency of skipping questions, is the number of active question skipping, is the number of logical jump questions, It is the total number of questions, logical jump questions: such as questions that are automatically skipped based on the answer to the previous question (such as "Do you smoke → No" will skip tobacco use details).

[0074] Early Termination Rate is defined as the proportion of incomplete questionnaires, combined with the completion progress to calculate the severity, targeting patient operational behavior.

[0075] formula:

[0076] in, is the early termination rate, It ends at Number of patients with the problem; is the total number of questions in the questionnaire; is the number of ending positions, is the total number of patients; Random forest algorithm training and weight determination: Feature engineering time window: extract behavioral data within the past 30 days and perform stratified sampling by role type (patient, doctor, data manager).

[0077] Derived features: delay time interval and operation sequence entropy; Delay interval: the average time between two modifications to the same field (the shorter the interval, the higher the ambiguity); Operation sequence entropy: The semantic entropy of the operation sequence is calculated using the word vector model in NLP. A higher entropy value indicates more disordered behavior.

[0078] Random forest training process label generation: Positive samples: manually labeled fields that need to be optimized (for example, the frequency of skipping questions decreases by >20% after modification).

[0079] Negative samples: Fields with stable operation and good character feedback.

[0080] Parameter optimization: 5-fold cross validation was used to determine the optimal parameters (tree depth = 10, subsampling ratio = 0.7) through grid search.

[0081] Feature Importance Matrix

[0082] in, is the number of random forest trees, It is the feature dimension (5 content degradation factors + derived features).

[0083] Content degradation factor weight:

[0084] in, is the final weight vector, each element of which corresponds to the weight of a content degradation factor; is the domain expert correction coefficient, which adjusts the weight to match clinical needs. , formal representation; is the number of trees in the random forest It is Tree ( ) Features ( ) is expressed as follows (e.g., information gain, Gini index, etc.), which is expressed as , , formal representation; Content degradation index calculation:

[0085] in, is the content degradation index, It is Content degradation factor value ( ); is the weight corresponding to the i-th content degradation factor, is the time decay coefficient; is the number of days since the last optimization, The value range is ,The higher the value is, the more serious the quality of the content has deteriorated, that is, the more profound the role has ,modified the content of this patient report.

[0086] Through matrix and vector operations, the transformation from the random forest model to domain adaptation weights is achieved, providing a quantitative and interpretable feature weight system for content degradation analysis, ensuring that the algorithm conforms to both data patterns and medical practice.

[0087] Step S4: Collect the field content of the patient report and the medical standard terminology library, input the field content in the patient report and the medical standard terminology library into the biomedical pre-training model, obtain the corresponding patient report term vector and medical standard term vector respectively, calculate the patient report term vector and medical standard term vector using the cosine distance formula, and obtain the medical compliance index.

[0088] Data Collection: Collect field content in patient reports, including but not limited to text data such as answers to various questionnaires filled out by patients, symptom descriptions, subjective scores, etc.

[0089] Synchronously obtain clinical guidelines and diagnosis and treatment standard texts (such as "Clinical Outcome Evaluation Standards" and "Disease Coding Standards") to build a medical standard terminology library.

[0090] Build a medical term mapping dictionary to solve the problem of synonyms and near-synonyms: Introduce authoritative medical dictionaries (such as UMLS), embed synonym dictionaries (such as UMLS semantic network) in the medical standard terminology database, and establish a mapping relationship between patient report terms and standard terms. For example: The patient term "headache" → standard term "headache (R42)"; the colloquial expression "confusion" → standard term "dizziness and vertigo (R42)".

[0091] Establish a domain-specific term extension mechanism: For rare disease or emerging therapy terms, clinical experts are allowed to manually add mappings to form a dynamically updated terminology library. For example, a new therapy commonly known as "targeted drug A" can be mapped to the standard term "targeted drug (ATC code)".

[0092] Data preprocessing: Text cleaning: Clean the field content in the patient report (such as "dizziness lasting for two days" and "obvious fatigue") and medical standard terminology (such as "R42 dizziness and vertigo" and "R53 fatigue" in ICD10) to remove special symbols, typos and other interfering information, and unify the text format.

[0093] Tokenization: Split the text into individual words or subwords. For example, "dizziness" is treated as a single word, while a complex term like "paroxysmal headache" may be split into "paroxysmal," "sexual," and "headache" (based on the model's tokenization rules).

[0094] Construct the input feature vector: Word embedding: Each segmented unit is assigned a unique word vector. For example, "dizziness" corresponds to an ID in the model vocabulary, which is mapped to an initial vector through the vocabulary.

[0095] Segment embeddings: Distinguish the source of the input text (patient report or standard medical terminology). For example, if both "dizziness" (patient report) and "R42 dizziness and vertigo" (standard terminology) are input, different segment embeddings identify them, making it easier for the model to distinguish semantic context.

[0096] Positional embedding: This takes into account the position of words in the text. For example, if the word "dizzy" appears at the beginning of a sentence or in the middle of a sentence, its positional encoding will be different, helping the model understand the contextual order.

[0097] Finally, the three embeddings are added together to form the model's input vector = word vector embedding + segment vector embedding + position vector embedding, which is then input into the BERT biomedical pre-training model. Model structure: BERT is based on a multi-layer, bidirectional Transformer architecture. Biomedical pre-trained models (such as BioBERT and ClinicalBERT) are further trained on a large amount of biomedical literature (such as PubMed papers and clinical records) based on the general BERT, and can more accurately capture the semantics of medical terminology.

[0098] Feature extraction: The input vector passes through multiple layers of Transformer blocks, each layer using a self-attention mechanism to calculate relationships between words. For example, when processing the word "fatigue," the model focuses on its co-occurrence with terms like "anemia" and "chronic disease" in medical texts, thereby learning more specialized semantic representations.

[0099] Output high-dimensional semantic vector Hidden state acquisition: Take the hidden state of the last layer of the model as the semantic vector of the term. Assuming the model output dimension is 768, then the terms such as "dizziness" and "fatigue" are finally converted into 768-dimensional vectors. (patient-reported term) and (Standard Medical Terminology).

[0100] Through the above steps, patient report terms and medical standard terms are converted into computer-computable high-dimensional semantic vectors. This lays the foundation for the subsequent calculation of the Medical Compliance Index (MCI), ensuring that the content of patient report fields is semantically consistent with medical standards.

[0101] Set field Include Terms }, each term generates a semantic vector through the BERT biomedical model , is the semantic vector representing the i-th term in the field.

[0102] Standard term set matching: Each field Corresponding to the reference term set in medical standards (For example, the "Symptom Description" field corresponds to the ICD10 symptom code set), each standard term The vector is , according to the field Whether the patient report term belongs to the domain-specific term (manually judged and labeled) depends on the mapping mechanism between the field F and the standard term set S. If it does, the domain-specific term expansion mechanism is used. If not, the medical term mapping dictionary is used for mapping and matching.

[0103] The average distance between the semantic vector of the term in the field and the standard term is calculated according to the cosine distance formula:

[0104] in, Semantic vectors representing terms in a field With the standard term vector The average semantic distance of is the number of standard terms, i.e., the same as the field Corresponding medical standard terminology set The number of terms in is the semantic vector representing the i-th term in the field; It is Semantic vectors of standard terms, Used to calculate the semantic vector of the i-th term in the field With the standard term vector The cosine similarity measures the semantic similarity between the two.

[0105] Medical Compliance Index (MCI) calculation: Introducing the standard term vector space radius (dynamically setting r based on the actual distribution of standard term vectors): )

[0106] in, Medical Compliance Index, The value range is ,The higher the value, the more the content complies with medical standards. Represents field semantic vector With the standard term vector The average semantic distance of is the radius of the introduced standard term vector space; for example, if a reported term completely matches a standard term, ,but , indicating that the field content is completely consistent with medical standards; if the deviation is extremely large, then d=r, , indicating that the field content does not comply with medical standards.

[0107] Step S5: Calculate a content optimization index by combining the content degradation index and the medical compliance index, compare the content optimization index with a preset threshold to obtain a comparison result, and classify the field content according to the comparison result to obtain a field content label; Determine the weights of content degradation index and medical compliance index through hierarchical analysis method Weight analysis method: Expert judgment matrix: Let the criteria be (interactive experience, corresponding to CDI) and (Medical compliance, corresponding to MCI), building a judgment matrix : ; in, express relatively The importance of The interactive experience is 20% more important). for relatively importance.

[0108] Weight calculation: Interaction experience weight and medical compliance weight They correspond to:

[0109] in, is the weight of the interactive experience, Medical compliance weight, k is the importance scaling factor set by the expert. By normalizing the matrix row sum, the expert's importance judgment is converted into a quantitative weight to ensure .

[0110] Combining the content degradation index and the medical compliance index, the content optimization index is calculated:

[0111] Among them, COI is the content optimization index; CDI is the content degradation index (range 0-100, reflecting interactive defects such as skipping topics and repeated revisions); MCI is the medical compliance index (range 0-100, reflecting semantic deviations from clinical standards); is the weight coefficient, determined by the Analytic Hierarchy Process (AHP), for example: clinical scenario (such as intensive care): =0.4 (prioritizing medical compliance); patient experience scenarios (such as chronic disease management): =0.6 (interaction optimization first).

[0112] Threshold rules and labels: Indicates a global threshold (e.g., 70, determined through statistical analysis of historical data to ensure that more than 90% of optimization requirements are covered).

[0113] Single threshold division: If COI , mark the field content as "content optimized item"; otherwise mark it as "content qualified item".

[0114] The weights are determined by AHP and labels are divided by a single threshold, thus achieving efficient and explainable determination of content optimization needs.

[0115] Step S6: When the field content is a content optimization item, the original field content is optimized to obtain the optimized field content, which is included in the optimization candidate pool. The optimized field content is tested by a comparative test method to obtain the experimental group indicators and the control group indicators. Based on the experimental group indicators and the control group indicators, the optimization effect index is calculated. Based on the optimization effect index, it is decided whether to extract the optimized field content from the optimization candidate pool and replace it with the original field content to obtain the final patient report.

[0116] Strategy matching: When field content is marked as a content optimization item, based on issues in the field content (such as interaction defects and medical compliance issues), preset strategies (such as terminology calibration and interface simplification) are called to generate optimized field content. , store it in the optimization candidate pool, record the strategy type (such as "interface simplification and folding of non-mandatory items") and expected indicators (such as reducing the skipping rate by 10%).

[0117] A / B testing implementation (comparison testing method) Traffic distribution: Experimental Group (Group B): 50% of roles (e.g., role ID hash modulo 049) use optimized field content .

[0118] Control group (Group A): 50% of roles (hashed modulo 5099) use the original field content , the hash algorithm is used to ensure that the grouping is unbiased.

[0119] Indicators of the experimental group and the control group: Data collection: role experience: skipping rate, interaction time, and submission success rate.

[0120] Data quality: field correction rate, medical compliance index.

[0121] System performance: page loading time, API response latency.

[0122] Optimization Effect Index (OEI) calculation:

[0123] in To optimize the effect index, It is The weight of each indicator (dynamic adjustment to meet , such as clinical scenarios higher); The experimental group (Group B) indicator values; : Control group (Group A) indicator values; : No. The relative rate of change of an indicator (positive means improvement, negative means decrease).

[0124] Version update decision: Replace: If (like , historical data calibration), will optimize the field content Extract the original field content from the optimization candidate pool Replace, update patient report versions, record logs (time, strategy, OEI).

[0125] Fallback: If , retain the original field content, adjust the strategy (such as modifying the prompt copy), and retest until it meets the requirements or is marked as an invalid strategy.

[0126] The present invention provides a full-process intelligent patient-reported outcome management method. Through role-based authority control, behavior chain analysis, multi-index fusion, and automated optimization and iteration, it realizes the security management, quality inspection, compliance verification, and autonomous evolution of PRO content, addressing the shortcomings of traditional methods in terms of authority, interaction, compliance, and iteration.

[0127] The present invention provides a patient-reported outcome management system, characterized in that the system comprises: The authority allocation module is used to allocate authority to each role in the patient report and obtain the information visibility and operation boundaries of each role; A role behavior chain generation module collects key operation behaviors of each role based on the information visibility range and operation boundaries of each role, obtains key operation behavior data, and obtains a role behavior chain by encoding the key operation behavior data in a unified format; a content degradation index generation module, configured to extract indicators from the character behavior chain to obtain a content degradation factor, perform weight analysis on the content degradation factor using a random forest algorithm to obtain a content degradation factor weight, and calculate a content degradation index by combining the original value of the content degradation factor and the content degradation factor weight; A medical compliance index generation module is configured to collect field contents from a medical standard terminology library and patient reports, input the field contents from the medical standard terminology library and patient reports into a biomedical pre-training model, obtain corresponding medical standard term vectors and patient report term vectors, respectively, and calculate the medical standard term vectors and patient report term vectors using a cosine distance formula to obtain a medical compliance index; a field content label generation module, configured to calculate a content optimization index by combining the content degradation index and the medical compliance index, compare the content optimization index with a preset threshold to obtain a comparison result, and classify the field content according to the comparison result to obtain a field content label; The decision-making module is used to optimize the original field content when the field content is a content optimization item, obtain the optimized field content, and include it in the optimization candidate pool. The optimized field content is tested by the comparative test method to obtain the experimental group indicators and the control group indicators. The optimization effect index is calculated based on the experimental group indicators and the control group indicators. Based on the optimization effect index, it is decided whether to extract the optimized field content from the optimization candidate pool and replace it with the original field content to obtain the final patient report.

[0128] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions. The sentence "including an element defined by..." does not exclude the presence of other identical elements in the process, method, article or device that includes the element."

[0129] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for managing patient-reported outcomes, characterized in that: The method comprises the following steps: Step S1: assign permissions to each role in the patient report to obtain the information visibility and operation boundaries of each role; Step S2: Based on the information visibility range and operation boundary of each role, collect the key operation behaviors of each role to obtain key operation behavior data, and obtain the role behavior chain by encoding the key operation behavior data in a unified format; Step S3: extracting indicators from the character behavior chain to obtain a content degradation factor, performing weight analysis on the content degradation factor using a random forest algorithm to obtain a content degradation factor weight, and calculating a content degradation index by combining the original value of the content degradation factor and the content degradation factor weight; Step S4: Collecting the field contents in the medical standard terminology library and the patient report, inputting the field contents in the medical standard terminology library and the patient report into the biomedical pre-training model, respectively obtaining corresponding medical standard term vectors and patient report term vectors, and calculating the medical standard term vectors and patient report term vectors using the cosine distance formula to obtain a medical compliance index; Step S5: Calculate a content optimization index by combining the content degradation index and the medical compliance index, compare the content optimization index with a preset threshold to obtain a comparison result, and classify the field content according to the comparison result to obtain a field content label; Step S6: When the field content is a content optimization item, the original field content is optimized to obtain the optimized field content, which is included in the optimization candidate pool. The optimized field content is tested by a comparative test method to obtain the experimental group indicators and the control group indicators. Based on the experimental group indicators and the control group indicators, the optimization effect index is calculated. Based on the optimization effect index, it is decided whether to extract the optimized field content from the optimization candidate pool and replace it with the original field content to obtain the final patient report.

2. A patient-reported outcome management method according to claim 1, characterized in that: The step S1 comprises: Assign permissions to each role in the patient report to obtain the information visibility and operation boundaries of each role, including patients, doctors, and managers; Patient role permissions: Information visibility: Subjective rating items: Patients can view their subjective ratings given in various questionnaires; Simple trend chart: Generate a simple trend chart for the patient to show the changing trend of certain key indicators of the patient over a period of time; Operational boundaries: Within the scope of their authority, patients can only confirm their subjective scores or re-score within the specified time. In addition, patients can view their own trend charts, but cannot modify the trend chart generation method or data; Doctor role permissions: Information visibility: Subjective scoring items and simple trend charts: Doctors can see all the subjective scoring items and simple trend charts that patients can see, which helps them understand the patient's initial assessment of their condition from the patient's perspective; Risk prediction score: Based on the patient's various questionnaire data, historical diagnosis and treatment records, and medical statistical models, the doctor generates a risk prediction score for the patient; Historical deviation range: The doctor is shown the deviation range of each patient's indicators from the normal range or the patient's own historical average value; Operational Boundaries: Doctors can modify labels on patient reports and annotate key data in patient reports. Doctors can view changes in patient data over different time periods and conduct comparative analysis. In addition, doctors can adjust the risk prediction model parameters in the system based on the patient's latest condition to improve the accuracy of risk prediction. Data manager role permissions: Information visibility: Data managers can view complete report data for all patients, including patients' subjective scores, trend charts, doctors' risk prediction scores, historical deviation intervals, and doctors' label modifications and annotations. They can access all data records stored in the system, including original questionnaire data, intermediate results during data processing, and various reports and analytical data generated in the end; Operational Boundaries: Data integrity review: Data management personnel are responsible for checking the integrity of patient report data and ensuring that all required fields have data entered; Problem field backtracking: When data anomalies or suspicious situations are found, data management personnel can backtrack the relevant fields.

3. A patient-reported outcome management method according to claim 2, characterized in that: The key operation behavior data is encoded in a unified format to obtain a role behavior chain, including: The system standardizes the coding of every click, question skipping, modification, delayed submission, and incomplete questionnaire behavior by patients during the questionnaire filling process. The frequency of doctors reviewing patient reports, label modifications, and annotations are also incorporated into the same behavioral record chain. Field corrections or suspicious flagging by data managers are incorporated into the behavioral chain as independent nodes. All behaviors are indexed by timestamps and stored in a chain structure, resulting in a role-behavior chain. Each chain can trace back to core elements such as the source role, operation intention, affected fields, and trigger time. The encoding format is standardized as <operation type, operation object, timestamp>.

4. A patient-reported outcome management method according to claim 3, characterized in that: The extracting of indicators based on the character behavior chain to obtain a content degradation factor includes: Extract indicators from the character behavior chain to obtain content degradation factors: behavior density, redundant operation frequency, topic skipping frequency, early termination rate, and field repeated revision rate; Behavior density: Definition: Without considering the time factor, behavior density can be defined as the ratio of the actual number of operations to the theoretical minimum number of operations, for patients, doctors, and data managers. ; in, is the behavior density, It refers to the actual number of operations performed by patients, doctors, and data managers during the operation process; The theoretical minimum number of operations is the theoretical minimum number of operations required based on the structure of the operation. It represents the number of operations required to complete the task under the most ideal and simplified circumstances. Operation redundancy frequency: Definition: The ratio of meaningless repeated operations to the total number of operations, quantifying invalid interactions, for patients, doctors, and data managers; formula: ; in, is the frequency of redundant operations, is the number of times you click the same option repeatedly. is the number of times the operation is repeated after undoing, It is the total number of operations; repeated confirmation of key operations in special scenarios is not counted as redundancy; Definition of field repetition revision rate: refers to the weighted average of the number of revisions of each field and its corresponding revision weight in a group of fields. This indicator is used to measure the frequency and importance of field revisions, targeting the modification behavior of patients, doctors, and data managers; formula: ; in, is the field repetition correction rate, It is Number of field corrections, It is Field correction weights; is the total number of fields; modify the weight More weight is given to recent revisions; Definition of question skipping frequency: It refers to the ratio of the number of questions that a character actively skips in a set of questions to the number of questions that can actually be skipped. It is used to measure the frequency of the patient's question skipping behavior; formula: ; in, is the frequency of skipping questions, is the number of active question skipping, is the number of logical jump questions, is the total number of questions; Definition of early termination rate: the proportion of incomplete questionnaires, combined with the completion progress to calculate the severity, targeting patient operation behavior; formula: ; in, is the early termination rate, It ends at Number of patients with the problem; is the total number of questions in the questionnaire; is the number of ending positions, is the total number of patients.

5. A patient-reported outcome management method according to claim 4, characterized in that: The content degradation index is calculated by combining the original value of the content degradation factor and the weight of the content degradation factor, including: Random forest algorithm training and weight determination: Feature Importance Matrix ; in, is the number of random forest trees, is the characteristic dimension; Content degradation factor weight: ; in, The final weight vector, each element of which corresponds to the weight of a content degradation factor; is the domain expert correction coefficient, which adjusts the weight to match clinical needs. , formal representation; is the number of trees in the random forest It is The first tree The importance value of the feature, , in the formula , , formal representation; Content degradation index calculation: ; in, is the content degradation index, It is Content degradation factor values: ; is the weight corresponding to the i-th content degradation factor, is the time decay coefficient; is the number of days since the last optimization, The value range is ,The higher the value is, the more serious the quality of the content has deteriorated, that is, the more profound the role has ,modified the content of this patient report.

6. A patient-reported outcome management method according to claim 5, characterized in that: The field content in the patient report and the medical standard terminology library are input into the biomedical pre-training model to obtain the medical standard term vector and the patient report term vector respectively, including: Construct the input feature vector: Word vector embedding: assign a unique word vector to each unit after word segmentation; Segment embeddings: Distinguish the source of input text: patient reports or standard medical terminology. For example, if both "dizziness" and "R42 dizziness and vertigo" are input, different segment embeddings are used to identify them, making it easier for the model to distinguish semantic context. Position vector embedding: Considering the position information of words in the text, such as "dizzy" at the beginning or in the middle of a sentence, its position encoding is different, helping the model understand the context order; Finally, the three embeddings are added together to form the model's input vector = word vector embedding + segment vector embedding + position vector embedding, which is then input into the BERT biomedical pre-training model. Model structure: BERT is based on a multi-layer bidirectional Transformer architecture. The biomedical pre-trained model is based on the general BERT and is further trained on a large amount of biomedical literature, which can more accurately capture the semantics of medical terminology. Feature extraction: The input vector passes through multiple layers of Transformer blocks in sequence, and each layer calculates the association between words through the self-attention mechanism; Output high-dimensional semantic vector: Take the hidden state of the last layer of the model as the semantic vector of the term; assuming the model output dimension is 768, then the terms "dizziness" and "fatigue" are finally converted into 768-dimensional vector patient report term vectors and Standard Medical Terminology .

7. A patient-reported outcome management method according to claim 6, characterized in that: The medical compliance index is obtained by calculating the medical standard term vector and the patient report term vector using the cosine distance formula, including: The average distance between the semantic vector of the term in the field and the standard term is calculated according to the cosine distance formula: ; in, Semantic vectors representing terms in a field With the standard term vector The average semantic distance of is the number of standard terms, i.e., the same as the field Corresponding medical standard terminology set The number of terms in is the semantic vector representing the i-th term in the field; It is Semantic vectors of standard terms, Used to calculate the semantic vector of the i-th term in the field With the standard term vector The cosine similarity of the two words is used to measure the semantic similarity between them; Medical Compliance Index Calculation: Introducing the standard term vector space radius : ) ; in, Medical Compliance Index, The value range is ,The higher the value, the more the content complies with medical standards. Represents field semantic vector With the standard term vector The average semantic distance of is the radius of the introduced standard term vector space; for example, if a reported term completely matches a standard term, ,but , indicating that the field content is completely consistent with medical standards; if the deviation is extremely large, then d=r, , indicating that the field content does not comply with medical standards; in, Medical Compliance Index, The value range is , the higher the value, the more the content complies with medical standards Represents field semantic vector With all standard term vectors The average semantic distance of is the radius of the introduced standard term vector space; for example, if a reported term completely matches a standard term, ,but , indicating that the field content is completely consistent with medical standards; if the deviation is extremely large, then , , indicating that the field content does not comply with medical standards.

8. A patient-reported outcome management method according to claim 7, characterized in that: The content optimization index is calculated by combining the content degradation index and the medical compliance index, including: Combining the content degradation index and the medical compliance index, the content optimization index is calculated: ; Among them, COI is the content optimization index; CDI is the content degradation index; MCI is the medical compliance index; is the weight coefficient, which is determined by the analytic hierarchy process.

9. A patient-reported outcome management method according to claim 8, characterized in that: The optimization effect index is calculated based on the experimental group index and the control group index, including: Optimization effect index calculation: ; in To optimize the effect index, It is The weight of each indicator; The experimental group indicator values; : Control group indicator values; : No. The relative rate of change of an indicator.

10. A patient-reported outcome management system, characterized in that: The system comprises: The authority allocation module is used to allocate authority to each role in the patient report and obtain the information visibility and operation boundaries of each role; A role behavior chain generation module collects key operation behaviors of each role based on the information visibility range and operation boundaries of each role, obtains key operation behavior data, and obtains a role behavior chain by encoding the key operation behavior data in a unified format; a content degradation index generation module, configured to extract indicators from the character behavior chain to obtain a content degradation factor, perform weight analysis on the content degradation factor using a random forest algorithm to obtain a content degradation factor weight, and calculate a content degradation index by combining the original value of the content degradation factor and the content degradation factor weight; A medical compliance index generation module is configured to collect field contents from a medical standard terminology library and patient reports, input the field contents from the medical standard terminology library and patient reports into a biomedical pre-training model, obtain corresponding medical standard term vectors and patient report term vectors, respectively, and calculate the medical standard term vectors and patient report term vectors using a cosine distance formula to obtain a medical compliance index; a field content label generation module, configured to calculate a content optimization index by combining the content degradation index and the medical compliance index, compare the content optimization index with a preset threshold to obtain a comparison result, and classify the field content according to the comparison result to obtain a field content label; The decision-making module is used to optimize the original field content when the field content is a content optimization item, obtain the optimized field content, and include it in the optimization candidate pool. The optimized field content is tested by the comparative test method to obtain the experimental group indicators and the control group indicators. The optimization effect index is calculated based on the experimental group indicators and the control group indicators. Based on the optimization effect index, it is decided whether to extract the optimized field content from the optimization candidate pool and replace it with the original field content to obtain the final patient report.

Citation Information

Cited By

  • Intelligent processing and analysis feedback system for gastric cancer PROs data

    CN122245576A

  • A smart processing and analysis feedback system for gastric cancer PROs data

    CN122245576B