Human genetic resource management full-process management and control risk assessment method and system
By integrating OCR, natural language processing, and machine learning technologies, an intelligent risk assessment system for human genetic resource management was constructed, which solved the problems of low data integration efficiency, insufficient interpretability, and poor model adaptability, and achieved efficient and accurate risk assessment and dynamic adjustment.
Patent Information
- Application Number
- CN202510879401.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies for human genetic resource management suffer from problems such as low data integration efficiency, insufficient interpretability, poor model adaptability, and difficulty in processing unstructured data, resulting in low accuracy and efficiency in risk assessment.
By employing OCR recognition, natural language processing, rule engine, structured indicator system, LightGBM integrated modeling, and SHAP interpretability analysis technology, an intelligent assessment mechanism is constructed to achieve structured output, visual presentation, and dynamic optimization of project risks.
It improves the scientific rigor, accuracy, and compliance of risk assessment, enhances data processing efficiency, strengthens the applicability and flexibility of the system, and provides stronger interpretability and predictive accuracy.
Smart Images

Figure CN121458024A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and in particular to a method and system for risk assessment and control of the entire process of human genetic resource management. Background Technology
[0002] Human genetic resource management involves multiple aspects, including the collection, preservation, use, and export of genetic resources. Each step involves complex administrative approvals and biosafety regulations, exhibiting high complexity and sensitivity. Under the current management model, administrative approvals for related operations rely heavily on manual processes, resulting in low efficiency and insufficient accuracy in risk identification. Manual approval not only suffers from high error rates and low efficiency but also struggles to comprehensively analyze and respond promptly to the impact of regulatory provisions and major events. Furthermore, the diverse and fragmented nature of the data involved in human genetic resource management, coupled with inherent quality issues, poses a significant challenge to comprehensive and effective risk assessment.
[0003] In various administrative approval activities, the existing "instant approval" model can significantly improve the efficiency of administrative approvals by integrating cross-departmental data through blockchain technology. This model relies on machine learning to dynamically optimize risk control thresholds, automatically intercepting high-risk applications, achieving rapid approval and intelligent verification, and greatly improving approval efficiency and accuracy. Although this model has made breakthroughs in other fields, due to the unique characteristics of human genetic resource management, including the complexity of data types and the high dependence of approval standards on policies and regulations, existing technologies have not effectively solved these risk assessment and approval bottlenecks, specifically in the following ways:
[0004] 1. Low data integration efficiency: Existing technologies mostly rely on manual summarization of historical data, which results in outdated information and cannot dynamically respond to the impact of regulatory changes or major events. This leads to delayed and inadequate risk control decisions in the approval process and makes it impossible to grasp potential risks in real time.
[0005] 2. Insufficient interpretability and feedback mechanisms: Most existing risk assessment systems in other fields are difficult to interpret risk conclusions, lack targeted intervention suggestions and result tracking methods, and are not convenient for applicant agencies to modify materials.
[0006] 3. Insufficient model adaptability: Most existing assessment models are based on weighted scoring methods, which cannot fully capture the nonlinear relationships and interaction effects between risk indicators, resulting in low accuracy of risk prediction and inability to adapt to changes in different business scenarios.
[0007] 4. Data sources are scattered and the types are complex: Traditional OCR technology has a high recognition error rate when processing unstructured data such as complex tables and handwriting, and still requires manual correction. It cannot efficiently process approval matters that include unstructured materials such as experimental records and review opinions.
[0008] Therefore, there is an urgent need for a system solution that integrates document parsing, intelligent rule modeling, risk prediction, and interpretability analysis to achieve intelligent evaluation, dynamic adjustment, and auxiliary judgment of approval items throughout the entire process. Summary of the Invention
[0009] This invention provides a method and system for risk assessment and control of the entire process of human genetic resource management. The system integrates technologies such as OCR recognition, natural language processing, rule engine, structured indicator system construction, LightGBM-based integrated modeling and SHAP interpretability analysis to build an intelligent assessment mechanism that can be used to automatically identify high-risk projects in administrative approval processes. It realizes the structured output, visualization and dynamic optimization of project risks, and improves the scientificity, accuracy and compliance of approval decisions.
[0010] To achieve the above objectives, the technical solution of the present invention includes the following:
[0011] A method for risk assessment and management of the entire process of human genetic resource management, the method comprising:
[0012] Extract key fields from the application report;
[0013] Based on the key fields, and according to the defined risk dimensions and the judgment rules for the indicators under the risk dimensions, a vector representation of the application report is generated;
[0014] The risk assessment result of the application report is obtained by classifying the application report based on its vector representation.
[0015] Furthermore, extract key fields from the application report, including:
[0016] OCR text recognition technology was used to extract the text content from the application report;
[0017] The method utilizes regular expression rules combined with dependency parsing to remove noise from redundant text content, including redundant headers, footers, and watermarks.
[0018] A domain-adaptive named entity recognition model is used to identify key fields in text content after noise removal. Specifically, for tables and semi-structured documents in the text content, key fields are extracted by cell label recognition, table structure reconstruction, and a combination of rule engine and field alignment rules.
[0019] Furthermore, the key fields include: project name, organization, ethics information, and approval number.
[0020] Furthermore, the risk dimensions include: administrative service risk, business processing risk, processing timeliness risk, deviation risk, and other compliance risks. Specifically, the primary indicators under administrative service risk include: research output risk, institutional credit, and institutional information quality; the primary indicators under business processing risk include: material completeness, material compliance, consistency of attachments, and cooperation information; the primary indicators under processing timeliness risk include: process delays and application lag; the primary indicators under deviation risk include: approval cycle and approval quality; and the primary indicators under other compliance risks include: ethical review, data security, consistency of purpose of use, and experimental protocol.
[0021] Furthermore, the secondary sub-indicators under the research output risk include: the proportion of professional titles, hardware facilities, and the number of unfinished projects; wherein, the determination rule for the proportion of professional titles is described as the proportion of senior professional titles being less than a first set value, the determination rule for hardware facilities is described as the number of equipment types being less than a second set value or the total number of equipment being less than a third set value, and the determination rule for the number of unfinished projects is described as the number of unfinished projects being greater than a fourth set value within a specified time range.
[0022] The secondary sub-indicators under the unit credit system include: administrative penalty records, notifications of scientific research misconduct, and records of dishonesty; wherein, the determination rule for administrative penalty records is described as any unit having a penalty announcement record, the determination rule for notifications of scientific research misconduct is described as units involved in notifications from the National Natural Science Foundation of China, and the determination rule for records of dishonesty is described as units being subject to enforcement by the court for dishonesty.
[0023] The secondary sub-indicators under the unit information quality include: incomplete information and inconsistent information. The judgment rule for incomplete information is described as missing fields in the system filing information, and the judgment rule for inconsistent information is described as the actual information not matching the filing.
[0024] The secondary sub-indicators under the material integrity include: missing core proof; wherein, the judgment rule for missing core proof is described as missing legal person authorization letter or missing ethics approval document;
[0025] The secondary sub-indicators under the material compliance include: naming / format anomalies; wherein, the judgment rule for naming / format anomalies is described as no version number or broken format;
[0026] The secondary sub-indicators under the attachment consistency include: incomplete protocol upload; wherein, the determination rule for incomplete protocol upload is described as the number that should be uploaded is not equal to the actual number;
[0027] The secondary sub-indicators under the cooperation information include: missing legal person / qualification; wherein, the judgment rule for missing legal person / qualification is described as lacking qualification certificates or not being stamped;
[0028] The secondary sub-indicators under the process delay include: system node lag; wherein, the judgment rule for system node lag is described as the duration of a certain approval node being greater than a fifth set value;
[0029] The secondary sub-indicators under the application delay include: abnormal submission cycle; wherein, the judgment rule for the abnormal submission cycle is described as the time from the initial registration of the project to the submission being greater than the sixth set value;
[0030] The secondary sub-indicators under the approval cycle include: deviation above the historical average; wherein, the judgment rule for deviation above the historical average is described as the current approval non-acceptance time being greater than the seventh set value;
[0031] The secondary sub-indicators under the approval quality include: significant differences from similar projects; wherein, the judgment rule for significant differences from similar projects is described as the only project that failed to pass among similar projects;
[0032] The secondary sub-indicators under the ethical review include: missing or abnormal content; wherein, the judgment rule for missing or abnormal content is described as not uploaded or the content not signed or stamped.
[0033] The secondary sub-indicators under the data security framework include: failure to report outbound travel; wherein, the determination rule for failure to report outbound travel is described as involving cross-border travel but without providing approval documentation;
[0034] The secondary sub-indicators under the purpose of use include: vague / incorrectly stated research purpose; wherein, the judgment rule for vague / incorrectly stated research purpose is described as the purpose being inconsistent with the ethical approval document;
[0035] The secondary sub-indicators under the consistency of the experimental scheme include: inconsistent sample information; wherein, the judgment rule for inconsistent sample information is described as the uploaded sample type not matching the scheme.
[0036] Furthermore, based on the aforementioned key fields, and according to the defined risk dimensions and the judgment rules for indicators under those risk dimensions, a vector representation of the application report is generated, including:
[0037] Based on the content of the key fields and the defined risk dimensions, extract the features from the application report;
[0038] The feature is vectorized by combining the corresponding judgment rules;
[0039] By combining the vectorization results of each feature, the vector representation of the application report is obtained.
[0040] Furthermore, based on the vector representation of the application report, a classification is performed to obtain the risk assessment result of the application report, including:
[0041] Build a risk assessment model based on the LightGBM model, XGBoost model, or CatBoost model;
[0042] Input the vector representation of the application report into the risk assessment model to obtain the overall risk score and the dimensional sub-risk scores;
[0043] The contribution of all features to the overall risk score is calculated using the SHAP algorithm to obtain the feature explanation. The feature explanation includes: the top few features that have the greatest impact on the overall risk score and the impact of these features.
[0044] The risk assessment result of the application report is obtained by combining the overall risk score, the dimensional sub-risk scores, and the feature interpretation.
[0045] Furthermore, after synthesizing the overall risk score, the dimensional sub-risk scores, and the feature interpretations to obtain the risk assessment result of the application report, the following additional steps are included:
[0046] If the risk assessment results of the application report differ from the expert review opinions, experts will be re-selected to review the risk assessment results of the application report.
[0047] If the review results show that the risk assessment results of the application report are incorrect, adjust the risk assessment model based on the review results and mark the solution in the system update log.
[0048] Furthermore, after classifying the application report based on its vector representation to obtain the risk assessment result, the process also includes:
[0049] A heat map is generated based on the risk assessment results of the application report, and the heat map is used to display the risk level;
[0050] A historical risk change trend is generated based on the risk assessment results of the application report.
[0051] A risk assessment system for the entire process of human genetic resource management, the system comprising:
[0052] The field extraction module is used to extract key fields from the application report;
[0053] The vector generation module is used to generate a vector representation of the application report based on the key fields and the judgment rules of the indicators under the set risk dimensions.
[0054] The result generation module is used to classify the application report based on its vector representation to obtain a risk assessment of the application report.
[0055] Compared with the prior art, the present invention has at least the following beneficial effects.
[0056] 1. Breakthrough in risk identification methods.
[0057] Existing technologies largely rely on static rule bases and human experience. Most risk assessment models employ scoring models or linear weighting methods, failing to effectively capture the non-linear relationships and interactions between indicators, resulting in assessment results lacking sensitivity and adaptability. This invention introduces a risk assessment model based on ensemble learning algorithms such as LightGBM, capable of automatically learning complex patterns in historical data and accurately identifying hidden high-risk factors. Its built-in feature importance assessment has demonstrated value when dealing with the cross-influence of multi-dimensional indicators, while the combination with model interpretation algorithms provides finer-grained interpretability for individual prediction results, thus possessing stronger interpretability and prediction accuracy.
[0058] 2. Improvement in data processing and integration capabilities.
[0059] Traditional risk assessment systems primarily process structured data, but have limited capabilities in handling unstructured materials such as PDFs, scanned documents, and handwritten forms, often requiring significant manual intervention. This invention integrates advanced natural language processing technologies such as OCR recognition, document parsing, named entity recognition, and relation extraction. It can automatically parse various unstructured application materials, including ethical review opinions, experimental record forms, and informed consent forms, enabling structured storage and verification of information. This effectively reduces the workload of manual proofreading and significantly improves data processing efficiency.
[0060] 3. It has a greater advantage in the design and flexibility of the indicator system.
[0061] Existing indicator systems generally suffer from limitations such as single-dimensionality and lack of flexibility, making it difficult to cover the complex scenarios involved in the approval of human genetic resources. In contrast, the indicator system constructed in this invention is based on regulations, expert opinions, and historical cases, covering multiple dimensions including administrative service quality, material completeness, ethical compliance, and historical behavioral deviations, and supports dynamic expansion and rule customization. The system can flexibly switch indicator weights and evaluation logic according to different approval types, significantly improving the system's applicability and stability in real-world business scenarios.
[0062] 4. It is more forward-looking in terms of dynamic adjustment and continuous learning mechanisms.
[0063] Traditional system evaluation models are outdated and lack feedback mechanisms, failing to optimize in a timely manner based on actual approval results or policy changes, leading to evaluation results that are out of touch with actual business operations. This system, however, introduces a scenario feature label library and a feedback learning mechanism, enabling continuous optimization of model parameters and weight settings based on data such as approver opinions and deviations in approval results, achieving true "self-evolution." Especially when faced with new legal policies or changes in regulatory standards, the system can quickly adjust its evaluation logic without completely retraining the model, effectively ensuring the system's long-term effectiveness and flexible responsiveness.
[0064] 5. It is more practical in terms of presenting risk assessment results and providing intervention support.
[0065] Existing systems typically only provide risk scores, lacking interpretability and business intervention suggestions, making it difficult to assist approvers in making practical judgments. This system features a visual risk dashboard and decision support module, which not only displays the scoring results but also analyzes the impact of each indicator on the risk score using model interpretation algorithms, and provides operational suggestions for different risk levels. This significantly enhances the system's business usability and transparency, shortens approval decision-making time, and improves the work efficiency of frontline approvers. Attached Figure Description
[0066] Figure 1 A flowchart of a risk assessment method for the entire process of human genetic resource management. Detailed Implementation
[0067] The present invention will now be described in further detail with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0068] The system first parses application materials using automatic document extraction technology. It then utilizes OCR and NLP modules to convert the format and extract semantics of unstructured text such as PDFs, scanned documents, and handwritten content, completing data cleaning, field mapping, and structured data storage. Based on historical approval data, legal provisions, expert knowledge, and external data, the system constructs a risk assessment indicator system covering five dimensions: administrative services, business processing, processing timeliness, approval deviation, and other compliance. A rule base is built upon this system to standardize risk expression. The risk assessment module trains and predicts input samples using the LightGBM model, generating multi-dimensional risk scores, including overall risk scores and dimensional sub-scores. The system further visualizes the scoring results and provides risk level determinations and explanations of key factor contributions based on the scores, assisting in manual review and decision-making. The system also includes a dynamic adjustment module that automatically adjusts feature weights and risk judgment thresholds based on feedback signals such as expert opinions, approval result deviations, and model misjudgment rates, improving model adaptability and accuracy. The quality verification module runs periodically, including an anomaly detection algorithm to automatically screen samples with extreme deviations, and a manual quality inspection mechanism to ensure system stability and evaluation credibility.
[0069] 1. Indicator construction.
[0070] The risk indicator system of this system originates from in-depth research and systematic study of human genetic resource management practices. Employing diverse methods such as expert feedback, empirical research, and literature review, it systematically identifies key risk factors that constrain the efficiency and compliance level of human genetic resource approval from multiple perspectives, including policies and regulations, business operations, technical materials, institutional qualifications, and ethical processes. This results in a risk rule base comprising five major risk dimensions and dozens of rules. The system supports modular expansion and can cover various application scenarios.
[0071] 1.1 Analysis of unstructured materials.
[0072] The system supports the parsing of diverse application materials, employing OCR text recognition technology to extract text content from images and PDFs, and utilizing a domain-adaptive named entity recognition model to identify key fields such as project names, organizations, ethical information, and approval numbers. The system uses regular expressions combined with dependency parsing to remove redundant headers, footers, watermarks, and other noise, and stores structured key information in the database. For tables and semi-structured documents, the system introduces cell label recognition and table structure reconstruction modules, combining a rule engine and field alignment rules to achieve refined data extraction.
[0073] 1.2 Evaluation index design.
[0074] By organizing relevant laws and regulations, historical approval data, expert review opinions, and other data, and using data crawling and feature extraction technologies, multiple evaluation indicators and measurement methods were designed.
[0075] For example, from expert opinion texts, the system can extract terms such as "ethical review," "missing version number," and "informed consent" through entity recognition, identify risk descriptive words (such as "unmarked," "unsigned," and "inconsistent") through a defect vocabulary database, and construct event semantic chains (such as "ethical approval document - missing - sample informed consent form") using relation extraction, ultimately transforming them into a structured risk label input model. To further improve the understanding of complex unstructured documents and the accuracy of risk extraction, especially when semantic relationship reasoning involves multiple parties, materials, and approval processes, the system preferentially uses knowledge graph construction to model and reason about the entity relationships involved in the approval process.
[0076] 1.3 Risk Dimension Classification and Rule Base Design.
[0077] The system's indicator framework is divided into five dimensions: (A) Administrative service risk, including research output risk, institutional credit, and institutional information quality; (B) Business processing risk, including material completeness, material compliance, consistency of attachments, and cooperation information; (C) Processing timeliness risk, including process delays and application lag; (D) Deviation risk, including approval cycle and approval quality; and (E) Other qualified risks, including ethical review, data security, purpose of use, and consistency of experimental protocols. Each dimension has several sub-indicators and corresponding rules, supporting structured rule judgment and model training. Table 1 lists some examples of dimensions and indicator rules.
[0078]
[0079]
[0080] Table 1
[0081] 2. Model training.
[0082] This system employs machine learning methods to construct a multi-dimensional risk prediction model for human genetic resource approvals. Based on historical approval project sample data and a structured indicator system, it achieves efficient and interpretable risk score output through an ensemble algorithm. This invention preferentially uses the LightGBM model, a high-performance gradient boosting tree algorithm, which has advantages such as fast modeling speed, high robustness to outliers, and support for feature importance analysis. It is suitable for approval data characterized by high feature dimensionality, uneven data distribution, and sparse rules.
[0083] The training process for a risk assessment model is completed through the following four steps:
[0084] 2.1 Data Preparation and Feature Construction: The system utilizes the results of structured data extraction to construct a sample dataset, including fields such as basic project information and rule matching labels. The target variable is the approval conclusion label (e.g., approved / rejected). The system automatically handles missing values, unifies feature distribution, and constructs training and testing sets for model development.
[0085] 2.2 Initial screening of indicators and dimensional modeling: Pearson correlation coefficient and feature importance assessment method were used to eliminate redundant and highly collinear variables and retain key factors; some strongly regulatory indicators were forcibly retained in the model to ensure regulatory traceability.
[0086] 2.3 Model Training and Optimization: Model input includes multi-dimensional structured indicators, rule matching results, business labels, etc., and hyperparameters such as tree depth, number of leaves, and learning rate are tuned using five-fold cross-validation. Model output includes an overall project risk score (0-100), scores for five risk dimensions, and the contribution of the main features behind each score. This multi-layered output structure supports automatic approval classification and also meets interpretability requirements.
[0087] 2.4 Dynamic Risk Adjustment Mechanism: The system utilizes predefined project type scenario tags (such as "data collection + overseas cooperation") to adjust the weight ratios of each dimension in the model based on scenario characteristics. Furthermore, a feedback learning mechanism is introduced to periodically summarize approval deviation data (such as high-risk assessments that are approved by experts). Weight configurations are automatically corrected through sliding window analysis and incremental training. The system supports hot model updates and records version information to ensure that the risk assessment logic remains consistent with the business status.
[0088] 3. Risk assessment.
[0089] After the risk assessment model is trained, such as Figure 1 As shown, the risk assessment of the application report is completed through the following three steps.
[0090] 3.1 Extract key fields from the application report.
[0091] 3.2 Based on the key fields, and according to the defined risk dimensions and the judgment rules of the indicators under the risk dimensions, a vector representation of the application report is generated.
[0092] 3.3 Classification is performed based on the vector representation of the application report to obtain the risk assessment result of the application report. The risk assessment result is divided into three levels:
[0093] Overall Risk Score: Represents the project's comprehensive risk level, ranging from 0 to 100. The system sets fixed risk thresholds, such as 0-39 for low risk, 40-69 for medium risk, and 70-100 for high risk, which are used for subsequent assessments.
[0094] Five-dimensional sub-risk scores: The system outputs sub-scores for five dimensions: "Administrative Service Risk," "Business Processing Risk," "Processing Timeliness Risk," "Approval Deviation Risk," and "Other Compliance Risk." Each sub-score ranges from 0 to 100, and its risk level (low, medium, high) can be determined based on preset threshold values for the dimension or based on the overall risk level ratio. Visual display and drill-down analysis are supported in the dashboard.
[0095] Key Feature Explanation and Risk Factor Contribution: The system calculates the contribution of all features to the score output based on the SHAP algorithm, outputs the top few features with the greatest impact on the total score and their explanatory direction (risk improvement or mitigation), and attaches them to the end of the risk report in the form of a list, supporting manual review and the generation of material correction suggestions.
[0096] Through a three-tiered structured output of "total score + sub-score + feature interpretation", the system can provide approvers with rapid judgment support and accurate risk identification capabilities, improving approval efficiency and accuracy, while also providing data support for subsequent risk control strategy adjustments and model optimization.
[0097] In another embodiment, the system may also select models such as XGBoost and CatBoost to replace LightGBM to adapt to different data scales and business complexity scenarios. However, based on the current requirements, this invention prioritizes the use of LightGBM to balance training efficiency, model size and interpretability.
[0098] 4. System Architecture
[0099] 4.1 Acquisition Layer
[0100] It includes modules for application material parsing, external data acquisition, and indicator construction. The application material parsing module is responsible for format conversion and text extraction of diverse application materials; the external data acquisition module is responsible for real-time capture of relevant legal and regulatory data, historical approval data, and third-party data; and the indicator construction module is responsible for maintaining the risk indicator system and generating and updating the risk rule base.
[0101] 4.2 Computation Layer
[0102] It includes modules for data storage and management and risk assessment. The data storage and management module is responsible for cleaning, storing and managing data to ensure its efficient use; the risk assessment module processes the input data based on machine learning models to generate risk scores and explanatory factors for each project.
[0103] 4.3 Application Layer.
[0104] The application layer includes risk dashboards, feedback adjustment, and quality verification modules.
[0105] 4.3.1 Risk Dashboard: Displays the risk level for each dimension and provides an overall risk report to support expert approval decisions.
[0106] A. Use a heat map to display the risk level, with a color gradient of green → yellow → red corresponding to low → medium → high risk, and generate a detailed risk report.
[0107] B. Displays historical risk trends, supports multi-dimensional filtering by time range and activity type, assists in analyzing the reasons for project failures, and provides prompts to applicant organizations.
[0108] C. Set tiered viewing permissions for administrators and experts.
[0109] 4.3.2 Feedback Adjustment: If the review comments and risk decisions are inconsistent, experts can be reassigned or feedback issues can be raised.
[0110] A. The system automatically performs conflict detection, comparing items that differ from the expert review opinions and risk assessment report conclusions. If there is a discrepancy, it triggers a process of re-selecting experts for review.
[0111] B. Experts can submit activities with controversial review results to the issue feedback section and mark the specific issues. The system will record and categorize these issues for system maintenance personnel or model optimization processes to handle.
[0112] C. If the model is modified, the solution will be marked in the system update log.
[0113] 4.3.3 Quality verification: anomaly detection and manual spot checks.
[0114] A. Random sampling checks were conducted to identify any data entry errors in the assessed projects.
[0115] B. Add the problem items to the training set for iterative model optimization.
[0116] In summary, this invention constructs a dynamic indicator system that covers a multi-dimensional risk assessment indicator system throughout the entire lifecycle and a rule base for implementing its logic, adapting to the risk assessment needs of different business scenarios.
[0117] This invention enables intelligent risk assessment by integrating intelligent document parsing, rule engines, and interpretable machine learning models to achieve automated extraction of key information, dynamic multi-dimensional risk scoring, and explanation of causes.
[0118] This invention features a feedback learning and dynamic adjustment mechanism, which automatically optimizes model parameters, feature weights, and rule thresholds based on approval feedback, result deviations, and changes in the external environment, ensuring the timeliness and accuracy of continuous evaluation.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail using examples, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for risk assessment and control throughout the entire process of human genetic resource management, characterized in that, The method includes: Extract key fields from the application report; Based on the key fields, and according to the defined risk dimensions and the judgment rules for the indicators under the risk dimensions, a vector representation of the application report is generated; The risk assessment result of the application report is obtained by classifying the application report based on its vector representation.
2. The method according to claim 1, characterized in that, Extract key fields from the application report, including: OCR text recognition technology was used to extract the text content from the application report; The method utilizes regular expression rules combined with dependency parsing to remove noise from redundant text content, including redundant headers, footers, and watermarks. A domain-adaptive named entity recognition model is used to identify key fields in text content after noise removal. Specifically, for tables and semi-structured documents in the text content, key fields are extracted by cell label recognition, table structure reconstruction, and a combination of rule engine and field alignment rules.
3. The method according to claim 1 or 2, characterized in that, The key fields include: project name, organization, ethics information, and approval number.
4. The method according to claim 1, characterized in that, The risk dimensions include: administrative service risk, business processing risk, processing timeliness risk, deviation risk, and other compliance risks. Specifically, the primary indicators under administrative service risk include: research output risk, institutional credit, and institutional information quality; the primary indicators under business processing risk include: material completeness, material compliance, consistency of attachments, and cooperation information; the primary indicators under processing timeliness risk include: process delays and application lag; the primary indicators under deviation risk include: approval cycle and approval quality; and the primary indicators under other compliance risks include: ethical review, data security, consistency of purpose of use, and experimental protocol.
5. The method according to claim 4, characterized in that, The secondary sub-indicators under the research output risk include: the proportion of professional titles, hardware facilities, and the number of unfinished projects; wherein, the determination rule for the proportion of professional titles is described as the proportion of senior professional titles being less than a first set value, the determination rule for hardware facilities is described as the number of equipment types being less than a second set value or the total number of equipment being less than a third set value, and the determination rule for the number of unfinished projects is described as the number of unfinished projects being greater than a fourth set value within a specified time range. The secondary sub-indicators under the unit credit system include: administrative penalty records, notifications of scientific research misconduct, and records of dishonesty; wherein, the determination rule for administrative penalty records is described as any unit having a penalty announcement record, the determination rule for notifications of scientific research misconduct is described as units involved in notifications from the National Natural Science Foundation of China, and the determination rule for records of dishonesty is described as units being subject to enforcement by the court for dishonesty. The secondary sub-indicators under the unit information quality include: incomplete information and inconsistent information. The judgment rule for incomplete information is described as missing fields in the system filing information, and the judgment rule for inconsistent information is described as the actual information not matching the filing. The secondary sub-indicators under the material integrity include: missing core proof; wherein, the judgment rule for missing core proof is described as missing legal person authorization letter or missing ethics approval document; The secondary sub-indicators under the material compliance include: naming / format anomalies; wherein, the judgment rule for naming / format anomalies is described as no version number or broken format; The secondary sub-indicators under the attachment consistency include: incomplete protocol upload; wherein, the determination rule for incomplete protocol upload is described as the number that should be uploaded is not equal to the actual number; The secondary sub-indicators under the cooperation information include: missing legal person / qualification; wherein, the judgment rule for missing legal person / qualification is described as lacking qualification certificates or not being stamped; The secondary sub-indicators under the process delay include: system node lag; wherein, the judgment rule for system node lag is described as the duration of a certain approval node being greater than a fifth set value; The secondary sub-indicators under the application delay include: abnormal submission cycle; wherein, the judgment rule for the abnormal submission cycle is described as the time from the initial registration of the project to the submission being greater than the sixth set value; The secondary sub-indicators under the approval cycle include: deviation above the historical average; wherein, the judgment rule for deviation above the historical average is described as the current approval non-acceptance time being greater than the seventh set value; The secondary sub-indicators under the approval quality include: significant differences from similar projects; wherein, the judgment rule for significant differences from similar projects is described as the only project that failed to pass among similar projects; The secondary sub-indicators under the ethical review include: missing or abnormal content; wherein, the judgment rule for missing or abnormal content is described as not uploaded or the content not signed or stamped. The secondary sub-indicators under the data security framework include: failure to report outbound travel; wherein, the determination rule for failure to report outbound travel is described as involving cross-border travel but without providing approval documentation; The secondary sub-indicators under the purpose of use include: vague / incorrectly stated research purpose; wherein, the judgment rule for vague / incorrectly stated research purpose is described as the purpose being inconsistent with the ethical approval document; The secondary sub-indicators under the consistency of the experimental scheme include: inconsistent sample information; wherein, the judgment rule for inconsistent sample information is described as the uploaded sample type not matching the scheme.
6. The method according to claim 1, characterized in that, Based on the aforementioned key fields, and according to the defined risk dimensions and the judgment rules for indicators under those risk dimensions, a vector representation of the application report is generated, including: Based on the content of the key fields and the defined risk dimensions, extract the features from the application report; The feature is vectorized by combining the corresponding judgment rules; By combining the vectorization results of each feature, the vector representation of the application report is obtained.
7. The method according to claim 6, characterized in that, Based on the vector representation of the application report, a classification is performed to obtain the risk assessment result of the application report, including: Build a risk assessment model based on the LightGBM model, XGBoost model, or CatBoost model; Input the vector representation of the application report into the risk assessment model to obtain the overall risk score and the dimensional sub-risk scores; The contribution of all features to the overall risk score is calculated using the SHAP algorithm to obtain the feature explanation. The feature explanation includes: the top few features that have the greatest impact on the overall risk score and the impact of these features. The risk assessment result of the application report is obtained by combining the overall risk score, the dimensional sub-risk scores, and the feature interpretation.
8. The method according to claim 7, characterized in that, After synthesizing the overall risk score, the dimensional sub-risk scores, and the feature interpretations to obtain the risk assessment result of the application report, the following further steps are included: If the risk assessment results of the application report differ from the expert review opinions, experts will be re-selected to review the risk assessment results of the application report. If the review results show that the risk assessment results of the application report are incorrect, adjust the risk assessment model based on the review results and mark the solution in the system update log.
9. The method according to claim 1, characterized in that, After classifying the application report based on its vector representation to obtain the risk assessment result, the process further includes: A heat map is generated based on the risk assessment results of the application report, and the heat map is used to display the risk level; A historical risk change trend is generated based on the risk assessment results of the application report.
10. A risk assessment system for the entire process of human genetic resource management, characterized in that, The system includes: The field extraction module is used to extract key fields from the application report; The vector generation module is used to generate a vector representation of the application report based on the key fields and the judgment rules of the indicators under the set risk dimensions. The result generation module is used to classify the application report based on its vector representation to obtain a risk assessment of the application report.