A method and device for generating experimental reports based on large models

By implanting the dual-unit channels of ethical supervision and scientific deduction into the large model and constructing a three-dimensional ethical constraint matrix and map, the problems of ethical compliance and scientific rationality in the generation of experimental reports are solved, and the dual precise supervision and correction of experimental reports are achieved, ensuring the compliance and logical rigor of the generated content.

CN120409433BActive Publication Date: 2025-10-14SHANGYU TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510465196.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-10-14
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Existing technologies lack dual supervision of ethical compliance and scientific rationality in the generation of experimental reports, making it difficult to comprehensively evaluate experimental compliance and safety. The generated content may contain compliance deviations or logical errors, making it difficult for users to trace the basis of the generated content and make effective corrections.

Method used

By implanting a dual-unit channel of ethical supervision and scientific deduction into the large model, a three-dimensional ethical constraint matrix is ​​constructed, the large model is jointly trained with compliant samples and non-compliant samples, a sample primary verification engine is designed, and ethical compliance and scientific rationality reports are generated. The dependency relationship is then inferred through attention weights to construct a three-dimensional map for correction.

Benefits of technology

It achieves dual and precise supervision in the process of generating experimental reports, ensures that the generated content complies with regulations and scientific principles, provides an ethical and scientific supervision channel, improves the compliance and logical rigor of the report, and avoids the deviation of a single model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409433B_ABST
    Figure CN120409433B_ABST
Patent Text Reader

Abstract

The application discloses a kind of experimental report generation method and device based on large model, it is related to natural language processing technical field, including, extracting experimental features from experimental record, constructs three-dimensional ethical restraint matrix;Ethical supervision and scientific derivation double-unit channel are implanted in large model, and large model is trained by compliance sample and violation sample jointly;Design sample primary verification engine, carry out sample preliminary screening, scoring and confirmation, output ethical compliance and scientific rationality report;The dependence relationship of ethical supervision and scientific derivation double-unit is deduced by attention weight, and three-dimensional atlas is constructed;High-risk sample node is positioned and corrected using three-dimensional atlas, generates correction log and is integrated to three-dimensional ethical restraint matrix, and outputs final correction report.The application realizes the double precision supervision of ethical compliance and scientific rationality in experimental report generation process by implanting ethical supervision and scientific derivation double-unit channel in BERT model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a method and device for generating an experimental report based on a large model. Background Art

[0002] In recent years, with the rapid development of artificial intelligence, especially breakthroughs in deep learning and natural language processing (NLP), text generation methods based on large models have shown significant application potential in multiple fields. In the field of laboratory report generation, traditional methods mostly rely on template-based automation tools, which can achieve the output of structured content to a certain extent. Pre-trained language models (such as BERT and GPT) have been introduced into the laboratory report generation task. By understanding the semantics of experimental records and generating them, the automation level and content quality of the reports have been improved. In addition, research combining ethical constraints and risk assessment has gradually emerged, aiming to ensure the compliance and scientific nature of the generated content. However, these technologies mostly focus on single-dimensional text generation or compliance testing, and lack modeling and comprehensive verification of the multidimensional characteristics of the experimental process.

[0003] Although existing technologies have made some progress in automated generation, there are still several shortcomings. First, traditional methods have limited ability to extract ethically sensitive attributes and potential risks in experimental records, making it difficult to comprehensively assess the compliance and safety of experiments, which can easily lead to the omission of key ethical issues in generated reports. Second, although existing large models have powerful generation capabilities, they lack a dual supervision mechanism for ethical and scientific logic, which may lead to compliance deviations or logical errors in the output content. In addition, existing technologies lack support for the interpretability of the model decision-making process, making it difficult for users to trace the basis of the generated content and make effective corrections. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a large-scale model-based experimental report generation method to solve the problems of ethical compliance and scientific rationality in experimental report generation.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In the first aspect, the present invention provides a method for generating an experimental report based on a large model, which includes extracting experimental features from experimental records, constructing a three-dimensional ethical constraint matrix, and outputting experimental risk ratings and rule mappings; implanting a dual-unit channel of ethical supervision and scientific deduction in the large model, and jointly training the large model with compliant samples and non-compliant samples to obtain a bimodal feature set; designing a sample primary verification engine to perform preliminary screening, scoring and confirmation of samples, and outputting an ethical compliance and scientific rationality report; inferring the dependency relationship between the dual units of ethical supervision and scientific deduction through attention weights, constructing a three-dimensional map, and dynamically linking the three-dimensional ethical constraint matrix; using the three-dimensional map to locate and correct high-risk sample nodes, generating a correction log and integrating it into the three-dimensional ethical constraint matrix, and outputting a final correction report.

[0008] As a preferred solution of the method for generating experimental reports based on a large model described in the present invention, the steps of extracting experimental features from experimental records, constructing a three-dimensional ethical constraint matrix, and outputting experimental risk ratings and rule mapping are as follows:

[0009] The experimental records include operation logs, instrument parameters and sample identification;

[0010] Use natural language processing tools to extract experimental features;

[0011] Construct a three-dimensional ethical constraint framework, analyze experimental characteristics based on predefined rules, and locate experimental characteristics within the three-dimensional ethical constraint framework;

[0012] Obtain multiple ethical guidelines from an external rule set, number each ethical guideline to form multiple rules and map them onto a three-dimensional ethical constraint framework;

[0013] The experimental characteristics and rule mapping are integrated through the priority rule method, the experimental risk rating is output, and a three-dimensional ethical constraint matrix is ​​generated.

[0014] As a preferred solution of the experimental report generation method based on the large model of the present invention, wherein: the dual-unit channel of ethical supervision and scientific deduction is implanted in the large model, and the large model is trained by jointly training the compliant samples and the non-compliant samples to obtain a dual-modal feature set. The specific steps are as follows:

[0015] Select the pre-trained BERT model as the basic framework of the large model;

[0016] Implanting ethical supervision units and scientific deduction units into the output layer of the BERT model;

[0017] Use NLP tools to extract experimental text data from operation logs and sample identifiers, and extract experimental operation data from operation logs and instrument parameters;

[0018] The three-dimensional ethical constraint matrix is ​​embedded into the fully connected layer of the ethics supervision unit through the attention mechanism for training, and the actual compliance score is output as the basis for evaluating the probability of sample violation;

[0019] Obtain scientific knowledge base through external scientific resources, load it into the scientific deduction unit, and output the logical consistency score as the basis for verifying the logical error rate of the sample;

[0020] Assign initial attention weights to ethical supervision units and scientific deduction units;

[0021] Obtain compliance samples and non-compliance samples from historical laboratory reports;

[0022] Divide the compliant samples and the noncompliant samples into training sets and validation sets in proportion;

[0023] The training set is input into the BERT model for training, and the attention weight is adjusted to achieve a balance between the ethical supervision unit and the scientific deduction unit;

[0024] The sample violation probability and attention weight distribution are output and combined into a bimodal feature set.

[0025] As a preferred solution of the large model-based experimental report generation method of the present invention, the design sample primary verification engine performs preliminary sample screening, scoring and confirmation, and outputs an ethical compliance and scientific rationality report. The specific steps are as follows:

[0026] Extract the sample violation probability from the bimodal feature set, set the violation threshold, and preliminarily screen out high-risk samples that exceed the violation threshold for first-level review;

[0027] Use the BERT model to perform text analysis on high-risk samples for the first-level review, input experimental text data and experimental operation data, and generate compliance scores and logical consistency scores for each high-risk sample as a second-level review;

[0028] Conduct a comprehensive review of high-risk samples from the second-level review, using a three-dimensional ethical constraint matrix to compare sample violation probabilities, compliance scores, and logical consistency scores, identify violation samples, and generate violation conclusions as a third-level review;

[0029] Comprehensive first-level review, second-level review and third-level review to output ethical compliance and scientific rationality reports.

[0030] As a preferred solution of the large-scale model-based experimental report generation method of the present invention, the following specific steps are used to reversely infer the dependency relationship between the ethical supervision and scientific deduction dual units through attention weight, construct a three-dimensional map, and dynamically link the three-dimensional ethical constraint matrix.

[0031] The dependency relationship between the ethical supervision unit and the scientific deduction unit in the BERT model is inferred using the attention weight distribution, compliance score, and logical consistency score.

[0032] Traverse the bimodal feature set and high-risk samples, integrate each high-risk sample, calculate the contribution weights of the ethical supervision unit and the scientific deduction unit, and generate a dependency table;

[0033] Sort the experimental text data and experimental operation data by time to generate the sequence of experimental steps;

[0034] Load the BERT model, input the dependency table and three-dimensional ethical constraint matrix, project the attention weight distribution, compliance score, logical consistency score and experimental step sequence into the three-dimensional coordinate system, dynamically link the numbering rules in the three-dimensional ethical constraint matrix, and output a three-dimensional map.

[0035] As a preferred solution of the large-model-based experimental report generation method described in the present invention, the three-dimensional coordinate system refers to the X-axis being the order of experimental steps, the Y-axis being the compliance score and the logical consistency score, and the Z-axis being the attention weight and the sample violation probability.

[0036] As a preferred solution of the large-scale model-based experimental report generation method of the present invention, the method uses a three-dimensional map to locate and correct high-risk sample nodes, generates a correction log and integrates it into a three-dimensional ethical constraint matrix, and outputs a final correction report. The specific steps are as follows:

[0037] Through the bimodal feature set of the BERT model and the three-dimensional ethical constraint matrix, we analyze the contribution of attention weight distribution to high-risk samples and mark high-risk sample nodes;

[0038] Locate high-risk sample nodes in the three-dimensional map, analyze the dependencies of high-risk sample nodes, make correction suggestions, and output correction logs;

[0039] Integrate the correction log into the three-dimensional ethical constraint matrix, update the rule mapping of the three-dimensional ethical constraint matrix, and output the final correction report.

[0040] In the second aspect, the present invention provides an experimental report generation device based on a large model, including a risk module for extracting experimental features from experimental records, constructing a three-dimensional ethical constraint matrix, and outputting experimental risk ratings and rule mappings; a training module for implanting a dual-unit channel of ethical supervision and scientific deduction in the large model, and jointly training the large model through compliant samples and non-compliant samples to obtain a bimodal feature set; a design module for designing a sample primary verification engine, performing preliminary screening, scoring and confirmation of samples, and outputting ethical compliance and scientific rationality reports; a construction module for inferring the dependency relationship between the dual units of ethical supervision and scientific deduction through attention weights, constructing a three-dimensional map, and dynamically linking the three-dimensional ethical constraint matrix; a correction module for using the three-dimensional map to locate and correct high-risk sample nodes, generate a correction log and integrate it into the three-dimensional ethical constraint matrix, and output a final correction report.

[0041] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the large model-based experimental report generation method as described in the first aspect of the present invention is implemented.

[0042] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the large model-based experimental report generation method as described in the first aspect of the present invention.

[0043] The beneficial effects of the present invention are as follows: the present invention realizes the dual precise supervision of ethical compliance and scientific rationality in the process of generating experimental reports by implanting a dual-unit channel of ethical supervision and scientific deduction into the BERT model. The ethical supervision unit generates a compliance score by embedding a three-dimensional ethical constraint matrix and combining experimental text data to effectively evaluate whether the sample complies with regulations and ethical requirements, ensure that the gene editing experiment has been approved by the state, and avoid generating illegal content. The scientific deduction unit loads the scientific knowledge base, analyzes the experimental operation data, outputs a logical consistency score, and verifies whether the operation complies with scientific principles, such as whether the protective measures for high-temperature experiments are reasonable. This dual-unit architecture provides an independent ethical and scientific supervision channel for the report content during the generation process, solving the problem that the traditional single model is difficult to balance compliance and logic. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 This is a flow chart of the method for generating an experimental report based on a large model in Example 1.

[0046] Figure 2 This is a flow chart of the sample review in Example 1. DETAILED DESCRIPTION

[0047] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0048] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0049] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0050] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides a method for generating an experimental report based on a large model, comprising the following steps:

[0051] S1. Extract experimental features from experimental records, construct a three-dimensional ethical constraint matrix, and output experimental risk rating and rule mapping.

[0052] Furthermore, experimental records include operation logs, instrument parameters, and sample identification;

[0053] Specifically, the operation log recorded by the operator during the experiment, including information such as timestamp, operation description, and operator identity;

[0054] Instrument parameters are the operating data of the equipment used in the experiment, such as temperature, pressure, radiation dose and other specific values;

[0055] The sample ID is the unique identifier of the experimental object and its description;

[0056] Use natural language processing tools to analyze experimental records line by line and extract experimental features;

[0057] Experimental characteristics include ethically sensitive attributes, experimental risks, and data source reliability markers;

[0058] Specifically, use natural language processing (NLP) tools, such as the BERT-based text analysis model, to parse the text content of the experimental records line by line;

[0059] The ethically sensitive attribute identifies descriptions involving ethics through keyword matching and semantic analysis;

[0060] Experimental risks identify potential risks based on the operation description and instrument parameters. For example, if the temperature reaches 500°C, it is marked as a high temperature risk.

[0061] Data source reliability marking is to check the qualification information of the record source;

[0062] Ethically sensitive attributes include human tissue, endangered animals, and gene editing, and experimental risks include high temperature risks, high pressure risks, and radioactive risks;

[0063] Data source reliability mark refers to the qualification information of the source of experimental records;

[0064] Construct a three-dimensional ethical constraint framework, analyze experimental characteristics through natural language processing tools combined with predefined rules, and map experimental characteristics to experimental risk type, experimental stage, and experimental impact severity, and locate them in the three-dimensional ethical constraint framework;

[0065] Specifically, the X-axis / horizontal axis (experimental risk type): risk classification based on ethical sensitivity attributes; the Y-axis / vertical axis (experimental stage): positioning based on log timestamps; the Z-axis / vertical axis (impact severity): scoring the impact severity based on risk judgment;

[0066] The predefined rules refer to a set of logic or conditions that have been formulated before the parsing process begins, which are used to map experimental features to framework dimensions;

[0067] Use data processing software (such as Python's Pandas) to generate a three-dimensional ethical constraint framework;

[0068] The horizontal axis of the three-dimensional ethical constraint framework is defined as the experimental risk type;

[0069] Types of experimental risks include biosafety, data privacy, and academic integrity;

[0070] Map ethically sensitive attributes and experimental risks to corresponding types, clarifying the classification basis of the X-axis / horizontal axis;

[0071] The vertical axis of the three-dimensional ethical constraint framework is defined as the experimental stage;

[0072] The experimental phase includes data collection, experimental execution, and result analysis;

[0073] The vertical axis of the three-dimensional ethical constraint framework is defined as the severity of the experimental impact;

[0074] The experimental impact severity is classified into five levels;

[0075] Specifically, a five-level scoring system is adopted, with level 1 being slight impact, level 2 being limited impact, level 3 being moderate impact, level 4 being high impact, and level 5 being serious consequences. The experimental impact severity is assessed based on the experimental risk, and the quantification standards for Z-axis / vertical axis are determined to measure the degree of experimental risk;

[0076] More than 300 ethical guidelines are obtained from external rule sets (such as OSHA safety standards, etc.) and directly converted into a structured rule base. Each ethical guideline is numbered to form a mapping of multiple rules to a three-dimensional ethical constraint framework;

[0077] Specifically, multiple ethical guidelines are collected, each guideline is assigned a unique number, and is converted into a structured format;

[0078] The multiple ethical guidelines are mapped to the three-dimensional ethical constraint framework, such as "rule number → X, experimental execution → Y, 4 levels → Z". The multiple ethical guidelines and the three-dimensional ethical constraint framework are associated to form a searchable three-dimensional ethical constraint matrix;

[0079] The structured file is JSON or CSV;

[0080] The experimental risk type, experimental stage, experimental impact severity, and rule mapping are integrated by the priority rule method to output the experimental risk rating and generate a three-dimensional ethical constraint matrix with experimental risk rating and rule mapping;

[0081] The experimental impact severity is taken as the initial value, and if the rule mapping specifies a specific rating, the rule is used as the reference;

[0082] The experimental risk rating = max (experimental impact severity, experimental risk rating specified by rule mapping);

[0083] If the rule specifies the experimental risk rating, the rule value is used, otherwise the experimental impact severity is used;

[0084] The experimental impact severity is taken as the reference, and if the rule mapping specifies the rating, it is used preferentially to output the experimental risk rating and generate a three-dimensional ethical constraint matrix with experimental risk rating and rule mapping;

[0085] It should be noted that by extracting experimental features from experimental records, constructing a three-dimensional ethical constraint matrix and outputting experimental risk rating and rule mapping, the evaluation of the ethical and experimental risk of the experimental process is realized. The NLP tool is used to analyze the operation log, instrument parameters and sample identification, and the features are mapped to the risk type, experimental stage and severity by combining the pre-defined rules to form a structured matrix, which provides compliance basis for subsequent generation. Its role is to comprehensively identify ethical issues and experimental risks in experiments, and integrate multi-dimensional information through priority rule method to ensure accurate rating. Finally, a matrix with experimental risk rating and rule mapping is generated, which improves the coverage of the report on ethical and safety hazards.

[0086] S2, implanting an ethical supervision and scientific derivation dual-unit channel in the large model, training the large model by combining compliant samples and non-compliant samples to obtain a dual-modal feature set.

[0087] Further, a pre-trained BERT model is selected as a large model basic framework;

[0088] Specifically, a pre-trained BERT-base-uncased model is selected from the HuggingFace model library as a large model basic framework;

[0089] The input maximum sequence length is 512 tokens, and the word embedding dimension is 768;

[0090] The output CLS (classification) captures the semantic information of the entire sequence. CLS is a special token in the BERT model used to represent the semantic information of the entire input sequence. The last layer output CLS vector is the pooling representation of the sequence, which is commonly used in downstream classification tasks;

[0091] Based on the output layer expansion of the BERT model, the ethical supervision unit and the scientific derivation unit are implanted, and the input and output of the ethical supervision unit and the scientific derivation unit are defined;

[0092] Specifically, a full connection layer and a Sigmoid activation function are added to the ethical supervision unit to generate compliance scores;

[0093] The scientific derivation unit adds a full connection layer and a Sigmoid activation function to generate logical consistency scores;

[0094] The outputs of the two units are fused through an attention mechanism, and the weight parameters w1 (ethics) and w2 (science) are defined

[0095] The CLS vector is input into the full connection layer of the two units, and the ethical supervision unit and the scientific derivation unit output a single scalar (0-1) respectively;

[0096] The input of the ethics supervision unit is experimental text data and a three-dimensional ethics constraint matrix, and the output is a compliance score;

[0097] The input of the scientific deduction unit is experimental operation data, and the output is logical consistency score;

[0098] Use NLP tools to extract experimental text data from operation logs and sample identifiers, and extract experimental operation data from operation logs and instrument parameters;

[0099] The input of the ethics supervision unit is experimental text data and a three-dimensional ethics constraint matrix, and the output is a compliance score;

[0100] Specifically, the ethics supervision unit inputs experimental text data and a three-dimensional ethics constraint matrix;

[0101] The BERT model generates CLS vectors, and the fully connected layer combines the three-dimensional ethical constraint matrix rules to output the compliance score;

[0102] The embedding layer is CLS vector extraction with a dimension of 768;

[0103] The fully connected layer changes from ReLU activation to Sigmoid activation;

[0104] The output is a compliance score;

[0105] It should be noted that this step is a static definition and does not explain how the compliance score is generated. It only specifies the input-output relationship without actually performing training or optimization, indicating that the ethics supervision unit will predict compliance based on these inputs;

[0106] The input of the scientific deduction unit is experimental operation data, and the output is logical consistency score;

[0107] Specifically, the scientific deduction unit inputs experimental operation data;

[0108] Generate CLS vectors through the BERT model, and use the fully connected layer to verify logical consistency with experimental operation data, and output a logical consistency score;

[0109] The embedding layer extracts dimensions through CLS vectors;

[0110] The fully connected layer changes from ReLU activation to Sigmoid activation;

[0111] The output is a logical consistency score;

[0112] The three-dimensional ethical constraint matrix is ​​embedded into the fully connected layer of the ethics supervision unit through the attention mechanism for training, and the actual compliance score is output as the basis for evaluating the probability of sample violation;

[0113] It should be noted that this step is implemented dynamically, which illustrates that the actual compliance score is generated by embedding the three-dimensional ethical constraint matrix and training;

[0114] Specifically, a compliance score is defined as 0-1, where the closer to 1, the more compliant, and closer to 0, the more non-compliant;

[0115] Logical consistency score, ranging from 0 to 1, the closer to 1, the more logically consistent, and closer to 0, the more logically incorrect;

[0116] For example, if the compliance score is 0.9 and the probability of violation is 0.1 (low violation), then the compliance score of 0.9 is close to 1 and the probability of violation is close to 0, which is a low violation, indicating that the sample is compliant.

[0117] Compliance score 0.2 → Violation probability 0.8 (high violation) → Then the compliance score 0.2 is close to 0 and the violation probability is close to 1, which is a high violation, indicating that the sample violates the rules;

[0118] Logical consistency score 0.8 → sample logical error rate 0.2 (low error) → logical consistency close to 1 (correct), indicating that the sample logic is correct;

[0119] Logical consistency score 0.4 → sample logical error rate 0.6 (high error) → logical consistency close to 0 (error), indicating sample logical errors;

[0120] The three-dimensional ethical constraint matrix is ​​flattened and encoded into a feature vector, which is then fused with the CLS vector through an attention mechanism and input into a fully connected layer. Compliance / violation samples are used to optimize the fully connected layer and attention weights to generate a compliance score.

[0121] This step generates the actual rating capability through the embedding matrix and finally outputs the result;

[0122] Obtain a scientific knowledge base through external scientific resources (physical law database), convert the scientific knowledge base into a structured format through NLP tools, load it into the scientific deduction unit, and output a logical consistency score as the basis for verifying the sample logical error rate (i.e., the updated scientific deduction unit parameters);

[0123] Assign initial attention weights to the ethics supervision unit and scientific deduction unit, such as w1(ethics) = 0.5, w2(science) = 0.5, satisfying w1+w2=1;

[0124] Obtain compliant samples (samples of experimental cases approved by national agencies, such as samples based on approved gene editing experiments) and noncompliant samples (samples of experimental cases that have not obtained national agency permission or violate regulations, such as samples based on unauthorized high-temperature experiments) from historical experimental reports (approximately 5,000 reports). Label compliant samples as positive samples and noncompliant samples as negative samples, and indicate the violation points.

[0125] Compliant samples and non-compliant samples are divided into training set (80%) and validation set (20%) in proportion;

[0126] The training set is input into the BERT model for training. The loss function is used to optimize the ethical supervision unit and the scientific deduction unit, and the attention weight is adjusted to achieve a balance between the ethical supervision unit and the scientific deduction unit.

[0127] Specifically, the ethics supervision unit optimizes the compliance score and generates the probability of sample violation (i.e., 1-compliance score);

[0128] Optimize logical consistency scoring in scientific deduction units;

[0129] Use the cross entropy loss function (sample compliance loss + sample consistency loss) to adjust the attention weight to balance;

[0130] The loss function is L = L c +L o , where L is the loss function, which represents the goal of BERT model optimization. BERT model parameters are usually adjusted by minimizing L. c is the loss of the ethical supervision unit, which measures the gap between the compliance score and the true compliance label (such as 1 or 0). For example, if the sample label is 1 for compliance and the prediction is 0.8, L o The loss of the scientific derivation unit is the loss of the scientific loss, which measures the gap between the logical consistency score and the true consistency label. For example, if the sample label is consistent, it is 1, and the prediction is 0.9, L o Calculate the error and dynamically adjust the total attention weight (e.g., to “0.6 for ethics, 0.4 for science”);

[0131] Output the sample violation probability and attention weight distribution and directly combine them into a bimodal feature set (including sample violation probability, experimental text data, experimental operation data, three-dimensional ethical constraint matrix and scientific knowledge base, that is, the trained BERT model), and save it as a JSON format file;

[0132] It should be noted that the sample violation probability is the result of the ethical judgment of each sample (sample level) and is directly used for screening;

[0133] The attention weight distribution is a trained global weight that is paired with each sample to provide explanatory and bimodal information;

[0134] Use the validation set to test the sample violation probability and sample logical error rate, and adjust the attention weight distribution to optimize the BERT model performance;

[0135] The sample logical error rate is an overall performance indicator (global level) of the validation set. It is not a sample-level feature and cannot provide a specific value for each sample. It is not suitable as graph content. Its role is to test the performance of the BERT model and optimize the attention weight distribution, rather than directly judging the sample.

[0136] Specifically, the probability of sample violation is: the ethics supervision unit score is compared with the compliance label to obtain the accuracy (e.g. ≥90%), and the accuracy is calculated as: the number of correct predictions / total number;

[0137] Sample logic error rate: Compare the scientific derivation unit score with the consistency label to obtain the error ratio (e.g. ≤5%). Error rate calculation: number of incorrect predictions / total number;

[0138] It should be noted that by implanting dual units of ethical supervision and scientific deduction into the BERT model, jointly training with compliance and violation samples, and outputting a dual-modal feature set, accurate dual supervision of ethics and science in experimental report generation is achieved. The ethical supervision unit combines the three-dimensional ethical constraint matrix to generate a compliance score, and the scientific deduction unit relies on the scientific knowledge base to output a logical consistency score to ensure that the content of the experimental report complies with regulations and scientific principles. Its role is to provide an intelligent compliance and rationality assessment tool for the scientific research field, which is suitable for high-demand scenarios such as gene editing and medical experiments. Finally, through dual-unit collaboration and attention weight optimization, the compliance and logical rigor of the experimental report are significantly improved, avoiding the deviation of a single model and generating a reliable dual-modal feature set.

[0139] S3. Design a sample primary verification engine to conduct preliminary sample screening, scoring and confirmation, and output ethical compliance and scientific rationality reports.

[0140] Furthermore, the sample violation probability is extracted from the bimodal feature set, and a violation threshold is set to preliminarily screen out high-risk samples that exceed the violation threshold for first-level review;

[0141] Specifically, the sample violation probability of about 5,000 samples is extracted from the bimodal feature set;

[0142] Using the rules in the three-dimensional ethical constraint matrix as a reference, set the violation threshold (for example, the severity of experimental risk or high-risk samples is set at 0.7, and samples below 0.3 are initially considered low risk) and determine the screening criteria;

[0143] Traverse each sample in the bimodal feature set and compare the sample violation probability of each sample with the violation threshold. If it is higher than the violation threshold (e.g. ≥0.7), it is marked as a high-risk sample that needs to be reviewed;

[0144] It should be noted that the violation threshold is set based on the severity classification of the three-dimensional ethical constraint matrix (level 1-5). The sample violation probability ≥ 0.7 is defined as the high-risk threshold, corresponding to samples of severity level 4 and above, indicating a high impact that requires immediate review; the violation probability ≤ 0.3 is defined as the low-risk threshold, corresponding to samples of severity level 2 and below, which are initially considered low risk. Samples between 0.3-0.7 are marked as potential risks and require further evaluation;

[0145] Output a list of high-risk samples as the first-level review result;

[0146] Use the BERT model to perform text analysis on high-risk samples for the first-level review, input experimental text data and experimental operation data, and generate compliance scores and logical consistency scores for each high-risk sample as a second-level review;

[0147] Specifically, read high-risk samples, load the BERT model, input experimental text data and experimental operation data, and splice the experimental text data and experimental operation data into a single sequence. The maximum sequence length is 512 tokens, and the number of tokens is 512.

[0148] Input the sequence into the BERT model and generate the CLS vector;

[0149] The CLS vectors are input into the fully connected layers of the ethics supervision unit and the scientific deduction unit respectively;

[0150] The ethics supervision unit outputs a compliance score, which indicates the probability that the sample complies with ethical rules;

[0151] The scientific deduction unit outputs a logical consistency score, which indicates the scientific rationality of the sample operation;

[0152] Example: Sample A: Compliance score 0.7, logical consistency score 0.9;

[0153] Sample B: Compliance score 0.6, logical consistency score 0.8;

[0154] Generate a compliance score and a logical consistency score for each high-risk sample;

[0155] Conduct a comprehensive review of high-risk samples from the second-level review, using a three-dimensional ethical constraint matrix to compare sample violation probabilities, compliance scores, and logical consistency scores, identify violation samples, and generate violation conclusions as a third-level review;

[0156] Specifically, the compliance score and logical consistency score of the high-risk samples are read from the high-risk samples with high scores in the second-level review;

[0157] Extracting the sample violation probability of the corresponding sample from the bimodal feature set;

[0158] Load the scientific knowledge base and the three-dimensional ethical constraint matrix. For each sample, combine the rules in the three-dimensional ethical constraint matrix with the scientific knowledge base, and compare the second-level review, such as the first-level sample violation probability, the second-level compliance score, and the second-level logical consistency score. If the compliance score is lower than the rule requirements, or the sample violation probability is high and violates the rules of the three-dimensional ethical constraint matrix, it is determined to be a violation. If the logical consistency score is low or violates the knowledge base rules, further confirm the violation.

[0159] For example:

[0160] Sample A: Sample violation probability is 0.8, compliance score is 0.7 (below the standard), logical consistency score is 0.9, and it was not approved and confirmed as a violation;

[0161] Sample B: Sample violation probability is 0.6, compliance score is 0.9, logical consistency score is 0.8, compliant and consistent, not confirmed as a violation;

[0162] Output generates a list of confirmed violation samples and violation conclusions as the third-level review results;

[0163] Comprehensively reviewing the first, second, and third levels of review, output an ethical compliance and scientific rationality report (including a list of high-risk samples, sample scoring details, conclusions on non-compliant samples, and statistics on compliant samples);

[0164] It should be noted that by designing a primary verification engine for samples, extracting violation probabilities from a bimodal feature set and conducting a three-level review, and outputting ethical compliance and scientific rationality reports, accurate verification of the generated content is achieved. Violation thresholds are used to screen high-risk samples, and the BERT model generates compliance and logical scores. Combined with review and a three-dimensional ethical constraint matrix for comprehensive evaluation, the ethical and scientific dual standards of the samples are ensured. Its role is to provide reliable quality control for experimental reports, and it is suitable for high-compliance fields such as medicine and biology. Ultimately, through multi-level verification, the accuracy and credibility of the reports are improved, and the deviation of single automated judgment is avoided.

[0165] S4. Use attention weights to infer the dependency relationship between ethical supervision and scientific deduction, construct a three-dimensional map, and dynamically link the three-dimensional ethical constraint matrix.

[0166] Furthermore, the attention weight distribution, compliance score, and logical consistency score are used to infer the dependency between the ethical supervision unit and the scientific deduction unit in the BERT model.

[0167] Specifically, the attention weight distribution in the bimodal feature set corresponds to the tokens input to the BERT model;

[0168] Load the pre-trained BERT model and extract the weight parameters of the ethical supervision unit and scientific derivation unit of the trained BERT model;

[0169] Use the SHAP method to calculate the contribution of attention weight distribution to compliance score and logical consistency score, and generate dependency relationships;

[0170] Traverse the bimodal feature set and high-risk samples, integrate the attention weight distribution, compliance score, and logical consistency score of each high-risk sample, calculate the contribution weights of the ethical supervision unit and scientific deduction unit to the sample violation probability, compliance score, and logical consistency score, and generate a dependency table;

[0171] Specifically, use Python script to traverse each high-risk sample;

[0172] Extract the attention weight distribution, compliance score, logical consistency score and sample violation probability of each sample;

[0173] Use gradient analysis (e.g., PyTorch's torch.autograd automatic differentiation) to calculate the contribution weights of the ethical supervision unit and the scientific deduction unit to the output, integrate the results, and generate a dependency table;

[0174] Sort the experimental text data and experimental operation data by time to generate the sequence of experimental steps;

[0175] Specifically, the experimental text data and experimental operation data are sorted by time to generate the sequence of experimental steps;

[0176] Load the BERT model, input the dependency table and the three-dimensional ethical constraint matrix, project the attention weight distribution, compliance score, logical consistency score, and experimental step sequence into the three-dimensional coordinate system, dynamically link the numbering rules in the three-dimensional ethical constraint matrix, and output a three-dimensional map;

[0177] The three-dimensional coordinate system means that the X-axis is the sequence of experimental steps;

[0178] The Y-axis is the compliance score and logical consistency score;

[0179] The Z axis is the attention weight and sample violation probability;

[0180] It should be noted that by inferring the dependency relationship between the two units of ethical supervision and scientific deduction through attention weights, constructing a three-dimensional map and dynamically linking the three-dimensional ethical constraint matrix, the visualization and traceability of the decision-making process are achieved. By using the SHAP method and gradient analysis, the scores and weights in the bimodal feature set are integrated to generate a dependency table and project it into a three-dimensional coordinate system, providing users with an intuitive analysis tool. Its role is to provide the basis for ethical and scientific judgments. Finally, by dynamically linking the matrix rules, the transparency and interpretability of the generation process are enhanced, making it easier to locate problem nodes.

[0181] S5. Use the three-dimensional map to locate and correct high-risk sample nodes, generate a correction log and integrate it into the three-dimensional ethical constraint matrix, and output the final correction report.

[0182] Furthermore, by using the BERT model's bimodal feature set and the three-dimensional ethical constraint matrix, we analyze the contribution of attention weight distribution to high-risk samples, locate the feature dependencies of high-risk samples in the three-dimensional graph, and mark high-risk sample nodes that require verification.

[0183] Specifically, the sample violation probability (e.g., sample A: 0.85), the ethical supervision unit attention weight (e.g., 0.7), and the scientific deduction unit attention weight (e.g., 0.3) are extracted from the bimodal feature set;

[0184] Extract the 3D coordinates of high-risk sample nodes from the 3D atlas:

[0185] Based on the dependency table (the contribution weights of the ethical supervision unit and the scientific deduction unit to the probability of sample violation), the ethical contribution ratio and the scientific contribution ratio are obtained;

[0186] If the ethical weight is 0.7 and the scientific weight is 0.3, then the ethical contribution ratio is 0.7 / (0.7+0.3)=70%, and the scientific contribution ratio is 30%;

[0187] According to the three-dimensional ethical constraint matrix, high-risk sample nodes are mapped to the rules in the three-dimensional ethical constraint matrix, such as mapping the sequence of experimental steps to the experimental stage, mapping the compliance score to the severity of the impact, and mapping the sample violation probability to the risk rating basis;

[0188] If the compliance score of a high-risk sample node is lower than the rule mapping requirement, it is marked as an ethical rule conflict;

[0189] If the logical consistency score of a high-risk sample node is lower than the scientific knowledge base standard, it is marked as a scientific logic conflict;

[0190] Mark ethical rule conflicts and scientific logic conflicts according to the ratio of ethical contributions and scientific contributions;

[0191] Locate high-risk sample nodes in the three-dimensional map, analyze the dependencies of high-risk sample nodes, make correction suggestions, and output correction logs;

[0192] Specifically, based on the rule mapping in the three-dimensional ethical constraint matrix, the causes of violations are matched and correction suggestions are generated;

[0193] Based on the loaded scientific knowledge base, match the violation type and generate correction suggestions;

[0194] Check the instrument parameters and experimental phases in the experimental operation data to confirm the inconsistency of the violation type;

[0195] Generate a remediation log for each high-risk sample node (including sample ID, experiment execution stage, violation type, dependency weight, rule mapping, remediation suggestions, and update status such as "pending" or "resolved");

[0196] Integrate the revision log into the three-dimensional ethical constraint matrix, update the rule mapping of the three-dimensional ethical constraint matrix, and output the final revision report;

[0197] Specifically, if the proposed amendment involves rules that are not covered by the three-dimensional ethical constraint matrix, new numbers are extracted from the rules in the three-dimensional ethical constraint matrix and assigned to them, which are mapped to the X-axis (experimental risk type), Y-axis (experimental stage), and Z-axis (impact severity) of the three-dimensional matrix;

[0198] Recalculate the experiment rating based on the revised suggestions and update the three-dimensional ethical constraint matrix according to the priority rule method (e.g., reduce the original rating from level 4 to level 3);

[0199] Output the final revision report (including a list of high-risk samples, revision details, updated records of the three-dimensional ethical constraint matrix, and compliance statistics, such as the proportion of compliant samples after revision (e.g., from 75% to 92%));

[0200] It should be noted that through the bimodal feature set and three-dimensional ethical constraint matrix of the BERT model, combined with the attention weight distribution, the feature dependencies of high-risk sample nodes can be accurately located, the causes of violations can be analyzed and correction suggestions can be generated. Based on rule mapping and scientific knowledge base, ethical rule conflicts or scientific logic conflicts can be marked, correction logs can be generated and the three-dimensional ethical constraint matrix can be updated. Finally, a correction report can be output to improve the compliance ratio and optimize the experimental risk rating and rule coverage.

[0201] This embodiment also provides an experimental report generation device based on a large model, including: a risk module, which is used to extract experimental features from experimental records, construct a three-dimensional ethical constraint matrix, and output experimental risk ratings and rule mappings; a training module, which is used to implant a dual-unit channel of ethical supervision and scientific deduction into the large model, and jointly train the large model through compliant samples and non-compliant samples to obtain a bimodal feature set; a design module, which is used to design a sample primary verification engine, perform preliminary sample screening, scoring and confirmation, and output an ethical compliance and scientific rationality report; a construction module, which is used to infer the dependency relationship between the dual units of ethical supervision and scientific deduction through attention weights, construct a three-dimensional map, and dynamically link the three-dimensional ethical constraint matrix; a correction module, which is used to use the three-dimensional map to locate and correct high-risk sample nodes, generate a correction log and integrate it into the three-dimensional ethical constraint matrix, and output a final correction report.

[0202] This embodiment also provides a computer device, which is suitable for the case of a large model-based experimental report generation method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the large model-based experimental report generation method proposed in the above embodiment.

[0203] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.

[0204] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating an experimental report based on a large model as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0205] In summary, the present invention implants a dual-unit channel of ethical supervision and scientific deduction into the BERT model, uses joint training of compliant samples and non-compliant samples, outputs compliance scores and logical consistency scores, and generates a bimodal feature set, thereby realizing dual precise supervision of ethical compliance and scientific rationality in the process of generating experimental reports. The ethical supervision unit embeds a three-dimensional ethical constraint matrix and combines experimental text data to generate a compliance score, effectively evaluating whether the sample complies with regulations and ethical requirements, ensuring that the gene editing experiment has been approved by the state and avoiding the generation of non-compliant content. The scientific deduction unit loads the scientific knowledge base, analyzes the experimental operation data, outputs a logical consistency score, and verifies whether the operation complies with scientific principles, such as whether the protective measures for high-temperature experiments are reasonable. This dual-unit architecture provides an independent ethical and scientific supervision channel for the report content during the generation process, solving the problem that the traditional single model is difficult to balance compliance and logic.

[0206] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for generating an experimental report based on a large model, characterized by: include, Extract experimental features from experimental records, construct a three-dimensional ethical constraint matrix, and output experimental risk rating and rule mapping; Embed the dual-unit channel of ethical supervision and scientific deduction into the large model, and jointly train the large model with compliant samples and non-compliant samples to obtain a bimodal feature set; Design a sample primary verification engine to conduct preliminary sample screening, scoring, and confirmation, and output ethical compliance and scientific rationality reports; By inferring the dependency relationship between ethical supervision and scientific deduction of dual units through attention weight, a three-dimensional map is constructed, and a three-dimensional ethical constraint matrix is ​​dynamically linked; Use the three-dimensional map to locate and correct high-risk sample nodes, generate correction logs and integrate them into the three-dimensional ethical constraint matrix, and output the final correction report; The specific steps of extracting experimental features from experimental records, constructing a three-dimensional ethical constraint matrix, and outputting experimental risk ratings and rule mapping are as follows: The experimental records include operation logs, instrument parameters and sample identification; Use natural language processing tools to extract experimental features; Construct a three-dimensional ethical constraint framework, analyze experimental characteristics based on predefined rules, and locate experimental characteristics within the three-dimensional ethical constraint framework; Obtain multiple ethical guidelines from an external rule set, number each ethical guideline to form multiple rules and map them onto a three-dimensional ethical constraint framework; The experimental characteristics and rule mapping are integrated through the priority rule method, the experimental risk rating is output, and a three-dimensional ethical constraint matrix is ​​generated.

2. The method for generating an experimental report based on a large model according to claim 1, wherein: The dual-unit channel of ethical supervision and scientific deduction is implanted in the large model, and the large model is trained by jointly training the compliant samples and the non-compliant samples to obtain the dual-modal feature set. The specific steps are as follows: Select the pre-trained BERT model as the basic framework of the large model; Implanting ethical supervision units and scientific deduction units into the output layer of the BERT model; Use NLP tools to extract experimental text data from operation logs and sample identifiers, and extract experimental operation data from operation logs and instrument parameters; The three-dimensional ethical constraint matrix is ​​embedded into the fully connected layer of the ethics supervision unit through the attention mechanism for training, and the actual compliance score is output as the basis for evaluating the probability of sample violation; Obtain scientific knowledge base through external scientific resources, load it into the scientific deduction unit, and output the logical consistency score as the basis for verifying the logical error rate of the sample; Assign initial attention weights to ethical supervision units and scientific deduction units; Obtain compliance samples and non-compliance samples from historical laboratory reports; Divide the compliant samples and the noncompliant samples into training sets and validation sets in proportion; The training set is input into the BERT model for training, and the attention weight is adjusted to achieve a balance between the ethical supervision unit and the scientific deduction unit; The sample violation probability and attention weight distribution are output and combined into a bimodal feature set.

3. The method for generating an experimental report based on a large model according to claim 2, wherein: The sample primary verification engine is designed to perform preliminary sample screening, scoring and confirmation, and output ethical compliance and scientific rationality reports. The specific steps are: Extract the sample violation probability from the bimodal feature set, set the violation threshold, and preliminarily screen high-risk samples as a first-level review; Use the BERT model to perform text analysis on high-risk samples for the first-level review, generating compliance scores and logical consistency scores for each high-risk sample as a second-level review; Conduct a comprehensive review of high-risk samples from the second-level review, identify violation samples, and generate violation conclusions as the third-level review; Comprehensive first-level review, second-level review and third-level review to output ethical compliance and scientific rationality reports.

4. The method for generating an experimental report based on a large model according to claim 3, wherein: The method of inferring the dependency relationship between ethical supervision and scientific deduction of dual units through attention weights, constructing a three-dimensional map, and dynamically linking the three-dimensional ethical constraint matrix is ​​as follows: The dependency relationship between the ethical supervision unit and the scientific deduction unit in the BERT model is inferred using the attention weight distribution, compliance score, and logical consistency score. Traverse the bimodal feature set and high-risk samples, integrate each high-risk sample, calculate the contribution weights of the ethical supervision unit and the scientific deduction unit, and generate a dependency table; Sort the experimental text data and experimental operation data by time to generate the sequence of experimental steps; Load the BERT model, input the dependency table and three-dimensional ethical constraint matrix, project the attention weight distribution, compliance score, logical consistency score and experimental step sequence into the three-dimensional coordinate system, dynamically link the numbering rules in the three-dimensional ethical constraint matrix, and output a three-dimensional map.

5. The method for generating an experimental report based on a large model according to claim 4, wherein: The three-dimensional coordinate system refers to the X-axis representing the order of experimental steps, the Y-axis representing the compliance score and the logical consistency score, and the Z-axis representing the attention weight and the sample violation probability.

6. The method for generating an experimental report based on a large model according to claim 5, wherein: The specific steps of using the three-dimensional map to locate and correct high-risk sample nodes, generate correction logs and integrate them into the three-dimensional ethical constraint matrix, and output the final correction report are as follows: Through the bimodal feature set of the BERT model and the three-dimensional ethical constraint matrix, we analyze the contribution of attention weight distribution to high-risk samples and mark high-risk sample nodes; Locate high-risk sample nodes in the three-dimensional map, analyze the dependencies of high-risk sample nodes, make correction suggestions, and output correction logs; Integrate the correction log into the three-dimensional ethical constraint matrix, update the rule mapping of the three-dimensional ethical constraint matrix, and output the final correction report.

7. A large-scale model-based experimental report generation device, based on the large-scale model-based experimental report generation method according to any one of claims 1 to 6, characterized in that: include, The risk module is used to extract experimental features from experimental records, construct a three-dimensional ethical constraint matrix, and output experimental risk ratings and rule mappings; The training module is used to embed the dual-unit channel of ethical supervision and scientific deduction into the large model. The large model is trained by jointly training the compliant and non-compliant samples to obtain a bimodal feature set. Design module, used to design a sample primary verification engine, conduct preliminary sample screening, scoring and confirmation, and output ethical compliance and scientific rationality reports; A building block is used to infer the dependency relationship between ethical supervision and scientific deduction through attention weights, construct a three-dimensional map, and dynamically link the three-dimensional ethical constraint matrix; The correction module is used to locate and correct high-risk sample nodes using the three-dimensional map, generate correction logs and integrate them into the three-dimensional ethical constraint matrix, and output the final correction report.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for generating an experimental report based on a large model according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating an experimental report based on a large model according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Image report generation method and model training method

    CN118072898A

  • System and methods for safe, scalable, artificial general intelligence (AGI)

    WO2024182285A2