Large language model-based ethical examination method and device
By performing secondary pre-training and fine-tuning of large language models, combined with search enhancement and review rule databases, optimized ethical review results are generated, and the existing model lacks understanding and analysis capabilities in ethical review are solved, and efficient ethical review is achieved.
Patent Information
- Application Number
- CN202510082900.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-06
AI Technical Summary
The existing general-purpose large language model is difficult to understand ethical norms in ethical review, cannot effectively conduct professional ethical analysis, and is limited in the search and analysis ability of cross-complex documents.
By performing secondary pre-training, secondary instruction fine-tuning and human feedback reinforcement learning on existing large language models, a well-trained domain model is obtained. Combining the search enhancement method and the review rule base, preliminary review results of project documents are generated, and preliminary review results are analyzed in-depth through the domain model to generate optimized review results.
It realizes rapid and efficient ethical review of project documents, provides in-depth ethical analysis and optimized review results, and assists reviewers to conduct fast and efficient ethical review.
Smart Images

Figure CN120106189A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the intersection of artificial intelligence and ethics, and specifically relates to an ethics review method and device based on a large language model. Background Art
[0002] In recent years, with the rapid advancement of scientific research and technological development, the demand for project ethics review has increased significantly. As an important means to ensure the legality, social responsibility and ethical standards of research, ethics review has become increasingly critical. However, the current ethics review model mainly relies on manual completion. This traditional method has exposed many shortcomings in the face of increasing demand and complexity.
[0003] At the same time, large language models have demonstrated excellent problem-solving capabilities in many fields, and their application range is becoming increasingly wide. Not only do they perform well in natural language processing-related tasks such as translation and sentiment analysis, they also show great potential in areas such as legal aid, medical diagnosis, and financial consulting.
[0004] Although the development of large language models has provided unprecedented technical possibilities for ethical review, there are still some technical challenges in practical application. On the one hand, due to the lack of professional knowledge for ethical review, existing general large language models are difficult to deeply understand ethical norms and cannot effectively conduct professional ethical analysis. On the other hand, the fine-tuned model is subject to certain limitations in its ability to retrieve and analyze across complex documents.
[0005] Therefore, how to build an efficient, accurate and comprehensive ethical review assistance system based on the large language model has become a technical challenge that needs to be solved urgently. Summary of the invention
[0006] The present invention is made to solve the above-mentioned problems, and its purpose is to provide an ethical review method and device based on a large language model.
[0007] The present invention provides an ethical review method based on a large language model, which is used to generate review results corresponding to project documents according to a review rule library, and has the following characteristics, including the following steps: Step S1, based on existing training data, the existing large language model is subjected to secondary pre-training, secondary instruction fine-tuning and human feedback reinforcement learning to obtain the trained large language model as a domain model; Step S2, the project document is input into the existing general model, and combined with the retrieval enhancement method and the review rule library, the preliminary review result corresponding to the project document is obtained; Step S3, the preliminary review result is input into the domain model, and combined with the review rule library, the review result is obtained.
[0008] The ethical review method based on the large language model provided by the present invention may also have the following features: wherein the review rule library includes a plurality of different review rule sub-libraries, and the general model is based on the review rule sub-libraries. Generate project documentation Corresponding preliminary review results The calculation expression is: , where It is a general model that adopts the retrieval enhanced generation method. The preliminary review results include the review rule-related information extracted from the project documents according to the review rules of each review rule sub-library, and the preliminary judgment results of each review rule-related information. The preliminary judgment result is whether the project document violates the corresponding review rule-related information.
[0009] The ethical review method based on the large language model provided by the present invention may also have the following features: wherein the domain model is based on the review rule sub-library and preliminary review results Generate corresponding review results The calculation expression is: , where For a domain model, the review results include analysis results of the relevant information of each review rule and the relevant content of its corresponding project documents, as well as improvement suggestions. The analysis result is whether the relevant content violates the relevant information of the corresponding review rule. When the analysis result is that the relevant content violates the relevant information of the corresponding review rule, the domain model generates corresponding improvement suggestions in combination with the relevant information of the review rule.
[0010] In the ethics review method based on the large language model provided by the present invention, it can also have the following characteristics: wherein, the existing training data includes existing ethics field data, the ethics field data includes data of ethics-related books, papers and policy and regulations, and general corpus data, and the secondary pre-training is to pre-train the large language model through the ethics field data and in combination with the autoregressive language modeling target, and the calculation expression of the autoregressive language modeling target is: , where is the length of the sequence, For the A word.
[0011] In the ethics review method based on the large language model provided by the present invention, it can also have the following characteristics: wherein, the existing training data includes an existing ethics instruction data set, the ethics instruction data set includes data of ethics-related examination questions, data of ethics case analysis, data of ethics review, and general instruction data, and the secondary instruction fine-tuning is to supervise and fine-tune the large language model completed by the secondary pre-training through the ethics instruction data set and in combination with the supervised learning objective function, and the calculation expression of the supervised learning objective function is: , where For the ethics instruction dataset, is the total number of training samples in the ethics instruction dataset, is the input instruction text in the training sample, Enter instruction text for training samples The corresponding model expected output.
[0012] In the ethics review method based on the large language model provided by the present invention, it can also have the following characteristics: wherein, the existing training data includes feedback data with manual annotations, and the human feedback reinforcement learning is to train the reward model through the feedback data, and adopt the trained reward model, and use the reinforcement learning method combined with the optimization objective function to optimize the parameters of the large language model of the secondary instruction fine-tuning to obtain the domain model, and the calculation expression of the training objective function of the training reward model is: , where Input data for the model in the feedback data, Input data for the model The corresponding high-quality model generation results, Input data for the model The corresponding suboptimal model generates results, Input data to the model for the reward model and high-quality model generation results The rating results, Input data to the model for the reward model and suboptimal models generate results The rating results of is the logistic function, and the calculation expression of the optimization objective function is: , , where Input data to the model for the reward model The large language model fine-tuned with the second instruction inputs data to the model Generated Output The rating results of is the trade-off factor, For the current strategy, For reference strategy, is the relative entropy.
[0013] The present invention also provides an ethical review device based on a large language model, which is used to generate review results corresponding to project documents according to a review rule library, and has the following characteristics, including: a data storage module, which stores a domain model, an existing general model and a review rule library; a preliminary review module, which is used to input the project document into the general model, and combine the retrieval enhancement method and the review rule library to obtain the preliminary review results corresponding to the project document; a review module, which is used to input the preliminary review results into the domain model, and combine the review rule library to obtain the review results, wherein the domain model is obtained by performing secondary pre-training, secondary instruction fine-tuning and human feedback reinforcement learning on the existing large language model using existing training data.
[0014] In the ethics review device based on the large language model provided by the present invention, it can also have such a feature, and also includes: a consultation judgment module, a first reply module and a second reply module, wherein the consultation judgment module is used to input the user's question into the general model to determine whether the question needs to obtain the corresponding project document related information. If so, the question is used as the first data, if not, the question is used as the second data. The first reply module is used to input the first data and the corresponding project document into the general model to obtain the corresponding project document related information, and input the first data and the project document related information into the domain model to obtain the answer to the question. The second reply module is used to input the second data into the domain model to obtain the answer to the question.
[0015] Functions and Effects of the Invention According to the ethics review method and device based on a large language model involved in the present invention, on the one hand, the general model and retrieval enhancement method are used to realize rapid parsing and retrieval of project documents and provide preliminary review results; on the other hand, the preliminary review results are deeply analyzed through a domain model with professional knowledge to generate optimized review results. Therefore, the ethics review method and device based on a large language model of the present invention can assist reviewers in conducting fast and efficient ethics reviews. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a block diagram of the ethics review device in an embodiment of the present invention.
[0017] Figure 2 Schematic diagram of domain model training and fine-tuning in an embodiment of the present invention.
[0018] Figure 3 It is a flowchart of an ethical review method based on a large language model in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and the accompanying drawings specifically illustrate the ethical review method and device based on a large language model of the present invention.
[0020] This embodiment provides an ethics review device based on a large language model, hereinafter referred to as an ethics review device, which is used to generate a review result corresponding to a project document according to a review rule library. The review rule library includes a plurality of different review rule sub-libraries, which are expressed as follows: , In the formula To review the rule base, For the A review rule sub-library.
[0021] Figure 1 It is a block diagram of the ethics review device in an embodiment of the present invention.
[0022] like Figure 1 As shown, the ethics review device 100 includes a data storage module 11, a preliminary review module 12, a review module 13, a consultation and judgment module 14, a first reply module 15, a second reply module 16 and a general control module 17 that controls the operation of the above modules.
[0023] The data storage module 11 stores domain models, existing general models and review rule bases.
[0024] Figure 2 Schematic diagram of domain model training and fine-tuning in an embodiment of the present invention.
[0025] like Figure 2 As shown, the domain model, namely the ethics domain model, is obtained by performing secondary pre-training, secondary instruction fine-tuning and human feedback reinforcement learning on the existing large language model, namely the base model, using existing training data. In this embodiment, the large language model is an open source large language model.
[0026] The existing training data includes existing ethics field data, existing ethics instruction data sets, and manually annotated feedback data. Ethics field data includes data on ethics-related books, papers, policies and regulations, and general corpus data. Ethics instruction data sets include data on ethics-related exam questions, data on ethics case analysis, data on ethics review, and general instruction data.
[0027] Secondary pre-training is to pre-train the large language model using ethics data combined with the autoregressive language modeling objective. The calculation expression of the autoregressive language modeling objective is: , In the formula is the length of the sequence, For the A word.
[0028] Secondary instruction fine-tuning is to fine-tune the large language model completed by secondary pre-training through the ethics instruction dataset and the supervised learning objective function. The calculation expression of the supervised learning objective function is: , In the formula For the ethics instruction dataset, is the total number of training samples in the ethics instruction dataset, is the input instruction text in the training sample, Enter instruction text for training samples The corresponding model expected output.
[0029] Human feedback reinforcement learning trains the reward model through feedback data, and uses the trained reward model to optimize the parameters of the large language model fine-tuned by secondary instructions using reinforcement learning methods combined with the optimization objective function to obtain the domain model.
[0030] In this embodiment, the reward model is trained using the large language model fine-tuned by the secondary instructions, and then the trained reward model is used to train the large language model, and then the trained large language model is used to train the reward model again, and this is repeated iteratively to obtain the final trained large language model as the domain model.
[0031] The calculation expression of the training objective function of the training reward model is: , In the formula Input data for the model in the feedback data, Input data for the model The corresponding high-quality model generation results, Input data for the model The corresponding suboptimal model generates results, Input data to the model for the reward model and high-quality model generation results The rating results of Input data to the model for the reward model and suboptimal models generate results The rating results of is the logistic function.
[0032] The calculation expression of the optimization objective function is: , , In the formula Input data to the model for the reward model And the large language model fine-tuned by the second instruction inputs data to the model Generated Output The rating results, is the trade-off factor, is the current strategy, which indicates the probability distribution currently generated by the model. is the reference strategy, an early version of the model training, is the relative entropy, i.e., KL divergence. This optimization objective is used to prevent the strategy Deviating too far from the reference strategy .
[0033] The preliminary review module 12 is used to input the project document into the general model, and combine the retrieval enhancement method and the review rule library to obtain the preliminary review results corresponding to the project document.
[0034] Among them, the general model is based on the review rule sub-library Generate project documentation Corresponding preliminary review results The calculation expression is: , In the formula A general model that adopts retrieval-augmented generation methods.
[0035] The preliminary review result includes the review rule related information extracted from the project document according to the review rules of each review rule sub-library, and the preliminary determination result of each review rule related information. The preliminary determination result is whether the project document violates the corresponding review rule related information.
[0036] For example, when the current review rule sub-library contains the review rule "In the collection of samples, or the personnel who carry out intervention measures need to have the corresponding professional qualifications", the general model extracts all the content related to "sample collection personnel" in the project document, such as "In the process of the study, professional researchers from the XX Institute need to draw blood from participants" as relevant information for the review rule. At the same time, the general model lacks the necessary professional knowledge training. Without knowing the specific sampling qualification requirements, it may determine that the project corresponding to the project document uses "blood collection by professional researchers" to meet the review point requirements, that is, the preliminary judgment result gives the result that "blood collection by professional researchers" meets the review.
[0037] The review module 13 is used to input the preliminary review results into the domain model and obtain the review results in combination with the review rule base.
[0038] Among them, the domain model is based on the review rule sub-library and preliminary review results Generate corresponding review results The calculation expression is: , In the formula For the domain model.
[0039] The review results include the analysis results of the relevant information of each review rule and the relevant content of the corresponding project documents, as well as improvement suggestions. The analysis result is whether the relevant content violates the relevant information of the corresponding review rule. When the analysis result is that the relevant content violates the relevant information of the corresponding review rule, the domain model combines the relevant information of the review rule to generate corresponding improvement suggestions.
[0040] For example, after the domain model has learned the professional knowledge content of "The Ethical Review Measures for Life Science and Medical Research Involving Humans 2023 issued by the National Science and Technology Ethics Committee: Article 19 (ii) Whether the qualifications, experience, and technical capabilities of the researcher meet the research requirements" and "Requirements for blood sample collection to be conducted by nurses from professionally qualified hospitals", the analysis results include that "blood collection by professional researchers" violates the review rules, and corresponding modification suggestions are given based on the above professional knowledge content.
[0041] The consultation judgment module 14 is used to input the user's question into the general model, and judge whether the question needs to obtain the corresponding project document related information. If so, the question is used as the first data, and if not, the question is used as the second data.
[0042] For example, when the question is "What privacy protection measures does this project use to protect personal privacy samples or data?", the general model will judge this question as the second data. When the question is "Is it necessary to clarify possible experimental risks in the informed consent form?", the general model will judge this question as the first data.
[0043] The first answering module 15 is used to input the first data and the corresponding project document into the general model to obtain the corresponding project document related information, and input the first data and the project document related information into the domain model to obtain the answer to the question.
[0044] The second answering module 16 is used to input the second data into the domain model to obtain an answer to the question.
[0045] The master control module 17 stores a control program for controlling the operation of each module.
[0046] The following describes the process of using the ethics review device 100 to perform an ethics review method based on a large language model in conjunction with the accompanying drawings.
[0047] Figure 3 It is a flowchart of an ethical review method based on a large language model in an embodiment of the present invention.
[0048] like Figure 3 As shown in FIG. 1 , the ethics review method based on the large language model includes the following steps: Step S1, based on the existing training data, perform secondary pre-training, secondary instruction fine-tuning and human feedback reinforcement learning on the existing large language model to obtain a trained large language model as a domain model.
[0049] Step S2, using the preliminary review module 12 to input the project document into the existing general model, and combining the retrieval enhancement method and the review rule library to obtain the preliminary review results corresponding to the project document.
[0050] Step S3, using the review module 13 to input the preliminary review results into the domain model, combined with the review rule base, to obtain the review results.
[0051] Functions and Effects of the Embodiments According to the ethics review method and device based on the large language model involved in this embodiment, on the one hand, the general model and the retrieval enhancement method are used to realize the rapid parsing and retrieval of project documents and provide preliminary review results; on the other hand, the preliminary review results are deeply analyzed through the domain model with professional knowledge to generate optimized review results. In short, this method can assist reviewers in conducting fast and efficient ethics review.
[0052] Furthermore, the user's questions are analyzed through the consultation judgment module, and the first reply module and the second reply module are combined to efficiently generate replies corresponding to the questions, assisting researchers to further improve the project plan to make it meet ethical requirements.
[0053] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. An ethical review method based on a large language model, used to generate review results corresponding to project documents according to a review rule library, characterized in that: The following steps are involved: Step S1, based on the existing training data, perform secondary pre-training, secondary instruction fine-tuning and human feedback reinforcement learning on the existing large language model to obtain a trained large language model as a domain model; Step S2, inputting the project document into the existing general model, and combining the search enhancement method and the review rule library to obtain a preliminary review result corresponding to the project document; Step S3, inputting the preliminary review result into the domain model, combining it with the review rule base, to obtain the review result.
2. The ethics review method based on a large language model according to claim 1, characterized in that: in, The review rule library includes a plurality of different review rule sub-libraries. The general model is based on the review rule sub-library Generate project documentation Corresponding preliminary review results The calculation expression is: , In the formula A general model that adopts the retrieval-augmented generation method. The preliminary review result includes the review rule related information extracted from the project document according to the review rules of each review rule sub-library, and the preliminary determination result of each review rule related information. The preliminary determination result is information related to whether the project document violates the corresponding review rule.
3. The ethics review method based on a large language model according to claim 2, characterized in that: in, The domain model is based on the review rule sub-library and the preliminary review results Generate corresponding review results The calculation expression is: , In the formula For the domain model, The review results include analysis results of the relevant information of each review rule and the relevant content of the corresponding project documents, as well as improvement suggestions. The analysis result is information related to whether the relevant content violates the corresponding review rules. When the analysis result is that the relevant content violates the corresponding information related to the review rules, the domain model generates the corresponding improvement suggestions in combination with the relevant information of the review rules.
4. The ethics review method based on a large language model according to claim 1, characterized in that: in, The existing training data includes existing ethics field data, The ethics field data includes data on ethics-related books, papers, policies and regulations, and general corpus data. The secondary pre-training is to pre-train the large language model using the ethics field data in combination with an autoregressive language modeling objective. The calculation expression of the autoregressive language modeling objective is: , In the formula is the length of the sequence, For the A word.
5. The ethics review method based on a large language model according to claim 1, characterized in that: in, The existing training data includes an existing ethics instruction dataset, The ethics instruction data set includes data on ethics-related examination questions, data on ethics case analysis, data on ethics review, and general instruction data. The secondary instruction fine-tuning is to perform supervised fine-tuning on the large language model completed by secondary pre-training by using the ethics instruction dataset and combining it with a supervised learning objective function. The calculation expression of the supervised learning objective function is: , In the formula For the ethics instruction dataset, is the total number of training samples in the ethics instruction dataset, is the input instruction text in the training sample, Enter instruction text for training samples The corresponding model expected output.
6. The ethics review method based on a large language model according to claim 1, characterized in that: in, The existing training data includes manually annotated feedback data. The human feedback reinforcement learning is to train the reward model through the feedback data, and use the trained reward model to optimize the parameters of the large language model fine-tuned by the secondary instruction using the reinforcement learning method combined with the optimization objective function to obtain the domain model. The calculation expression of the training objective function of training the reward model is: , In the formula Input data for the model in the feedback data, Input data for the model The corresponding high-quality model generation results, Input data for the model The corresponding suboptimal model generates results, Input data to the model for the reward model and high-quality model generation results The rating results, Input data to the model for the reward model and suboptimal models generate results The rating results, is the logistic function, The calculation expression of the optimization objective function is: , , In the formula Input data to the model for the reward model And the second instruction fine-tunes the large language model to the model input data Generated Output The rating results of is the trade-off factor, For the current strategy, For reference strategy, is the relative entropy.
7. An ethics review device based on a large language model, used to generate review results corresponding to project documents according to a review rule library, characterized in that: include: A data storage module storing a domain model, an existing general model and the review rule library; A preliminary review module, used to input the project document into the general model, and combine the search enhancement method and the review rule library to obtain a preliminary review result corresponding to the project document; A review module is used to input the preliminary review result into the domain model and combine it with the review rule base to obtain the review result. The domain model is obtained by performing secondary pre-training, secondary instruction fine-tuning and human feedback reinforcement learning on the existing large language model using existing training data.
8. The ethics review device based on a large language model according to claim 7 is characterized in that: Also includes: Consultation judgment module, first reply module and second reply module, The consultation judgment module is used to input the user's question into the general model, and judge whether the question needs to obtain the corresponding project document related information. If so, the question is used as the first data, and if not, the question is used as the second data. The first answer module is used to input the first data and the corresponding project document into the general model to obtain the corresponding project document related information, and input the first data and the project document related information into the domain model to obtain the answer to the question, The second answering module is used to input the second data into the domain model to obtain an answer to the question.
Citation Information
Cited By
Functional security auxiliary authentication method and system based on large language model
CN121234904A
A biosafety risk assessment intelligent agent construction method
CN122734740A