AI evaluation auxiliary system based on RLHF

The AI ​​evaluation assistance system based on RLHF solves the problem of inaccuracy in behavior evaluation after AI system deployment, achieves accurate assessment and controllability of AI system behavior, ensures the accuracy and standardization of evaluation data, and improves the accuracy of subsequent model evaluations.

CN120952095APending Publication Date: 2025-11-14SHENZHEN WANGAN COMP SECURITY DETECTION TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511160317.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies lack mature and systematic RLHF (Real-Time Risk Detection and Handling) mechanisms for behavioral evaluation after AI system deployment, making it difficult to identify complex behavioral patterns in real time, affecting the controllability and security of AI systems. Furthermore, traditional methods lack accurate measurement of behavioral risks.

Method used

Design an AI assessment assistance system based on RLHF, including on-site assessment module, unit assessment module, level assessment and analysis module, overall assessment module, report compilation module, and report review module. Through speech transcription, data archiving, feedback processing, multi-level analysis, and report generation, the system optimizes model behavior by combining the RLHF mechanism.

Benefits of technology

It enables precise evaluation and controllability of AI system behavior, ensures the accuracy and standardization of evaluation data, and improves the accuracy and reliability of the model in subsequent evaluation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952095A_ABST
    Figure CN120952095A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of evaluation, and particularly discloses an RLHF-based AI evaluation auxiliary system, and the system comprises a field evaluation module which is used for recording key data of each evaluation item, and carrying out the data archiving and processing, and obtaining field evaluation data; the unit evaluation module is used for processing and feeding back based on the evaluation text; the level evaluation analysis module is used for acquiring a plurality of unit evaluation records, performing classification analysis according to the safety control level to which the unit evaluation records belong, and performing evaluation feedback on each level; the overall evaluation module is used for carrying out feedback, modification and adoption or regeneration based on overall evaluation result analysis; the report compiling module is used for combining the overall evaluation results to generate an evaluation report first draft and auditing the evaluation report first draft; and the report auditing module is used for carrying out learning adjustment based on an RLHF mechanism to generate an evaluation report, strengthening the precision and controllability of a model generation behavior under the RLHF mechanism, reversely optimizing an AI behavior through the RLHF mechanism, and improving the accuracy of the AI model in a subsequent evaluation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of evaluation technology, and more specifically to an AI evaluation assistance system based on RLHF. Background Technology

[0002] In recent years, RLHF (Reinforcement Learning with Human Feedback), which combines Reinforcement Learning (RL) with Human Feedback (HF), has gradually become one of the core technologies in Large Language Model (LLM) training. RLHF introduces human ratings or preference feedback to construct a reward function, guiding the model to optimize its behavioral strategies, significantly improving the rationality and alignment of the model's generated results. During the training phase, RLHF technology has been widely used to enhance the AI ​​model's responsiveness to human intentions.

[0003] However, in the post-deployment evaluation and behavioral assessment of models, there is a lack of mature and systematic RLHF (Real-Time Risk Assessment) mechanisms to evaluate and monitor the actual interactive behavior of the models. Furthermore, during the actual operation of AI systems, current behavioral assessment schemes often rely on manually set rules or static detection mechanisms, making it difficult to identify and score complex behavioral patterns in real time. This affects the controllability and security of AI systems. In behavioral risk assessment, traditional methods lack weighting mechanisms for the degree of risk of the behavior itself, the severity of the behavior type, and the criticality of the actor, thus failing to accurately measure the impact of behavior on the overall operational stability of the system. Summary of the Invention

[0004] The purpose of this invention is to provide an AI assessment assistance system based on RLHF to solve the problems mentioned above.

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] The AI-assisted assessment system based on RLHF includes the following modules:

[0007] On-site assessment module: Based on the assessment manual, key data of each assessment item is recorded during the execution process, and the key assessment data is transcribed and recognized by speech. At the same time, the data is archived and processed to obtain on-site assessment data.

[0008] Unit assessment module: acquires on-site assessment data, generates prompt words based on predefined templates, generates assessment text based on prompt words and sends it to assessors, processes and provides feedback based on the assessment text, and obtains unit assessment records;

[0009] Level-based assessment and analysis module: Acquire multiple unit assessment records, classify and analyze them according to their respective security control levels, and then analyze them based on the prompts of the unit assessment results at each level, providing evaluation feedback for each level;

[0010] Overall evaluation module: It records and inputs the unit evaluation results into the AI ​​big model according to the level, outputs "overall evaluation result analysis", and provides feedback, modification and adoption or regeneration based on the overall evaluation result analysis;

[0011] Report preparation module: Generates an assessment report based on the results of all unit assessments, merges the overall assessment results to generate a draft assessment report, and reviews the draft assessment report;

[0012] Report review module: The initial draft of the evaluation report is input into the AI ​​big model for quality review, outputting format and logic issues. Combined with human review issues, it is returned for revision, and finally, the evaluation report is generated based on the RLHF mechanism for learning and adjustment.

[0013] As a further aspect of the present invention, the steps of the on-site evaluation module include:

[0014] Step 1: Obtain key assessment data based on assessment items, perform speech-to-text recognition on key assessment data, and extract the standard number and content of assessment items;

[0015] Step 2: Search for content that matches the assessment item based on the pre-set internal template library and present it as a prompt card;

[0016] Step 3: Annotate the assessment based on the content matching the assessment items on the prompt card;

[0017] Step 4: Record assessments and perform operations based on assessment annotations, including: dynamically matching assessment item prompts, completing assessment records, providing language prompts, and outputting assessment results.

[0018] As a further aspect of the present invention: the key evaluation data obtained from the evaluation manual includes, but is not limited to, access control policies, audit data, and key data; the recording of key evaluation data can adopt a multi-segment expression method, such as a three-segment format: evaluation item name, performance status, and supporting evidence.

[0019] As a further aspect of the present invention, the data archiving and processing steps include:

[0020] Based on the on-site evaluation content, terminal configurations, security policies, audit configurations, etc., are uploaded by taking photos and scanning; the system automatically labels and categorizes them, and matches the images with the corresponding evaluation items to form archived data;

[0021] Then, based on the archived data, the image content is identified, keywords are extracted, and a structured transformation is performed.

[0022] As a further aspect of the present invention: in the unit evaluation module, the feedback processing based on the evaluation text is as follows:

[0023] The evaluators modify and provide feedback on semantic biases based on the evaluation text, forming a feedback reinforcement learning loop. The methods of modification and feedback include: editing feedback, selective feedback, and labeled feedback.

[0024] Among them, editorial feedback refers to directly editing and modifying the test text online, comparing the modified record directly with the original text, and then feeding back to the system for recording and archiving;

[0025] Selective feedback refers to selecting the inaccurate reason and regenerating based on the inaccurate reason;

[0026] Labeled feedback means that the record meets the requirements and actual needs; at the same time, it can be used as an item in the sample training set.

[0027] As a further aspect of the present invention: the security control layer in the layer assessment and analysis module includes: technical layer, management layer and physical layer;

[0028] The categorization and analysis of unit assessment records includes:

[0029] First, the unit test records are categorized based on the test item standards, which include the test item number, test text, and security control level.

[0030] Then, semantic extraction and identification are performed on each unit evaluation record. The extraction technology identifies the following information: control measures, implementation effectiveness, problem level, and risk identification.

[0031] As a further aspect of the present invention, the working steps of the overall evaluation module include:

[0032] Obtain multi-level evaluation results, including safety control measures analysis and safety problem analysis, and input the multi-level evaluation results into the large value model;

[0033] The AI ​​big model summarizes the multi-layered evaluation results, then generates structured prompts, and combines the structured prompts to generate the overall evaluation results.

[0034] The structured prompts include overall security control capabilities, overall security issues, and overall risks.

[0035] Feedback is provided based on the overall assessment results, and then the RLHF mechanism is used to learn and adjust based on the feedback information.

[0036] The feedback loop mechanism includes visual traceability, review label feedback, and polishing and regeneration.

[0037] As a further aspect of the present invention, the working steps of the report preparation module include:

[0038] The unit evaluation records, the summary of the level evaluation, the overall evaluation analysis and the extrusion are fitted together. Then, based on the template structure, the AI ​​large model is called to generate text, and the text is formatted in a unified way. Finally, an evaluation report is generated.

[0039] The template content structure includes evaluation information, system overview, evaluation methods and processes, system security control, analysis of major problems, evaluation conclusions, and rectification suggestions;

[0040] The methods for reviewing and approving evaluation reports include: editing, regenerating, and restructuring the evaluation report.

[0041] As a further aspect of the present invention, the quality review in the report review module specifically includes: content authenticity check, expression standardization check, structural integrity check, logical consistency check, data accuracy check, and professional review.

[0042] As a further aspect of the present invention, the method for learning and adjusting using the RLHF mechanism specifically includes the following steps:

[0043] Collect feedback signals: Acquire and automatically record the feedback behavior of evaluators and reviewers on each AI-generated content, including: direct adoption without modification, modification, editing, marking and regeneration;

[0044] Building a reward model: Collect feedback information groups to form the core data for training the "reward model" and then train it;

[0045] Policy fine-tuning: Fine-tuning the original AI model based on the reward model.

[0046] The beneficial effects of this invention are:

[0047] In this invention, key content is acquired and processed based on on-site information, ensuring the accuracy of the data structure and the integrity of the semantic content. This ensures that the evaluation content is based on real conditions, solves the problems of inaccurate or unfounded content in the evaluation process, and guarantees the credibility and standardization of the evaluation data.

[0048] By using AI to globally abstract and aggregate multi-layered assessment content, assessors can quickly construct a "system-level security status judgment" based on authenticity and compliance, while enhancing the accuracy and controllability of model generation behavior under the RLHF mechanism.

[0049] By employing an AI-guided generation and evaluation personnel review mechanism, we ensure that the content is fluent, the language is standardized, the source is authentic, and there is no fabrication, thus guaranteeing the authenticity and accuracy of the evaluation. Then, based on the AI ​​big model, we conduct language compliance review, logical consistency verification, and content source tracing check. Human review opinions are fed back into the AI ​​model to optimize the accuracy and reliability of its subsequent output. The RLHF mechanism is used to optimize the AI ​​behavior in reverse, improving the accuracy of the AI ​​model in the subsequent evaluation process. Attached Figure Description

[0050] The invention will now be further described with reference to the accompanying drawings.

[0051] Figure 1 This is a schematic diagram of the system of the present invention;

[0052] Figure 2 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] Example 1

[0055] Please see Figure 1 As shown, this invention is an AI assessment assistance system based on RLHF, comprising the following modules:

[0056] On-site assessment module: Based on the assessment manual, key data of each assessment item is recorded during the execution process, and the key assessment data is transcribed and recognized by speech. At the same time, the data is archived and processed to obtain on-site assessment data.

[0057] Unit assessment module: acquires on-site assessment data, generates prompt words based on predefined templates, generates assessment text based on prompt words and sends it to assessors, processes and provides feedback based on the assessment text, and obtains unit assessment records;

[0058] Level-based assessment and analysis module: Acquire multiple unit assessment records, classify and analyze them according to their respective security control levels, and then analyze them based on the prompts of the unit assessment results at each level, providing evaluation feedback for each level;

[0059] Overall evaluation module: It records and inputs the unit evaluation results into the AI ​​big model according to the level, outputs "overall evaluation result analysis", and provides feedback, modification and adoption or regeneration based on the overall evaluation result analysis;

[0060] Report preparation module: Generates an assessment report based on the results of all unit assessments, merges the overall assessment results to generate a draft assessment report, and reviews the draft assessment report;

[0061] Report review module: The initial draft of the evaluation report is input into the AI ​​big model for quality review, outputting format and logic issues. Combined with human review issues, it is returned for revision, and finally, the evaluation report is generated based on the RLHF mechanism for learning and adjustment.

[0062] Example 2

[0063] like Figure 2 As shown, based on the above embodiments, this embodiment provides an AI assessment assistance method based on RLHF, including the following steps:

[0064] Step 1: The evaluators record key data for each evaluation item during the execution of the evaluation according to the evaluation manual, and perform voice transcription and recognition of the key evaluation data. At the same time, the data is archived and processed to obtain the on-site evaluation data.

[0065] Specifically, the steps of the on-site assessment module include:

[0066] Step 1: The evaluators obtain key evaluation data based on the evaluation items, perform speech-to-text recognition on the key evaluation data, and extract the standard number and content of the evaluation items;

[0067] Step 2: Based on the pre-set internal template library, find content that matches the assessment item and present it to the assessor in the form of a prompt card;

[0068] Step 3: The assessors provide assessment annotations based on the content matching the assessment items on the prompt cards;

[0069] Step 4: Record assessments and perform operations based on assessment annotations, including: dynamically matching assessment item prompts to execute methods, completing assessment records, providing language prompts, and outputting assessment results;

[0070] The key assessment data obtained by assessors based on the assessment manual includes, but is not limited to, access control policies, audit data, and key data; the recording of key assessment data can adopt a multi-segment expression method, such as a three-segment format: assessment item name, performance status, and supporting evidence;

[0071] In addition, during the process of speech transcription and semantic recognition, the voice recordings of on-site assessment and communication are extracted using a natural language understanding model to extract core content, such as keyword extraction and extraction of key assessment data.

[0072] In addition, the steps for data archiving and processing include:

[0073] Based on the on-site evaluation content, terminal configurations, security policies, audit configurations, etc., are uploaded by taking photos and scanning; the system automatically labels and categorizes them, and matches the images with the corresponding evaluation items to form archived data;

[0074] Then, based on the archived data, the image content is identified, keywords are extracted, and structural transformation is performed;

[0075] To ensure semantic accuracy in subsequent unit assessments, this step requires adherence to a specific structured format.

[0076] Each assessment item's record is organized into the same structure:

[0077] For example: for the evaluation unit number, technical control, access control, and identity recognition; for the evaluation items, user identity should be authenticated; for supporting materials, attached pictures, screenshots, and documents; and for data sources, manual shorthand, audio transcription, and photography.

[0078] In this step, key content is acquired and processed based on records from the on-site information exchange process, such as notes, audio recordings, and photos. This ensures the accuracy of the data structure, the integrity of the semantic content, and that the evaluation content is based on the real situation. It also resolves inaccurate or unfounded content in the evaluation process and ensures the credibility and standardization of the evaluation data.

[0079] Step 2: Obtain on-site assessment data, generate prompt words based on predefined templates, generate assessment text based on prompt words and send it to assessors, who then process and provide feedback on the assessment text to obtain unit assessment records.

[0080] In this step, it should be noted that the predefined templates are set by the evaluators, such as "evaluation status, control test, judgment result", etc. At the same time, the prompt information also needs to meet certain data requirements during processing, such as the description of the evaluation information, the status of on-site implementation, and the requirements for generating the evaluation report.

[0081] The process by which evaluators process feedback based on the evaluation text is as follows:

[0082] The evaluators modify and provide feedback on semantic biases based on the evaluation text, forming a feedback reinforcement learning loop. The methods of modification and feedback include: editing feedback, selective feedback, and labeled feedback.

[0083] Among them, editorial feedback refers to directly editing and modifying the test text online, comparing the modified record directly with the original text, and then feeding back to the system for recording and archiving;

[0084] Selective feedback refers to selecting the inaccurate reason and regenerating based on the inaccurate reason;

[0085] Labeled feedback means that the record meets the requirements and actual needs; at the same time, it can be used as an item in the sample training set.

[0086] It should be noted that in this step, all the content generated by the unit assessments is stored in the database. The system uses the comparison function to summarize and learn from the differences between its own assessment text and assessment feedback, thus achieving reinforcement learning based on human feedback.

[0087] This step involves using AI models to verbally express, technically standardize, and evaluate the security of each assessment item based on on-site evaluation data. This crucial step of "generating compliant language content from real data" is also the core of the entire intelligent assessment document writing process.

[0088] Step 3: Obtain multiple unit evaluation records, classify and analyze them according to their respective security control levels, and then analyze them based on the prompts of the unit evaluation results at each level, and provide evaluation feedback for each level;

[0089] The security control layer includes: technical layer, management layer and physical layer;

[0090] The categorization and analysis of unit assessment records includes:

[0091] First, the unit test records are categorized based on the test item standards, which include the test item number, test text, and security control level.

[0092] Then, semantic extraction and identification are performed on each unit evaluation record. The extraction technology identifies the following information: control measures, implementation effectiveness, problem level, and risk identification.

[0093] Examples include: control measures (e.g., whether control measures exist); implementation effectiveness (e.g., whether implementation is sufficient); problem level (e.g., the severity of the problem, vulnerabilities, deviations, etc.); and risk identification (e.g., whether risks exist).

[0094] The method for analyzing prompts based on the unit evaluation results at each level includes: analyzing prompts based on safety control measures and safety issues.

[0095] The analysis of security control measure prompts includes: deploying multiple control measures, including a bastion host system for unified login and command auditing for operation and maintenance users, centralized authentication for unified account management, and the adoption of a two-factor authentication mechanism based on identity authentication and access control management;

[0096] The analysis of security question prompts also includes: adopting strategy configuration to increase the length and complexity of assessment texts and setting up session control mechanisms;

[0097] In addition, the evaluation feedback methods for each layer include: editing, regeneration, marking and recording, and referencing. In this step, editing refers to manual editing and optimization; regeneration refers to negating the content so that the system can regenerate the evaluation content; marking and recording essentially involves recording and marking the content, feeding it back to the system, and using it for RLHF training to improve the accuracy of subsequent evaluations; referencing refers to referencing the content. In this step, the system records each evaluation feedback as training material for continuous model optimization.

[0098] Step 4: Input the unit evaluation results into the AI ​​big model according to the level, and output "overall evaluation result analysis". The evaluators will make feedback, modification and adoption or regeneration based on the overall evaluation result analysis;

[0099] Specifically, the overall assessment module's workflow includes:

[0100] Obtain multi-level evaluation results, including safety control measures analysis and safety problem analysis, and input the multi-level evaluation results into the large value model;

[0101] The AI ​​big model summarizes the multi-layered evaluation results, then generates structured prompts, and combines the structured prompts to generate the overall evaluation results.

[0102] Specifically: structured prompts include overall security control capabilities, overall security issues, and overall risks;

[0103] The evaluators provide feedback based on the overall evaluation results, and then use the RLHF mechanism to learn and adjust based on the feedback information;

[0104] Specifically, the feedback loop mechanism includes visual traceability, review label feedback, and polishing and regeneration;

[0105] Visual traceability marks the source unit evaluation record for each generated content in reverse. Clicking on it allows you to view the related original text, helping the evaluator determine whether the content exceeds the basis.

[0106] Review tags are used to label entire paragraphs or sentences, and the next round of model optimization is based on the content of these tags.

[0107] Polishing and regeneration refers to the ability of evaluators to manually modify or regenerate based on the evaluation results. The purpose of the modification or regeneration can be annotated to enhance the ability to guide the large AI model.

[0108] This step uses AI to globally abstract and aggregate multi-layered assessment content, enabling assessors to quickly construct a "system-level security status judgment" based on authenticity and compliance, while strengthening the accuracy and controllability of model generation behavior under the RLHF mechanism;

[0109] Step 5: Generate an assessment report based on all unit assessment results records, merge the overall assessment results to generate a draft assessment report, and have the assessment personnel review the draft assessment report;

[0110] Specifically: The unit evaluation records, the summary of the level evaluation, the overall evaluation analysis and the extrusion are fitted together, and then based on the template structure, the AI ​​large model is called to generate text, and the text is formatted in a unified manner, and finally an evaluation report is generated.

[0111] The template content structure includes, but is not limited to, evaluation information, system overview, evaluation methods and processes, system security control, analysis of major issues, evaluation conclusions, and rectification suggestions.

[0112] The methods used by evaluators to review and approve evaluation reports include: editing, regenerating, and restructuring the evaluation report;

[0113] This step organizes scattered structured assessment content into a standardized assessment result text. Through AI-guided generation and assessment personnel review mechanism, it ensures that the content is fluent, the language is standardized, the source is authentic, and there is no fabrication, thus guaranteeing the authenticity and accuracy of the assessment.

[0114] Step Six: The reviewers input the initial draft of the evaluation report into the AI ​​model for quality review, outputting format and logic issues. Combined with the issues raised by the human reviewers, the report is returned to the evaluators for revision. Finally, the evaluation report is generated based on the RLHF mechanism for learning and adjustment.

[0115] The quality review specifically includes: checking the authenticity of the content, checking the standardization of expression, checking the integrity of the structure, checking the logical consistency, checking the accuracy of the data, and reviewing the professionalism.

[0116] Specifically: Content authenticity check: Check whether there is any generated content with an unclear source or that exceeds the logic of the original data, and suppress content that is fabricated out of thin air or subjectively conjectured;

[0117] Standardization check: Check whether the terminology is consistent with the template of the graded protection assessment report, whether the sentence structure is coherent and the grammar is correct;

[0118] Structural integrity check: Check if the report is missing any core content;

[0119] Logical consistency check: Check whether there are self-contradictions in the descriptions of each section, and whether the descriptions at multiple levels are consistent.

[0120] Data accuracy check: Check whether the date, evaluation period, and number of systems tested are correct;

[0121] Professional review: Check whether the language of the report is fluent and professional.

[0122] Specifically, the RLHF mechanism's learning and adjustment methods include the following steps:

[0123] Collect feedback signals: Acquire and automatically record the feedback behavior of evaluators and reviewers on each AI-generated content, specifically including: direct adoption without modification, modification, editing, marking, and regeneration;

[0124] Building a reward model: Collect feedback information groups to form the core data for training the "reward model" and then train it;

[0125] Policy fine-tuning: Fine-tuning the original AI model based on the reward model, specifically including the following methods: common policy optimization methods in reinforcement learning to control the model from deviating too far from the original semantics; learning the preference function directly based on the human preference distribution, without relying on the simulation environment.

[0126] This step utilizes a large AI model to conduct language compliance reviews, logical consistency checks, and content source tracing. Human review feedback is then input into the AI ​​model to optimize the accuracy and reliability of its subsequent outputs. Furthermore, the RLHF mechanism is used to reverse-engineer the AI's behavior, improving its accuracy in subsequent evaluation processes.

[0127] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. An AI assessment assistance system based on RLHF, characterized in that, Includes the following modules: On-site assessment module: Based on the assessment manual, key data of each assessment item is recorded during the execution process, and the key assessment data is transcribed and recognized by speech. At the same time, the data is archived and processed to obtain on-site assessment data. Unit assessment module: acquires on-site assessment data, generates prompt words based on predefined templates, generates assessment text based on prompt words and sends it to assessors, processes and provides feedback based on the assessment text, and obtains unit assessment records; Level-based assessment and analysis module: Acquire multiple unit assessment records, classify and analyze them according to their respective security control levels, and then analyze them based on the prompts of the unit assessment results at each level, providing evaluation feedback for each level; Overall evaluation module: It records and inputs the unit evaluation results into the AI ​​big model according to the level, outputs "overall evaluation result analysis", and provides feedback, modification and adoption or regeneration based on the overall evaluation result analysis; Report preparation module: Generates an assessment report based on the results of all unit assessments, merges the overall assessment results to generate a draft assessment report, and reviews the draft assessment report; Report review module: The initial draft of the evaluation report is input into the AI ​​big model for quality review, outputting format and logic issues. Combined with human review issues, it is returned for revision, and finally, the evaluation report is generated based on the RLHF mechanism for learning and adjustment.

2. The AI ​​assessment assistance system based on RLHF according to claim 1, characterized in that, The steps of the on-site assessment module include: Step 1: Obtain key assessment data based on assessment items, perform speech-to-text recognition on key assessment data, and extract the standard number and content of assessment items; Step 2: Search for content that matches the assessment item based on the pre-set internal template library and present it as a prompt card; Step 3: Annotate the assessment based on the content matching the assessment items on the prompt card; Step 4: Record assessments and perform operations based on assessment annotations, including: dynamically matching assessment item prompts, completing assessment records, providing language prompts, and outputting assessment results.

3. The AI ​​assessment assistance system based on RLHF according to claim 1, characterized in that, The key data obtained from the evaluation manual includes, but is not limited to, access control policies, audit data, and key data; the recording of key evaluation data can be expressed in a multi-segment manner, such as a three-segment format: evaluation item name, performance status, and supporting evidence.

4. The AI ​​assessment assistance system based on RLHF according to claim 1, characterized in that, The steps of data archiving and processing include: Based on the on-site evaluation content, terminal configurations, security policies, audit configurations, etc., are uploaded by taking photos and scanning; the system automatically labels and categorizes them, and matches the images with the corresponding evaluation items to form archived data; Then, based on the archived data, the image content is identified, keywords are extracted, and a structured transformation is performed.

5. The AI ​​assessment assistance system based on RLHF according to claim 1, characterized in that, In the unit assessment module, the feedback processing based on the assessment text is as follows: The evaluators modify and provide feedback on semantic biases based on the evaluation text, forming a feedback reinforcement learning loop. The methods of modification and feedback include: editing feedback, selective feedback, and labeled feedback. Among them, editorial feedback refers to directly editing and modifying the test text online, comparing the modified record directly with the original text, and then feeding back to the system for recording and archiving; Selective feedback refers to selecting the reason for inaccuracy and regenerating based on that reason; labeled feedback refers to the record meeting the requirements and actual needs; at the same time, it can be used as an item in the sample training set.

6. The AI ​​assessment assistance system based on RLHF according to claim 1, characterized in that, The security control layers in the layer assessment and analysis module include: technical layer, management layer, and physical layer. The categorization and analysis of unit assessment records includes: First, the unit test records are categorized based on the test item standards, which include the test item number, test text, and security control level. Then, semantic extraction and identification are performed on each unit evaluation record. The extraction technology identifies the following information: control measures, implementation effectiveness, problem level, and risk identification.

7. The AI ​​assessment assistance system based on RLHF according to claim 1, characterized in that, The overall assessment module's workflow includes: Obtain multi-level evaluation results, including safety control measures analysis and safety problem analysis, and input the multi-level evaluation results into the large value model; The AI ​​big model summarizes the multi-layered evaluation results, then generates structured prompts, and combines the structured prompts to generate the overall evaluation results. The structured prompts include overall security control capabilities, overall security issues, and overall risks. Feedback is provided based on the overall assessment results, and then the RLHF mechanism is used to learn and adjust based on the feedback information. The feedback loop mechanism includes visual traceability, review label feedback, and polishing and regeneration.

8. The AI ​​assessment assistance system based on RLHF according to claim 1, characterized in that, The work steps of the report preparation module include: The unit evaluation records, the summary of the level evaluation, the overall evaluation analysis and the extrusion are fitted together. Then, based on the template structure, the AI ​​large model is called to generate text, and the text is formatted in a unified way. Finally, an evaluation report is generated. The template content structure includes evaluation information, system overview, evaluation methods and processes, system security control, analysis of major problems, evaluation conclusions, and rectification suggestions; The methods for reviewing and approving evaluation reports include: editing, regenerating, and restructuring the evaluation report.

9. The AI ​​assessment assistance system based on RLHF according to claim 1, characterized in that, The quality review in the report review module specifically includes: content authenticity check, expression standardization check, structural integrity check, logical consistency check, data accuracy check, and professional review.

10. The AI ​​assessment assistance system based on RLHF according to claim 1, characterized in that, The specific steps involved in the learning and adjustment process of the RLHF mechanism are as follows: Collect feedback signals: Acquire and automatically record the feedback behavior of evaluators and reviewers on each AI-generated content, including: direct adoption without modification, modification, editing, marking and regeneration; Building a reward model: Collect feedback information groups to form the core data for training the "reward model" and then train it; Policy fine-tuning: Fine-tuning the original AI model based on the reward model.

Citation Information

Cited By

  • Intelligent agent-based automatic verification and rule matching system for insurance-waiting evaluation report

    CN121304091A