Multi-agent personnel examination scoring method based on end-to-end

By employing an end-to-end multi-agent scoring method, utilizing a heterogeneous OCR engine and multi-dimensional scoring agents, the difficulties in text analysis and logical consistency in personnel examinations are resolved, achieving efficient and fair automated scoring and improving the efficiency and credibility of the evaluation process.

CN120975642APending Publication Date: 2025-11-18SHANDONG NUOMAXIN INFORMATION TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511121767.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing artificial intelligence technologies in personnel examinations suffer from difficulties in text analysis, insufficient comprehension, inability to ensure logical consistency between multiple dimensions and the total score, and failure to meet the requirements for fairness and appealability of results.

Method used

An end-to-end multi-agent scoring method is adopted, which uses heterogeneous OCR engines to collaboratively recognize the candidate's answer text, and combines multi-dimensional scoring agents and reinforcement learning models to achieve high-precision automated scoring of candidate's answers, ensuring the fairness and interpretability of the scoring.

Benefits of technology

It has achieved a closed-loop scoring process from paper-based answer sheets to high-precision, automated, and interpretable evaluation, significantly improving the efficiency, fairness, and credibility of personnel examinations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975642A_ABST
    Figure CN120975642A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of personnel examination scoring, in particular to an end-to-end-based multi-agent personnel examination scoring method, which comprises the steps of generating fusion features based on image data of an examinee answer sheet, inputting the fusion features into a heterogeneous OCR engine for processing, and outputting an examinee answer text; establishing a knowledge base, analyzing a question requirement text, and screening out a reference answer text from the knowledge base; constructing a plurality of dimension scoring agents, determining an original state based on the question requirement text and the reference answer text, training each dimension scoring agent through the original state, inputting the examinee answer text into the dimension scoring agents, outputting dimension scoring scores, and forming dimension scoring vectors; carrying out self-consistency analysis by integrating the dimension score vector and the original state, and outputting a final score; the whole-process closed-loop processing from paper answer sheet scanning to high-precision, automatic and interpretable scoring is realized, and the marking efficiency, fairness and credibility are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of personnel examination scoring technology, specifically to an end-to-end multi-agent personnel examination scoring method. Background Technology

[0002] In the field of personnel examinations, such as civil service recruitment and professional qualification examinations, the grading of large-scale paper-based answer sheets has long faced severe challenges. These examinations typically involve a large number of candidates, and the question types include both objective questions and subjective questions that heavily rely on subjective judgment, such as essay writing and case analysis. Traditional manual grading methods, when dealing with massive volumes of essay-type answer sheets, have revealed significant bottlenecks, including high manpower requirements, lengthy processing times, inconsistent scoring standards, and an inability to systematically provide detailed feedback on complex answers based on multi-dimensional scoring criteria. These bottlenecks ultimately reduce the credibility and fairness of the exam answers. Early automated technologies based on Optical Mark Recognition (OMR) could only handle multiple-choice questions and were ineffective for core subjective question types such as essay writing.

[0003] To change this situation, the rapid development of artificial intelligence technology in recent years, especially the breakthroughs in deep learning in the fields of Natural Language Processing (NLP) and Computer Vision (CV), has provided a technological foundation for building intelligent marking systems for personnel examinations. Deep learning-based models can learn complex scoring patterns from massive amounts of labeled data, mimicking expert cognitive processes to perform semantic understanding, logical structure analysis, policy terminology recognition, grammatical checks, and even writing style evaluation on essay answers. Theoretically, this offers hope for automating and objectively scoring subjective questions such as essays, improving efficiency, consistency, and scalability, and generating structured comments as a basis for grade review.

[0004] However, the application of existing technologies in personnel examination scenarios still faces unique challenges: Answer sheet images often suffer from accumulated OCR (Optical Character Recognition) errors due to illegible handwriting and poor scanning quality, severely impacting subsequent text analysis; the precise quantification and understanding of abstract dimensions such as the depth of policy understanding, the effectiveness of countermeasures, and the rigor of logic contained in essay answers remains insufficient; the robustness and generalization ability of scoring models in the face of complex, non-standard answers need improvement; and more importantly, ensuring the logical consistency between multi-dimensional scoring and the final total score, as well as the transparency and interpretability of the model's decision-making process, to meet the extremely high requirements of personnel examinations for fairness and appealability, are all key challenges that urgently need to be overcome. Summary of the Invention

[0005] To address the technical problems of existing artificial intelligence technologies in scoring subjective questions, such as difficulties in text analysis, insufficient understanding, inability to ensure logical consistency between multiple dimensions and the total score, and inadequate fairness and appealability, this invention aims to provide an end-to-end multi-agent personnel examination scoring method. The specific technical solution adopted is as follows:

[0006] The system acquires image data of the candidates' answer sheets and constructs a heterogeneous OCR engine. Based on the image data of the candidates' answer sheets, it generates fusion features, inputs the fusion features into the heterogeneous OCR engine for processing, and outputs the candidates' answer text.

[0007] Collect data from multiple sources, build a knowledge base, parse the question requirement text, and filter out the reference answer text from the knowledge base;

[0008] Construct multi-dimensional scoring agents. Determine the original state based on the question requirement text and the reference answer text. Train each dimension scoring agent through the original state. Input the candidate's answer text into the dimension scoring agent and output the dimension scoring score to form a dimension scoring vector.

[0009] A self-consistency analysis is performed on the comprehensive dimension scoring vector and the original state to output the final score.

[0010] Preferably, image data of the candidate's answer sheet is acquired, and a heterogeneous OCR engine is constructed. Fusion features are generated based on the image data of the candidate's answer sheet, and the fusion features are input into the heterogeneous OCR engine for processing to output the candidate's answer text, including:

[0011] The image data of the candidate's answer sheet is processed in grayscale to obtain the corresponding grayscale image. Then, it is successively processed through a multi-scale Gabor filter bank, CLAHE dynamic contrast enhancement, and combined with image gradient information and adaptive channel feature fusion to generate fused features.

[0012] The heterogeneous OCR engine includes architectures based on CNN-RNN-CTC, Transformer, and LSTM and attention mechanisms. The fused features are input into each heterogeneous OCR engine in parallel to activate each engine, obtain the original output sequence, analyze the character similarity, fuse the weight-aware confidence, determine the final characters, and combine word-level N-Gram semantic optimization for coherence to output the candidate's answer text.

[0013] Preferably, the fused features are input in parallel into each heterogeneous OCR engine to activate each engine, obtain the original output sequence, analyze the character similarity, fuse the weighted confidence, determine the final characters, and combine word-level N-gram semantic optimization for coherence to output the candidate's answer text, including:

[0014] The fused features are input into each heterogeneous OCR engine in parallel. Quality parameters are extracted based on the fused features. A state vector is constructed using the quality parameters to form a state set.

[0015] The expected reward for each heterogeneous OCR engine is calculated using the Contextual Bandits model for the set of states, and then processed in conjunction with the ε-greedy strategy. The accuracy and processing time of the heterogeneous OCR engines are recorded, and the Contextual Bandits model is updated.

[0016] The combination of heterogeneous OCR engines activated based on the updated Contextual Bandits model output is determined, and the weight of each heterogeneous OCR engine is determined.

[0017] By using the original output sequences of heterogeneous OCR engines, a cumulative distance matrix is ​​constructed by combining the weights of each heterogeneous OCR engine, the character similarity is derived, the weight-aware confidence is fused, the recognition result is obtained, and corrections are made to determine the final character.

[0018] By combining word-level N-Gram semantics and the Viterbi algorithm to optimize coherence and correcting errors in heterogeneous OCR engines, the test taker's answer text is output.

[0019] Preferably, the weighted confidence level is fused to obtain the recognition result, which is then corrected to determine the final character, specifically as follows:

[0020] The character prediction probability of any heterogeneous OCR engine is determined based on weight-aware confidence, the recognition result is obtained, a judgment threshold is set, and a confusion matrix is ​​preset. If the recognition result is less than the judgment threshold, the error correction mechanism is triggered, the recognition result is corrected using the confusion matrix, and the final character is determined by the maximum a posteriori probability.

[0021] Preferably, multi-source data is collected, a knowledge base is established, and the question requirement text is parsed. Reference answer texts are then selected from the knowledge base, including:

[0022] Collect multi-source data, including test papers and answers crawled from any of the following channels: personnel examination resource websites, officially released examination outlines, and past civil service recruitment and professional qualification examination question banks, to form a raw dataset. Then, preprocess the dataset, establish a knowledge base, and determine the true quality based on the knowledge base.

[0023] A semi-supervised learning model is established, which includes a discriminator. The discriminator is trained, and the prediction quality is output based on the knowledge base.

[0024] A crawling strategy optimization module is built. Based on the text requirements of the question, the original crawling strategy is determined by combining the actual quality and the predicted quality. The error between the predicted quality and the actual quality is calculated and fed back to the crawling strategy optimization module to adjust the crawling priority, determine the new crawling strategy, and update the knowledge base.

[0025] A deep reinforcement learning model is established, an adaptive crawling strategy is designed, the new crawling strategy is analyzed, the corresponding actions are executed, and reference answer texts are selected from the updated knowledge base.

[0026] Preferably, the preprocessing includes unstructured data parsing, question type classification, and labeling.

[0027] Preferably, a deep reinforcement learning model is established, an adaptive crawling strategy is designed, the new crawling strategy is analyzed, corresponding actions are executed, and reference answer texts are selected from the updated knowledge base, including:

[0028] Based on the new crawling strategy, define the action space and state space, determine the constraints, and design the reward function in combination with the actual quality.

[0029] A deep reinforcement learning-based model is established and trained. The optimal deep reinforcement learning-based model is obtained by combining the reward function with the experience replay and target network stable training process.

[0030] The optimal deep reinforcement learning model is used to select reference answer texts from the updated knowledge base.

[0031] Preferably, a multi-dimensional scoring agent is constructed. The initial state is determined based on the question requirement text and the reference answer text. Each dimension scoring agent is trained using the initial state. The candidate's answer text is input into the dimension scoring agent, which outputs a dimension score, forming a dimension scoring vector, including:

[0032] The dimensional agents include topic relevance, content completeness, and language fluency. Each dimensional agent adopts an Actor-Critic architecture. Each dimensional agent is trained with the original state and optimized using temporal difference and proximal policy optimization algorithms to obtain the trained Actor-Critic architecture.

[0033] The candidate's answer text is input into the Actor-Critic framework, which outputs dimensional score scores and integrates them to form a dimensional score vector.

[0034] Preferably, a self-consistency analysis is performed on the comprehensive dimensional scoring vector and the original state to output the final score, including:

[0035] A comprehensive decision-making agent is constructed, and an attention mechanism is adopted. The attention weight of the score of each dimension in the dimension score vector is determined through learnable parameters. The context vector is determined and concatenated with the original state and input into a multi-layer fully connected network to obtain the discretized total score.

[0036] Set up a score reward function to enhance scoring consistency and adherence to scoring rules;

[0037] The comprehensive decision-making agent is trained collaboratively using a collaborative training strategy of at least two stages. In the first stage, the comprehensive decision-making agent is frozen and the dimensional agents are optimized independently. In the second stage, a consistent reward and gradient backpropagation mechanism are generated. Based on the discretized total score, the goal of coordinating local optima and global logical self-consistency is achieved to obtain the trained comprehensive decision-making agent.

[0038] The trained integrated decision-making agent outputs the final score corresponding to the candidate's answer text.

[0039] Preferably, a consistent reward and gradient backpropagation mechanism is generated, based on the discretized total score coordinating local optima and global logical self-consistency objectives, to obtain the trained comprehensive decision-making agent, including:

[0040] Set a control threshold to determine the error of the dimensional agent. If the error is less than the control threshold, freeze all parameters of the dimensional agent and fine-tune them together with the comprehensive decision agent. Generate a consistent reward based on the discretized total score, forcing the comprehensive decision agent to adjust its strategy.

[0041] By adding gradient backpropagation paths and combining them with consistency rewards, we can coordinate local optima and global logical self-consistency objectives to achieve multi-objective collaborative optimization.

[0042] The present invention has the following beneficial effects:

[0043] This application addresses the bottleneck in recognizing test paper answers by employing adaptive recognition based on heterogeneous OCR collaboration and enhanced routing; it ensures the authority and timeliness of scoring criteria by constructing an examination knowledge base based on quality judgment and dynamic crawling strategies; and it achieves multi-dimensional accurate scoring and logical consistency of the total score through a global consistency scoring method based on multi-dimensional intelligent agent reinforcement learning, thereby determining the final score. This realizes a closed-loop process from scanning paper answer sheets to high-precision, automated, and interpretable scoring, significantly improving the efficiency, fairness, and credibility of personnel examination evaluation. Attached Figure Description

[0044] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 A flowchart illustrating the steps of an end-to-end multi-agent personnel examination scoring method provided in one embodiment of the present invention;

[0046] Figure 2 This is a schematic diagram of the framework of an end-to-end multi-agent personnel examination scoring method provided in one embodiment of the present invention. Detailed Implementation

[0047] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of an end-to-end multi-agent personnel examination scoring method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0049] The following description, in conjunction with the accompanying drawings, details a specific scheme for an end-to-end multi-agent personnel examination scoring method provided by the present invention.

[0050] Please combine Figure 1 and Figure 2 The diagrams show a flowchart and a framework diagram of an end-to-end multi-agent personnel examination scoring method provided in the first embodiment of the present invention. The method includes:

[0051] Step S1: Acquire image data of the candidate's answer sheet, construct a heterogeneous OCR engine, generate fusion features based on the image data of the candidate's answer sheet, input the fusion features into the heterogeneous OCR engine for processing, and output the candidate's answer text;

[0052] Step S2: Collect multi-source data, build a knowledge base, parse the question requirement text, and filter out the reference answer text from the knowledge base;

[0053] Step S3: Construct multi-dimensional scoring agents. Determine the original state based on the question requirement text and the reference answer text. Train each dimension scoring agent through the original state. Input the candidate's answer text into the dimension scoring agent and output the dimension scoring score to form a dimension scoring vector.

[0054] Step S4: Perform a self-consistency analysis on the comprehensive dimension scoring vector and the original state, and output the final score.

[0055] To better illustrate this, intelligent scoring for subjective questions currently faces three major bottlenecks. First, the massive volume of paper-based answer sheets suffers from significant differences in handwriting and fluctuating scanning quality, leading to accumulated OCR recognition errors that severely impact the accurate extraction of textual information. Second, the exam knowledge base supporting the scoring suffers from issues such as limited data sources, uncontrollable quality, and outdated updates, making it difficult to dynamically adapt to authoritative policy statements and scoring standards. Third, existing automated scoring schemes struggle to simultaneously and accurately quantify multi-dimensional scoring standards such as topic relevance, content completeness, and language fluency, and there is a lack of logical consistency between the scores of each dimension and the total score, failing to meet the extremely high requirements of personnel examinations for fairness and traceability in scoring results.

[0056] This application proposes an end-to-end multi-agent personnel examination scoring method. "End-to-end" means that the input data and output results are obtained based on a unified model, requiring no human intervention. This forms a closed-loop process from scanning paper answer sheets to high-precision, automated, and interpretable scoring, reducing human error. "Multi-agent" refers to multiple independent and collaborative agents, each responsible for a corresponding task or function, working together to achieve the overall goal and ensuring the comprehensiveness and fairness of the scoring.

[0057] As an optional implementation method, the entire scoring method is applicable to the automated scoring of subjective questions in civil service recruitment, public institution recruitment, and various professional qualification examinations, ensuring that the scoring of each candidate's answer sheet is fair, efficient, and transparent.

[0058] Understandably, step S1 is an adaptive recognition method based on heterogeneous OCR collaboration and enhanced routing, aiming to solve the core problem in personnel examinations where image quality fluctuates greatly due to illegible handwriting, poor scanning quality, binding misalignment, etc., resulting in insufficient representation ability of single image features or insufficient robustness of single OCR engine recognition. Addressing the issue of varying quality in massive amounts of answer sheet images, it innovatively designs a multi-parameter evaluation system based on ambiguity, contrast, and geometric deviation angle, and adopts a contextual bandits reinforcement learning framework, combined with... The -greedy strategy dynamically schedules the optimal combination of OCR engines and their weights. Finally, to address the temporal misalignment and confidence conflict in the output of multiple engines, a weight-aware progressive DTW (Dynamic Time Warping) sequence alignment and confidence fusion strategy is proposed. Combined with pre-trained character embeddings and word-level N-Gram semantic optimization, the accuracy and coherence of the final output text are ensured. The entire method achieves end-to-end adaptive processing from complex image input to high-quality text recognition.

[0059] Further, step S1 includes:

[0060] Step S11: Grayscale processing of the candidate's answer sheet image data to obtain the corresponding grayscale image, followed by multi-scale Gabor filter bank, CLAHE dynamic contrast enhancement, and fusion of image gradient information and adaptive channel features to generate fused features.

[0061] The process involves scanning the candidate's answer sheet image data and processing it to obtain a grayscale image. A multi-scale, multi-directional Gabor filter bank is used to extract texture information sensitive to cursive features. The CLAHE algorithm is used to enhance the contrast of local writing areas and suppress noise, combined with image gradient information. Then, an adaptive weighting function is designed to group and weight the extracted features, perform fusion and cross-channel normalization, and finally generate a fused feature. This fused feature is then used in parallel as input to a heterogeneous OCR engine, fully utilizing the advantages of different engines in feature extraction and sequence modeling, laying the foundation for subsequent dynamic routing and fusion.

[0062] Specifically, grayscale images are defined as... Multi-scale Gabor filter banks pass through 8 directions and 3 scales Texture feature extraction is performed using a filter bank. The direction is defined by eight fixed, evenly spaced directions at 22.5° intervals, covering all basic stroke directions of Chinese characters, including common horizontal, vertical, and diagonal strokes, ensuring robustness to any writing style. The scale is defined by three fixed scales to match the scanning resolution of the answer sheet. Scale 4 is used to determine pen tip details and capture fine-grained stroke transitions and other fine-grained connecting strokes. Scale 8 is used to determine single-character structure and extract medium-scale structural features such as character outlines. Scale 16 is used for paragraph texture analysis, capturing macro-level text line textures such as paragraph layout. The directional handwriting features have the best response, and a 24-channel Gabor texture map is generated by multiplying 8 directions by 3 scales. ,in, These correspond to the height, width, and number of channels of the texture map, respectively, to significantly improve the sensitivity to the characteristics of handwriting strokes.

[0063] It can be explained that each Gabor filter generates a texture feature channel, reflecting the texture response of the image at a specific direction and scale, such as the stroke trajectory and stroke thickness, which is independent of color. Multiple Gabor channels are determined because of the multi-scale directionality of handwritten handwriting. That is, Chinese handwriting contains rich directions, such as the 45° and 135° of the left and right strokes, the 90° of the vertical stroke, and the scale of single strokes or paragraph structures. Using a filter with a single direction or scale cannot fully capture the features. Moreover, since grayscale images only contain brightness information, the Gabor filter can expand a single brightness channel into 24-dimensional texture features, highlighting the directionality and structure of handwritten notes.

[0064] Next, the CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithm is used to enhance the contrast of the writing region through local histogram equalization. A contrast enhancement threshold is limited to prevent excessive noise amplification. An adaptive weighting function is then set based on image gradient information. The corresponding calculation formula is as follows:

[0065]

[0066]

[0067]

[0068] in, , , These represent the Gabor weight function, the CLAHE weight function, and the gradient weight function, respectively. This represents the slope of the Sigmoid function; Indicates the image sharpness value; Indicates the sharpness threshold; This indicates the maximum image sharpness. Represents the natural constant; Indicates the standard deviation of local contrast; Indicates the maximum local contrast. This represents the image gradient.

[0069] Then, a grouped weighted fusion was performed, grouping the 24 channels of the Gabor feature group by scale. Each group generates a weighted feature map, namely a 3-scale Gabor feature, a CLAHE enhancement map, and a gradient magnitude map. After each feature is normalized using Min-Max, cross-channel weighting is performed to obtain the final fused feature. The corresponding calculation formula is:

[0070]

[0071]

[0072] in, Indicates Gabor characteristics; Indicates the first Gabor weights in each direction, express The first direction Gabor scale characteristics; Indicates 3-scale Gabor features; Indicates the sharpness weight; , , These represent the normalized Gabor feature, the normalized CLAHE enhancement image, and the normalized grayscale image, respectively.

[0073] Understandably, given the vast amount of answer sheets generated in personnel examinations and the resulting variations in image quality due to significant differences in handwriting and scanning equipment, a single OCR engine struggles to dynamically adapt. This system, based on a multi-parameter evaluation system considering ambiguity, contrast, and geometric deviation angles, specifically addresses common bottlenecks in personnel examination answer sheets such as illegible handwriting, scanning shadows, and binding misalignment. Furthermore, it dynamically schedules the optimal engine combination and weights using a contextual bandits reinforcement learning framework. - The greedy strategy balances exploration and utilization to achieve adaptive routing decisions in complex personnel examination scenarios.

[0074] Step S12: The heterogeneous OCR engine includes architectures based on CNN-RNN-CTC, Transformer, and LSTM and attention mechanisms. The fused features are input into each heterogeneous OCR engine in parallel to activate each heterogeneous OCR engine, obtain the original output sequence, analyze the character similarity, fuse the weight-aware confidence, determine the final characters, combine word-level N-Gram semantic optimization for coherence, and output the candidate's answer text.

[0075] To clarify, each heterogeneous OCR engine corresponds to a different feature extraction paradigm. Among them, the CNN-RNN-CTC (Convolutional Neural Network - Recurrent Neural Network - Connectionist Temporal Classification) model extracts local gradient orientation histogram (HOG) features through convolutional layers, models temporal context associations through recurrent neural networks, and solves the character alignment problem through the CTC algorithm. The Transformer architecture-based model establishes global feature associations through a self-attention mechanism, and its position encoding module generates position vectors to model long-distance dependencies. The LSTM (Long Short-Term Memory) and attention mechanism-based model uses convolutional layers to extract local features, combines LSTM to process sequence features, and uses the attention mechanism to assist in character localization.

[0076] Further, step S12 includes:

[0077] Step S121: Input the fused features into each heterogeneous OCR engine in parallel, extract quality parameters based on the fused features, construct a state vector through the quality parameters, and form a state set.

[0078] It can be explained that the quality parameters are derived from the fusion characteristics. The image is directly extracted from the image. Blur level is calculated using the Tenengrad gradient method; contrast is obtained based on the standard deviation formula; and the geometric deviation angle is obtained by the absolute difference between the actual tilt angle of the text line in the detected grayscale image and the standard angle of the template. The corresponding calculation formula is as follows:

[0079]

[0080]

[0081]

[0082] in, , , These represent ambiguity, contrast, and geometric deviation angle, respectively. , These represent the pixel coordinates of the fused features; , These represent the total number of pixel values ​​on the x and y coordinates in the fused feature, i.e., the size of the fused feature; Indicates the mean of the fusion features; This represents the actual tilt angle, which is obtained by detecting the text line direction using the Hough transform. This indicates the standard angle of the template.

[0083] Specifically, a state vector is constructed using mass parameters, denoted as... Integrate each state vector to form a state set, denoted as . The quality parameters are discretized into a finite state space. Each state consists of three dimensions: ambiguity, contrast, and geometric deviation angle. Ambiguity is divided into high, medium, and low; contrast is divided into strong, medium, and weak; and geometric deviation angle is divided into none, slight, and severe.

[0084] Step S122: Calculate the expected reward for each heterogeneous OCR engine using the Contextual Bandits model, process it in conjunction with the ε-greedy policy, record the accuracy and processing time of the heterogeneous OCR engines, and update the Contextual Bandits model.

[0085] This paper explains that a dynamic routing strategy is designed and analyzed based on a heterogeneous OCR engine to activate the engine. The Contextual Bandits model is a reinforcement learning algorithm that selects an action based on the current context information at each time step and updates its strategy based on the reward obtained from performing that action. The ε-greedy strategy is a common explore-exploit strategy where, at each decision point, an action is randomly selected for exploration with probability ε, and the currently considered optimal action is selected for exploitation with probability 1-ε. Preferably, in this embodiment, the strategy is based on probability... The engine combination with the highest expected reward, and through probability. Randomly select other engine combinations to obtain the dynamic routing strategy that needs to be executed, and then execute the strategy.

[0086] Specifically, by using quality parameters, an initial weight allocation function is designed to determine the optimal operating parameters for each heterogeneous OCR engine. The corresponding calculation formula is as follows:

[0087]

[0088] in, Indicates the initial weights; This represents the prior weight allocation coefficient, used to control the smoothness of weight allocation; Indicates quality parameters; , They represent the first The and the first The optimal quality threshold for a heterogeneous OCR engine.

[0089] It can be explained that the initial weights The initial state representation, used as prior knowledge, is then employed for training the reinforcement learning model. A dynamic routing strategy is trained and optimized using the Contextual Bandits framework, and the action space is defined, i.e., the engine set is determined based on heterogeneous OCR engines. In the middle, select the list of engines to activate. and its corresponding normalized weight vector Engine 1 is based on a CNN-RNN-CTC architecture, Engine 2 is based on a Transformer architecture, and Engine 3 is based on an LSTM and attention mechanism architecture. It also allows for hybrid strategies, meaning a strategy can include one independent engine or multiple hybrid engines, for example... A linear contextual bandits model is adopted. Input any state vector, and process it through the model Output the expected reward for each heterogeneous OCR engine, denoted as the reward function, and the corresponding calculation formula is:

[0090]

[0091] in, Indicates the first The expected reward corresponding to each state vector; Indicates the first Nonlinear mapping of state features of a state vector; Indicates the first Transpose of the parameter vectors of a heterogeneous OCR engine.

[0092] After implementing the strategy, the accuracy and processing time of the heterogeneous OCR engine are recorded. Accuracy represents the correctness of the heterogeneous OCR engine's output compared to the reference answer text; processing time represents the time from receiving image data to obtaining the correct result, i.e., the corresponding delay processing time, as a penalty. Then, the model parameters of the Contextual Bandits model are updated. The corresponding calculation formula is as follows:

[0093]

[0094]

[0095]

[0096]

[0097]

[0098] in, Representation based on state vector Execution strategy The rewards received; This represents the accuracy-based reward weighting coefficient; This represents the weighting coefficient for the delay penalty; This indicates the degree of matching between the output of the heterogeneous OCR engine and the reference answer text; This represents the total latency in completing a single image data, which includes the processing time of a single heterogeneous OCR engine and the processing time of parallel heterogeneous OCR engines. Indicates the first The time consumption of each heterogeneous OCR engine; This represents the updated Contextual Bandits model. parameter; Indicates the learning rate; Indicates in the parameter Next, strategy The gradient of the logarithmic probability; Representation based on state vector Execution strategy The probability distribution; This indicates the currently selected activation strategy, i.e., the selected heterogeneous OCR engine. Represents a subset of all possible strategies; Indicates the first A heterogeneous OCR engine in the state vector The expected reward.

[0099] Step S123: Output the activated heterogeneous OCR engine combination based on the updated Contextual Bandits model, and determine the weight of each heterogeneous OCR engine.

[0100] Specifically, based on the trained linear Contextual Bandits model Then, select to activate the heterogeneous OCR engine and normalize the corresponding heterogeneous OCR engine weights. The corresponding calculation formula is as follows:

[0101]

[0102]

[0103] in, Indicates the currently selected activation strategy; The output of the linear Contextual Bandits model represents the first... A heterogeneous OCR engine in the state vector The activation probability is calculated using the softmax function. Indicates the activation threshold; used to control the sensitivity of activating heterogeneous OCR engines; Represents based on the current state vector The maximum activation probability of all heterogeneous OCR engines is used as a dynamic benchmark. Indicates the first The weights of each heterogeneous OCR engine are used to represent the proportion of contribution among multiple heterogeneous OCR engines; This represents the total activation probability of the currently selected heterogeneous OCR engines.

[0104] Understandably, processing with multiple heterogeneous OCR engines can lead to issues such as output temporal misalignment, confidence conflicts, and semantic breaks. Therefore, a weight-aware progressive DTW sequence alignment and confidence fusion strategy is adopted. Similarity is calculated through pre-trained character embeddings to eliminate temporal biases, and word-level N-Gram semantic optimization is combined to correct conflicting characters, ensuring the coherence and accuracy of the output text.

[0105] Step S124: Using the original output sequence of the heterogeneous OCR engine, construct a cumulative distance matrix by combining the weights of each heterogeneous OCR engine, derive the character similarity, fuse the weight-aware confidence, obtain the recognition result, make corrections, and determine the final character.

[0106] The explanation is that the original output sequence of the heterogeneous OCR engines is processed and fused based on the parallel processing of features from three types of heterogeneous OCR engines. The original OCR recognition result sequence is output, i.e., the original output sequence, denoted as... The weights of the heterogeneous OCR engines are mapped to the three types of correspondences in this embodiment, and these correspondences are denoted as follows: .

[0107] Specifically, the original output sequence is preprocessed and standardized to eliminate format differences and enhance alignment; among which, For the output of the CNN-RNN-CTC engine, the blank label is removed and duplicate characters are compressed using the CTC decoding rules; For Transformer engine output, words are separated by spaces; For the output of LSTM and attention mechanism, take the character with the largest attention weight to generate a deterministic sequence; then complete the preprocessing of each heterogeneous OCR engine, and then perform progressive DTW alignment to eliminate temporal misalignment, that is, optimize the matching between character sequences by gradually refining the alignment process; construct the cumulative distance matrix.

[0108] To better illustrate this, we will construct a cumulative distance matrix based on any two engines, assuming... For heterogeneous OCR engines The character sequence, For heterogeneous OCR engines The character sequence, where, , Similarly, it is denoted as Furthermore, pre-trained character embeddings are introduced, which are character-level vector representations pre-trained based on character sequences. These vector representations are used to capture the semantic features of characters in different contexts, making similar characters close to each other in the vector space. This allows for the derivation of character similarity to understand the relationships between characters and improve the performance of text processing tasks. The corresponding calculation formula is as follows:

[0109]

[0110]

[0111] in, Represents a character sequence and The Middle The character and the first The cumulative distance matrix of each character; Represents the weighting factor, i.e. , , These represent heterogeneous OCR engines. and The weights; Represents a character sequence and The Middle The character and the first Character similarity of 1 character; Indicates cosine similarity; , Representing characters respectively and The pre-trained character embedding vectors.

[0112] Next, based on the original output sequences from the heterogeneous OCR engines, three sets of alignment paths are generated by pairwise alignment. A weighted voting mechanism is used to determine the reference sequence, that is, the original output sequence with higher alignment with the image data of the candidate's answer sheet. Specifically, if two original output sequences have the same character at any position, they are retained; otherwise, the original output sequence of the heterogeneous OCR engine with higher weight is selected. If the original output sequences of the three alignment paths are all inconsistent, that is, there are no identical characters or symbols, the character with the highest weighted average confidence score is selected to complete the alignment. Here, confidence score is used to measure the reliability and accuracy of each character in the alignment process, that is, in this embodiment, it is represented by the number of times the character appears in the alignment path. Then, the weight-aware confidence score fusion stage is entered.

[0113] Further, in step S124, the weighted perceptual confidence is fused to obtain the recognition result, which is then corrected to determine the final character, specifically as follows:

[0114] The character prediction probability of any heterogeneous OCR engine is determined based on weight-aware confidence, the recognition result is obtained, a judgment threshold is set, and a confusion matrix is ​​preset. If the recognition result is less than the judgment threshold, the error correction mechanism is triggered, the recognition result is corrected using the confusion matrix, and the final character is determined by the maximum a posteriori probability.

[0115] It can be explained that the character prediction probability of any heterogeneous OCR engine determined based on the weight-aware confidence score is denoted as the fused weight-aware confidence score. This means that the average character prediction probability is obtained based on the confidence score for further analysis. The corresponding calculation formula is:

[0116]

[0117] in, Indicates that it is identified as the first The average character prediction probability of each character; This indicates that the recognition result of the heterogeneous OCR engine is the first... One character; Indicates the first The weight of each heterogeneous OCR engine; No. The output of the heterogeneous OCR engine is the first Character prediction probability for each character; It represents the set of all possible characters, that is, the set of all characters that could appear in the exam scenario.

[0118] Optionally, in this embodiment, the judgment threshold is: .

[0119] when When this occurs, an error correction mechanism is triggered, which calculates the corrected character prediction probability using the confusion matrix. The confusion matrix is ​​used to evaluate classification performance, and its calculation formula is as follows:

[0120]

[0121] in, This indicates the corrected character prediction probability; Character Identified as frequency, Represents the confusion matrix; It represents the set of all possible characters.

[0122] Then, the final character is determined using the maximum posterior probability, and the corresponding calculation formula is:

[0123]

[0124] Among them, represents the final character; that is, if the recognition result is greater than or equal to the judgment threshold, it means that the prediction probability of the current character meets the standard, and the final character is directly determined by the corresponding predicted character; if the error correction mechanism is triggered, the corrected character prediction probability is used to determine the final character through the maximum a posteriori probability.

[0125] Step S125: Optimize the coherence by combining word-level N-Gram semantics and the Viterbi algorithm, correct the errors of heterogeneous OCR engines, and output the candidate's answer text.

[0126] Specifically, after the processing of step S124 is completed, each recognition result in the character sequence is integrated to obtain a fused character sequence, denoted as , and processing is performed based on the fused character sequence, that is, the output coherence is enhanced by optimizing the word-level N-Gram semantics. Among them, word-level N-Gram semantics refers to a method in natural language processing for capturing and understanding the meaning and context of language by analyzing the combinations of multiple consecutive words in the text; then a character transition probability matrix is defined, which represents the probability that the character is followed by to reflect the transition relationship between different characters. The maximum probability path is found through the Viterbi algorithm, that is, the maximum probability that the current character is followed by a relevant character. Among them, the Viterbi algorithm is a dynamic programming algorithm that can efficiently find the state sequence that is most likely to generate the given sequence; correct the semantic breaks caused by OCR errors, and finally output the candidate's answer text, denoted as .

[0127] For better illustration, in practical applications, for the connected strokes, there are two cases of standard folded strokes such as the character "弓" and scribbled connected strokes. Specifically, in this embodiment, the contrast of the stroke edges is enhanced by the CLAHE algorithm and the image gradient information. Among them, the edges of the standard folded strokes are sharp, and the edges of the scribbled folded strokes are blurred to assist in suppressing the misjudgment of the inherent connected strokes; then when the blur degree in the quality parameters is low but the contrast is high, it means that it conforms to the standard folded stroke writing method. Furthermore, for the enhanced features, an OCR engine with a Transformer architecture is selected, which has a stronger understanding of the glyph structure and is not likely to misjudge the standard folded strokes as blurred; if both of the above two steps fail to detect, the character probability is corrected in combination with the error correction mechanism, that is, the semantic anomaly of misidentifying the character "弓" as an uncommon character combination such as "口" + "刁" can be captured during the cosine similarity calculation, and it is forced to be corrected by referring to common collocations such as "弓箭" and "弓形" through the word-level N-Gram semantics to adjust the characters of the standard folded strokes and improve the recognition accuracy.

[0128] Understandably, existing knowledge bases that rely on intelligent scoring generally suffer from problems such as insufficient coverage due to single data sources, uncontrollable quality affecting the authority of scoring, and lagging updates that make it difficult to adapt to policy changes. Therefore, it is proposed to establish a knowledge base to solve the corresponding technical deficiencies in personnel examinations; and to build a high-quality, timely knowledge base to provide authoritative and reliable knowledge support for intelligent scoring.

[0129] Further, step S2 includes:

[0130] Step S21: Collect multi-source data, which includes test papers and answers crawled from any of the following channels: personnel examination resource websites, officially released examination outlines, and past civil service recruitment and professional qualification examination question banks, to form a raw dataset. This dataset is then preprocessed, a knowledge base is established, and the actual quality is determined based on the knowledge base.

[0131] To clarify, multi-source data refers to datasets acquired from diverse sources, resulting in data formats such as scanned PDF versions of past exam questions, Word documents containing case analysis answers, and webpage texts of policy interpretations, with significant differences. Therefore, preprocessing is performed on the resulting raw dataset, denoted as [original dataset name missing]. .

[0132] Furthermore, in step S21, preprocessing includes unstructured data parsing, question type classification, and labeling.

[0133] Specifically, unstructured data parsing involves: extracting text from PDF scanned exam papers using PP-OCRv4 (Paddle Optical Character Recognition v4); matching questions and answers using regular expressions for webpage text; and directly parsing paragraph structures for Word documents. This process converts different types of data into processable text formats. Question type classification and tagging involve: accurately identifying core subjective question types in personnel examinations, such as essay writing and case analysis, through rule matching and a BERT (Bidirectional Encoder Representations from Transformers) classifier. Each question is then tagged with unique metadata such as policy areas, difficulty levels, and core scoring dimensions relevant to the personnel examination. This enhances understanding of the question's background and requirements and improves its manageability.

[0134] It can be noted that the preprocessed original dataset forms the established knowledge base, and the true quality is determined through this knowledge base. The corresponding calculation formula is as follows:

[0135]

[0136]

[0137]

[0138]

[0139] in, Indicates true quality; Indicates coverage rate; Indicates redundancy; Indicates the matching rate; express , , The corresponding weight coefficients are obtained by using a grid search method based on historical data to find the optimal combination for optimization, that is, by traversing all possible weight combinations to determine the optimal weight coefficients.

[0140] Provide an explanation, coverage This indicates the proportion of exam knowledge points covered in the knowledge base. This represents the set of knowledge points covered in the knowledge base. Represents the set of all knowledge points in multi-source data; redundancy. This indicates the percentage of repeated questions. This represents a set of repeated questions in the knowledge base; Represents the total number of questions in the knowledge base; matching rate. This indicates the degree to which the answer matches the requirements of the question. This represents the set of questions in the knowledge base whose answers meet the requirements of the question.

[0141] Step S22: Establish a semi-supervised learning model, including a discriminator, and train the discriminator to output prediction quality based on the knowledge base.

[0142] The semi-supervised learning model is based on a semi-supervised Transformer model and is used to distinguish between high-quality and low-quality personnel examination data, and to evaluate data quality in real time. Specifically, it is based on a discriminator to design a dynamic data quality evaluation mechanism, trains the discriminator with a small amount of labeled data and a large amount of unlabeled data, introduces conditional entropy loss to enhance the consistency of evaluation of unlabeled data in dimensions such as policy fit and countermeasure feasibility, and improves the robustness of the semi-supervised Transformer model to noise in personnel examination data through adversarial training, so as to output prediction quality.

[0143] Specifically, a semi-supervised learning model is established, which includes a discriminator. The discriminator is trained using a small amount of labeled data with either high-quality or low-quality labels and a large amount of unlabeled data, based on the original dataset corresponding to the knowledge base. The labeled dataset is denoted as... , Represents the dataset Quality label, 1 represents poor quality, and 1 represents high quality; unlabeled datasets are denoted as poor quality. Furthermore, it contains candidate data from the original dataset that are yet to be evaluated, meaning that data of the appropriate quality has not yet been output.

[0144] Define the discriminator loss function, and the corresponding calculation formula is as follows:

[0145]

[0146]

[0147] in, Represents the loss function; Indicates the first in the labeled dataset Quality labels for each data sample; This represents the regularization strength coefficient of the loss in a semi-supervised Transformer model. Indicates the discriminator; Represents the conditional entropy function; This indicates that the discriminator operates on an unlabeled dataset. Predicted as category The probability of; , These represent the total number of data samples in the labeled dataset and the unlabeled dataset, respectively.

[0148] By minimizing the loss function, the semi-supervised Transformer model learns to distinguish between high-quality and low-quality data on labeled datasets, thereby enhancing prediction consistency on unlabeled datasets; the prediction quality output by the trained discriminator is denoted as... .

[0149] Step S23: Construct a crawling strategy optimization module. Based on the text requirements of the question, determine the original crawling strategy by combining the actual quality and the predicted quality. Calculate the error between the predicted quality and the actual quality, feed it back to the crawling strategy optimization module, adjust the crawling priority, determine a new crawling strategy, and update the knowledge base.

[0150] Specifically, a crawling strategy optimization module is constructed, which includes a strategy for crawling multi-source data based on actual quality or predicted quality as the standard; that is, in this embodiment, the actual quality is obtained based on the text of the question requirements currently being analyzed. Set a crawling threshold, i.e., a threshold for stopping crawling based on knowledge base quality, denoted as . ;like If crawling stops, it means the currently established knowledge base meets the scoring requirements; otherwise, it means... The crawling strategy optimization module is optimized based on the predicted quality output of the discriminator.

[0151] First, the probability distribution of crawling actions is defined based on the crawling strategy optimization module, denoted as... ,in This indicates adding crawling actions such as capturing a certain type of question or switching data sources; then, based on the predicted quality and the actual quality, the error is calculated, and the corresponding calculation formula is:

[0152]

[0153] in, This represents the error; it is then used to adjust the gradient of the crawling strategy and fed back to the crawling strategy optimization module to adjust the crawling priority. The corresponding calculation formula is:

[0154]

[0155] in, This means that it is proportional to, that is, the probability distribution of crawling actions is proportional to... They are directly proportional; This represents the gradient of the crawling action.

[0156] A threshold is preset and denoted as . ,like A positive result indicates a large discriminator error. In this case, the crawling strategy needs to be optimized, a new crawling strategy determined, and corresponding crawling actions executed to obtain new data. This new data is then added to the knowledge base, and the true quality is recalculated to ensure the rebuilt knowledge base meets the scoring requirements. Conversely, a negative result indicates a large discriminator error. We will maintain the current crawling strategy.

[0157] Preferably, the discriminator incorporates an adversarial training mechanism to generate adversarial examples, thereby enhancing the robustness of the discriminator. These adversarial examples are constructed using the gradient ascent method, and the corresponding calculation formula is as follows:

[0158]

[0159] in, Indicates adversarial examples; Represents a knowledge base; Indicates the disturbance intensity parameter; This represents the gradient of the knowledge base; This represents the loss function.

[0160] By simultaneously training a semi-supervised Transformer model to distinguish between real data and adversarial examples, the discriminator's tolerance to noise and anomalous data is improved. In other words, through a self-supervised learning mechanism, the model can discover potential structures and patterns in unlabeled datasets to better identify adversarial examples and more accurately identify subtle perturbations, ensuring the stability and reliability of the model in practical applications.

[0161] Step S24: Establish a deep reinforcement learning-based model, design an adaptive crawling strategy, analyze the new crawling strategy, execute the corresponding actions, and filter out the reference answer text from the updated knowledge base.

[0162] Understandably, if the new crawling strategy still does not meet the scoring requirements, continuing to use the traditional crawling strategy will result in high resource consumption and a lack of quality feedback. Therefore, an adaptive crawling strategy is proposed to analyze the new crawling strategy, dynamically adjust the crawling strategy, and efficiently and accurately obtain reference answer text from the knowledge base.

[0163] Further, step S24 includes:

[0164] Step S241: Define the action space and state space according to the new crawling strategy, determine the constraints, and design the reward function based on the actual quality.

[0165] The explanation is as follows: a reward function is designed with data quality improvement and budget consumption as constraints. Data quality improvement refers to the improvement in data quality obtained by the new crawling strategy; budget consumption represents the total amount of computing resources and time used by the model when performing machine learning or data processing tasks.

[0166] Specifically, define the action space, that is ={Stop, add specific question type, switch data source}; State space Includes the data quality corresponding to the current crawling strategy. Remaining budget timestamp The reward function is designed based on the true quality of the data, and the corresponding calculation formula is as follows:

[0167]

[0168] in, Represents the reward function; , Representing timestamps and The corresponding data quality; This represents the cost penalty coefficient; Indicates action The corresponding cost; when the new strategy improves quality and saves budget, a positive reward is given; if the quality decreases and the budget is wasted, a negative reward is given; otherwise, there is no reward; that is, the feasibility of the current new crawling strategy is determined by the reward function.

[0169] Step S242: Establish and train a deep reinforcement learning-based model, and combine the reward function with the experience replay and target network stable training process to obtain the optimal deep reinforcement learning-based model.

[0170] To clarify, the deep reinforcement learning model refers to the Deep Q-Network (DQN). By using a convolutional neural network as a function approximator, DQN can handle high-dimensional input data. The crawling strategy is modeled based on the state space, denoted as... , Representing the state space Any state in the representation; Let the state-space value function be represented, where, Indicates model parameters.

[0171] Specifically, for the training process based on deep reinforcement learning models, experience replay and target network stabilization are adopted, and experience replay stores historical trajectories. This involves defining a replay buffer to store various states, actions, and reward information from the data collected during training. This facilitates random sampling and utilization in subsequent training, breaking down correlations between data and improving training efficiency and stability. (Target network) Periodically update model parameters This reduces training fluctuations and avoids frequent changes in the learning objective caused by frequent parameter updates. Specifically, the model parameter updates follow the Bellman equation, which describes how, in a Markov decision process, the value of the current state can be recursively calculated using the value of the future state. This ensures that the model can progressively optimize its crawling strategy during training to maximize the cumulative reward. The corresponding calculation formula is as follows:

[0172]

[0173] in, Indicates the discount factor; This represents the experience replay buffer.

[0174] Understandably, based on the current state... Choose the optimal action, that is The data obtained based on the current crawling strategy is input into the aforementioned trained discriminator to obtain the corresponding prediction quality, and is then analyzed in conjunction with the actual quality to execute the corresponding action in the action space; specifically, if and If so, the "stop" action will be executed, where, This represents the current budget for the experience buffer replay area; Indicates the budget threshold; if And the coverage rate of a certain type of question Then, the action of "adding a specific question type" will be executed, where, This represents the coverage threshold; if a data source has redundancy... Then, the "Switch Data Source" action will be executed, where, This indicates the threshold for redundant segments.

[0175] And introduce budget constraints on the cost of the action, namely ,in, Indicates action The cost per crawl, i.e., the crawling fee; Let represent the total budget. To ensure the crawling strategy maximizes data quality within the budget, the budget constraint is embedded into the objective function using the Lagrange multiplier method. This is then used to train the deep reinforcement learning model. The corresponding calculation formula is:

[0176]

[0177] in, This represents the objective function during the training process; This represents the original loss function of the DQN model, used to update model parameters during training; It represents the Lagrange multiplier.

[0178] Step S243: Select reference answer texts from the updated knowledge base using the optimal deep reinforcement learning model; that is, determine the corresponding crawling strategy using the optimal deep reinforcement learning model and execute the relevant actions. Throughout the process, adjust the model parameters in real time to adaptively select reference answer texts from the established knowledge base and accurately match the requirements of the question.

[0179] Understandably, given the current poor consistency and low efficiency of multi-dimensional scoring for subjective questions, a multi-dimensional agent reinforcement learning approach is adopted to achieve globally consistent scoring of questions. This unifies the scoring accuracy with the logical consistency of the scoring results across all dimensions, significantly improving the reliability and efficiency of automated scoring in personnel examinations.

[0180] Furthermore, step S3 includes:

[0181] Step S31: The dimensional agents include topic relevance, content completeness, and language fluency. Each dimensional agent adopts the Actor-Critic architecture. Each dimensional agent is trained with the original state and optimized using temporal difference and proximal policy optimization algorithms to obtain the trained Actor-Critic architecture.

[0182] The explanation is that a dimensional agent is constructed to address the issues of dimensional independence and insufficient semantic understanding in the scoring of subjective questions in personnel examinations. Specifically, semantic features of questions and answers are extracted using BERT-wwm (BERT with Whole Word Masking). An Actor-Critic architecture is used to train agents that independently output scores based on topic relevance, content completeness, and language fluency. The PPO (Proximal Policy Optimization) algorithm is then used to ensure the stability of this scoring strategy.

[0183] Specifically, the original state is determined based on the question requirements text and the reference answer text, that is, the semantic features of the question requirements are determined separately. semantic features of the reference answer The original state is obtained by splicing the parts together, and is denoted as . At the same time, it perceives the matching degree between the question requirements and the reference answer, avoiding evaluation bias caused by a single feature; then, it captures the deep semantic information of the original features through a large-scale pre-trained language model (BERT-wwm), that is, it understands complex language structures and contextual relationships, and extracts more accurate semantic features.

[0184] In this context, the dimensional agent adopts an Actor-Critic architecture, where Actor represents the policy network used to process the initial state. It outputs continuous scoring actions through multiple fully connected layers and a sigmoid activation function. , to reflect the prediction score of the agent in the corresponding dimension, and the policy function is denoted as . The parameters are ; Critic represents a value network used to share the initial state and output a single-valued state, denoted as . , indicating the current original state Below, follow the scoring strategy The expected cumulative discount rewards that can be obtained.

[0185] It can be explained that the training objective for each dimension of the agent is to maximize the cumulative reward. , This represents the cumulative reward discount factor for the dimensional agent. Indicates the first The reward for the agent is calculated using a multi-dimensional algorithm; the reward function is designed based on the features of the reference answer, and the corresponding calculation formula is:

[0186]

[0187] in, Indicates the first The reward of an agent in its original state; This indicates the corresponding number based on the reference answer. The pseudo-labels for the intelligent agent are generated through automated rules. Specifically, topic relevance is calculated by mapping the semantic similarity between the question and the reference answer to a score range proportionally; content completeness is calculated by extracting the key knowledge points of the question, checking the coverage ratio of the reference answer, and calculating the coverage score; and language fluency is scored by weighting sentence fluency, the proportion of repeated words, and the density of connecting words. This represents the Critic value-guided weighting coefficient. This means guiding the corresponding dimension's intelligent agent to weigh the local scoring error against the global policy benefit.

[0188] Specifically, the Actor-Critic architecture is optimized using temporal difference and proximal policy optimization algorithms, respectively. Optionally, the data used to train the dimensional agents is used to form answer pairs containing a fixed number of batches using randomly sampled test papers obtained from the knowledge base and reference text.

[0189] Critic optimization involves learning and updating through temporal difference (TD) to calculate the TD objective, with the corresponding formula as follows:

[0190]

[0191] in, Indicates TD target; Indicates the discount factor; This represents the value of the previous state in the original state.

[0192] Then minimize the TD error, i.e. This allows for an accurate value assessment of the Critic and provides the Actor with directions for strategic improvement.

[0193] The Actor is optimized using a proximal policy optimization algorithm, which ensures stability by limiting the update step size of the scoring policy. This is achieved by calculating the advantage function estimate and maximizing the shearing objective function. The corresponding calculation formula is as follows:

[0194]

[0195]

[0196]

[0197] in, This represents the estimation of the advantage function; Represent the objective function for shearing; Indicates the importance sampling ratio; Indicates time step Expectations; This indicates an operation to limit the range of ratio changes, used to prevent excessively large single-step time-step updates from compromising the stability of the scoring strategy; This represents the shearing hyperparameters; then, corresponding optimization training is performed for agents of all dimensions to obtain the trained Actor-Critic architecture.

[0198] Step S32: Input the candidate's answer text into the Actor-Critic architecture, output the dimension score, and integrate them to form a dimension score vector; that is, input the candidate's answer text into the trained Actor-Critic architecture, and output the scoring policy network for each dimension agent. The corresponding scoring results are obtained and integrated to form a dimensional scoring vector, denoted as . ,in, This represents the scoring results of the agent regarding topic relevance; This represents the scoring result of the agent indicating content integrity; The score given to the agent indicates its fluency in language.

[0199] Understandably, even with scores obtained from multiple dimensions, there are still issues such as logical inconsistencies between subjective question dimension scores and the total score, making it difficult to meet the strictness of the scoring procedures. Therefore, a comprehensive decision-making intelligent agent is constructed, namely a comprehensive decision-making system based on hierarchical reward mechanisms and attention-driven approaches. This system dynamically fuses dimension score vectors using attention weights activated by LeakyReLU (Leaky Rectified Linear Unit), and explicitly constrains the logical relationship between the total score and the sub-item scores through a consistent reward function. This effectively punishes local abnormal deviations and ensures that the final scoring results conform to the multi-dimensional scoring standards of personnel examinations. Specifically, it addresses the different emphases of different question types in personnel examinations, such as essay writing emphasizing logic and case analysis emphasizing countermeasures, by dynamically allocating the score weights of each dimension. Furthermore, it establishes logical consistency between the discretized total score and the dimension score scores, avoiding the contradictory phenomenon of "high total score but abnormally low scores in some dimensions."

[0200] Further, step S4 includes:

[0201] Step S41: Construct a comprehensive decision-making agent, adopt an attention mechanism, determine the attention weight of each dimension score in the dimension score vector through learnable parameters, determine the context vector, and concatenate it with the original state to input it into a multi-layer fully connected network to obtain the discretized total score.

[0202] To clarify, based on the aforementioned multi-dimensional intelligent agent after training and stabilization, a comprehensive decision-making intelligent agent is constructed, namely a comprehensive decision-making system based on hierarchical reward mechanisms and attention-driven approaches, which includes multi-layer fully connected networks; specifically, the comprehensive decision-making intelligent agent adopts the Dueling DQN architecture, decomposing the Q-value into state values. With advantage function This explicitly distinguishes between the inherent value of a state and the relative advantage of an action, improving the learning efficiency of the scoring strategy; while the target network and priority experience replay alleviate the overestimation problem of Q-learning, and the experience replay pool is based on TD error. By weighting samples based on historical experience, the comprehensive decision-making agent is made to focus on samples with large prediction errors, such as complex policy analysis questions or cases with contradictory scoring tendencies, in order to accelerate the learning and optimization of scoring samples for difficult points in personnel examinations.

[0203] Specifically, based on the comprehensive decision-making intelligent agent, the dimensional scoring vectors are input respectively. Used to reflect local quality and original state Used to provide global context; a scoring policy function for a comprehensive decision-making agent. An attention mechanism is employed, in which, This represents the discretized total score; using learnable parameters, the attention weights for each dimension's score are calculated, and the context vector is determined. The corresponding calculation formula is as follows:

[0204]

[0205]

[0206]

[0207] in, Indicates intermediate representation; Indicates the first The weights of the scores for each dimension; , All represent learnable parameters; This represents the context vector.

[0208] The context vector is concatenated with the original state and used as the input to a multi-layer fully connected network, which outputs a discretized total score. .

[0209] Preferably, to prevent the scoring strategy function of the comprehensive decision-making agent from converging prematurely, the probability distribution of the Top-5 actions is retained for entropy regularization of the scoring strategy function, thereby encouraging the exploration of diverse scoring schemes.

[0210] Step S42: Set the score reward function to enhance scoring consistency and adherence to scoring rules. The corresponding calculation formula is as follows:

[0211]

[0212] in, Represents the score reward function; , All represent the reward weight coefficients of the comprehensive decision-making intelligent agent; This represents a pseudo-total score based on the reference answer text, expressed through dimensional pseudo-labels. Weighted generation; This represents the maximum difference threshold between dimension score ranges, used to ensure normalization of error calculations. Indicates an indicator function.

[0213] It can be explained that, Used to penalize dimension scores that deviate from pseudo-labels beyond a threshold. This is to reduce the risk of logical inconsistency in the generated discretized total score due to abnormal dimensional scoring scores in personnel examinations.

[0214] Understandably, in personnel examinations such as civil service exams (essay writing) and professional qualification exams (case analysis), the multi-dimensional scoring results must strictly adhere to the scoring procedures, the rigid requirement of a high degree of self-consistency between the total score and the sub-items, and the potential for inter-dimensional conflicts that may arise from independent optimization by multiple agents. This is to coordinate the goal of local optima with global logical consistency, ensuring that the subsequent output scores conform to the scoring rules and logic.

[0215] Step S43: Perform collaborative training on the comprehensive decision-making agent, using a collaborative training strategy of at least two stages. In the first stage, freeze the comprehensive decision-making agent and independently optimize the dimensional agents. In the second stage, generate a consistent reward and gradient backpropagation mechanism, based on the discretized total score coordinating local optimality and global logical self-consistency objective, to obtain the trained comprehensive decision-making agent.

[0216] The two-stage collaborative training strategy is explained below. In the first stage, the overall decision-making agent is frozen, while the dimensional agents are optimized independently. Specifically, the focus is on independently optimizing the dimensional scores output by each dimensional agent. In the first N rounds of training, the overall decision-making agent is frozen, and only the dimensional agents are optimized, focusing on their independent scoring capabilities for each dimension: topic relevance, content completeness, and language fluency. This is achieved by maximizing the cumulative reward. The first stage aims to converge the dimensional scores of each dimension to a local optimum, ensuring that each dimension agent fully learns the scoring pattern based on the personnel examination reference answers and scoring standards, thus avoiding optimization difficulties caused by excessive initial bias in subsequent collaborative training. The second stage further optimizes the output of the first stage by integrating the decision-making agent with the strategies and value networks of the dimensional agents, combined with consistency rewards based on the inverse model, to ensure the internal logical consistency between the final score and the scores of each dimension in accordance with the personnel examination scoring rules.

[0217] Further, step S43 includes:

[0218] Step S431: Set a control threshold and determine the error of the dimensional agent. If the error is less than the control threshold, freeze all parameters of the dimensional agent and fine-tune them together with the comprehensive decision agent. Generate a consistent reward based on the discretized total score to force the comprehensive decision agent to adjust its strategy.

[0219] The explanation is as follows: based on the corresponding candidate's answer text, the dimensional score obtained from any original state and the predicted score output by the dimensional agent in the ideal original state are respectively compared. The difference between the two is used to determine the error of the dimensional agent, and a control threshold is set according to the actual situation.

[0220] When the error is less than the control threshold, all parameters of the dimensional agent are frozen to ensure that it remains stable in the current state. The M rounds are then fine-tuned in conjunction with the comprehensive decision agent to improve its performance, accuracy, and response speed through repeated iterations.

[0221] Consistency rewards are generated based on discretized total scores, i.e., globally consistent rewards driven by personnel examination scenarios, to satisfy scoring consistency constraints. The corresponding calculation formula is as follows:

[0222]

[0223]

[0224] in, Indicates a consistency reward; Represents the dimension score vector; This represents the inverse model, which is achieved by discretizing the total score. Compared with the original state Predicted ideal dimension score Inverse model is a mathematical model used for reverse reasoning or reverse engineering, which can infer input conditions based on output results.

[0225] It can be explained that by setting a globally consistent reward, logical consistency between the dimensional score vector and the discretized total score is forced. If the discretized total score is high, but any dimension's score is abnormally low, then... Triggering penalties forces the integrated decision-making agent to adjust its scoring strategy.

[0226] Step S432: Add a gradient backpropagation path, combine it with consistency rewards, coordinate local optimum and global logical self-consistency objectives, and achieve multi-objective collaborative optimization.

[0227] To achieve consistent rewards, i.e., global consistency, a gradient backpropagation path is added to the Critic network in the Actor-Critic architecture for dimensional agents. This gradient backpropagation path refers to the path used to change model parameter updates by adjusting the gradient direction during the optimization process of the integrated decision-making agent. The corresponding calculation formula is:

[0228]

[0229] in, Indicates the coordination coefficient; Indicates global error; Indicates local error; Indicates global consistency reward regarding The gradient; that is, by balancing local errors and global consistency objectives through coordination coefficients, it is ensured that while pursuing the accuracy of dimensional scores, the dimensional agent is also subject to the adjustment of the comprehensive decision agent, ultimately achieving multi-objective collaborative optimization, and thus obtaining the comprehensive decision agent that has completed training.

[0230] Step S44: Output the final score corresponding to the candidate's answer text through the trained integrated decision-making agent.

[0231] Specifically, a scoring dimension vector is generated in parallel based on each dimension of the intelligent agent, i.e. Then, the trained integrated decision-making agent outputs the final score, i.e., the optimized discretized score, based on the attention fusion strategy. .

[0232] Understandably, this application addresses the bottleneck of recognizing candidates' answers by using adaptive recognition based on heterogeneous OCR collaboration and enhanced routing; it ensures the authority and timeliness of the scoring criteria by constructing an examination knowledge base based on quality judgment and dynamic crawling strategies; and it achieves multi-dimensional accurate scoring and logical consistency of the total score through a global consistency scoring method based on multi-dimensional intelligent agent reinforcement learning, thereby determining the final score. This realizes a closed-loop process from scanning paper answer sheets to high-precision, automated, and interpretable scoring, significantly improving the efficiency, fairness, and credibility of personnel examination evaluation.

[0233] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0234] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A multi-agent personnel examination scoring method based on end-to-end, characterized in that, The method includes: The system acquires image data of the candidates' answer sheets and constructs a heterogeneous OCR engine. Based on the image data of the candidates' answer sheets, it generates fusion features, inputs the fusion features into the heterogeneous OCR engine for processing, and outputs the candidates' answer text. Collect data from multiple sources, build a knowledge base, parse the question requirement text, and filter out the reference answer text from the knowledge base; Construct multi-dimensional scoring agents. Determine the original state based on the question requirement text and the reference answer text. Train each dimension scoring agent through the original state. Input the candidate's answer text into the dimension scoring agent and output the dimension scoring score to form a dimension scoring vector. A self-consistency analysis is performed on the comprehensive dimension scoring vector and the original state to output the final score.

2. The end-to-end multi-agent personnel examination scoring method according to claim 1, characterized in that, Image data of the candidate's answer sheet is acquired, and a heterogeneous OCR engine is constructed. Fusion features are generated based on the image data of the candidate's answer sheet, and these fusion features are input into the heterogeneous OCR engine for processing. The output text of the candidate's answer sheet is then generated, including: The image data of the candidate's answer sheet is processed in grayscale to obtain the corresponding grayscale image. Then, it is successively processed through a multi-scale Gabor filter bank, CLAHE dynamic contrast enhancement, and combined with image gradient information and adaptive channel feature fusion to generate fused features. The heterogeneous OCR engine includes architectures based on CNN-RNN-CTC, Transformer, and LSTM and attention mechanisms. The fused features are input into each heterogeneous OCR engine in parallel to activate each engine, obtain the original output sequence, analyze the character similarity, fuse the weight-aware confidence, determine the final characters, and combine word-level N-Gram semantic optimization for coherence to output the candidate's answer text.

3. The end-to-end multi-agent personnel examination scoring method according to claim 2, characterized in that, The fused features are input in parallel into each heterogeneous OCR engine, activating each engine to obtain the original output sequence. Character similarity is then analyzed, weighted confidence is fused, and the final characters are determined. Word-level N-gram semantic optimization is combined to improve coherence, outputting the candidate's answer text, including: The fused features are input into each heterogeneous OCR engine in parallel. Quality parameters are extracted based on the fused features. A state vector is constructed using the quality parameters to form a state set. The expected reward for each heterogeneous OCR engine is calculated using the Contextual Bandits model for the set of states, and then processed in conjunction with the ε-greedy strategy. The accuracy and processing time of the heterogeneous OCR engines are recorded, and the Contextual Bandits model is updated. The combination of heterogeneous OCR engines activated based on the updated Contextual Bandits model output is determined, and the weight of each heterogeneous OCR engine is determined. By using the original output sequences of heterogeneous OCR engines, a cumulative distance matrix is ​​constructed by combining the weights of each heterogeneous OCR engine, the character similarity is derived, the weight-aware confidence is fused, the recognition result is obtained, and corrections are made to determine the final character. By combining word-level N-Gram semantics and the Viterbi algorithm to optimize coherence and correcting errors in heterogeneous OCR engines, the test taker's answer text is output.

4. The end-to-end multi-agent personnel examination scoring method according to claim 3, characterized in that, By fusing weighted confidence scores, the recognition result is obtained, and then corrected to determine the final character. Specifically: The character prediction probability of any heterogeneous OCR engine is determined based on weight-aware confidence, the recognition result is obtained, a judgment threshold is set, and a confusion matrix is ​​preset. If the recognition result is less than the judgment threshold, the error correction mechanism is triggered, the recognition result is corrected using the confusion matrix, and the final character is determined by the maximum a posteriori probability.

5. The end-to-end multi-agent personnel examination scoring method according to claim 1, characterized in that, Collect data from multiple sources, build a knowledge base, parse the question requirement text, and filter out the reference answer text from the knowledge base, including: Collect multi-source data, including test papers and answers crawled from any of the following channels: personnel examination resource websites, officially released examination outlines, and past civil service recruitment and professional qualification examination question banks, to form a raw dataset. Then, preprocess the dataset, establish a knowledge base, and determine the true quality based on the knowledge base. A semi-supervised learning model is established, which includes a discriminator. The discriminator is trained, and the prediction quality is output based on the knowledge base. A crawling strategy optimization module is built. Based on the text requirements of the question, the original crawling strategy is determined by combining the actual quality and the predicted quality. The error between the predicted quality and the actual quality is calculated and fed back to the crawling strategy optimization module to adjust the crawling priority, determine the new crawling strategy, and update the knowledge base. A deep reinforcement learning model is established, an adaptive crawling strategy is designed, the new crawling strategy is analyzed, the corresponding actions are executed, and reference answer texts are selected from the updated knowledge base.

6. The end-to-end multi-agent personnel examination scoring method according to claim 5, characterized in that, The preprocessing includes unstructured data parsing, question type classification, and labeling.

7. The end-to-end multi-agent personnel examination scoring method according to claim 5, characterized in that, A deep reinforcement learning model is established, an adaptive crawling strategy is designed, the new crawling strategy is analyzed, corresponding actions are executed, and reference answer texts are selected from the updated knowledge base, including: Based on the new crawling strategy, define the action space and state space, determine the constraints, and design the reward function in combination with the actual quality. A deep reinforcement learning-based model is established and trained. The optimal deep reinforcement learning-based model is obtained by combining the reward function with the experience replay and target network stable training process. The optimal deep reinforcement learning model is used to select reference answer texts from the updated knowledge base.

8. The end-to-end multi-agent personnel examination scoring method according to claim 1, characterized in that, Construct a multi-dimensional scoring agent. Determine the initial state based on the question requirements text and the reference answer text. Train each dimension's scoring agent using the initial state. Input the examinee's answer text into the dimension's scoring agent, and output the dimension's score to form a dimension scoring vector, including: The dimensional agents include topic relevance, content completeness, and language fluency. Each dimensional agent adopts an Actor-Critic architecture. Each dimensional agent is trained with the original state and optimized using temporal difference and proximal policy optimization algorithms to obtain the trained Actor-Critic architecture. The candidate's answer text is input into the Actor-Critic framework, which outputs dimensional score scores and integrates them to form a dimensional score vector.

9. The end-to-end multi-agent personnel examination scoring method according to claim 1, characterized in that, A self-consistency analysis is performed on the comprehensive dimensional scoring vector and the original state to output the final score, including: A comprehensive decision-making agent is constructed, and an attention mechanism is adopted. The attention weight of the score of each dimension in the dimension score vector is determined through learnable parameters. The context vector is determined and concatenated with the original state and input into a multi-layer fully connected network to obtain the discretized total score. Set up a score reward function to enhance scoring consistency and adherence to scoring rules; The comprehensive decision-making agent is trained collaboratively using a collaborative training strategy of at least two stages. In the first stage, the comprehensive decision-making agent is frozen and the dimensional agents are optimized independently. In the second stage, a consistent reward and gradient backpropagation mechanism are generated. Based on the discretized total score, the goal of coordinating local optima and global logical self-consistency is achieved to obtain the trained comprehensive decision-making agent. The trained integrated decision-making agent outputs the final score corresponding to the candidate's answer text.

10. The end-to-end multi-agent personnel examination scoring method according to claim 9, characterized in that, A consistent reward and gradient backpropagation mechanism is generated. Based on the discretized total score, a local optimum and a globally logically consistent objective are coordinated to obtain a trained comprehensive decision-making agent, including: Set a control threshold to determine the error of the dimensional agent. If the error is less than the control threshold, freeze all parameters of the dimensional agent and fine-tune them together with the comprehensive decision agent. Generate a consistent reward based on the discretized total score, forcing the comprehensive decision agent to adjust its strategy. By adding gradient backpropagation paths and combining them with consistency rewards, we can coordinate local optima and global logical self-consistency objectives to achieve multi-objective collaborative optimization.

Citation Information

Cited By

  • Generation processing method based on multi-dimensional feature matching screening and confidence coefficient verification

    CN121960505A

  • A generation and processing method based on multidimensional feature matching and confidence verification

    CN121960505B

  • Text data quality evaluation method and system

    CN122197856A

  • A method and system for assessing the quality of text data

    CN122197856B