Public examination interview simulation system based on VR and large language model
The public examination interview simulation system constructed through VR and large language models solves the problems of high cost and inconsistent evaluation of traditional public examination interviews, realizes low-cost and efficient simulation training and objective evaluation, and improves the simulation training effect of candidates.
Patent Information
- Application Number
- CN202510507939.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
AI Technical Summary
The traditional public examination interview simulation training is high cost and has low evaluation consistency, and lacks unified standards, which makes it difficult for candidates to train frequently and obtain objective feedback.
VR and large language models are used to build a virtual examination room, combine digital examiners and multi-modal evaluation, analyze the candidate's answer content through large language models, generate unified scores and improvement suggestions, integrate environmental interference factors and pressure adjustment modules, and provide an immersive simulation interview experience.
Significantly reduce interview costs, improve evaluation consistency and objectivity, help candidates obtain standardized feedback in the virtual environment, shorten the training cycle and improve simulation results.
Smart Images

Figure CN120339009A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of interview virtual digital simulation, and specifically to a public examination interview simulation system based on VR and large language models. Background Art
[0002] In the traditional mode of civil servant interview training, it is crucial to conduct on-site simulation sessions for public examination interviews in advance. The main purpose of this session is to enable candidates to become familiar with the examination room environment in advance, so as to effectively reduce nervousness during the actual interview and improve their coping abilities. However, currently, this training method relies on manual simulation and has several problems: First of all, the cost of the simulation session is relatively high. Generally, for a simulated interview, 5 to 7 simulated examiners, 2 supervisors, 1 timekeeper, and 1 scorekeeper need to be configured. This manpower allocation results in significant labor costs, making it difficult for candidates to conduct low-cost and frequent simulation training, thus limiting the opportunities for candidates to improve their interview skills. Secondly, the levels of simulated examiners vary. Since different examiners may have different evaluation criteria for candidates' answers, the consistency of the evaluation of candidates' performance is relatively low. In addition, due to the lack of a unified evaluation standard, the feedback is often highly subjective, making it difficult for candidates to obtain standardized and objective evaluation bases, thus affecting the effect of simulation training.
[0003] In view of this, in response to the above problems, in-depth research has been carried out, and thus this case has emerged. Summary of the Invention
[0004] The purpose of the present invention is to provide a public examination interview simulation system based on VR and large language models, so as to solve the problems of high cost and low evaluation consistency existing in the traditional manual simulation of public examination interviews proposed in the above background art and help candidates achieve digital public examination interview simulation.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A public examination interview simulation system based on VR and large language models, the simulation system includes: a VR interaction module, a VR scene construction module, a data processing module, a digital human examiner interaction module, a pressure dynamic adjustment module, a large language model analysis module, a multi-modal evaluation module, and a feedback output module; The VR interaction module includes a head-mounted virtual reality display, a gesture sensor, a pressure sensor, and a voice acquisition port, and is used to cooperate with the VR scene construction module to construct a virtual examination room scene; The VR scene construction module is equipped with a 360° panoramic rendering engine, which is used to generate interview venue scenes of different styles, support customizing the scale of examiners, and at the same time integrate an environmental interference factor library to trigger interactive scenes such as noise and item dropping through the VR interaction module; The data processing module is connected to the VR interaction module and is used to collect the examinee's voice, stress indication and movement data; The digital human examiner interaction module builds an intelligent examiner group based on a large language model, recognizes the answer content in real time, and dynamically generates questioning questions or interruptions in combination with the pressure dynamic adjustment module. At the same time, it uses an emotional computing algorithm to adjust the digital human's expression and body language according to the quality of the user's answer; The pressure dynamic adjustment module establishes a three-level pressure index model, calibrates the pressure level through the dual channels of heart rate detection and speech coherence detection, and regulates the questioning or interruption behavior of the digital human examiner and the superposition of environmental interference in the environmental interference factor library; The large language model analysis module is trained based on the civil servant interview question bank, analyzes the answers of candidates in real time, and generates scores and improvement suggestions in combination with the multimodal evaluation module; The large language model analysis module adopts dynamic scoring model optimization for civil servant interviews, introduces hierarchical analysis method to reconstruct indicator weights, establishes differentiated scoring matrices for different civil servant examinations and different positions, including comprehensive management, administrative law enforcement, and professional and technical position examinations, and provides position matching analysis, recommending suitable departments for applicants based on interview performance data; At the same time, a context-sensitive scoring algorithm is adopted, including focusing on theoretical depth for policy questions for administrative posts and focusing on response speed for questions for emergency management posts. In addition, special improvements are made to the construction of policy knowledge graphs in the big model, and the central and local policy documents within the scope of the civil service examination are stored in a structured manner to achieve automatic comparison of the answer content with the latest policies. The multimodal assessment module generates a three-dimensional ability map: content logic, language fluency, and body expression, and generates a progress curve by comparing historical data; The feedback output module combines the scoring results with the dynamic lip shape and gestures of the virtual examiner and presents them to the examinee in real time through the VR scene; The simulation process includes the following steps: candidate sign-in, candidate draw, candidate waiting, candidate interview answer, and examiner feedback and evaluation. The specific interactive process is as follows: After the candidate puts on the VR headset, the system loads the virtual examination scene. The candidate experiences the check-in, lottery, and waiting stages in an immersive manner before entering the interview examination room. The virtual examiner in the center of the interview scene is driven by a large language model. When the candidate answers questions, the candidate's voice data is converted into text and input into the large language model analysis module. The model generates a score based on preset evaluation rules, and the score is fed back to the candidate through the virtual examiner's voice synthesis and facial expression animation.
[0006] Preferably, the VR interaction module is connected to the data processing module server via a wireless network; The head-mounted virtual reality display has built-in headphones and a screen display, and is responsible for outputting audio and video information of the VR scene; The voice acquisition port carrier is a microphone for picking up the voices emitted by students, cooperating with a gesture sensor and external video recording to capture micro-expressions and body behaviors, and converting the collected information into digital signals and inputting them to the data processing module server.
[0007] Preferably, the server is equipped with a database for cooperating with the large language model analysis module to provide training data. The database includes a question bank corresponding to the applied job responsibilities, answer keywords closely related to the questions, and key points for attention in body behaviors and language abilities.
[0008] Preferably, the database establishes information communication with cooperative public examination units through the cloud network, obtains the actual interview data and question banks of the cooperative units, and stores and supplements them.
[0009] Preferably, the database stores the information fed back by the VR interaction module and the corresponding virtual scene information, combines the audio-visual and behavioral information into a simulated interview record based on the time axis, and reproduces the scene through the VR interaction module combined with the feedback output module, and gives corresponding opinions and suggestions for the interview simulation.
[0010] Preferably, the preset evaluation rule parameters of the large language model analysis module are divided into two major modules: content quality and presentation ability, and the weight ratios of content quality and presentation ability are 70% and 30% respectively.
[0011] Preferably, the weight distribution logic of the content quality module covers the core abilities of civil servant interviews, highlights relevance, hierarchy, and content depth, analyzes the content of the candidates' answers converted from speech to text, and generates scores and improvement suggestions; The weight distribution of the content quality module includes 20% for comprehensive analysis ability, 20% for answering logic and framework, 15% for content matching, 10% for practice and organization ability, and 5% for values and motivation.
[0012] Preferably, the weight distribution logic of the presentation ability module includes using VR to capture micro-expressions, voice features, and body language, quantifying the implicit standards of traditional interviews, and being a targeted training item for pre-examination interview simulations.
[0013] Preferably, the weight distribution of the presentation ability module includes a 15% weight ratio for language expression and a 5% weight ratio for time control. This part calculates and comprehensively scores the fluency, speech rate, volume, and rationality of the answering duration of the audio information collected by the voice acquisition port.
[0014] Preferably, the weight distribution of the presentation ability module includes a 10% proportion for confidence and emotion management. This part conducts body language analysis on the videos and action information collected by the gesture sensor and external video recording, and comprehensively scores according to the criteria of whether the interviewee's gestures are natural without small movements and whether the attention of the facial eyes is concentrated.
[0015] Compared with the prior art, the public examination interview simulation system based on VR and large language models involved in the present invention has the following advantages: 1. The interview cost of this simulation system is significantly reduced. Through the immersive simulation environment provided by virtual reality technology, candidates can independently complete the simulated interview without artificial teachers acting as examiners. This not only saves human resources but also improves the operability of the simulated interview. 2. The consistency in evaluating candidates' performance has been significantly improved. Through learning with large language models, digital virtual examiners can accurately master and apply unified evaluation criteria to provide feedback to candidates. Since these criteria are strictly set and precisely executed by the model, the evaluations given by digital examiners have low subjectivity, ensuring the consistency and objectivity of the evaluations, which is an important advantage for candidates. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic diagram of the simulation process of the present invention; Figure 2 is a schematic diagram of the specific interaction process of the present invention; Figure 3 is a schematic block diagram of the system of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Please refer to Figures 1 - 3 , the present invention provides a technical solution: a public examination interview simulation system based on VR and large language models, the simulation system includes: a VR interaction module, a VR scene construction module, a data processing module, a digital human examiner interaction module, a pressure dynamic adjustment module, a large language model analysis module, a multimodal evaluation module, and a feedback output module; The VR interaction module includes a head-mounted virtual reality display, a gesture sensor, a pressure sensor, and a voice collection port, and is used to construct a virtual examination room scene; The data processing module is connected to the VR interaction module and is used to collect the examinee's voice, heart rate stress indicators, and motion data; The VR scene construction module is equipped with a 360° panoramic rendering engine and can generate interview venue scenes in more than 20 different styles such as auditoriums and meeting rooms, supporting customizing the scale of examiners from 5 to 10 people, as well as supervisors, timekeepers, and scorers, etc.
[0019] The integrated environmental interference factor library includes noise, item dropping, etc., which are triggered through the VR interaction module.
[0020] The digital human examiner interaction module constructs an intelligent examiner group based on a large language model, performs real-time speech recognition on the speech content, and dynamically generates skeptical questions such as "Is your data source reliable?" or interruption behaviors.
[0021] Adopt an emotion computing algorithm to adjust the digital human's facial expressions and body languages according to the quality of the user's answer, such as applauding, skeptical gestures, and whispering to each other, etc.
[0022] The pressure dynamic adjustment module establishes a three-level pressure index model: Primary level: 5% of the examiners ask random questions at intervals of >2 minutes; Intermediate level: 10% of the examiners continuously follow up with questions, including semantic trap questions; Advanced level: 20% of the examiners synchronously interrupt + environmental interference superposition; Calibrate the pressure level through the dual channels of heart rate detection and speech coherence detection.
[0023] The multi-modal evaluation module generates a three-dimensional ability map: evaluate the content logic, language fluency, and body performance through large language model scoring, voiceprint analysis, and motion capture, and generate a progress curve by comparing historical data.
[0024] The large language model analysis module is trained based on the civil servant interview question bank, analyzes the examinee's answer content in real time, and generates scores and improvement suggestions; The large language model analysis module optimizes the dynamic scoring model for civil servant interviews, introduces the analytic hierarchy process to reconstruct the index weights, establishes a differentiated scoring matrix for different civil servant examinations and different positions, including comprehensive management, administrative law enforcement, and professional and technical position examinations, and provides an analysis of the position matching degree. According to the interview performance data, recommend suitable departments for application, such as the emergency management position requires a reaction speed in the top 20%; At the same time, adopt a context-sensitive scoring algorithm, including that for administrative positions, policy-related questions focus on theoretical depth, while questions for emergency management positions focus on reaction speed and set a time weight coefficient of 0.5 - 1.5 times accordingly. And for the large model, there is a special improvement in constructing a policy knowledge graph, which structurally stores the central and local policy documents within the scope of the civil servant examination, and realizes the automatic comparison of the answer content with the latest policies; The feedback output module combines the scoring results with the dynamic lip movements and gestures of the virtual examiner and presents them to the candidates in real time through the VR scenario; The VR interaction module is connected to the server of the data processing module through a wireless network; The head-mounted virtual reality display is equipped with built-in headphones and a screen display, which is responsible for outputting the audio and video information of the VR scenario; The voice acquisition port is carried by a microphone to pick up the voice emitted by the student, and cooperates with the gesture sensor and external image recording to capture micro-expressions and body behaviors, and converts the collected information into digital signals and inputs them to the server of the data processing module.
[0025] The server is equipped with a database for cooperating with the large language model analysis module to provide training data. The database includes a question bank corresponding to the job responsibilities of the application, answer keywords closely related to the questions, and key points of attention for body behaviors and language abilities. Further, the database establishes information communication with cooperative public examination units through the cloud network, obtains the real interview data and question banks of the cooperative units, and stores and supplements them. At the same time, the database stores the information fed back by the VR interaction module and the corresponding virtual scenario information, combines the audio and video and behavior information to simulate the interview record based on the time axis, and reproduces the scenario through the VR interaction module combined with the feedback output module, and gives corresponding opinions and suggestions for the interview simulation.
[0026] The preset evaluation rule parameters of the large language model analysis module are divided into two major modules: content quality and presentation ability, and the weight ratios of content quality and presentation ability are 70% and 30% respectively.
[0027] The weight distribution logic of the content quality module covers the core abilities of civil servant interviews, highlighting the relevance, hierarchy and content depth, analyzes the content of the candidates' answers converted from speech to text, and generates scores and improvement suggestions; The weight distribution of the content quality module includes 20% for comprehensive analysis ability, 20% for answering logic and framework, 15% for content matching, 10% for practice and organization ability, and 5% for values and motivation. The specific content includes: 1. Comprehensive analysis ability - Evaluation dimensions: depth of view, grasp of the essence of the problem, comprehensiveness of multi-angle analysis, such as at the policy, social, and personal levels.
[0028] 2. Answering logic and framework - Sub-dimensions: - Structural clarity, such as total score - total score / progressive / causal chain; - Hierarchy of sub-arguments, whether vertically stratified clearly; - Logical coherence, whether there are no jumps or contradictions.
[0029] 3. Content matching - Sub-dimensions: - Relevance, whether closely related to the topic keywords; - Job fit, whether combined with the responsibilities of the applied position; - Accuracy of policy theory citation.
[0030] 4. Practical and organizational capabilities - Evaluation dimensions: Feasibility of solutions, resource coordination strategies, rationality of contingency measures.
[0031] 5. Values and motivation - Evaluation dimensions: Public service awareness, job stability, organizational culture recognition.
[0032] The weight allocation logic of the demonstrated ability module includes using VR to capture microexpressions, voice features, and body language, quantifying the implicit criteria of traditional interviews, and providing targeted training items for pre-examination interview simulations. The specific content includes: 1. Language expression - Sub-dimensions: - Fluency, whether there are no stutters or repetitions, and the AI detects the stutter frequency; - Speech rate and volume, whether stable and moderate, and VR voiceprint analysis is performed; - Intonation infectivity, with cadence and natural emotional transmission; - Filler words, dynamic deduction of points, 0.5 points deducted for each occurrence.
[0033] 2. Confidence and emotion management - Sub-dimensions: - Body language, whether gestures are natural and there are no small movements, and VR motion capture is performed; - Stress resistance, such as the degree of smooth answering when faced with follow-up questions or sudden interferences.
[0034] 3. Time control - Evaluation dimensions: - Rationality of answering duration, 1 point deducted for every 10 seconds of overtime, and 0.5 points deducted for being too short; - Rhythm compactness, whether the key points are evenly distributed without redundancy.
[0035] The weight allocation of the demonstrated ability module includes a 15% weight for language expression and a 5% weight for time control. This part calculates and comprehensively evaluates the audio information collected through the voice acquisition port for fluency, speech rate and volume, and rationality of answering duration.
[0036] The weight distribution of the demonstration ability module includes 10% for self-confidence and emotional management. This part conducts body language analysis on the video and motion information collected by gesture sensors and external image records, and extracts standards such as whether the interviewee's gestures are natural and free of small movements, and whether the facial and eye direction is focused, for comprehensive scoring.
[0037] The simulation process includes the following steps: candidate sign-in, candidate draw, candidate waiting, candidate interview answer, and examiner feedback and evaluation. The specific interactive process is as follows: After the candidate puts on the VR headset, the system loads the virtual examination scene. The candidate experiences the check-in, lottery, and waiting stages in an immersive manner before entering the interview examination room. The virtual examiner in the center of the interview scene is driven by a large language model. When the candidate answers questions, the candidate's voice data is converted into text and input into the large language model analysis module. The model generates a score based on preset evaluation rules, and the score is fed back to the candidate through the virtual examiner's voice synthesis and facial expression animation.
[0038] This improved system has been verified in a pilot project for civil service examination training, shortening the average training period of candidates by 37%, improving the consistency of examiners' scoring by 28%, and achieving 92% satisfaction of candidates with the fairness of scoring. Through the immersive simulation environment provided by virtual reality technology, candidates can independently complete the simulated interview without the need for human teachers to act as examiners, effectively saving human resources. At the same time, digital virtual examiners learn large language models and apply unified evaluation standards to provide feedback to candidates, ensuring the consistency and objectivity of the evaluation, which is an important advantage for candidates.
[0039] Contents not described in detail in this specification belong to the prior art known to professional and technical personnel in the field. Although the embodiments of the present invention have been shown and described, it is understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the attached claims and their equivalents.
Claims
1. A mock interview system for civil service examinations based on VR and large language models, characterized in that: The simulation system includes: a VR interaction module, a VR scene construction module, a data processing module, a digital human examiner interaction module, a pressure dynamic adjustment module, a large language model analysis module, a multimodal evaluation module and a feedback output module; The VR interaction module includes a head-mounted virtual reality display, a gesture sensor, a pressure sensor, and a voice acquisition port, which are used to cooperate with the VR scene construction module to build a virtual examination room scene; The VR scene construction module is equipped with a 360° panoramic rendering engine, which is used to generate interview venue scenes of different styles, supports custom examiner scale, and integrates an environmental interference factor library, and triggers noise and item drop interaction scenes through the VR interaction module; The data processing module is connected to the VR interaction module and is used to collect the examinee's voice, stress indication and movement data; The digital human examiner interaction module builds an intelligent examiner group based on a large language model, recognizes the answer content in real time, and dynamically generates questioning questions or interruptions in combination with the pressure dynamic adjustment module. At the same time, it uses an emotional computing algorithm to adjust the digital human's expression and body language according to the quality of the user's answer; The pressure dynamic adjustment module establishes a three-level pressure index model, calibrates the pressure level through the dual channels of heart rate detection and speech coherence detection, and regulates the questioning or interruption behavior of the digital human examiner and the superposition of environmental interference in the environmental interference factor library; The large language model analysis module is trained based on the civil servant interview question bank, analyzes the answers of candidates in real time, and generates scores and improvement suggestions in combination with the multimodal evaluation module; The large language model analysis module adopts dynamic scoring model optimization for civil servant interviews, introduces hierarchical analysis method to reconstruct indicator weights, establishes differentiated scoring matrices for different civil servant examinations and different positions, including comprehensive management, administrative law enforcement, and professional and technical position examinations, and provides position matching analysis, recommending suitable departments for applicants based on interview performance data; At the same time, a context-sensitive scoring algorithm is adopted, including focusing on theoretical depth for policy questions for administrative posts and focusing on response speed for questions for emergency management posts. In addition, special improvements are made to the construction of policy knowledge graphs in the big model, and the central and local policy documents within the scope of the civil service examination are stored in a structured manner to achieve automatic comparison of the answer content with the latest policies. The multimodal assessment module generates a three-dimensional ability map: content logic, language fluency, and body expression, and generates a progress curve by comparing historical data; The feedback output module combines the scoring results with the dynamic lip shape and gestures of the virtual examiner, and presents them to the examinee in real time through a VR scene.
2. The public examination interview simulation system based on VR and large language model according to claim 1, characterized in that: The simulation system process includes the following steps: candidate sign-in, candidate draw, candidate waiting, candidate interview answer, examiner feedback and evaluation. The specific interaction process is as follows: After the candidate puts on the VR headset, the system loads the virtual examination scene. The candidate experiences the check-in, lottery, and waiting stages in an immersive manner before entering the interview examination room. The virtual examiner in the center of the interview scene is driven by a large language model. When the candidate answers questions, the candidate's voice data is converted into text and input into the large language model analysis module. The model generates a score based on preset evaluation rules, and the score is fed back to the candidate through the virtual examiner's voice synthesis and facial expression animation.
3. The public examination interview simulation system based on VR and large language model according to claim 1, characterized in that: The VR interaction module is connected to the data processing module server via a wireless network; The head-mounted virtual reality display is built-in with headphones and a screen display, responsible for outputting audio and video information of the VR scene; The carrier of the voice collection port is a microphone for picking up the voices emitted by students, cooperating with the gesture sensor and external image recording to capture micro-expressions and body behaviors, and converting the collected information into digital signals and inputting them to the data processing module server; The server is equipped with a database for cooperating with the large language model analysis module to provide training data. The database includes a question bank corresponding to the job responsibilities of the application, answer keywords closely related to the questions, and key points for attention in body behavior and language ability.
4. The public examination interview simulation system based on VR and large language model according to claim 3, wherein: The database establishes information communication with cooperative public examination units through the cloud network, obtains the actual interview data and question banks of the cooperative units, and stores and supplements them.
5. The public examination interview simulation system based on VR and large language model according to claim 3, characterized in that: The database stores the information fed back by the VR interaction module and the corresponding virtual scene information, combines the audio-visual and behavior information into a simulated interview record based on the time axis, and reproduces the scene through the VR interaction module combined with the feedback output module, and gives corresponding opinions and suggestions for the interview simulation.
6. The public examination interview simulation system based on VR and large language model according to claim 1, characterized in that: The preset evaluation rule parameters of the large language model analysis module are divided into two major modules: content quality and presentation ability, and the weight ratios of content quality and presentation ability are 70% and 30% respectively.
7. The public examination interview simulation system based on VR and large language model according to claim 6, wherein: The weight distribution logic of the content quality module covers the core capabilities of civil servant interviews, highlighting relevance, hierarchy and content depth, analyzing the content of the candidates' answers converted from speech to text, and generating scores and improvement suggestions; The weight distribution of the content quality module includes 20% for comprehensive analysis ability, 20% for answering logic and framework, 15% for content matching, 10% for practice and organization ability, and 5% for values and motivation.
8. The public examination interview simulation system based on VR and large language model according to claim 6, characterized in that: The weight distribution logic of the presentation ability module includes using VR to capture micro-expressions, voice features, and body language, quantifying the implicit standards of traditional interviews, and targeting training items for pre-examination interview simulations.
9. The public examination interview simulation system based on VR and large language model according to claim 8, characterized in that: The weight distribution of the presentation ability module includes a weight ratio of 15% for language expression and 5% for time control. This part calculates and comprehensively scores the fluency, speech rate and volume of the audio information collected by the voice collection port, as well as the rationality of the answering duration.
10. A mock system for public examination interviews based on VR and large language models according to claim 8, characterized in that: The weight distribution of the presentation ability module includes a weight ratio of 10% for confidence and emotion management. This part conducts body language analysis on the video and action information collected by the gesture sensor and external image recording, and extracts the criteria of whether the interviewee's gestures are natural without small movements and whether the facial eye orientation and attention are concentrated for comprehensive scoring.
Citation Information
Cited By
Dynamic hierarchical knowledge graph construction method and system for public operator examination
CN121436135A