Voice interview method and device based on artificial intelligence, equipment and medium
Through the artificial intelligence model, the voice data of interview users are processed, interview questions are generated and scored, and standardized voice interview reports are provided, which solves the problem of inconsistent judgment standards in intelligent interviews, and achieves fast and efficient interview results analysis.
Patent Information
- Application Number
- CN202510675450.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing intelligent interview technology, the interview evaluation standards are inconsistent, and the subjective judgment of real-person interviewers leads to inconsistent judgments.
The interview user's voice interview instructions are obtained through the artificial intelligence model, the question audio is generated, the answer audio is converted into text and preprocessed, the answer scoring model is used to score, the interview score data is generated, and the interview analysis is conducted to provide standardized voice interview reports.
It realizes fast and efficient voice interview based on artificial intelligence, standardizes the analysis of interview results, reduces the human subjective influence, and improves the consistency of judgments.
Smart Images

Figure CN120340489A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a voice interview method, device, equipment and medium based on artificial intelligence. Background Technique
[0002] With the continuous development of electronic technology, artificial intelligence technology has gradually become active in various technical fields, such as intelligent vehicle driving, intelligent customer service, intelligent floor cleaning robots, intelligent interviews, etc. Among them, intelligent interviews are conducted by artificial intelligence interviewers instead of traditional interviewers to interview candidates, solving the low-efficiency problem of traditional manual interviews restricted by time, venue, manpower, etc.
[0003] For example, intelligent interview means can be applied in the interview procedures for positions such as bank account managers or fund managers in the financial field, and can also be applied in the interview procedures for positions such as hospital doctors and nurses in the medical field, or sales representatives and procurement managers in pharmaceutical companies, to conduct interviews on candidates' human resource requirements and / or professional skill requirements with artificial intelligence interviews instead of traditional interviewers.
[0004] However, in intelligent interviews, the existing technology usually simply processes the audio data of the interviewees and provides it to real interviewers for manual review, relying on the real interviewers' understanding ability of the language expressions of different interview users, which is prone to subjective influences brought by human judgment, resulting in inconsistent interview evaluation criteria. Summary of the Invention
[0005] The present invention provides a voice interview method, device, computer equipment and medium of artificial intelligence to solve the technical problem of inconsistent interview evaluation criteria in intelligent voice interviews.
[0006] In a first aspect, a voice interview method based on artificial intelligence is provided, including:
[0007] Obtain a voice interview instruction sent by an interview user, where the voice interview instruction includes a position type and a voice type;
[0008] Obtain a plurality of interview question texts from a preset question bank according to the position type, and generate a question audio according to the voice type and the interview question texts through an artificial intelligence model;
[0009] Obtain an answer audio of the interview user based on the question audio, convert the answer audio into an interview answer text, and extract text features after preprocessing the interview answer text;
[0010] Obtain an answer scoring model corresponding to the interview question text, and input the text features into the answer scoring model to obtain interview score data;
[0011] Perform an interview analysis through an artificial intelligence model based on the text features and the interview score data to obtain a voice interview analysis report.
[0012] In a second aspect, a voice interview device based on artificial intelligence is provided, including:
[0013] An instruction acquisition module, configured to acquire a voice interview instruction sent by an interview user, where the voice interview instruction includes a position type and a voice type;
[0014] An audio generation module, configured to obtain a plurality of interview question texts from a preset question bank according to the position type, and generate a question audio according to the voice type and the interview question texts through an artificial intelligence model;
[0015] A text acquisition module, configured to acquire a response audio of the interview user based on the question audio, convert the response audio into an interview response text, and extract text features after preprocessing the interview response text;
[0016] A scoring module, configured to obtain an answer scoring model corresponding to the interview question text, and input the text features into the answer scoring model to obtain interview score data;
[0017] An analysis module, configured to perform an interview analysis through an artificial intelligence model based on the text features and the interview score data to obtain a voice interview analysis report.
[0018] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned voice interview method based on artificial intelligence is implemented.
[0019] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned voice interview method based on artificial intelligence is implemented.
[0020] In the solutions implemented by the above-mentioned voice interview method, device, computer device, and storage medium based on artificial intelligence, multiple interview question texts can be obtained from a preset question bank according to a voice interview instruction sent by an interview user through an artificial intelligence model and a question audio can be generated. A response audio of the interview user is acquired based on the question audio, and the response audio is scored through an answer scoring model corresponding to the interview question text to obtain interview score data. An artificial intelligence model analyzes the text features extracted from the response audio and the interview score data to obtain a voice interview analysis report, realizing fast and efficient voice interviews based on artificial intelligence and standardized analysis of interview results. Description of the Drawings
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 It is a schematic diagram of an application environment of a voice interview method based on artificial intelligence in an embodiment of the present invention;
[0023] Figure 2 It is a schematic flowchart of a voice interview method based on artificial intelligence in an embodiment of the present invention;
[0024] Figure 3 It is a schematic structural diagram of a voice interview device based on artificial intelligence in an embodiment of the present invention;
[0025] Figure 4 It is a schematic structural diagram of a computer device in an embodiment of the present invention;
[0026] Figure 5 It is another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0028] The voice interview method based on artificial intelligence provided by the embodiments of the present invention can be applied, for example, in Figure 1In the application environment, the client communicates with the server through the network. The server can receive the voice interview instruction of the interviewed user sent by the client, obtain multiple interview question texts from a preset question bank according to the position type of the voice interview instruction, and generate a question audio according to the voice type of the voice interview instruction and the interview question texts through an artificial intelligence model; obtain the answer audio of the interviewed user sent by the client based on the question audio, convert the answer audio into an interview answer text, and extract text features after preprocessing the interview answer text; obtain the answer scoring model corresponding to the interview question text, and input the text features into the answer scoring model to obtain interview score data; perform interview analysis through an artificial intelligence model according to the text features and the interview score data to obtain a voice interview analysis report, realizing fast and efficient voice interviews based on artificial intelligence and standardized interview result analysis. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0029] Please refer to Figure 2 as shown in Figure 2 a flowchart of a method for voice interviews based on artificial intelligence provided by an embodiment of the present invention, including the following steps:
[0030] S1: Obtain the voice interview instruction sent by the interviewed user, where the voice interview instruction includes a position type and a voice type.
[0031] The voice interview method provided by the present invention can be applied to a voice interview system based on artificial intelligence. The voice interview system based on artificial intelligence can be an interview application program constructed based on a grid recruitment platform, combined with an interview question bank, intelligent speech recognition, and a neural network model, such as a bank interview program in the financial field, a hospital interview program in the medical field, or a pharmaceutical company interview program.
[0032] In one embodiment, when the interviewed user accesses the home page, the voice interview system plays a preset voice generated based on the default voice type. The voice content of the preset voice can be a welcome message pre-set by the voice interview system, such as "Welcome to the interview of XX Company. Wish you a successful interview." This welcome message can be set according to the application scenario and actual needs of the voice interview system.
[0033] In this embodiment, the obtaining of the voice interview instruction sent by the interviewed user includes:
[0034] Obtain the user information of the interviewed user, and provide corresponding position type options in the interface of the voice interview system according to the user information;
[0035] Receive the job types selected by the interviewed user and provide a plurality of preset voice type options in the interface of the voice interview system;
[0036] Receive the voice type selected by the interviewed user to obtain a voice interview instruction composed of the job type and the voice type.
[0037] For example: after the voice interview system receives the instruction for the interviewed user to enter the interview, it receives the user information input by the user through the user information acquisition interface. The user information can be the information input in real time by the interviewed user obtained through the information input box and selection box of the user interface, or the user information obtained by identifying the resume attachment uploaded by the interviewed user through an artificial intelligence model.
[0038] In this factual example, providing corresponding job type options on the interface of the voice interview system according to the user information includes:
[0039] Obtain all the current job types in the voice interview system and the job requirements corresponding to each job type;
[0040] Match the user information with each job requirement respectively to obtain the job type corresponding to the job requirement that the user information matches;
[0041] Generate a multiple-choice question based on the matched job type and provide the multiple-choice question and the corresponding job type options on the interface of the voice interview system.
[0042] For example: the job types that the voice interview system can currently interview include front-end developers, back-end developers, testers, and business docking personnel. The job requirements corresponding to each job type include relevant technical requirements and years of experience requirements, etc. According to the job requirements corresponding to each job type, match the obtained user information in turn to determine whether the user information meets the relevant technical requirements in the job requirements (such as front-end developers are required to master technologies such as HTML, CSS, JS, Unix, and Linux operating systems, and Web server configuration, and determine whether the technologies mastered by the user in the user information meet the above requirements) and years of experience requirements (such as front-end developers are required to have 3 years of front-end development work experience, and determine whether the work experience of the user in the user information meets the above requirements) to screen out the job types that the user information meets the job requirements. In other application scenarios, the job requirements of some job types may also include relevant certificates. For example, in the field of fintech, accounting positions require accounting certificates, and in the field of healthcare, medical nursing positions require nursing certificates. The voice interview system can identify and verify the certificate pictures uploaded by the user to determine whether the user information meets the relevant certificates in the job requirements.
[0043] S2: Obtain multiple interview question texts from a preset question bank according to the job type, and generate question audio according to the voice type and the interview question text through an artificial intelligence model.
[0044] In this embodiment, before obtaining multiple interview question texts from a preset question bank according to the position type, the following steps are included:
[0045] Analyze multiple pre-set job types to obtain the job requirement keywords corresponding to each job type;
[0046] Using an artificial intelligence language model, generate corresponding interview question texts, standard answers to the interview question texts, and question weights according to the job requirement keywords;
[0047] Generate multiple derived answers based on the interview question text through an artificial intelligence language model, and score each derived answer based on the standard answer;
[0048] Based on the standard answer, the derived answer and the corresponding score, a pre-built neural network model is trained to obtain an answer scoring model corresponding to the interview question text;
[0049] The interview question text is associated with the corresponding question weight, the answer scoring model and the position type and stored in a preset question bank.
[0050] In this embodiment, the pre-set multiple job types are analyzed to obtain job requirement keywords corresponding to each job type, including: analyzing the job requirements and work content corresponding to each job type through an artificial intelligence model to obtain corresponding job requirement keywords, and the job requirement keywords include competency type keywords, personality type keywords, ability type keywords, etc.
[0051] Competence refers to the characteristics that an individual possesses and consistently uses in an appropriate manner to achieve ideal performance. These characteristics include knowledge, skills, self-image, social motivation, traits, thinking patterns, mental sets, and ways of thinking, perceiving, and acting. Competence includes three dimensions: career, behavior, and strategic integration. The career dimension refers to the skills for handling specific, daily tasks, the behavior dimension refers to the skills for handling non-specific, arbitrary tasks, and the strategic integration dimension refers to management skills combined with organizational contexts.
[0052] In the application scenario of recruitment interviews, companies can use competencies to identify whether the applicant's behavior can achieve the company's predetermined development goals. The impact of competencies on predetermined goals is measurable, and companies can use the measurability of competencies to evaluate the current gaps in applicants' competencies and the direction and degree of improvement needed in the future.
[0053] For example, the competencies of an enterprise include the following three modules: The self-management module, including achievement-oriented qualities and learning and innovation qualities. Among them, the achievement-oriented qualities include the following dimensions: self-vision, courage to challenge, stress tolerance, and pursuit of excellence; the learning and innovation qualities include the following dimensions: learning willingness, learning strategies, application of what is learned, and innovation awareness. The others management module, including team management qualities and communication and coordination qualities. Among them, the team management qualities include the following dimensions: teamwork, subordinate cultivation, effective motivation, and culture shaping; the communication and coordination qualities include the following dimensions: effective expression, attentive listening, positive feedback, and conflict resolution. The task management module, including customer-oriented qualities and plan management qualities. Among them, the customer-oriented qualities include the following dimensions: service awareness, need exploration, effective response, and sustainable win-win; the plan management qualities include the following dimensions: plan formulation, time management, execution ability, and result orientation. In this embodiment, the competency dimensions required for the position type can be selected as the competency type keywords included in the position requirement keywords according to the position requirements and work content of the position type for interview recruitment.
[0054] In one embodiment, a large amount of data is obtained based on the position type for big data analysis, descriptions regarding the qualification and ability parts are found, and statistics and classification extraction are performed on the qualification and ability parts, and the position requirements of the position type can be roughly obtained. The long text of the extracted description part is segmented, function words are removed, and the job responsibilities and requirements are specified and generalized. According to the statistical results, all the words related to the competency requirements of the position type are marked to complete the preliminary determination of the competency words, and some words very important for a specific position are obtained, such as mining, information search, consulting, customer, understanding, coordination, knowledge, transaction, product, policy, risk, negotiation, business trip, overtime, learning, contract, procedures, etc. Through big data analysis, the competency indicators corresponding to the position type are selected according to the word frequency of the words of the position type. In this embodiment, the competency dimension indicators of this position can also be manually reviewed and confirmed.
[0055] For example, for related positions in the fintech field, such as bank tellers, the corresponding competency words include customer, understanding, coordination, consulting, product, procedures, etc. For positions in the medical and health field, such as hospital nurses, the corresponding competency words include carefulness, patience, communication, coordination, etc.
[0056] In this embodiment, using the artificial intelligence language model to generate the interview question text, the standard answer of the interview question text, and the question weight according to the position requirement keywords includes:
[0057] Each position requirement keyword is respectively input into the artificial intelligence language model to generate multiple interview question texts respectively corresponding to each position requirement keyword and the corresponding standard answer that meets the position requirement keyword;
[0058] Obtain the importance levels of each job requirement keyword in the job type, and set corresponding question weights for the interview question texts corresponding to each job requirement keyword according to the importance levels.
[0059] For example: Input the ability type keyword "HTML" in the job requirement keywords into the artificial intelligence language model. The artificial intelligence language model extracts relevant technical points from the preset data corpus according to this keyword and generates corresponding technical questions and answers. Use this technical question as the interview question text and the corresponding answer as the standard answer.
[0060] In this factual example, the importance levels of each job requirement keyword in the job type can be set according to the recruitment requirements of the job type. For example: In the recruitment requirements for front-end developers, the technical requirements are HTML (required item), CSS (required item), JS (required item), Web server configuration (bonus item), Unix and Linux operating systems (bonus item). Then, the importance level of the job requirement keyword corresponding to the required item is higher than that of the job requirement keyword corresponding to the bonus item. Another example: In the recruitment requirements for accounting positions in the fintech field, the importance level of accounting professional skills is higher than that of office software. In the medical and health field, the importance level of the medical and nursing experience of medical staff is higher than that of communication skills.
[0061] In this embodiment, the artificial intelligence language model generates multiple derivative answers according to the interview question text, and scores each derivative answer based on the standard answer, including:
[0062] Identify the keywords in the interview question text, and perform synonym replacement or antonym replacement on the keywords to obtain multiple derivative question texts;
[0063] Generate derivative answers corresponding to each derivative question text through the artificial intelligence language model;
[0064] Convert the standard answer into a standard text vector, convert each derivative answer into a corresponding derivative text vector, calculate the vector distance between each derivative text vector and the standard text vector, and score the derivative answer corresponding to each derivative text vector according to the vector distance.
[0065] In this embodiment, training the pre-constructed neural network model based on the standard answer, the derivative answer, and the corresponding scores to obtain the answer scoring model corresponding to the interview question text includes:
[0066] Convert the standard answer and the derivative answer into corresponding answer text vectors;
[0067] Inputting the answer text vector into a pre-built neural network model to obtain a prediction score for the answer text vector;
[0068] The loss value of the predicted score and the actual score corresponding to the answer text vector is calculated, and the neural network model is optimized according to the loss value to obtain the answer scoring model.
[0069] In this embodiment, the step of generating a question audio according to the voice type and the interview question text by using an artificial intelligence model includes:
[0070] Modeling, extracting and encoding the interview question text to obtain a text matrix;
[0071] Extracting spectral features of the text matrix through an acoustic model to obtain spectral feature information;
[0072] The timbre features of the voice type are extracted, and an artificial intelligence model is used to perform sound synthesis processing according to the timbre features and the spectrum feature information to obtain the question audio.
[0073] In this embodiment, modeling, extracting and encoding the interview question text includes:
[0074] Convert each character in the interview question text into a modeling unit; use onehot encoding to convert the modeling unit corresponding to the character into a character vector; concatenate the character vectors according to the order of the characters in the corresponding interview question text to obtain a text matrix. For example: convert "我" into the initial and final format "w o3".
[0075] S3: Obtain the answer audio of the interview user based on the question audio, convert the answer audio into an interview answer text, and extract text features after preprocessing the interview answer text.
[0076] In this embodiment, obtaining the answer audio of the interview user based on the question audio includes:
[0077] Providing the question audio through the voice interview system, and providing a re-listen button and an answer button in the system interface;
[0078] After receiving a re-listening instruction sent by the interview user through the re-listening button, replaying the question audio;
[0079] After receiving the answer instruction sent by the interview user through the answer button, the audio data of the interview user within a preset time period is obtained as the answer audio of the interview user.
[0080] In this embodiment, converting the response audio into an interview response text, and extracting text features after preprocessing the interview response text includes:
[0081] Inputting the response audio into a preset speech-to-text model to obtain the text information corresponding to the response audio;
[0082] Using an artificial intelligence language model to correct the grammar and semantics of the text information to obtain the interview response text;
[0083] Performing word segmentation on the interview response text, and identifying and removing stop words from the segmented interview response text to obtain a preprocessed text;
[0084] Converting the preprocessed text into a text vector to obtain the text features of the interview response text.
[0085] In this embodiment, the response audio of the interview user can be converted into corresponding text information through real-time ASR audio transcoding technology.
[0086] Automatic Speech Recognition (ASR) is an important technology in the fields of artificial intelligence and natural language processing, aiming to convert human speech signals into corresponding texts. A typical ASR system transcribes sound into text through a series of steps, including preprocessing, feature extraction, acoustic model calculation, language model application, and decoding output, etc.:
[0087] Preprocessing: Operations such as noise reduction, silence segment detection, and pre-emphasis filtering are performed on the input speech to improve the quality of the speech signal. This step can reduce the influence of environmental noise and segment the audio into frames suitable for processing.
[0088] Feature extraction: Converting the original audio into a feature representation convenient for machine processing, such as Mel Frequency Cepstral Coefficients (MFCC) or spectrogram. Feature extraction aims to compress the audio data volume and extract acoustic features useful for distinguishing speech content.
[0089] Acoustic model calculation: The acoustic model predicts the probabilities of corresponding speech units (such as phonemes, syllables, or characters) based on the extracted features. In traditional systems, the acoustic model usually uses Hidden Markov Model (HMM) combined with an observation probability model to model speech sequences; modern systems mostly use deep neural networks to directly output the probability distribution of each speech unit at each moment.
[0090] Function of language model: Language model provides a priori probability scores for candidate transcription results based on the statistical laws of language, in order to prefer word sequences that are more in line with language habits. In the early days, the n-gram model based on frequency statistics was commonly used; nowadays, neural network language models are increasingly used to capture long-distance dependencies and improve the ability to handle complex contexts.
[0091] Decoding and output: The decoder combines the acoustic model probability and the language model probability to find the most likely recognition result in the search space of all possible text sequences. The Viterbi algorithm or beam search algorithm is usually used to efficiently complete this step and output the final transcription text. The pronunciation dictionary is also used in the decoding process to map the output units of the acoustic model (such as phonemes) to specific words.
[0092] Post-processing: Correct spelling, add punctuation, and restore capitalization on the decoded text to make the output text easy to read and apply. For example, a separate model can be trained to add punctuation and correct capitalization to the transcription results to obtain a complete and readable sentence.
[0093] S4. Obtain an answer scoring model corresponding to the interview question text, and input the text features into the answer scoring model to obtain interview score data.
[0094] In this embodiment, the step of obtaining the answer scoring model corresponding to the interview question text and inputting the text features into the answer scoring model to obtain the interview score data includes:
[0095] Obtain the answer scoring model corresponding to each interview question text, and input the text features of the interview answer text corresponding to each interview question text into the corresponding answer scoring model to obtain the initial score corresponding to each interview question text;
[0096] According to the question weights and initial scores corresponding to the interview question texts, the weighted score of each question is calculated by weighted average;
[0097] The weighted scores of all questions are aggregated to obtain the interview score data of the interview user.
[0098] For example: Obtain interview question texts a, b, and c from a preset question bank according to the position type, as well as the answer scoring model A and weight value 1.2 corresponding to the interview question text a, the answer scoring model B and weight value 1 corresponding to the interview question text b, and the answer scoring model C and weight value 0.7 corresponding to the interview question text c. Input the text features of the interview answer text corresponding to the interview question text a into the answer scoring model A to obtain an initial score of 80, input the text features of the interview answer text corresponding to the interview question text b into the answer scoring model B to obtain an initial score of 92, and input the text features of the interview answer text corresponding to the interview question text c into the answer scoring model C to obtain an initial score of 75. Calculate the weighted average according to the weight value and the initial score, and round the calculation result of (80×1.2 + 92×1 + 75×0.7)÷3 to the nearest integer to obtain 80 as the interview score data.
[0099] S5: Perform interview analysis through an artificial intelligence model based on the text features and the interview score data to obtain a voice interview analysis report.
[0100] In this embodiment, the performing interview analysis through an artificial intelligence model based on the text features and the interview score data to obtain a voice interview analysis report includes:
[0101] Input the text features into the artificial intelligence model for personality type analysis and ability level analysis, and generate corresponding intelligent analysis conclusions according to the interview score data;
[0102] Generate second-round interview suggestions through the artificial intelligence model according to the analysis conclusions, and generate a voice interview analysis report based on the personality type analysis, ability level analysis, intelligent analysis conclusions, and second-round interview suggestions.
[0103] It can be seen that in the above solution, through the artificial intelligence model, multiple interview question texts are obtained from the preset question bank according to the voice interview instruction sent by the interview user and a question audio is generated, the answer audio of the interview user is obtained based on the question audio, the answer audio is scored through the answer scoring model corresponding to the interview question text to obtain the interview score data, and the text features extracted from the answer audio and the interview score data are analyzed through the artificial intelligence model to obtain a voice interview analysis report, realizing fast and efficient voice interviews based on artificial intelligence and standardized interview result analysis.
[0104] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0105] In one embodiment, a voice interview device based on artificial intelligence is provided. The voice interview device based on artificial intelligence corresponds one-to-one with the voice interview method based on artificial intelligence in the above embodiment. As Figure 3 shown, the voice interview device includes an instruction acquisition module 101, an audio generation module 102, a text acquisition module 103, a scoring module 104, and an analysis module 105. The detailed description of each functional module is as follows:
[0106] The instruction acquisition module 101 is configured to acquire a voice interview instruction sent by an interview user, where the voice interview instruction includes a position type and a voice type;
[0107] The audio generation module 102 is configured to obtain multiple interview question texts from a preset question bank according to the position type, and generate a question audio according to the voice type and the interview question texts through an artificial intelligence model;
[0108] The text acquisition module 103 is configured to acquire an answer audio of the interview user based on the question audio, convert the answer audio into an interview answer text, and extract text features after preprocessing the interview answer text;
[0109] The scoring module 104 is configured to obtain an answer scoring model corresponding to the interview question text, and input the text features into the answer scoring model to obtain interview score data;
[0110] The analysis module 105 is configured to perform an interview analysis through an artificial intelligence model according to the text features and the interview score data to obtain a voice interview analysis report.
[0111] In one embodiment, the acquisition of the voice interview instruction sent by the interview user in the instruction acquisition module 101 is specifically used for:
[0112] Acquire the user information of the interview user, and provide corresponding position type options in the interface of the voice interview system according to the user information;
[0113] Receive the position type selected by the interview user, and provide a plurality of preset voice type options in the interface of the voice interview system;
[0114] Receive the voice type selected by the interview user to obtain a voice interview instruction composed of the position type and the voice type.
[0115] In one embodiment, the voice interview device further includes a question bank generation module, which is specifically used before the audio generation module 102 obtains multiple interview question texts from a preset question bank:
[0116] Analyze multiple pre-set job types to obtain the job requirement keywords corresponding to each job type;
[0117] Use an artificial intelligence language model to generate interview question texts, the standard answers to the interview question texts, and question weights according to the job requirement keywords;
[0118] Generate multiple derivative answers through the artificial intelligence language model according to the interview question texts, and score each derivative answer based on the standard answers;
[0119] Train a pre-built neural network model based on the standard answers, the derivative answers, and the corresponding scores to obtain an answer scoring model corresponding to the interview question texts;
[0120] Associate and store the interview question texts with the corresponding question weights, the answer scoring model, and the job types in a preset question bank.
[0121] In one embodiment, in the question bank generation module, training the pre-built neural network model based on the standard answers, the derivative answers, and the corresponding scores to obtain an answer scoring model corresponding to the interview question texts is specifically used for:
[0122] Convert the standard answers and the derivative answers into corresponding answer text vectors;
[0123] Input the answer text vectors into the pre-built neural network model to obtain the predicted scores of the answer text vectors;
[0124] Calculate the loss value between the predicted scores and the actual scores corresponding to the answer text vectors, and optimize the neural network model according to the loss value to obtain the answer scoring model.
[0125] In one embodiment, in the text acquisition module 103, obtaining the answer audio of the interview user based on the question audio is specifically used for:
[0126] Provide the question audio through a voice interview system, and provide a replay button and an answer button on the system interface;
[0127] After receiving the replay instruction sent by the interview user through the replay button, replay the question audio;
[0128] After receiving the answer instruction sent by the interview user through the answer button, obtain the audio data of the interview user within a preset time period as the answer audio of the interview user.
[0129] In one embodiment, the text acquisition module 103 converts the answer audio into an interview answer text, and extracts text features after preprocessing the interview answer text, specifically for:
[0130] Input the answer audio into a preset speech-to-text model to obtain text information corresponding to the answer audio;
[0131] Using an artificial intelligence language model to perform grammatical and semantic correction on the text information to obtain the interview answer text;
[0132] Performing word segmentation processing on the interview answer text, and identifying and removing stop words from the interview answer text after word segmentation to obtain a preprocessed text;
[0133] The preprocessed text is converted into a text vector to obtain text features of the interview answer text.
[0134] In one embodiment, the scoring module 103 obtains an answer scoring model corresponding to the interview question text, and inputs the text features into the answer scoring model to obtain interview score data, which is specifically used for:
[0135] Obtain the answer scoring model corresponding to each interview question text, and input the text features of the interview answer text corresponding to each interview question text into the corresponding answer scoring model to obtain the initial score corresponding to each interview question text;
[0136] According to the question weights and initial scores corresponding to the interview question texts, the weighted score of each question is calculated by weighted average;
[0137] The weighted scores of all questions are aggregated to obtain the interview score data of the interview user.
[0138] The present invention provides a voice interview device based on artificial intelligence. The artificial intelligence model obtains multiple interview question texts from a preset question bank and generates question audio according to the voice interview instruction sent by the interview user, obtains the interview user's answer audio based on the question audio, scores the answer audio through the answer scoring model corresponding to the interview question text to obtain interview score data, and analyzes the text features extracted from the answer audio and the interview score data through the artificial intelligence model to obtain a voice interview analysis report, thereby realizing fast and efficient voice interview and standardized interview result analysis based on artificial intelligence.
[0139] For the specific limitations of the voice interview device, reference can be made to the limitations on the intelligent question and answer method described above, which will not be elaborated here. Each module in the above voice interview device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0140] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The storage media include flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, disks, optical discs, etc. The non-volatile storage media stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The memory can be the internal storage unit of the computer device in some embodiments, such as the mobile hard disk of the computer device. The memory can also be an external storage device of the computer device in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Further, the storage can also include both the internal storage unit and the external storage device of the computer device. The memory can be used not only to store application software installed on the computer device and various types of data, such as the code of the voice interview program, etc., but also to temporarily store data that has been output or will be output. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a voice interview method based on artificial intelligence.
[0141] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Optionally, in some embodiments, the display screen may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display screen can also be appropriately referred to as a display or a display unit, which is used to display the information processed in the computer device and to display a visual user interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The storage medium includes flash memory, a mobile hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The memory may be an internal storage unit of the computer device in some embodiments, such as the mobile hard disk of the computer device. The memory may also be an external storage device of the computer device in some other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Further, the storage may also include both an internal storage unit and an external storage device of the computer device. The memory can be used not only to store application software installed on the computer device and various types of data, such as the code of a voice interview program, etc., but also to temporarily store data that has been output or will be output. The network interface of the computer device is used to communicate with an external server via a network connection. The computer program, when executed by the processor, implements the functions or steps on the client side of a voice interview method based on artificial intelligence.
[0142] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0143] Obtain a voice interview instruction sent by an interview user, where the voice interview instruction includes a position type and a voice type;
[0144] Obtain multiple interview question texts from a preset question bank according to the position type, and generate a question audio according to the voice type and the interview question texts through an artificial intelligence model;
[0145] Obtain the response audio of the interviewee based on the question audio, convert the response audio into an interview response text, and extract text features after preprocessing the interview response text;
[0146] Obtain the answer scoring model corresponding to the interview question text, and input the text features into the answer scoring model to obtain interview score data;
[0147] Conduct interview analysis through an artificial intelligence model based on the text features and the interview score data to obtain a voice interview analysis report.
[0148] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Obtain a voice interview instruction sent by an interviewee, where the voice interview instruction includes a position type and a voice type;
[0149] Obtain multiple interview question texts from a preset question bank according to the position type, and generate question audio according to the voice type and the interview question texts through an artificial intelligence model;
[0150] Obtain the response audio of the interviewee based on the question audio, convert the response audio into an interview response text, and extract text features after preprocessing the interview response text;
[0151] Obtain the answer scoring model corresponding to the interview question text, and input the text features into the answer scoring model to obtain interview score data;
[0152] Conduct interview analysis through an artificial intelligence model based on the text features and the interview score data to obtain a voice interview analysis report.
[0153] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0154] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0155] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0156] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use.
[0157] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0158] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, device, article or method comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, device, article or method comprising that element.
[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0160] The above-described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention and should all be included within the protection scope of the present invention.
Claims
1. A voice interview method based on artificial intelligence, characterized in that Including: Obtain a voice interview instruction sent by an interview user, where the voice interview instruction includes a position type and a voice type; Obtain multiple interview question texts from a preset question bank according to the position type, and generate a question audio according to the voice type and the interview question texts through an artificial intelligence model; Obtain the answer audio of the interview user based on the question audio, convert the answer audio into an interview answer text, and extract text features after preprocessing the interview answer text; Obtain an answer scoring model corresponding to the interview question text, and input the text features into the answer scoring model to obtain interview score data; Conduct an interview analysis through an artificial intelligence model according to the text features and the interview score data to obtain a voice interview analysis report.
2. The voice interview method based on artificial intelligence according to claim 1, wherein The obtaining of the voice interview instruction sent by the interview user includes: Obtain the user information of the interview user, and provide corresponding position type options in the interface of the voice interview system according to the user information; Receive the position type selected by the interview user, and provide a plurality of preset voice type options in the interface of the voice interview system; Receive the voice type selected by the interview user to obtain a voice interview instruction composed of the position type and the voice type.
3. The voice interview method based on artificial intelligence according to claim 1, characterized in that, Before obtaining multiple interview question texts from a preset question bank according to the position type, it includes: Analyze a plurality of preset position types to obtain position requirement keywords corresponding to each position type; Use an artificial intelligence language model to generate interview question texts, standard answers to the interview question texts, and question weights according to the position requirement keywords; Generate multiple derivative answers according to the interview question texts through an artificial intelligence language model, and score each derivative answer based on the standard answer; Train a pre-constructed neural network model based on the standard answer, the derivative answers, and the corresponding scores to obtain an answer scoring model corresponding to the interview question text; Associate and store the interview question text with the corresponding question weight, the answer scoring model, and the position type in a preset question bank.
4. The voice interview method based on artificial intelligence according to claim 3, wherein The training of the pre-constructed neural network model based on the standard answer, the derivative answers, and the corresponding scores to obtain an answer scoring model corresponding to the interview question text includes: Convert the standard answer and the derivative answers into corresponding answer text vectors; Input the answer text vectors into a pre-constructed neural network model to obtain the predicted scores of the answer text vectors; Calculate the loss value between the predicted score and the actual score corresponding to the answer text vector, and optimize the neural network model according to the loss value to obtain the answer scoring model.
5. The voice interview method based on artificial intelligence according to claim 1, characterized in that, The obtaining of the answer audio of the interview user based on the question audio includes: Provide the question audio through the voice interview system, and provide a replay button and an answer button in the system interface; After receiving the replay instruction sent by the interview user through the replay button, replay the question audio; After receiving the response instruction sent by the interviewed user through the response button, obtain the audio data of the interviewed user within a preset time period as the response audio of the interviewed user.
6. The speech interview method based on artificial intelligence according to claim 1, wherein Converting the response audio into an interview response text, and extracting text features after preprocessing the interview response text, including: Input the response audio into a preset speech-to-text model to obtain the text information corresponding to the response audio; Use an artificial intelligence language model to correct the grammar and semantics of the text information to obtain the interview response text; Perform word segmentation on the interview response text, and identify and remove stop words from the word-segmented interview response text to obtain a preprocessed text; Convert the preprocessed text into a text vector to obtain the text features of the interview response text.
7. The voice interview method based on artificial intelligence according to claim 1, characterized in that Obtain the answer scoring model corresponding to the interview question text, and input the text features into the answer scoring model to obtain the interview score data, including: Obtain the answer scoring model corresponding to each interview question text, and input the text features of the interview response text corresponding to each interview question text into the corresponding answer scoring model to obtain the initial score corresponding to each interview question text; According to the question weights and initial scores corresponding to each interview question text, use weighted average to calculate the weighted score of each question; Summarize the weighted scores of all questions to obtain the interview score data of the interviewed user.
8. A voice interview device based on artificial intelligence, characterized in that, Including: An instruction acquisition module for acquiring a voice interview instruction sent by an interviewed user, where the voice interview instruction includes a position type and a voice type; An audio generation module for obtaining multiple interview question texts from a preset question bank according to the position type, and generating a question audio according to the voice type and the interview question texts through an artificial intelligence model; A text acquisition module for obtaining the response audio of the interviewed user based on the question audio, converting the response audio into an interview response text, and extracting text features after preprocessing the interview response text; A scoring module for obtaining the answer scoring model corresponding to the interview question text, and inputting the text features into the answer scoring model to obtain the interview score data; An analysis module for performing interview analysis through an artificial intelligence model according to the text features and the interview score data to obtain a voice interview analysis report.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the artificial intelligence-based voice interview method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the artificial intelligence-based voice interview method according to any one of claims 1 to 7.