An AI Interviewer-based Intention Recognition Method, Device, Electronic Device and Medium
By combining machine learning, basic meaning and multi-level intention recognition methods with knowledge graph engines, the problem of low accuracy in intention recognition of AI interviewers is solved, achieving higher matching and evaluation accuracy.
Patent Information
- Application Number
- CN202510457572.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The current AI interviewers have low accuracy and accuracy when identifying candidate intentions, resulting in low matching between interviewers and companies' recruitment portraits.
Using a combination method of machine learning engine, basic meaning engine and knowledge graph engine, through text similarity matching, keyword disassembly and association intention recognition, the matching score of intention is comprehensively calculated to determine the winning intention.
It significantly improves the accuracy and reliability of intention recognition, and provides AI interviewers with efficient and accurate assessment capabilities.
Smart Images

Figure CN119989062B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent interviews. Specifically, it relates to a method, device, electronic device and medium for intention recognition based on an AI interviewer. Background Art
[0002] In today's digital age, virtual assistants have been widely used in many fields such as AI (Artificial Intelligence) interviews, customer service, and intelligent office. Nowadays, enterprises rely more and more on virtual assistants and have increasingly strict requirements for their performance and accuracy. For example, in the scenario of an AI interview, the virtual interviewer needs to accurately grasp the enterprise's recruitment profile and accurately judge whether the candidate's answer matches the recruitment profile, which requires the virtual interviewer to be able to accurately recognize the intention of the candidate's answer.
[0003] However, currently, the AI interviewer can only recognize the candidate's intention from the surface meaning of the candidate's answer, resulting in low accuracy and precision of intention recognition, and further resulting in a low matching degree between the interviewee and the enterprise's recruitment profile. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, device, electronic device and medium for intention recognition based on an AI interviewer to improve the accuracy of intention recognition and further improve the matching degree between the interviewee and the enterprise.
[0005] In a first aspect, a method for intention recognition based on an AI interviewer is provided, which is applied to an intention recognition engine. The intention recognition engine includes a machine learning engine, a basic meaning engine and a knowledge graph engine. The method includes:
[0006] Obtain the interview speech text of the interview candidate;
[0007] Perform text similarity matching on the interview speech text and the example speech set of the preset intention labels through the machine learning engine, and identify at least one initial intention and determine the first matching score of each initial intention based on the similarity matching result; the preset intention labels are predefined based on the enterprise's recruitment requirements;
[0008] Decompose the interview speech text through the basic meaning engine to obtain keywords associated with each initial intention, and calculate the second matching score of each initial intention based on the lexical meaning of the keywords;
[0009] Identify the associated intention of the interview speech text through the knowledge graph engine; determine the third matching score of each initial intention based on the relevance between the associated intention and each initial intention;
[0010] Determine a comprehensive score based on the first matching score, the second matching score, and the third matching score for each initial intention, and determine the winning intention based on the comprehensive score.
[0011] Optionally, when the number of intentions identified based on the similarity matching result is multiple, the method further includes:
[0012] Filter the initial intentions based on a preset intention exclusion rule to obtain the filtered initial intentions.
[0013] Optionally, the preset intention exclusion rule includes at least a first exclusion rule and a second exclusion rule:
[0014] The first exclusion rule is to preferentially exclude the intentions based on entity value matching; the entity value includes at least date and number;
[0015] The second exclusion rule is to determine the intention matching type based on the first matching score of the initial intention, and the intention matching type includes possible matching and definite matching;
[0016] When two types are simultaneously included in multiple initial intentions, exclude the initial intentions with the intention matching type of possible matching.
[0017] Optionally, calculating the second matching score for each initial intention based on the lexical meaning of the keyword includes:
[0018] Determine the original word role score of the keyword, and the original word role represents the role or function of the word in its original form;
[0019] Determine the role score of the keyword in the sentence, and the role in the sentence represents the role of the word in the entire sentence structure;
[0020] Determine the word state score of the keyword after preprocessing; the preprocessing is synonym replacement, lemmatization, and normalization;
[0021] Calculate the second matching score for each initial intention based on the original word role score of the keyword, the role score in the sentence, the word state score after preprocessing, the weight value corresponding to each score, as well as the preset bonus items and penalty items.
[0022] Optionally, identify the associated intentions of the interview discourse text through a knowledge graph engine; determining the third matching score for each initial intention based on the relevance between the associated intentions and each initial intention includes:
[0023] Extract the term set of the interview discourse text;
[0024] Based on the term set, find the set of paths in the pre-constructed knowledge graph of the target industry domain that can associate these terms;
[0025] Determine the associated intentions based on the direction of the set of paths;
[0026] Determine the third matching score of each initial intention based on the similarity between the associated intention and each initial intention.
[0027] Optionally, determining a comprehensive score based on the first matching score, the second matching score, and the third matching score of each initial intention, and determining the winning intention based on the comprehensive score includes:
[0028] Determine the first confidence level of each initial intention based on a machine learning engine;
[0029] Determine the second confidence level of each initial intention based on a basic meaning engine;
[0030] Determine the third confidence level of each initial intention based on a knowledge graph engine;
[0031] Calculate a comprehensive score based on the first matching score, the second matching score, the third matching score, the first confidence level, the second confidence level, the third confidence level, and the preset weights of each engine;
[0032] Compare the comprehensive score with a preset score threshold, and if it exceeds the preset score threshold, determine it as the winning intention.
[0033] Optionally, after determining the winning intention, the method further includes:
[0034] Determine the matching type of the winning intention based on the comprehensive score of each winning intention, and the matching type includes possible match and definite match;
[0035] When the number of winning intentions is 1, if the matching type of the winning intention is a definite match, directly output the feedback information of the winning intention; if the matching type of the winning intention is a possible match, output the winning intention in the form of a preset dialog box for the user to confirm;
[0036] When the number of winning intentions is 2 or more, if the matching type of the winning intention is multiple definite matches, output all the winning intentions for the user to select; if the matching type of the winning intention is multiple possible matches, output the winning intentions in the form of a preset dialog box for the user to confirm; if the matching type of the winning intention includes possible matches and definite matches, output the feedback information of the winning intention corresponding to the definite match.
[0037] In a second aspect, there is provided an intention recognition system based on an AI interviewer, the system including:
[0038] An acquisition unit for acquiring the interview speech text of an interview candidate;
[0039] A machine learning engine for performing text similarity matching between an interview speech text and a set of example speeches with preset intent tags, and identifying at least one initial intent and determining a first matching score for each initial intent based on the similarity matching result; the preset intent tags are predefined based on enterprise recruitment requirements;
[0040] A basic meaning engine for disassembling the interview speech text to obtain keywords associated with each initial intent, and calculating a second matching score for each initial intent based on the lexical meaning of the keywords;
[0041] A knowledge graph engine for identifying associated intents of the interview speech text; determining a third matching score for each initial intent based on the relevance between the associated intents and each initial intent;
[0042] A determination unit for determining a comprehensive score based on the first matching score, the second matching score, and the third matching score of each initial intent, and determining a winning intent based on the comprehensive score.
[0043] In a third aspect, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0044] The memory is used for storing a computer program;
[0045] The processor, when executing the program stored on the memory, implements the method steps described in any one of the first aspect.
[0046] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of the first aspect are implemented.
[0047] An intent recognition method, device, electronic device, and medium based on an AI interviewer provided by an embodiment of the present invention obtain the interview speech text of an interview candidate; perform text similarity matching on the interview speech text and a set of example speeches of preset intent tags through a machine learning engine, and identify at least one initial intent and determine the first matching score of each initial intent based on the similarity matching result; the preset intent tags are predefined based on enterprise recruitment requirements; disassemble the interview speech text through a basic meaning engine to obtain keywords associated with each initial intent, and calculate the second matching score of each initial intent based on the lexical meaning of the keywords; identify the associated intent of the interview speech text through a knowledge graph engine; determine the third matching score of each initial intent based on the relevance between the associated intent and each initial intent; determine the comprehensive score based on the first matching score, second matching score, and third matching score of each initial intent, and determine the winning intent based on the comprehensive score. Through the close cooperation of three engines in the AI interviewer scenario, the present invention is associated with each other and progresses layer by layer, processes the user's answer from different levels and different angles, gradually deeply mines the user's intent, significantly improves the accuracy and reliability of intent recognition, and lays a solid foundation for the efficient and accurate evaluation of the AI interviewer.
[0048] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.
[0050] Figure 1 Shows a flowchart of an intent recognition method based on an AI interviewer provided by an embodiment of the present invention;
[0051] Figure 2 Shows a structural schematic diagram of an intent recognition device based on an AI interviewer provided by an embodiment of the present invention;
[0052] Figure 3 Shows a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. The components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents the selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0054] Considering that the current AI interviewers can only identify the intentions of candidates from the surface meaning of their answers, resulting in low accuracy and precision of intention recognition, and further leading to a low matching degree between the interviewees and the enterprise's recruitment portraits.
[0055] Based on this, the embodiments of the present invention provide an intention recognition method and system based on an AI interviewer, which will be described below through embodiments.
[0056] The embodiments of the present invention provide an intention recognition method based on an AI interviewer. This method is applied to an intention recognition engine, which consists of three parts: a machine learning engine, a basic meaning engine, and a knowledge graph engine. After pre-testing and optimizing these three engines to meet the required intention recognition accuracy, they are then deployed and applied.
[0057] The following explains the test and optimization process of this intention recognition engine:
[0058] First, log in to the pre-built test and tuning platform. On this test and tuning platform, select the intention recognition engine to be tested. In the left menu, click the Test - Discourse Test option in sequence. And select the corresponding engine for the test discourse as needed.
[0059] Taking the machine learning engine as an example, first select the machine learning engine in the drop-down menu of the engine selection menu, and type the user discourse to be tested in the input user discourse box, such as "Are you familiar with the Java language?". This machine learning engine will output a test result. This test result will be presented in several situations: single, multiple, or no matching intention.
[0060] When the test result is a single matching intention, display the single intention that matches the user discourse below the input user discourse field. If the tester determines that the match is correct, the test can continue to improve the intention matching score. If there is an error, mark and select the correct intention and then continue the test.
[0061] When the test results show multiple matching intents, the tester can click on the radio button of one of the multiple matching intents to perform training.
[0062] When the test results show no matching intent, at this time, an intent needs to be selected and trained to fit the user's utterance.
[0063] Through a large number of test utterances, continuously test and optimize the machine learning model to continuously improve the matching score of the output intent until the preset matching score is reached.
[0064] Based on the same method, the basic meaning engine and the knowledge graph engine can be tested.
[0065] Among them, when training the knowledge graph engine, for the AI interviewer, common interview questions and their answers can be used as part of the knowledge graph. For example, for the common question "Please introduce your strengths and weaknesses", by setting terms (such as "strengths", "weaknesses"), term configurations or categories from the common question answer page, train the knowledge graph. When the interviewee answers a similar question, the intent can be better matched and a reasonable evaluation can be made. Or add the interviewee's answer as an alternative question to the common questions selected on the knowledge graph page to further train the knowledge graph and improve the model's intent recognition ability for various answers in the interview scenario.
[0066] In addition, when testing the user's utterance, an intent analysis box is also set on the test tuning platform. In this analysis box, the shortlisted intents, the engines called, the corresponding intent matching scores, and the final winning intent can be quickly overviewed.
[0067] Next, the intent recognition is performed through the intent recognition engine completed by test optimization. An embodiment of the present invention provides an intent recognition method based on an AI interviewer, as Figure 1 shown, the method includes:
[0068] Step S101: Obtain the interview utterance text of the interview candidate.
[0069] In this step, in the AI interview scenario, in a virtual interview room, the candidate connects through video conferencing, and the virtual image of the AI interviewer is displayed on the screen. The interface is simple and intuitive, with a timer, a question display area, an answer input box, etc. The interview starts, and the AI interviewer starts: "Welcome to participate in the interview for the software R & D engineer position. Next, we will start the professional skills assessment session. Please get ready." The interview candidate starts answering the questions according to the questions of the AI interviewer.
[0070] In an example, the voice or directly input text information of the interview candidate can be obtained. If it is voice, the voice can be further converted into text information to obtain the interview utterance text.
[0071] To improve the quality of the interview speech text, before intent recognition, the interview speech text can be preprocessed, such as removing irrelevant information: First, all information irrelevant to the interview content needs to be removed, such as background noise, non-verbal communication (such as coughing, laughing), etc. If it is in text form, any unnecessary annotations or tags need to be deleted. Another example is to standardize the text format: unify the text format, including font, case specification, etc. For example, convert all text to lowercase to avoid duplicate counting of words due to different cases.
[0072] Step S102: Use a machine learning engine to perform text similarity matching between the interview speech text and a set of example utterances with preset intent tags, and identify at least one initial intent and determine the first matching score for each initial intent based on the similarity matching result.
[0073] In this step, the preset intent tags are predefined based on the enterprise recruitment requirements. For example, it includes intent tags such as "introduce work experience" and "elaborate on project details".
[0074] In the embodiment of the present invention, first, the TF-IDF algorithm is used to vectorize the interview speech text and the set of example utterances, and then cosine similarity is used for similarity matching, and the first matching score for each initial intent is determined based on the cosine similarity score.
[0075] In a specific example, assume that the user input utterance is , and the set of preset intent tags is ; the set of example utterances is . First, all the example utterances in the set of example utterances and the user input utterance are vectorized. Let be the TF-IDF vector of the user input utterance , and be the TF-IDF vector of the th example utterance in the preset intent tag .
[0076] Calculate the cosine similarity scores between the user input utterance and all the example utterances under one of the preset intent tags, and take the maximum similarity score as the matching score for the preset intent recognition tag :
[0077] (1);
[0078] Among them, is the formula for calculating the cosine similarity between two vectors, It is the code name of the machine learning engine.
[0079] After calculating the scores of each preset intention label, by comparing with the preset first matching score threshold, the intention corresponding to the preset intention label greater than the first matching score threshold is determined as the initial intention.
[0080] In a specific example, when the AI interviewer asks "Please describe a product you are most familiar with and analyze its core value and target user group", and the interviewee answers with Douyin as an example, the machine learning engine uses the TF-IDF algorithm to vectorize the user's answer and calculate the similarity with the example sentences corresponding to a large number of predefined intention labels. It quickly scans through the massive example sentence data, preliminarily filters out the intention directions that may match the user's answer, obtains the preliminary recognition intentions, such as intentions like "product description", "core value analysis", "target user positioning", etc., and gives the corresponding matching scores.
[0081] Step S103: The interview speech text is disassembled by the basic meaning engine to obtain keywords associated with each initial intention, and the second matching score of each initial intention is calculated based on the lexical meaning of the keywords.
[0082] The machine learning engine can only judge the surface similarity of the text and it is difficult to deeply explore the complex relationships and meanings behind the semantics. Therefore, in this step, the basic meaning engine further analyzes the semantics of the user's speech to more accurately judge the user's intention and improve the accuracy of intention recognition.
[0083] For example, when the machine learning engine preliminarily judges that the user's answer is related to the "core value analysis" intention, the basic meaning engine further analyzes the close connection between the original words such as "convenient, rich, creative" mentioned by the user and the description of the core value of "short video creation and sharing platform" to confirm the accuracy of the core value analysis intention.
[0084] The specific implementation process will be described in the following embodiments and will not be elaborated here.
[0085] Step S104: The associated intentions of the interview speech text are identified by the knowledge graph engine; the third matching score of each initial intention is determined based on the relevance between the associated intentions and each initial intention.
[0086] In this step, a knowledge graph related to interviews is constructed, including company information, job requirements, industry knowledge, etc. When the interviewee mentions "I am familiar with big data technology and used Hadoop to process data in previous projects", terms (such as "big data technology", "Hadoop") can be extracted from the answer and mapped to the nodes in the knowledge graph (such as the company's requirements for big data skills, Hadoop-related knowledge modules). Find the knowledge graph path that matches the interviewee's answer to obtain the associated intention, and through the associated intention Figure 1 On the one hand, it can verify the accuracy of the initial intention recognized by the machine learning engine, and on the other hand, it can make the initial intention more explicit.
[0087] Step S105: Determine the comprehensive score based on the first matching score, the second matching score, and the third matching score of each initial intention, and determine the winning intention based on the comprehensive score.
[0088] In this step, by balancing the scores of the three dimensions, it is possible to avoid a situation where the score of a certain dimension is too high or the score of a dimension is too low, resulting in overfitting of the intention recognition result. For example, for an interviewee's answer to a question about technical capabilities, the machine learning engine may give a relatively high matching score based on technical keyword matching, the basic meaning engine may give a relatively low score based on sentence structure and lexical accuracy, and the knowledge graph may give a scoring rate between the two based on the association between technology and job requirements. By determining the final intention matching degree through the comprehensive score, the intention matching can be made more accurate, providing an accurate judgment for the AI interviewer and better assisting the interview evaluation.
[0089] The present invention identifies the user's intention through three dimensions: the machine learning engine, the basic meaning engine, and the knowledge graph engine. They are related to each other and progress step by step, jointly contributing to the accurate understanding and intention judgment of the user's answer. Through this progressive and collaborative method, the three engines cooperate closely in the scenario of the AI interviewer, process the user's answer from different levels and angles, gradually dig deeper into the user's intention, significantly improving the accuracy and reliability of intention recognition, and laying a solid foundation for the efficient and accurate evaluation of the AI interviewer.
[0090] Based on the above embodiments, when the number of intentions recognized based on the similarity matching result is multiple, the method further includes:
[0091] Step S106: Screen the initial intentions based on a preset intention exclusion rule to obtain the screened initial intentions.
[0092] Among them, the preset intention exclusion rule at least includes a first exclusion rule and a second exclusion rule:
[0093] The first exclusion rule is to preferentially exclude the intention based on entity value matching; the entity value at least includes a date and a number.
[0094] Since information such as dates and numbers is a relatively clear intention without much ambiguity, it can be directly excluded without further intention recognition, saving computing resources. Only those unclear intentions need to be further recognized.
[0095] The second exclusion rule is to determine the intention matching type based on the first matching score of the initial intention. The intention matching types include possible matching and certain matching.
[0096] In this step, the first matching score is compared with a preset threshold range. For example, the preset threshold range is 60 - 90 points. If it is between 60 - 90, it is a possible matching; if it is above 90, it is a certain matching. Of course, according to specific computing requirements, this threshold range can also be mapped to 0 - 1.
[0097] Possible matching means that the score is relatively high but does not meet the confidence requirement of exact matching. Certain matching means that it has a high confidence and is judged to be precisely matched with the user's utterance.
[0098] When both types are included in multiple initial intentions, exclude the initial intentions with the intention matching type of possible matching.
[0099] If there are both possible matching and certain matching at the same time, then the possible matching can be directly excluded, and the intention of certain matching can be retained. On the one hand, it reduces the processing volume of the subsequent engine, saves computing resources, and improves computing efficiency. On the other hand, it helps to quickly screen out the intentions that are closer to the user's true thoughts.
[0100] Of course, other exclusion rules can also be set according to requirements. For example, if the user's utterance contains two intentions (let intention i and j be the two intentions in the user's utterance), and a certain matching has been found previously (let intention i be the previously found certain matching intention), then exclude the subsequent certain matching intention j.
[0101] Based on the above embodiments, calculating the second matching score of each initial intention based on the lexical meaning of keywords includes the following steps:
[0102] Step S103A: Determine the original word role score of the keyword. The original word role represents the role or function of the word in its original form.
[0103] In one example, for instance, the original word role of a certain keyword is a noun, a verb, an adjective, etc.
[0104] In this step, keywords associated with each intention are first extracted. For example, if the response content of an interview candidate includes multiple intentions such as "product description", "core value analysis", "target user positioning", etc., then for the intention associated with "product description", keywords such as "beautiful, practical" are included. These keywords are analyzed by the basic meaning engine to determine their original word roles, roles in the sentence, etc., to judge their relevance to the intention of "product description". The greater the relevance, the higher the second matching score, and the smaller the relevance, the lower the second matching score.
[0105] Step S103B: Determine the role score of the keyword in the sentence. The role in the sentence represents the role of the word in the entire sentence structure.
[0106] In this step, the roles in the sentence, such as subject, object, and predicate, can reflect their positions and functions in the entire sentence structure.
[0107] Step S103C: Determine the word status score of the keyword after preprocessing; the preprocessing is synonym replacement, lemmatization, and normalization.
[0108] Step S103D: Calculate the second matching score of each initial intention based on the original word role score of the keyword, the role score in the sentence, the word status score after preprocessing, the weight values corresponding to each score, as well as the preset bonus items and penalty items.
[0109] Among them, the bonus item is the score calculated based on the sentence structure, word position, sequential reward, role reward, extended reward, etc., and the penalty item is the score calculated based on the number of conjunctions in the user's speech, etc.
[0110] In a feasible implementation, the ranges of the original word role score, the role score in the sentence, and the word status score after preprocessing can be normalized and limited to [0, 1].
[0111] In an example, for example, the score of the intention The calculation formula is:
[0112] (2);
[0113] Among them, is the original word role score of the keyword; is the role score of the keyword in the sentence; is the word status score of the keyword after preprocessing; is the reward item score; is the penalty item score; is the weight of the original word role; is the role score in the sentence; is the weight of the word state after preprocessing; is the code name of the basic meaning engine.
[0114] Based on the above embodiments, the associated intent of the interview discourse text is identified through the knowledge graph engine; determining the third matching score of each initial intent based on the relevance between the associated intent and each initial intent includes:
[0115] Step S104A: Extract the term set of the interview discourse text.
[0116] In this step, a pre-trained entity recognition model (such as the BERT model) can be used to extract the term set of the interview discourse text. For example, terms such as "Douyin", "short video creation", "algorithm recommendation", "young people", and "middle-aged and elderly users" are extracted from the user's answer.
[0117] Step S104B: Based on the term set, find the set of paths in the pre-constructed knowledge graph of the target industry field that can be associated with these terms.
[0118] In this step, the knowledge graph engine searches for associated paths in its own network according to these term sets, and then takes the associated paths with the number of terms exceeding the preset threshold as the set of paths.
[0119] In a specific example, assume the user input discourse is , from the extracted term set is ;
[0120] The set of associated paths of the term in the knowledge graph is ;
[0121] Based on the preset threshold , filter out the set of paths with the number of terms exceeding :
[0122] (3);
[0123] Among them, is one of the associated paths in the set of paths ; is the preset term number threshold; is one of the terms in the term set.
[0124] Step S104C: Determine the associated intent based on the direction of the set of paths.
[0125] For example, there is a path in the knowledge graph that connects "Douyin" with key information such as "The core value of a short video creation platform lies in satisfying users' desire to express and creativity" and "Attracts users of all ages, mainly young people", and infers the associated intention of the interviewee's answer from this path.
[0126] Step S104D: Determine the third matching score of each initial intention based on the similarity between the associated intention and each initial intention.
[0127] In a feasible implementation, the cosine similarity can be used to calculate the similarity between the associated intention and the initial intention, and the third matching score is determined based on the similarity score.
[0128] Based on the above embodiments, determining the comprehensive score based on the first matching score, the second matching score, and the third matching score of each initial intention, and determining the winning intention based on the comprehensive score includes:
[0129] Determine the first confidence level of each initial intention based on the machine learning engine.
[0130] In this step, the confidence level can be determined by calculating the standard deviation or variance of the first matching scores of each initial intention. Taking the standard deviation as an example, the lower the standard deviation, the higher the consistency of the machine learning engine in identifying a specific intention, and the higher the confidence level can be considered. To convert it into a more intuitive confidence level, one of the following conversion methods can be used: Confidence level = - σ ; If you want to obtain a confidence level value between 0 and 1, you can also use a normalization function to limit it to between 0 and 1.
[0131] Determine the second confidence level of each initial intention based on the basic meaning engine.
[0132] In this step, the second confidence level can also be determined based on the standard deviation or variance of the second matching score, which will not be elaborated here.
[0133] Determine the third confidence level of each initial intention based on the knowledge graph engine.
[0134] In this step, the second confidence level can also be determined based on the standard deviation or variance of the third matching score, which will not be elaborated here.
[0135] Calculate the comprehensive score based on the first matching score, the second matching score, the third matching score, the first confidence level, the second confidence level, the third confidence level, and the preset weight of each engine.
[0136] In a specific example, the comprehensive score is calculated as follows:
[0137] (4);
[0138] Among them, is the weight of each engine, and ; is the first matching score of the intent ; is the second matching score of the intent ; is the third matching score of the intent ; is the first confidence level of the intent ; is the second confidence level of the intent ; is the third confidence level of the intent ;
[0139] Compare the comprehensive score with a preset score threshold. If it exceeds the preset score threshold, it is determined as the winning intent.
[0140] In this step, the number of winning intents may be 1 or multiple.
[0141] Based on the above embodiments, after determining the winning intent, the method further includes:
[0142] Step S107: Determine the matching type of the winning intent based on the comprehensive score of each winning intent. The matching types include possible match and definite match.
[0143] In this step, compare the comprehensive score with a preset threshold range. The preset threshold range is, for example, 60 - 90, or the range is limited between 0 - 1. The preset threshold range is 0.6 - 0.9. If the comprehensive score is between 0.6 - 0.9, it is a possible match. If it exceeds 0.9, it is a definite match.
[0144] Step S108: When the number of winning intents is 1, if the matching type of the winning intent is a definite match, directly output the feedback information of the winning intent; if the matching type of the winning intent is a possible match, output the winning intent in the form of a preset dialog box for the user to confirm.
[0145] In one example, the preset dialog box is, for example, "Do you mean XX?", so as to let the interview candidate confirm, in order to determine the user's intent and give corresponding feedback according to the determined intent.
[0146] Step S109: When the number of winning intentions is two or more, if the matching types of the winning intentions are multiple definite matches, then all the winning intentions are output for the user to select; if the matching types of the winning intentions are multiple possible matches, then the winning intentions are output in the form of a preset dialog box for the user to confirm; if the matching types of the winning intentions include possible matches and definite matches, then the feedback information of the winning intentions corresponding to the definite matches is output.
[0147] By further confirming the less certain winning intentions, the interviewer's intentions can be more precisely matched, and corresponding more accurate answers or feedback can be given according to the interviewer's intentions.
[0148] Based on the same inventive concept, an intention recognition system based on an AI interviewer is provided, as Figure 2 shown. This system includes:
[0149] An acquisition unit 201, configured to acquire the interview speech text of the interview candidate.
[0150] A machine learning engine 202, configured to perform text similarity matching on the interview speech text and an example speech set of preset intention tags, and identify at least one initial intention and determine a first matching score for each initial intention based on the similarity matching result; the preset intention tags are predefined based on the enterprise recruitment requirements.
[0151] A basic meaning engine 203, configured to disassemble the interview speech text to obtain keywords associated with each initial intention, and calculate a second matching score for each initial intention based on the lexical meaning of the keywords.
[0152] A knowledge graph engine 204, configured to identify the associated intentions of the interview speech text; determine a third matching score for each initial intention based on the relevance between the associated intentions and each initial intention.
[0153] A determination unit 205, configured to determine a comprehensive score based on the first matching score, the second matching score, and the third matching score of each initial intention, and determine the winning intention based on the comprehensive score.
[0154] Based on the same technical concept, an embodiment of the present invention further provides an electronic device, as Figure 3 shown, including a processor 301, a communication interface 302, a memory 303, and a communication bus 304. Among them, the processor 301, the communication interface 302, and the memory 303 complete communication with each other through the communication bus 304.
[0155] The memory 303 is used to store a computer program;
[0156] The processor 301 is configured to implement the steps of the AI interviewer-based intention recognition method when executing the program stored in the memory 303.
[0157] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0158] The communication interface is used for communication between the above electronic device and other devices.
[0159] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0160] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0161] The computer program product for implementing the AI interviewer-based intention recognition method provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, which will not be elaborated herein.
[0162] The device for intent recognition based on an AI interviewer provided by the embodiments of the present invention can be specific hardware on a device or software or firmware installed on the device, etc. For the device provided by the embodiments of the present invention, its implementation principle and the resulting technical effects are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the foregoing-described systems, devices, and units can all refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0163] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0164] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] In addition, the functional units in the embodiments provided by the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0166] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.
[0167] It should be noted that like reference numerals and letters denote like items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.
[0168] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments or can easily conceive of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All of them should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An intention recognition method based on AI interviewer, characterized in that: Applied to an intention recognition engine, the intention recognition engine includes a machine learning engine, a basic meaning engine and a knowledge graph engine, the method includes: Obtain the interview speech text of the interview candidate; The interview speech text is matched with a set of example speech with preset intent labels by a machine learning engine for text similarity, and at least one initial intent is identified based on the similarity matching result and a first matching score of each initial intent is determined; the preset intent labels are predefined based on the recruitment needs of the enterprise; The initial intent determination process is: Calculate the cosine similarity score between the user input utterance and all example utterances under one of the preset intent labels; The maximum similarity score is taken as the matching score of the preset intent recognition tag. After calculating the scores of each preset intent tag, the scores are compared with the preset first matching score threshold, and the intent corresponding to the preset intent tag greater than the first matching score threshold is determined as the initial intent. Decomposing the interview speech text by a basic meaning engine to obtain keywords associated with each of the initial intentions, and calculating a second matching score for each initial intention based on the lexical meaning of the keywords; The second matching score calculation process is: Determine the original role score of the keyword, the original role represents the role or function of the word in its original form; Determine the role score of the keyword in the sentence. The role in the sentence represents the role of the word in the entire sentence structure. Determine the word status score of the keyword after preprocessing; the preprocessing is synonym replacement, word form restoration and standardization; Calculate the second matching score of each initial intent based on the original word role score of the keyword, the role score in the sentence, the preprocessed word status score, the weight value corresponding to each score, and the preset bonus item and penalty item; Identify the associated intention of the interview speech text through the knowledge graph engine; determine the third matching score of each initial intention based on the relevance of the associated intention to each initial intention; The third matching score determination process is: Extracting a term set from the interview discourse text; Based on the term set, searching for a path set that can associate these terms in a pre-built knowledge graph of the target industry field; Determining the association intention based on the orientation of the path set; determining a third matching score for each initial intent based on a similarity between the associated intent and each initial intent; A comprehensive score is determined based on the first matching score, the second matching score, and the third matching score of each initial intent, and a winning intent is determined based on the comprehensive score.
2. The method according to claim 1, characterized in that When the number of intentions identified based on the similarity matching results is multiple, the method further includes: The initial intent is filtered based on the preset intent exclusion rule to obtain the filtered initial intent.
3. The method according to claim 2, characterized in that The preset intention exclusion rule includes at least a first exclusion rule and a second exclusion rule: The first exclusion rule is to prioritize the exclusion of intents based on entity value matching; the entity value includes at least a date and a number; The second exclusion rule is to determine the intent match type based on the first match score of the initial intent, and the intent match type includes possible match and confirmed match; When multiple initial intents include both types, exclude the intent matching type as the initial intent that may be matched.
4. The method according to claim 1, characterized in that The step of determining a comprehensive score based on the first matching score, the second matching score, and the third matching score of each initial intent, and determining a winning intent based on the comprehensive score includes: Determining a first confidence level for each initial intent based on a machine learning engine; determining a second confidence level for each initial intent based on the base meaning engine; Determine the third confidence level of each initial intent based on the knowledge graph engine; Calculating a comprehensive score based on the first matching score, the second matching score, the third matching score, the first confidence level, the second confidence level, the third confidence level, and a preset weight of each engine; The comprehensive score is compared with a preset score threshold, and if it exceeds the preset score threshold, it is determined as a winning intention.
5. The method according to claim 4, characterized in that After determining the winning intention, the method further includes: Determine a match type of the winning intention based on the comprehensive score of each winning intention, where the match type includes a possible match and a confirmed match; When the number of winning intentions is 1, if the matching type of the winning intention is a confirmed match, the feedback information of the winning intention is directly output; if the matching type of the winning intention is a possible match, the winning intention is output in the form of a preset dialog box for confirmation by the user; When the number of winning intentions is 2 or more, if the matching type of the winning intention is multiple confirmed matches, all the winning intentions are output for user selection; if the matching type of the winning intention is multiple possible matches, the winning intention is output in the form of a preset dialog box for user confirmation; if the matching type of the winning intention includes possible matches and confirmed matches, feedback information of the winning intention corresponding to the confirmed matches is output.
6. An intention recognition system based on AI interviewer, characterized in that: The system comprises: An acquisition unit, used for acquiring the interview speech text of the interview candidate; A machine learning engine is used to perform text similarity matching between the interview utterance text and a set of example utterances with preset intent labels, and to identify at least one initial intent based on the similarity matching result and determine a first matching score for each of the initial intents; the preset intent labels are predefined based on the recruitment needs of the enterprise; The initial intent determination process is: Calculate the cosine similarity score between the user input utterance and all example utterances under one of the preset intent labels; The maximum similarity score is taken as the matching score of the preset intent recognition tag. After calculating the scores of each preset intent tag, the scores are compared with the preset first matching score threshold, and the intent corresponding to the preset intent tag greater than the first matching score threshold is determined as the initial intent. A basic meaning engine, used for disassembling the interview speech text to obtain keywords associated with each of the initial intentions, and calculating a second matching score for each initial intention based on the lexical meaning of the keywords; The second matching score calculation process is: Determine the original role score of the keyword, the original role represents the role or function of the word in its original form; Determine the role score of the keyword in the sentence. The role in the sentence represents the role of the word in the entire sentence structure. Determine the word status score of the keyword after preprocessing; the preprocessing is synonym replacement, word form restoration and standardization; Calculate the second matching score of each initial intent based on the original word role score of the keyword, the role score in the sentence, the preprocessed word status score, the weight value corresponding to each score, and the preset bonus item and penalty item; A knowledge graph engine, configured to identify the associated intent of the interview utterance text; and determine a third matching score of each initial intent based on the association between the associated intent and each initial intent; The third matching score determination process is: Extracting a term set from the interview discourse text; Based on the term set, searching for a path set that can associate these terms in a pre-built knowledge graph of the target industry field; Determining the association intention based on the orientation of the path set; determining a third matching score for each initial intent based on a similarity between the associated intent and each initial intent; The determination unit is used to determine a comprehensive score based on the first matching score, the second matching score and the third matching score of each initial intention, and determine a winning intention based on the comprehensive score.
7. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 5 when executing a program stored in a memory.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Intention recognition method and device, computer processing equipment and storage medium
CN116126994A
Method and apparatus for ai interview recognition, computer device and storage medium
WO2021217866A1