Intelligent Question Pushing Method, Device, Equipment and Storage Medium

By obtaining and analyzing the video and audio data of the collected personnel, combining the judgment of emotional characteristics and answer content, the second question is intelligently determined and pushed to the auditor, which solves the problem of low audit quality in the existing technology and improves the accuracy and efficiency of audits.

CN114724072BActive Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210428378.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-22
Publication Date
2025-06-24
Estimated Expiration
2042-04-22

AI Technical Summary

Technical Problem

The audit quality of audited personnel in the prior art is low, mainly because there are differences in the rating of question selection and answer content of auditors, and the prior art only scores for voice data, and lacks intelligent support for question selection.

Method used

By obtaining the video data and audio data when the collected person answers the first question, the emotional feature extraction model is used to extract the characteristic parameters of the facial data, and the audio data is judged by the judgment model, the second question is determined based on the emotional state, and pushed to the auditor.

Benefits of technology

The audit quality and processing efficiency of audited personnel are improved. Through the comprehensive analysis of emotional characteristics and answer content, we ensure the intelligence and personalization of question selection, and improve the accuracy and efficiency of audits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724072B_ABST
    Figure CN114724072B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology, and discloses an intelligent question pushing method, device, equipment and storage medium. The method includes: acquiring video data and audio data of the person being collected when answering the first question; using an emotion feature extraction model to extract features from the facial data of the person being collected in the video data to obtain emotion feature parameters, and using a judgment model to judge the answer content of the audio data to obtain a judgment result; determining the emotional state of the person being collected according to the emotion feature parameters; determining the second question of the person being collected based on the judgment result and the emotional state; and pushing the second question to the auditing personnel. This application improves the auditing quality and processing efficiency of the person being audited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to an intelligent question-pushing method, device, equipment and storage medium. Background Art

[0002] Currently, when dealing with audited personnel, the auditors often select questions and score the answers. Due to the different experiences of the auditors, there are certain differences in the question selection and answer scoring for the audited personnel. In the prior art, a machine learning model is introduced to score the answers of the audited personnel to the questions, and the scoring is only for the voice data of the answers. For the questions given to the audited personnel, they are still selected by the auditors, resulting in a low audit quality. Therefore, how to solve the problem of low audit quality for audited personnel has become an urgent problem to be solved. Summary of the Invention

[0003] This application provides an intelligent question-pushing method, device, equipment and storage medium to solve the problem of low audit quality for audited personnel in the prior art.

[0004] To solve the above problems, this application provides an intelligent question-pushing method, including:

[0005] Obtain the video data and audio data of the person being collected when answering the first question;

[0006] Use an emotion feature extraction model to extract features from the facial data of the person being collected in the video data to obtain emotion feature parameters, and use a judgment model to judge the answer content of the audio data to obtain a judgment result;

[0007] Determine the emotional state of the person being collected according to the emotion feature parameters;

[0008] Based on the judgment result and the emotional state, determine the second question for the person being collected;

[0009] Push the second question to the auditor.

[0010] Further, before obtaining the video data and audio data of the person being collected when answering the first question, it further includes:

[0011] Obtain the basic information of the person being collected;

[0012] Process the basic data by using a recommendation model to obtain the first question of the person being collected, and the recommendation model is trained based on the Apriori model.

[0013] Further, the feature extraction of the facial data of the person being collected in the video data by the emotion feature extraction model includes:

[0014] Intercept the video of the position where the head portrait of the person being collected is located in the video data to obtain the facial data;

[0015] Extract the facial data of the person being collected through the emotion feature extraction model, and the emotion feature extraction model is trained based on an enhanced long-term recurrent convolutional network model.

[0016] Further, the judgment of the answer content of the audio data by the judgment model to obtain the judgment result includes:

[0017] Convert the audio data into corresponding text data through the translation sub-model under the judgment model;

[0018] Use the matching sub-model under the judgment model to match the text data with the preset answer of the first question to obtain the corresponding matching degree, and the matching sub-model is trained based on the Bimpm model;

[0019] Determine the judgment result according to the matching degree.

[0020] Further, the matching of the text data with the preset answer of the first question by using the matching sub-model under the judgment model includes:

[0021] Use the feature extraction model to extract features from the text data to obtain the corresponding keywords, and the feature extraction model is trained based on the LDA model;

[0022] Use the matching sub-model to match the keywords with the preset answer.

[0023] Further, the determination of the emotional state of the person being collected according to the emotion feature parameters includes:

[0024] Based on the emotion feature parameters, query the preset state comparison table to determine the emotional state of the person being collected.

[0025] Further, the determination of the second question of the person being collected based on the judgment result and emotional state includes:

[0026] Use the classification model to perform hierarchical judgment on the judgment result and emotional state to determine the response level of the person being collected when answering the first question, and the classification model is trained based on the decision tree model;

[0027] Determine the second question corresponding to the corresponding level according to the response level.

[0028] To solve the above problems, the present application also provides an intelligent question-pushing device, which includes:

[0029] An acquisition module, configured to acquire video data and audio data of the person being collected when answering the first question;

[0030] An audio-video processing module, configured to extract feature parameters of the facial data of the person being collected in the video data by using an emotion feature extraction model, and obtain a judgment result by judging the answer content of the audio data through a judgment model;

[0031] A state confirmation module, configured to determine the emotional state of the person being collected according to the emotion feature parameters;

[0032] A question determination module, configured to determine the second question of the person being collected based on the judgment result and the emotional state;

[0033] A push module, configured to push the second question to the auditing personnel.

[0034] To solve the above problems, the present application also provides a computer device, including:

[0035] At least one processor; and,

[0036] A memory communicatively connected to the at least one processor; wherein,

[0037] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the intelligent question-pushing method as described above.

[0038] To solve the above problems, the present application also provides a non-volatile computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the intelligent question-pushing method as described above is implemented.

[0039] According to an intelligent question-pushing method, device, equipment and storage medium provided by an embodiment of the present application, compared with the prior art, it has at least the following beneficial effects:

[0040] By acquiring the video data and audio data of the person being collected when answering the first question, using an emotion feature extraction model to extract features from the facial data of the person being collected in the video data to obtain emotion feature parameters, and using a judgment model to judge the answer content of the audio data to obtain a judgment result, where the judgment result is the score level of the person being collected when answering the first question; processing the audio and video data respectively to obtain effective parameters; then determining the emotional state of the person being collected according to the emotion feature parameters; determining the second question of the person being collected based on the judgment result and the emotional state; by scoring the audio data of the user's answer and determining the emotional state according to the video data, the second question of the person being collected is finally determined, thereby improving the audit quality and processing efficiency of the person being audited. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 The overall flowchart of the intelligent question pushing method provided by an embodiment of the present application;

[0043] Figure 2 For Figure 1 A schematic flowchart of a specific implementation manner of step S2 in

[0044] Figure 3 For Figure 1 A schematic flowchart of a specific implementation manner of step S4 in

[0045] Figure 4 The module schematic diagram of the intelligent question pushing device provided by an embodiment of the present application;

[0046] Figure 5 The structural schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0048] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and is not necessarily referring to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those of ordinary skill in the art will explicitly or implicitly understand that the embodiments described herein can be combined with other embodiments.

[0049] This application provides an intelligent question-pushing method, which is mainly used in a remote interview system and is commonly used in scenarios such as remote interviews and promotion reviews. Refer to Figure 1 as shown Figure 1 is a schematic flowchart of the intelligent question-pushing method provided by an embodiment of this application.

[0050] In this embodiment, the intelligent question-pushing method includes:

[0051] S1. Obtain the video data and audio data of the person being collected when answering the first question;

[0052] Specifically, obtain the complete video data and audio data of the person being collected when answering the first question, mainly through the front end using a camera and a microphone for collection and real-time storage in a database or memory; when processing, obtain the video data and audio data from the database or memory.

[0053] Further, before obtaining the video data and audio data of the person being collected when answering the first question, it further includes:

[0054] Obtain the basic information of the person being collected;

[0055] By using a recommendation model to process the basic data, obtain the first question of the person being collected, and the recommendation model is trained based on the Apriori model.

[0056] Specifically, the first question of the interviewee is determined by the basic information of the interviewee. First, the basic information of the interviewee is obtained from the database, or the interviewee uploads his / her basic data. Since each question has a corresponding label in advance, the basic data is associated with the recommendation model to obtain the most relevant question, i.e., the first question. In the remote interview scenario, the basic data includes the identity information, work experience information, award information, and specialty information of the interviewee.

[0057] The Apriori (association rule mining) model is based on the principle that if an item set is frequent, then all its subsets are also frequent. If an item set is infrequent, then all its supersets are also infrequent. Based on this, the Apriori model starts with a single-element item set and forms a larger set by combining item sets that meet the minimum support. In fact, the Apriori model selects frequent item sets and association rules by elimination.

[0058] By adopting the recommendation model, the basic information of the collected personnel is processed to obtain the first question with certain relevance, thereby improving work efficiency and the relevance of the questions.

[0059] S2, using the emotion feature extraction model to extract features from the facial data of the collected person in the video data to obtain emotion feature parameters, and using the judgment model to judge the answer content of the audio data to obtain a judgment result;

[0060] Specifically, the emotional feature extraction model is used to extract features from facial data, that is, some expressions of the person being collected are extracted to obtain emotional feature parameters; and the translation sub-model and matching sub-model under the judgment model are used to process the data to obtain the score of the person being collected's answer, and based on the score, the corresponding judgment result is obtained.

[0061] Further, such as Figure 2 As shown, the feature extraction of the facial data of the collected person in the video data using the emotion feature extraction model includes:

[0062] S21, obtaining the facial data by intercepting the video at the location of the head portrait of the collected person in the video data;

[0063] S22. Extracting the facial data of the collected person through an emotional feature extraction model, wherein the emotional feature extraction model is trained based on an enhanced long-term recursive convolutional network model.

[0064] Specifically, by locating the head of the person being captured in the video data, the video at the location of the located head is intercepted to obtain the facial data; then, an emotion feature extraction model is used to extract the facial data of the person being captured to obtain emotion feature parameters, and the emotion feature parameters include values corresponding to the entire facial contour, mouth movements, eyebrow movements, cheek movements, etc. For example, the emotion feature parameters corresponding to a certain video frame are that the facial contour is 1, the mouth movement is 2, the eyebrow movement is 2, and the cheek movement is 0, etc.

[0065] For the enhanced long short-term recurrent convolutional network model, first, each micro-expression frame is encoded into a feature vector through a CNN module, and then the feature vector is passed through a long short-term memory (LSTM) module. This framework contains two different network variants:

[0066] Channel stacking of spatially enriched input data; functional stacking of features for temporal enrichment.

[0067] By intercepting the corresponding facial data from the complete video data and inputting the facial data containing only the human face into the emotion feature extraction model for processing, the accuracy is improved, and by processing through the emotion feature extraction model, not only the processing efficiency is improved but also the recognition rate is improved.

[0068] Further, the judgment model judges the answer content of the audio data to obtain judgment results including:

[0069] The translation sub-model under the judgment model converts the audio data into corresponding text data;

[0070] The matching sub-model under the judgment model matches the text data with the preset answer of the first question to obtain the corresponding matching degree, and the matching sub-model is trained based on the Bimpm model;

[0071] Determine the judgment result according to the matching degree.

[0072] Specifically, since the obtained audio data cannot be directly processed, the audio data needs to be translated into text data, and the translation sub-model directly calls ready-made tools of Baidu or iFlytek to execute. The obtained text data is matched with the preset answer of the corresponding first question through the matching sub-model to obtain the corresponding matching degree, and the matching degree is also the corresponding score, such as 0.8 corresponding to 80 points; finally, the judgment result is determined according to the matching degree, that is, the level is determined according to the interval to which the matching degree belongs. For example, when the matching degree is 0.8, the corresponding judgment result is good, and when the matching degree is 0.85 or above, the corresponding judgment result is excellent.

[0073] BiMPM (Bilateral Multi-perspective Matching) is a semantic matching model based on the "matching-aggregation" framework. For two input sentences P and Q, after adopting the pre-trained language model embedding, the model uses bidirectional LSTMs in the representation layer to perform representation on P and Q respectively. Then, matching is performed in two directions, namely P-->Q and Q-->P, while combining four matching methods, which is the multi-perspective matching. The matching results are input into a bidirectional LSTM for aggregation, and then the result is obtained through a fully connected layer softmax.

[0074] The audio data is converted into text data through the translation sub-model, and the text data is matched with the preset answer through the matching sub-model to obtain the matching degree, and finally the judgment result is determined, improving the processing efficiency.

[0075] Furthermore, the matching of the text data with the preset answer of the first question by using the matching sub-model under the judgment model includes:

[0076] The feature extraction model is used to extract features from the text data to obtain corresponding keywords, and the feature extraction model is trained based on the LDA model;

[0077] The matching sub-model is used to match the keywords with the preset answer.

[0078] Specifically, before matching, the feature extraction model is used to extract features from the text data to obtain corresponding keywords, and the matching sub-model is used to match the keywords with the preset answer.

[0079] LDA (Latent Dirichlet Allocation) is a document topic generation model, also known as a three-layer Bayesian probability model, including three layers of structure: words, topics, and documents. The so-called generation model means that we think that each word in an article is obtained through a process of "the article selects a certain topic with a certain probability and selects a certain word from this topic with a certain probability".

[0080] By first performing a keyword extraction step and then matching the keywords with the preset answer, the accuracy of the result is improved.

[0081] S3. Determine the emotional state of the person being collected according to the emotional feature parameters;

[0082] Specifically, based on the emotional feature parameters, query a preset status comparison table to determine the emotional state of the person being collected.

[0083] Further, determining the emotional state of the person being collected according to the emotional feature parameters includes:

[0084] Based on the emotional feature parameters, query a preset status comparison table to determine the emotional state of the person being collected.

[0085] Specifically, the preset status report stores the corresponding emotional states of each sub-parameter combination under the emotional feature parameters. By using the obtained emotional feature parameters to query the preset status comparison table, the corresponding emotional state can be obtained. For example, when the emotional feature parameters are that the facial contour is 1, the mouth movement is 2, the eyebrow movement is 2, and the cheek movement is 0, the corresponding emotional state is nervous.

[0086] The corresponding emotional state can be obtained quickly and accurately in the form of looking up a table.

[0087] S4. Based on the judgment result and the emotional state, determine the second question of the person being collected;

[0088] Specifically, use a grading model to judge the coping level of the judgment result and the emotional state. According to the coping level, determine the second question of the person being collected, and the second question is one or more questions.

[0089] Further, as Figure 4 shown, determining the second question of the person being collected based on the judgment result and the emotional state includes:

[0090] S41. Use a grading model to perform a grading judgment on the judgment result and the emotional state to determine the coping level to which the person being collected belongs when answering the first question. The grading model is trained based on a decision tree model;

[0091] S42. According to the coping level, determine the second question corresponding to the level.

[0092] Specifically, perform a grading judgment by using the corresponding judgment result and emotional state of the person being collected to determine the coping level of the person being collected when answering the first question. For example, when the judgment result is excellent and the emotional state is relaxed, through the grading judgment of the grading model, the coping level is obtained as level 1, that is, the highest level. According to this coping level of level 1, determine the second question corresponding to the level.

[0093] The decision tree model is a simple and easy-to-use non-parametric classifier. It does not require any prior assumptions about the data, has a relatively fast calculation speed, the results are easy to interpret, and it is robust.

[0094] By performing a hierarchical judgment based on the judgment result and the emotional state to determine the level of the person being collected's answer to the first question, it is beneficial to determine the corresponding-level questions subsequently, thereby improving the audit quality.

[0095] S5. Push the second question to the auditor.

[0096] Specifically, push the second question to the auditor, and the auditor sends one of the questions in the second question to the person being collected as needed, in the form of directly sending text or voice.

[0097] It should be emphasized that in order to further ensure the privacy and security of data, all the video data and audio data can also be stored in a node of a blockchain.

[0098] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.

[0099] By obtaining the video data and audio data of the person being collected when answering the first question, using the emotional feature extraction model to extract features from the facial data of the person being collected in the video data to obtain emotional feature parameters, and using the judgment model to judge the answer content of the audio data to obtain a judgment result, the judgment result is the score level of the person being collected's answer to the first question; process the audio and video data respectively to obtain effective parameters; then determine the emotional state of the person being collected according to the emotional feature parameters; based on the judgment result and the emotional state, determine the second question of the person being collected; by scoring the audio data of the user's answer and determining the emotional state according to the video data, to finally determine the second question of the person being collected, thereby improving the audit quality and processing efficiency of the person being audited.

[0100] This embodiment also provides an intelligent question-pushing device, as Figure 4 shown, which is the functional module diagram of the intelligent question-pushing device of this application.

[0101] The intelligent question-pushing device 100 described in this application can be installed in an electronic device. According to the functions achieved, the intelligent question-pushing device 100 may include an acquisition module 101, an audio-video processing module 102, a status confirmation module 103, a question determination module 104, and a push module 105. The modules described in this application may also be referred to as units, which refer to a series of computer program segments that can be executed by the processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0102] In this embodiment, the functions of each module / unit are as follows:

[0103] The acquisition module 101 is used to acquire video data and audio data when the person being collected answers the first question;

[0104] Furthermore, the intelligent question-pushing device 100 further includes an information acquisition module and a recommendation module;

[0105] The acquisition module is used to acquire the basic information of the person being collected;

[0106] The recommendation module is used to process the basic data by using a recommendation model to obtain the first question of the person being collected, and the recommendation model is trained based on the Apriori model.

[0107] Through the cooperation of the information acquisition module and the recommendation module, and by using the recommendation model, the basic information of the person being collected is processed to obtain the first question with a certain correlation, so as to improve work efficiency and the relevance of the questions.

[0108] The audio-video processing module 102 is used to extract features from the facial data of the person being collected in the video data by using an emotion feature extraction model to obtain emotion feature parameters, and to judge the answer content of the audio data through a judgment model to obtain a judgment result;

[0109] Furthermore, the audio-video processing module 102 includes a truncation sub-module and a parameter extraction sub-module;

[0110] The truncation sub-module is used to truncate the video at the position of the head portrait of the person being collected in the video data to obtain the facial data;

[0111] The parameter extraction sub-module is used to extract the facial data of the person being collected through an emotion feature extraction model, and the emotion feature extraction model is trained based on an enhanced long-term recurrent convolutional network model.

[0112] Through the cooperation of the intercepting sub-module and the parameter extraction sub-module, the corresponding facial data is intercepted from the complete video data. By inputting the facial data containing only the human face into the emotion feature extraction model for processing, the accuracy is improved. And by processing through the emotion feature extraction model, not only the processing efficiency is improved but also the recognition rate is increased.

[0113] Further, the audio-video processing module 102 includes a conversion sub-module, a matching sub-module, and a result determination sub-module;

[0114] The conversion sub-module is used to convert the audio data into corresponding text data through the translation sub-model under the judgment model;

[0115] The matching sub-module is used to match the text data with the preset answer of the first question by using the matching sub-model under the judgment model to obtain the corresponding matching degree, and the matching sub-model is trained based on the Bimpm model;

[0116] The result determination sub-module is used to determine the judgment result according to the matching degree.

[0117] Through the cooperation of the conversion sub-module, the matching sub-module, and the result determination sub-module, the audio data is converted into text data by using the translation sub-model and the text data is matched with the preset answer by using the matching sub-model to obtain the matching degree, and finally the judgment result is determined, which improves the processing efficiency.

[0118] Still further, the matching sub-module further includes a keyword extraction unit and a corresponding matching unit;

[0119] The keyword extraction unit is used to extract features from the text data by using the feature extraction model to obtain the corresponding keywords, and the feature extraction model is trained based on the LDA model;

[0120] The corresponding matching unit is used to match the keywords with the preset answer by using the matching sub-model.

[0121] Through the cooperation of the keyword extraction unit and the corresponding matching unit, a keyword extraction step is first performed, and then the keywords are matched with the preset answer, which improves the accuracy of the result.

[0122] The status confirmation module 103 is used to determine the emotional state of the collected person according to the emotion feature parameters;

[0123] Further, the status confirmation module 103 includes a query sub-module;

[0124] The query sub-module is used to query the preset status comparison table based on the emotion feature parameters to determine the emotional state of the collected person.

[0125] The corresponding emotional state can be obtained quickly and accurately in the form of querying a table by the query sub-module.

[0126] The problem determination module 104 is used to determine the second problem of the person being collected based on the judgment result and the emotional state.

[0127] Furthermore, the problem determination module 104 includes a grading sub-module and a corresponding determination sub-module.

[0128] The grading sub-module is used to perform a grading judgment on the judgment result and the emotional state by using a grading model to determine the coping level to which the person being collected belongs when answering the first question. The grading model is trained based on a decision tree model.

[0129] The corresponding determination sub-module is used to determine the second question corresponding to the corresponding level according to the coping level.

[0130] Through the cooperation of the grading sub-module and the corresponding determination sub-module, a grading judgment is made according to the judgment result and the emotional state to determine the level of the person being collected when answering the first question, which is beneficial to subsequent determination of questions corresponding to the corresponding level, thereby improving the audit quality.

[0131] The push module 105 is used to push the second question to the auditor.

[0132] By adopting the above device, the intelligent question-pushing device 100 uses the cooperation of the acquisition module 101, the audio-video processing module 102, the status confirmation module 103, the problem determination module 104, and the push module 105. By acquiring the video data and audio data of the person being collected when answering the first question, the facial data of the person being collected in the video data is feature-extracted by using an emotional feature extraction model to obtain emotional feature parameters, and the audio data is judged for the answer content by using a judgment model to obtain a judgment result. The judgment result is the score level of the person being collected when answering the first question; the audio and video data are processed respectively to obtain effective parameters; then, according to the emotional feature parameters, the emotional state of the person being collected is determined; based on the judgment result and the emotional state, the second question of the person being collected is determined; by scoring the audio data of the user's answer and determining the emotional state according to the video data, the second question of the person being collected is finally determined, thereby improving the audit quality and processing efficiency of the person being audited.

[0133] The embodiment of the present application also provides a computer device. For details, please refer to Figure 5 , Figure 5 which is the basic structural block diagram of the computer device in this embodiment.

[0134] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 4 with components 41 - 43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0135] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc.

[0136] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the intelligent question-pushing method. In addition, the memory 41 can also be used to temporarily store various data that have been output or will be output.

[0137] In some embodiments, the processor 42 may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the computer-readable instructions stored in the memory 41 or process data, such as running the computer-readable instructions of the intelligent question-pushing method.

[0138] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0139] In this embodiment, when the processor executes the computer-readable instructions stored in the memory, the steps of the intelligent question-pushing method in the above embodiment are implemented. By acquiring the video data and audio data of the person being collected when answering the first question, using the emotion feature extraction model to extract features from the facial data of the person being collected in the video data to obtain emotion feature parameters, and using the judgment model to judge the answer content of the audio data to obtain a judgment result, the judgment result is the score level of the person being collected answering the first question; the audio and video data are processed respectively to obtain effective parameters; then, according to the emotion feature parameters, the emotion state of the person being collected is determined; based on the judgment result and the emotion state, the second question of the person being collected is determined; by scoring the audio data of the user's answer and determining the emotion state according to the video data, the second question of the person being collected is finally determined, thereby improving the audit quality and processing efficiency of the person being audited.

[0140] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor, so that the at least one processor executes the steps of the intelligent question-pushing method as described above. By acquiring the video data and audio data of the person being collected when answering the first question, using the emotion feature extraction model to extract features from the facial data of the person being collected in the video data to obtain emotion feature parameters, and using the judgment model to judge the answer content of the audio data to obtain a judgment result, the judgment result is the score level of the person being collected answering the first question; the audio and video data are processed respectively to obtain effective parameters; then, according to the emotion feature parameters, the emotion state of the person being collected is determined; based on the judgment result and the emotion state, the second question of the person being collected is determined; by scoring the audio data of the user's answer and determining the emotion state according to the video data, the second question of the person being collected is finally determined, thereby improving the audit quality and processing efficiency of the person being audited.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0142] The intelligent question-pushing device, computer device, and computer-readable storage medium of the above embodiments of the present application have the same technical effects as the intelligent question-pushing method of the above embodiments, and will not be elaborated here.

[0143] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure made by using the specification and drawings of the present application, directly or indirectly applied in other related technical fields, is equally within the scope of the patent protection of the present application.

Claims

1. An intelligent question pushing method, characterized in that, The method includes: Obtaining the basic information of the person to be collected, processing the basic data by using a recommendation model to obtain the first question of the person to be collected, where the recommendation model is trained based on the Apriori model, and obtaining the video data and audio data of the person to be collected when answering the first question; Intercepting the video of the position where the head of the person to be collected is located in the video data to obtain facial data, extracting emotional feature parameters from the facial data of the person to be collected by using an emotional feature extraction model, where the emotional feature extraction model is trained based on an enhanced long-term recurrent convolutional network model, and judging the answer content of the audio data by using a judgment model to obtain a judgment result; Determining the emotional state of the person to be collected according to the emotional feature parameters; Using a grading model to judge the coping level of the judgment result and the emotional state, and determining the second question of the person to be collected according to the coping level; Pushing the second question to the auditor.

2. The intelligent question pushing method according to claim 1, wherein The judging the answer content of the audio data by using a judgment model to obtain a judgment result includes: Converting the audio data into corresponding text data by using a translation sub-model under the judgment model; Using a matching sub-model under the judgment model to match the text data with the preset answer of the first question to obtain a corresponding matching degree, where the matching sub-model is trained based on the Bimpm model; Determining the judgment result according to the matching degree.

3. The intelligent question pushing method according to claim 2, wherein The using a matching sub-model under the judgment model to match the text data with the preset answer of the first question includes: Extracting features from the text data by using a feature extraction model to obtain corresponding keywords, where the feature extraction model is trained based on the LDA model; Using the matching sub-model to match the keywords with the preset answer.

4. The intelligent question pushing method according to claim 1, characterized in that, The determining the emotional state of the person to be collected according to the emotional feature parameters includes: Querying a preset state comparison table based on the emotional feature parameters to determine the emotional state of the person to be collected.

5. The intelligent question pushing method according to any one of claims 1 to 4, characterized in that The determining the second question of the person to be collected based on the judgment result and the emotional state includes: Using a grading model to perform grading judgment on the judgment result and the emotional state to determine the coping level to which the person to be collected belongs when answering the first question, where the grading model is trained based on a decision tree model; Determining the second question corresponding to the level according to the coping level.

6. An intelligent question pushing device, characterized in that, The device includes: An acquisition module, configured to obtain the basic information of the person to be collected, process the basic data by using a recommendation model to obtain the first question of the person to be collected, where the recommendation model is trained based on the Apriori model, and obtain the video data and audio data of the person to be collected when answering the first question; An audio-video processing module, configured to intercept the video of the position where the head of the person to be collected is located in the video data to obtain facial data, and extract emotion feature parameters from the facial data of the person to be collected through an emotion feature extraction model, wherein the emotion feature extraction model is trained based on an enhanced long-term recurrent convolutional network model, and judge the answer content of the audio data through a judgment model to obtain a judgment result; A status confirmation module, configured to determine the emotional state of the person to be collected according to the emotion feature parameters; A question determination module, configured to use a grading model to judge the response level of the judgment result and the emotional state, and determine the second question of the person to be collected according to the response level; A push module, configured to push the second question to the auditor.

7. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the intelligent question-pushing method as described in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the intelligent question-pushing method as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Question and answer process optimization method and device, computer device and storage medium

    CN109767321A

  • Article recommendation method and related device

    CN110648170A

  • Method and device for assisting user in talking, computer equipment and storage medium

    CN114220055A