Scoring Method, Device, Terminal Device and Storage Medium Based on Semantic Analysis

Through the scoring method based on semantic analysis, the trained neural network model uses semantic analysis and text classification of interviewee speech information, which solves the problems of low interview efficiency and inaccurate scoring caused by the large amount of language model parameters, and achieves rapid and accurate interview scoring and improves interview efficiency.

CN111695352BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010469517.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-28
Publication Date
2025-05-27
Estimated Expiration
2040-05-28

AI Technical Summary

Technical Problem

In the prior art, the large amount of parameters of the language model makes it difficult to support the terminal processor memory, slow training and reasoning speed, and difficult to judge the accuracy, which increases the cost of interviews and reduces the accuracy of judging capabilities in various dimensions, affecting the efficiency of intelligent interviews.

Method used

A scoring method based on semantic analysis is adopted. By obtaining the voice information of the target user and converting it into text information, inputting the first neural network model after training for semantic analysis, the text classification results are obtained and the interview score results are calculated. This method uses the training sample set and the first neural network model trained by the second neural network model to achieve rapid and accurate scoring of the interviewer's ability points in each dimension.

Benefits of technology

It improves the efficiency of smart interviews and the accuracy of ratings, reduces the cost of interviews, and enhances the accuracy of judging capabilities in various dimensions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111695352B_ABST
    Figure CN111695352B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of computer technology, and provides a scoring method, device, terminal device, and storage medium based on semantic analysis. The method includes: obtaining the voice information of a target user and converting the voice information into text information; inputting the text information into a trained first neural network model to perform semantic analysis on the text information to obtain the output text classification result of the first neural network model; wherein, the text classification result includes a scoring label corresponding to the text information, and according to the scoring label, calculating the interview scoring result of the target user. By means of this application, the problems of slow precision inference speed of the language model increasing the interview cost, low accuracy of interview dimension determination, and low interview efficiency are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and particularly relates to a scoring method, device, terminal device, and storage medium based on semantic analysis. Background Art

[0002] With the expansion of enterprise scale, the number of recruited employees also increases; in the case of a large number of recruitments, intelligent interviews can be used for ability scoring. In the intelligent interview ability scoring scenario, each dimension ability point of the user is scored according to the user's answer.

[0003] However, currently, the number of parameters of the language model is very large, and it is difficult for the terminal processor memory to support, resulting in slow training and inference speeds of the language model, and it is difficult to judge the accuracy of the language model, which not only increases the interview cost but also reduces the accuracy of the determination of each dimension ability, thus directly affecting the efficiency of intelligent interviews. Summary of the Invention

[0004] The embodiments of this application provide a scoring method, device, terminal device, and storage medium based on semantic analysis, which can solve the problems of slow inference speed of language model accuracy increasing the interview cost, low accuracy of interview dimension determination, and low interview efficiency.

[0005] In a first aspect, the embodiments of this application provide an application program resource update method, including:

[0006] Obtain the voice information of the target user and convert the voice information into text information;

[0007] Input the text information into the trained first neural network model to perform semantic analysis on the text information and obtain the output text classification result of the first neural network model; wherein, the text classification result includes the scoring label corresponding to the text information, the first neural network model is trained based on a training sample set and a second neural network model, the second neural network model is trained based on the training sample set and the output result of the first neural network model, the output result of the first neural network model is obtained by taking the training sample set as the input, and the training sample set includes multiple interview corpus texts;

[0008] Calculate the interview scoring result of the target user according to the scoring label.

[0009] Through this embodiment, a scoring method based on semantic analysis is adopted. After training the language model, the conversation content of the target user during the interview is collected, semantic analysis is performed on the conversation content, the conversation content is classified into text categories based on the semantic analysis results, and the scores of the conversation content in the corresponding text categories are calculated, realizing the rapid and accurate scoring of the target user's various dimensional ability points according to the target user's answers in the intelligent interview scenario, improving the interview efficiency and the accuracy of interview scoring.

[0010] In a possible implementation manner of the first aspect, obtaining the voice information of the target user and converting the voice information into text information includes:

[0011] Identifying the voice information through a speech recognition algorithm and extracting the acoustic features in the voice information;

[0012] Converting the voice information into text information according to the acoustic features.

[0013] Exemplarily, establishing the corresponding relationship between the text information and the current conversation topic provides a more accurate and reliable basis for subsequent classification of the text information, making the scoring of the interviewee more accurate according to the voice information during the intelligent interview process.

[0014] In a possible implementation manner of the first aspect, before inputting the text information into the trained first neural network model, it includes:

[0015] Dividing the text information according to the preset number of word segments to obtain at least one short sentence text that meets the preset number of word segments;

[0016] Alternatively, during the process of converting the voice information into the text information, setting the maximum number of short sentences, dividing the voice information into at least one voice short sentence that is less than or equal to the maximum number of short sentences, and converting the at least one voice short sentence into the text information.

[0017] In a possible implementation manner of the first aspect, before inputting the text information into the trained first neural network model, it includes:

[0018] Obtaining a training sample set, where the training sample set includes multiple interview corpus texts;

[0019] Dividing the sentence texts in the training sample set into a short sentence set with a preset number of word segments, and encoding the word segments in the short sentence set to obtain a word segment matrix;

[0020] Performing convolution calculation on the word segment matrix to obtain a target matrix, and taking the dot product of the target matrix and the parameter matrix as the output matrix of the first neural network;

[0021] Obtain the predicted vectors corresponding to the masked word segments in the output matrix, and calculate the cross-entropy loss between the predicted vectors and the true vectors corresponding to the masked words actually, as the first loss.

[0022] In a possible implementation manner of the first aspect, before inputting the text information into the trained first neural network model, it includes:

[0023] Input the output matrix into a second neural network model, and the second neural network model performs bidirectional convolution calculation on the output matrix to output the probability that each word segment in the output matrix is masked;

[0024] Calculate the cross-entropy loss corresponding to all the masked word segments in the probability matrix, as the second loss.

[0025] In a possible implementation manner of the first aspect, the method includes:

[0026] After completing the training of the first neural network according to the preset number of iterative training times, based on the output matrix of the first neural network and the training sample set, perform iterative training on the second neural network model according to the preset number of training times of the second neural network model, and adjust the parameter matrix of the second neural network model.

[0027] In a possible implementation manner of the first aspect, the method includes:

[0028] Perform interactive training on the first neural network model and the second neural network model, adjust the parameter matrix, and obtain the first target parameter matrix of the first neural network model and the second target parameter matrix of the second neural network model respectively.

[0029] In a second aspect, an embodiment of the present application provides a scoring device based on semantic analysis, including:

[0030] An acquisition unit, configured to acquire the voice information of the target user and convert the voice information into text information;

[0031] A processing unit, configured to input the text information into the trained first neural network model, perform semantic analysis on the text information, and obtain the output text classification result of the first neural network model; wherein, the text classification result includes a scoring label corresponding to the text information, the first neural network model is trained based on a training sample set and a second neural network model, the second neural network model is trained based on the training sample set and the output result of the first neural network model, the output result of the first neural network model is obtained by using the training sample set as the input, and the training sample set includes multiple interview corpus texts;

[0032] A scoring unit, configured to calculate an interview scoring result of the target user according to the scoring tags.

[0033] In a third aspect, an embodiment of the present application provides a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the scoring method based on semantic analysis according to any one of the above first aspects is implemented.

[0034] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the scoring method based on semantic analysis according to any one of the above first aspects is implemented.

[0035] In a fifth aspect, an embodiment of the present application provides a computer program product, which when running on a terminal device, causes the terminal device to execute the scoring method based on semantic analysis according to any one of the above first aspects.

[0036] It can be understood that the beneficial effects of the above second to fifth aspects can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here.

[0037] The beneficial effects of the embodiment of the present application compared with the prior art are as follows: Through the embodiment of the present application, the voice information of the target user is obtained, and the voice information is converted into text information; the text information is input into the trained first neural network model, and semantic analysis is performed on the text information to obtain the output text classification result of the first neural network model; wherein, the text classification result includes the scoring tags corresponding to the text information; according to the scoring tags, the interview scoring result of the target user is calculated; realizing the rapid and accurate scoring of each dimension ability point of the target user according to the answer in the intelligent interview scenario, improving the interview efficiency and the accuracy of the interview scoring; having strong usability and practicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0040] Figure 2It is a schematic flowchart of a scoring method based on semantic recognition provided by an embodiment of the present application;

[0041] Figure 3 It is a schematic flowchart of speech model training provided by another embodiment of the present application;

[0042] Figure 4 It is a schematic structural diagram of a scoring device based on semantic analysis provided by an embodiment of the present application;

[0043] Figure 5 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0044] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0045] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0046] It should also be understood that the term "and / or" as used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0047] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" depending on the context.

[0048] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0049] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0050] Currently, in the intelligent interview session scenario, especially in application scenarios with a large recruitment volume, the voice information during the interviewee's session is received through the microphone of the terminal device, and based on the semantic analysis of the voice information, the interviewee's answers are scored to evaluate the interviewee's capabilities in various dimensions and improve the interview efficiency.

[0051] As Figure 1 shown, the interviewee is the user, and the terminal device can ask the user questions in multiple feature dimensions in the form of text or voice, receive the user's answers, and score the user's answers based on semantic analysis to finally obtain the ability scores of the user in various feature dimensions.

[0052] Among them, the terminal device can be a mobile phone, a laptop computer, a super personal computer 30 (ultra-mobile personal computer, UMPC), etc.; it can also include but is not limited to tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, netbooks, personal digital assistants (PDAs), etc. The specific type of the carrier of the client, that is, the terminal device, is not limited in any way in the embodiments of this application.

[0053] See Figure 2 is a schematic flowchart of the implementation process of the scoring method based on semantic analysis provided by the embodiments of this application. The method includes:

[0054] Step S201, obtain the voice information of the target user and convert the voice information into text information.

[0055] In this embodiment, the target user can be the interviewee, and the terminal device can act as the role of the interviewer to ask the target user questions in multiple aspects; the terminal device receives the voice information of the target user to implement the intelligent interview session scenario.

[0056] In some embodiments, obtaining the voice information of the target user and converting the voice information into text information includes:

[0057] A1. Identifying the voice information through a speech recognition algorithm and extracting the acoustic features in the voice information;

[0058] A2. Converting the voice information into text information according to the acoustic features.

[0059] In the embodiment of the present application, in the conversation scenario of an intelligent interview, the terminal device can receive the voice information during the conversation of the target user through a microphone, identify the voice information through a speech recognition algorithm, extract the acoustic features of the voice, obtain the phoneme information of the voice information, and convert the voice information into text information by corresponding the phoneme information with the characters or words in the dictionary.

[0060] In some embodiments, before inputting the text information into the trained first neural network model, it includes:

[0061] Dividing the text information according to a preset number of word segments to obtain at least one short sentence text that meets the preset number of word segments;

[0062] Or, during the process of converting the voice information into the text information, setting the maximum number of short sentences, dividing the voice information into at least one voice short sentence less than or equal to the maximum number of short sentences, and converting the at least one voice short sentence into the text information.

[0063] Specifically, the terminal device divides the text information according to the preset number of word segments to obtain multiple short sentence texts that meet the preset number of word segments; or during the process of converting the voice information into the text information, sets the maximum number of short sentences, divides the voice information into multiple voice short sentences less than or equal to the maximum number of short sentences, and converts the multiple voice short sentences into the corresponding text information. So as to keep the size of the target parameter matrix consistent before and after when performing semantic recognition on the text information subsequently, which is convenient for the data processing of the terminal device.

[0064] It should be noted that in the application scenario of the actual conversation process, establishing the correspondence relationship between the text information and the current conversation topic provides a more accurate and reliable basis for the subsequent classification of the text information, making the scoring of the interviewee more accurate during the intelligent interview process.

[0065] Step S202: Input the text information into the trained first neural network model, perform semantic analysis on the text information, and obtain the output text classification result of the first neural network model; wherein, the text classification result includes the scoring label corresponding to the text information.

[0066] In this embodiment, the first neural network model is a language model, which performs semantic recognition on text information, and classifies the text information according to the recognized semantics to obtain a scoring label corresponding to the classification result of the text information.

[0067] Specifically, during the process of performing semantic recognition on text information by the terminal device, the sentence corresponding to the text information is divided into short sentences, which are divided into multiple words or characters; the divided words or characters are converted into a vector matrix representation, and semantic understanding is performed through a semantic recognition algorithm; the text information is classified according to the semantics, and a text classification result corresponding to the text information is output.

[0068] Among them, the first neural network model is trained based on a training sample set and a second neural network model, the second neural network model is trained based on the training sample set and the output result of the first neural network model, the output result of the first neural network model is obtained by using the training sample set as input, and the training sample set includes multiple interview corpus texts.

[0069] See Figure 3 , the schematic flowchart of the training method of the speech recognition model provided by the embodiment of the present application. Before inputting the text information into the trained first neural network model, the training process of the model includes:

[0070] Step S301, obtain a training sample set, where the training sample set includes multiple interview corpus texts;

[0071] Specifically, the training sample set includes interview corpus texts in multiple dimensions, and the first neural network model is trained in multiple dimensions to facilitate multi-dimensional classification of the speech information input by the target user, so as to realize scoring of the multi-dimensional capabilities of the target user.

[0072] Step S302, divide the sentence text in the training sample set into a short sentence set with a preset number of word segments, and encode the word segments in the short sentence set to obtain a word segment matrix;

[0073] The terminal device divides the sentence text in the training sample set according to the preset number of segmentations to obtain a set of short sentences less than or equal to the preset number of segmentations. For example, "The weather has been bad in the past few days, but it is rare that the weather is good today, which is very suitable for outing" is divided into {"before", "a few days", "weather", "always", "bad", "," "rarely", "today", "weather", "good", "," "very", "suitable", "outing"}, plus punctuation marks, a total of 14 segmentations, then the preset number of segmentations can be 14, and different segmentation number thresholds can be set according to the model size. Encode each segmentation to obtain an encoded segmentation matrix, each row of the matrix identifies the representation vector of each segmentation, for example, the above sentence text includes 14 segmentations, then the segmentation matrix includes 14 rows. Specifically, taking the above sentence text as an example, the segmentation matrix M of 14*100 dimensions is obtained by encoding the segmentations in the short sentence set, and Mi is denoted as the i-th row of the segmentation matrix M.

[0074] Step S303, performing convolution calculation on the word segmentation matrix to obtain a target matrix, and taking the dot product of the target matrix and the parameter matrix as the output matrix of the first neural network;

[0075] Specifically, before the convolution calculation of the word segmentation matrix, one or more word segmentations in the short sentence set are randomly masked, that is, one of the word segmentations is encoded as an unknown quantity. Taking the above-mentioned word segmentation matrix M as an example, the fifth word "bad" and the ninth word "good" are masked and used as the input of the first neural network model. The input word segmentation matrix is ​​convoluted. Taking the first row of the word segmentation matrix M as an example, M1 is vector dot producted with M1 to M14 to obtain r1 to r14, where r1 to r14 are scalar values; then r1*M1+r2*M2+......+r4*M14=P1, P1 is a 100-dimensional vector. Each row of the word segmentation matrix M is calculated according to the operation process of the first row, M1 to M14 are updated to P1 to P14, and vectors P1 to P14 are combined into a 14*100-dimensional matrix P. In order to make the first neural network model learn more semantics, perform convolution calculation on matrix P again according to the operation of matrix M to obtain matrix S, and perform convolution calculation on matrix S again according to the operation of matrix M to obtain matrix K, and the size of matrix K is 14*100. According to the dictionary size and the preset number of word segmentations of the first neural network model, set the parameter matrix; for example, for the matrix K obtained after the above convolution calculation, the dictionary size of the first neural network model is 2000, then set the size of parameter matrix Q to 100*2000, set K*Q=T, and obtain a matrix T of size 14*2000, and use matrix T as the output matrix of the first neural network.

[0076] Step S304: Obtain the predicted vectors corresponding to the masked word segments in the output matrix, and calculate the cross-entropy loss between the predicted vectors and the true vectors corresponding to the actually masked words as the first loss.

[0077] Specifically, for example, calculate the cross-entropy loss between the predicted vectors corresponding to the 5th and 9th rows in matrix T and the true vectors corresponding to the masked words "not good" and "not bad" as the first loss Loss1.

[0078] In some embodiments, before inputting the text information into the trained first neural network model, it includes:

[0079] B1: Input the output matrix into a second neural network model, and let the second neural network model perform bidirectional convolution calculation on the output matrix to output the probability that each word segment in the output matrix is masked.

[0080] Specifically, the second neural network model is a sequence labeling model. Taking the output matrix output by the first neural network model as the input, calculate the probability that the word segment corresponding to each row vector in the output matrix is masked and the probability that it is not masked, so as to realize the recognition and labeling of each word segment in the output matrix, making the semantic analysis of the first neural network model more accurate.

[0081] In the bidirectional LSTM layer of the second neural network model, perform convolution calculation, splice the results of the bidirectional calculation and input them into the output layer of the second neural network model; let the output layer perform a linear transformation on the vector corresponding to each word segment in the bidirectional LSTM layer; for example, taking the above output matrix T as an example, after the linear transformation of the bidirectional LSTM layer and the output layer, the output of the first word segment is a 100-dimensional vector Y1. Set a parameter matrix G of size 100*2, and obtain the output of the first word segment of the output layer through Y1*G = C1; where C1 is a 2-dimensional vector, and the first element in the 2-dimensional vector represents the probability that the word segment is masked, and the second element represents the probability that the word segment is not masked. Based on the same operation, 2-dimensional vectors C1 to C14 corresponding to all word segments can be obtained, and the probability matrix C corresponding to all masked word segments is output.

[0082] B2: Calculate the cross-entropy loss corresponding to all masked word segments in the probability matrix as the second loss.

[0083] Specifically, the second loss Loss2 = sum{cross-entropy loss(whether the i-th word is masked, Ci)}, i = 1, 2, 3,..., 14.

[0084] In one embodiment, the loss of the first neural network model is defined as Loss1 - Loss2. The better the recognition effect of the second neural network model, the easier it is for the second neural network model to find which words in the output matrix of the first neural network model are masked. That is to say, the greater the gap between the word segmentation or semantics analyzed by the first neural network model and the true semantics.

[0085] In one embodiment, the first neural network model and the second neural network model are interactively trained. The parameter matrices of the first neural network model and the second neural network model are randomly initialized respectively, that is, a parameter matrix of a preset size is defined, and a predetermined initial value is set for the parameter matrix. The first neural network model and the second neural network model are trained in rounds according to the number of iterative training times. In the first round, the first neural network model is iteratively trained to adjust the parameter matrix of the first neural network model. The second neural network model is not iteratively trained. Only the probability that each word segmentation in the output matrix of the first neural network model is masked is calculated through the second neural network model, and the second loss is calculated. The first neural network model is iteratively trained according to the second loss and the first loss to adjust the parameter matrix of the first neural network model.

[0086] In one embodiment, after completing the training of the first neural network according to the preset number of iterative training times, based on the output matrix of the first neural network and the training sample set, the second neural network model is iteratively trained according to the preset number of training times of the second neural network model to adjust the parameter matrix of the second neural network model.

[0087] The first neural network model and the second neural network model are interactively trained to adjust the parameter matrix, and the first target parameter matrix of the first neural network model and the second target parameter matrix of the second neural network model are obtained respectively.

[0088] Among them, the number of iterative training times can be set according to the data volume. For example, if there are a total of L sentence data and N data are set for each training, the number of iterative training times is L / N. Generally, N is set to 128.

[0089] In one embodiment, after the first neural network model is iteratively trained, for the matrix output by the output layer of the first neural network model, a scoring parameter matrix is set according to the scoring levels. For example, for the output matrix T, a scoring parameter matrix of size 2000*5 is set. The output matrix T is multiplied by the scoring parameter matrix to obtain a predicted scoring label S corresponding to the input statement text (T*U = S). The cross-entropy loss between the predicted scoring label and the true scoring label is calculated, and the first neural network model is continuously iteratively trained through the cross-entropy loss to adjust the scoring parameter matrix to obtain a target scoring parameter matrix. The first neural network model using the target scoring parameter matrix is used as a model for semantic recognition and text classification of the target user's input voice information. By multiplying the output matrix of the first neural network model by the target predicted scoring label, the probability of each score level in the scoring label level is obtained, and the score level with the highest probability is used as the scoring result of this session.

[0090] Specifically, the scoring label is a label that sets score levels for the results of text classification, so that the ability level of the target user can be determined according to the scoring label. The scoring label can be set, for example, as scores in five levels: 1, 2, 3, 4, and 5. The scoring result of this session scenario is determined according to the scoring label corresponding to the text classification result.

[0091] Through the embodiments of the present application, the first neural network model and the second neural network model are interactively trained. The second neural network model is used to judge whether the output of the first neural network model is true and reasonable, and the loss of the second neural network model is added to the first neural network model as a reference index for iteratively training the first neural network model; the closer the output of the first neural network model is to the true semantics, the more difficult it is for the second neural network model to accurately judge whether the semantics of the output of the first neural network model is incorrect, which further promotes the iterative training of the second neural network model. After the iterative training of the second neural network model, the authenticity of the output result in the first neural network model can be judged more accurately, and thus the output of the first neural network model will also be closer to the true semantics. During the iterative training process of the two models, the semantic recognition and sequence annotation capabilities become stronger and stronger, improving the disadvantage that the first neural network model must output a specified word or short sentence to be considered as recognizing the semantics, making the output semantics of the first neural network model more flexible and variable, and thus the classification of different input text information is more accurate. In addition, the two models are trained simultaneously during the training process, and only the trained first neural network model is used in the actual application process. Therefore, when deploying the semantic analysis unit on the terminal device, the number of parameters is greatly reduced, the inference speed of the model is greatly improved, the storage space occupied by the model is reduced, and the processing performance of the terminal device is improved.

[0092] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0093] Corresponding to the scoring method based on semantic analysis described in the above embodiments, Figure 4 FIG. shows a structural block diagram of a scoring device based on semantic analysis provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0094] Referring to Figure 4 , the device includes:

[0095] An acquisition unit 41, configured to acquire voice information of a target user and convert the voice information into text information;

[0096] A processing unit 42, configured to input the text information into a trained first neural network model, perform semantic analysis on the text information, and obtain an output text classification result of the first neural network model; wherein, the text classification result includes a scoring label corresponding to the text information, the first neural network model is trained based on a training sample set and a second neural network model, the second neural network model is trained based on the training sample set and an output result of the first neural network model, the output result of the first neural network model is obtained by taking the training sample set as an input, and the training sample set includes a plurality of interview corpus texts;

[0097] A scoring unit 43, configured to calculate an interview scoring result of the target user according to the scoring label.

[0098] It should be noted that for the information interaction, execution process, etc. between the above devices / units, since they are based on the same concept as the method embodiments of the present application, the specific functions and the technical effects brought thereby can be specifically referred to the method embodiment part, and will not be elaborated here.

[0099] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.

[0100] Figure 5 FIG. is a schematic structural diagram of a terminal device provided in an embodiment of the present application. As Figure 5 shown, the terminal device 5 in this embodiment includes: at least one processor 50 ( Figure 5 only one is shown in the figure), a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50. When the processor 50 executes the computer program 52, the steps in any of the above-mentioned embodiments of the semantic analysis-based scoring method are implemented.

[0101] The terminal device 5 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art can understand that Figure 5 merely an example of the terminal device 5, which does not constitute a limitation on the terminal device 5, and may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0102] The so-called processor 50 may be a Central Processing Unit (CPU), and the processor 50 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0103] In some embodiments, the memory 51 may be an internal storage unit of the terminal device 5, such as the hard disk or memory of the terminal device 5. In other embodiments, the memory 51 may also be an external storage device of the terminal device 5, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal device 5. Further, the memory 51 may also include both the internal storage unit and the external storage device of the terminal device 5. The memory 51 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 51 may also be used to temporarily store data that has been output or is to be output.

[0104] The embodiments of the present application further provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0105] The embodiments of the present application provide a computer program product, and when the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned various method embodiments when executed.

[0106] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0107] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0108] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0109] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0110] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0111] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A scoring method based on semantic analysis, characterized in that, it includes: Obtain the voice information of the target user, and convert the voice information into text information; Input the text information into the trained first neural network model, perform semantic analysis on the text information, and obtain the output text classification result of the first neural network model; wherein, the text classification result includes the scoring label corresponding to the text information, and the first neural network model is trained based on the training sample set and the second neural network model, and the second neural network model is trained based on the training sample set and the output result of the first neural network model, and the output result of the first neural network model is obtained by taking the training sample set as the input, and the training sample set includes multiple interview corpus texts; Calculate the interview scoring result of the target user according to the scoring label; The process of training the first neural network model based on the training sample set and the second neural network model includes: Divide the training samples into a short sentence set with a preset number of word segments, randomly mask one or more word segments in the short sentence set, and encode the masked short sentence set to obtain a word segment matrix, and use the word segment matrix as the input of the first neural network model; the first neural network model performs convolution calculation on the word segment matrix, and calculates an output matrix based on the parameter matrix; input the output matrix into the second neural network model, and the second neural network model performs bidirectional convolution calculation on the output matrix to output the probability matrix of each word segment masked in the output matrix; Take the cross-entropy loss between the predicted vector corresponding to the masked word segment in the output matrix and the true vector actually corresponding to the masked word as the first loss; take the cross-entropy loss corresponding to all masked word segments in the probability matrix as the second loss; based on the first loss and the second loss, perform interactive training on the first neural network model and the second neural network model according to the number of iterative training times, and adjust the parameter matrix of the first neural network model and the parameter matrix of the second neural network model to obtain the first target parameter matrix of the first neural network model and the second target parameter matrix of the second neural network model.

2. The method according to claim 1, characterized in that, the obtaining the voice information of the target user and converting the voice information into text information includes: Identify the voice information through a speech recognition algorithm, and extract the acoustic features in the voice information; Convert the voice information into text information according to the acoustic features.

3. The method according to claim 1, characterized in that, before inputting the text information into the trained first neural network model, it includes: Divide the text information according to the preset number of word segments to obtain at least one short sentence text that meets the preset number of word segments; Alternatively, during the process of converting the voice information into the text information, set the maximum number of short sentences, divide the voice information into at least one voice short sentence less than or equal to the maximum number of short sentences, and convert the at least one voice short sentence into the text information.

4. The method according to claim 1, wherein, before inputting the text information into the trained first neural network model, it includes: obtaining a training sample set, the training sample set including a plurality of interview corpus texts; dividing the sentence texts in the training sample set into a short sentence set with a preset number of word segments, and encoding the word segments in the short sentence set to obtain a word segment matrix; performing convolution calculation on the word segment matrix to obtain a target matrix, and taking the dot product of the target matrix and a parameter matrix as the output matrix of the first neural network; obtaining a predicted vector corresponding to the masked word segment in the output matrix, and calculating the cross-entropy loss between the predicted vector and the true vector actually corresponding to the masked word as the first loss.

5. The method according to claim 4, wherein, the method includes: after completing the training of the first neural network according to a preset number of iterative training times, based on the output matrix of the first neural network and the training sample set, perform iterative training on the second neural network model according to a preset number of training times for the second neural network model, and adjust the parameter matrix of the second neural network model.

6. A scoring device based on semantic analysis, wherein, it includes: an acquisition unit, configured to acquire the voice information of a target user and convert the voice information into text information; a processing unit, configured to input the text information into the trained first neural network model, perform semantic analysis on the text information, and obtain the output text classification result of the first neural network model; wherein, the text classification result includes a scoring label corresponding to the text information, the first neural network model is trained based on a training sample set and a second neural network model, the second neural network model is trained based on the training sample set and the output result of the first neural network model, the output result of the first neural network model is obtained by taking the training sample set as the input, and the training sample set includes a plurality of interview corpus texts; a scoring unit, configured to calculate the interview scoring result of the target user according to the scoring label. The processing unit is further configured to train the first neural network model based on the training sample set and the second neural network model. The training process includes: dividing the training samples into a short sentence set with a preset number of word segments, randomly masking one or more word segments in the short sentence set, and encoding the masked short sentence set to obtain a word segment matrix, and using the word segment matrix as the input of the first neural network model; the first neural network model performs convolution calculation on the word segment matrix and calculates an output matrix based on a parameter matrix; inputting the output matrix into the second neural network model, and the second neural network model performs bidirectional convolution calculation on the output matrix to output a probability matrix of the masking of each word segment in the output matrix. Taking the cross-entropy loss between the predicted vector corresponding to the masked word segment in the output matrix and the true vector corresponding to the actually masked word as the first loss; taking the cross-entropy loss corresponding to all masked word segments in the probability matrix as the second loss; based on the first loss and the second loss, performing interactive training on the first neural network model and the second neural network model according to the number of iterative training times, and adjusting the parameter matrix of the first neural network model and the parameter matrix of the second neural network model to obtain the first target parameter matrix of the first neural network model and the second target parameter matrix of the second neural network model.

7. A terminal device Characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the method described in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium storing a computer program Characterized in that when the computer program is executed by a processor, the method described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • AI model privacy protection method for resisting member reasoning attack based on adversarial sample

    CN110516812A

  • Interview answer text classification method and device, electronic equipment and storage medium

    CN110717023A