Data analysis method and device in answering scenario, electronic equipment and storage medium
By constructing a semantic network and causal influence graph of the response text, static, dynamic, and logical features are determined, solving the problem of incomplete evaluation in single-person interviews and realizing a comprehensive assessment of the respondent's abilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2023-08-31
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies are insufficient for efficiently, comprehensively, and fairly assessing a candidate's cognitive, thinking, and expressive abilities in a one-on-one interview, resulting in an inadequate comprehensive evaluation of interview performance.
By constructing a semantic network and causal influence diagram of the response text, static, dynamic, and logical features are determined. Combined with the index values of the analysis indicators, a comprehensive quantitative and qualitative assessment of the respondent's key abilities and performance is achieved.
It enables specific and hierarchical identification and assessment of the respondent's abilities, providing a more comprehensive evaluation of interview performance.
Smart Images

Figure CN117350273B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, and in particular to a data analysis method, apparatus, electronic device, and storage medium for answering scenarios. Background Technology
[0002] In individual interviews, candidates elaborate on the topics discussed, comprehensively assessing their multifaceted abilities. Past interviews often relied on scoring scales and expert / interviewer evaluations. This approach failed to efficiently, comprehensively, and fairly assess candidates' cognitive, thinking, and communication skills, making it difficult to provide a more holistic evaluation of their interview performance. Summary of the Invention
[0003] To address the problems existing in the prior art, the present invention provides a data analysis method, apparatus, electronic device, and storage medium for answering scenarios.
[0004] In a first aspect, the present invention provides a data analysis method for answering questions, comprising:
[0005] Obtain the response text, and determine the semantic network and causal influence graph based on the response text. The semantic network is a node network with words screened out from the response text as nodes and mutual information between words as connecting edges. The causal influence graph is a node graph with words and sentences screened out from the response text as nodes and causal relationships between words and sentences as connecting edges.
[0006] Static features are determined based on the semantic network and a preset reference semantic network. The static features include information characterizing the number of nodes in the semantic network and the connections between nodes.
[0007] Dynamic features are determined based on the semantic network, and the dynamic features include information characterizing directed paths between nodes in the semantic network;
[0008] Logical features are determined based on the causal influence diagram and a preset reference causal influence diagram. The logical features include information characterizing the directed connections between nodes in the causal influence diagram.
[0009] The index values of the analysis indicators are determined based on the static features, the dynamic features, and the logical features. The analysis results are determined based on the index values, and the analysis results include the results of multiple analysis items.
[0010] In one embodiment, determining static features based on the semantic network and a preset reference semantic network includes:
[0011] Count the number of nodes in the semantic network;
[0012] The first similarity distance between nodes is determined based on the mutual information between words in the semantic network and the reference semantic network, and the first similarity distance is the semantic similarity between words.
[0013] The semantic network is clustered to determine node clusters, and the modularity of the network is calculated based on the distance between each node in the node cluster and the first central node; the first central node is the center of the node cluster.
[0014] Calculate the PMI matrix distance between the semantic network and the reference semantic network;
[0015] Calculate the skewness of the node degree distribution in the semantic network; node degree is the number of edges connected to a node, and skewness is a statistic that characterizes the distribution.
[0016] In one embodiment, determining dynamic features based on the semantic network includes:
[0017] Multiple traversal paths are identified based on the semantic network, and the distance between adjacent nodes in the path is determined based on the shortest path among the multiple traversal paths.
[0018] The number of steps between nodes that meet the continuity condition, the number of times the current node is stopped, and the number of steps between nodes that meet the jump condition are determined based on the movement distance.
[0019] Based on the movement direction of the nodes in the walk path, determine the number of nodes that satisfy the regressivity condition and the number of nodes that satisfy the divergence condition.
[0020] In one embodiment, determining the logical features based on the causal influence diagram and a preset reference causal influence diagram includes:
[0021] Multiple causal paths are determined based on the aforementioned causal influence diagram;
[0022] Determine the words and phrases corresponding to the nodes in the causal path that match the words and phrases in the reference causal influence map, and determine the corresponding causal path in the reference causal image map based on the matching words and phrases;
[0023] Based on the reference causal influence diagram, determine the causal relationship information between the words and phrases in the corresponding two causal paths;
[0024] The error rate of determining the causal path and the effective argumentation rate of the causal path for the answer topic are determined based on the causal relationship information.
[0025] In one embodiment, the method further includes:
[0026] Obtain the audio file corresponding to the answer text;
[0027] The audio file is analyzed and divided into multiple audio segments;
[0028] Determine the emotional characteristics of each speech segment, and determine the positive / negative results and the persuasiveness of each speech segment based on the emotional characteristics.
[0029] In one embodiment, the method further includes:
[0030] Get the corresponding scene image of the answer text;
[0031] The scene of the response is identified to obtain the morphological features of the respondent in the scene;
[0032] The degree to which the respondent in the picture conforms to etiquette norms is determined based on the morphological characteristics described.
[0033] In one embodiment, determining the analysis result based on the indicator value includes:
[0034] Obtain the weights of each analysis indicator for each sub-item;
[0035] The analysis results of the analysis items are determined based on the index values and weights of the analysis indicators.
[0036] Secondly, the present invention provides a data analysis device for answering questions, comprising:
[0037] A construction module is used to obtain the response text and determine the semantic network and causal influence graph based on the response text. The semantic network is a node network with words screened out from the response text as nodes and the connection edges are determined based on the mutual information between words. The causal influence graph is a node graph with words and sentences screened out from the response text as nodes and the causal relationship between words and sentences as connection edges.
[0038] The first acquisition module is used to determine static features based on the semantic network and a preset reference semantic network. The static features include information characterizing the number of nodes in the semantic network and the connections between nodes.
[0039] The second acquisition module is used to determine dynamic features based on the semantic network, wherein the dynamic features include information characterizing directed paths between nodes in the semantic network;
[0040] The third acquisition module is used to determine logical features based on the causal influence diagram and the preset reference causal influence diagram. The logical features include information characterizing the directed connections between nodes in the causal influence diagram.
[0041] The analysis module is used to determine the index values of the analysis indicators based on the static features, the dynamic features, and the logical features, and to determine the analysis results based on the index values. The analysis results include the results of multiple analysis items.
[0042] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the data analysis method in the answering scenario as described in any one of the first aspects.
[0043] Fourthly, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the data analysis method in the answering scenario as described in any one of the first aspects.
[0044] The data analysis method, apparatus, electronic device, and storage medium provided by this invention for answering scenarios determine semantic networks and causal influence diagrams based on the answer text, and obtain different features based on the semantic networks and causal influence diagrams formed by the answer, as well as reference semantic networks and causal influence diagrams. Based on the different features, the corresponding analysis index values are determined, and the analysis results are determined based on the index values. This enables specific and hierarchical identification of the respondent's key abilities and performance, and completes a comprehensive quantitative and qualitative assessment of the respondent's performance. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0046] Figure 1 This is a flowchart illustrating the data analysis method for answering questions provided by the present invention.
[0047] Figure 2 This is a schematic diagram of the data analysis device for answering questions provided by the present invention;
[0048] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0050] The following is combined Figures 1 to 3This invention describes the data analysis method, apparatus, electronic device, and storage medium provided for answering scenarios.
[0051] Figure 1 This diagram illustrates a data analysis method for a response scenario provided by the present invention. (See attached diagram.) Figure 1 The method includes:
[0052] 11. Obtain the response text, and determine the semantic network and causal influence graph based on the response text. The semantic network is a node network with words screened out from the response text as nodes and the connection edges are determined based on the mutual information between words. The causal influence graph is a node graph with words and sentences screened out from the response text as nodes and the causal relationship between words and sentences as connection edges.
[0053] 12. Determine static features based on the semantic network and the preset reference semantic network. Static features include information representing the number of nodes in the semantic network and the connections between nodes.
[0054] 13. Determine dynamic features based on the semantic network. Dynamic features include information representing directed paths between nodes in the semantic network.
[0055] 14. Determine the logical features based on the causal influence diagram and the preset reference causal influence diagram. The logical features include information that characterizes the directed connections between nodes in the causal influence diagram.
[0056] 15. Determine the index values of the analysis indicators based on static, dynamic, and logical characteristics, and determine the analysis results based on the index values. The analysis results include the results of multiple analysis items.
[0057] Regarding steps 11 to 15, it should be noted that in this invention, in reality, a person's various aspects are tested and observed to determine whether that person meets the organization's talent requirements. For example, a company conducts a question-and-answer session with an interviewee, and the interviewee provides answers.
[0058] In this invention, it is necessary to analyze the respondent's answers to evaluate their abilities in various aspects. The respondent's answers can be obtained from various sources, such as directly edited text, audio recordings of the answer, and video files of the answer.
[0059] In this invention, respondents will provide their own interpretations of certain topics and viewpoints raised by the organizers. The content of these statements can be used by the organizers to assess their abilities. Therefore, the statements can serve as the text for analysis. This text can be directly edited text or extracted from a recording.
[0060] In this invention, the response text is a statement of the proposed topic viewpoint. Therefore, the response text may be a rigorous statement around the topic viewpoint, or it may be a divergent statement that doesn't quite grasp the topic viewpoint, and so on. Different statements are related to the respondent's understanding of the topic viewpoint and their ability to perform under pressure. Therefore, some words and phrases in the response text can reflect the relationship between the topic viewpoint and the respondent's performance. Thus, a semantic network can be constructed. The response text is preprocessed, converted into a word segmentation format, and key words are identified. Each word is mapped to a node in the semantic network, and the nodes in the semantic network are connected based on the mutual information between words (here, the connection edge corresponds to the connection information between words). This forms a semantic network.
[0061] In this invention, the semantic network only contains contextual co-occurrence information and not causal information. However, when a respondent says a certain word or phrase (i.e., a sentence or paragraph), it may be laying the groundwork for one or more subsequent words or phrases. That is, a certain word or phrase has a causal relationship with subsequent words or phrases. Therefore, multiple words or phrases with this causal relationship need to be analyzed to construct a causal influence graph. This causal influence graph is a node graph with words or phrases screened from the response text as nodes and causal relationships between words or phrases as connecting edges.
[0062] In this invention, static features are determined based on a semantic network and a pre-defined reference semantic network. These static features are obtained by analyzing the topological characteristics of the semantic network and can reflect the overall global semantic features of the response text. Analyzing the topological characteristics of the semantic network and the pre-defined reference semantic network yields information such as the number of nodes in the semantic network and the connections between nodes. The pre-defined reference semantic network is constructed by analyzing a large number of standard response texts (i.e., reference answers). Specifically, a text corpus with a large number of argumentative texts on various topics is established, and the TF-IDF statistics of words corresponding to semantic network nodes are calculated as the importance of the word in the response context. Texts in the response texts where important nodes (words) appear comprehensively are selected as the content / opinion base for candidate reference answers. Reference answers can be obtained semi-automatically through expert review. The reference answers are preprocessed in a similar manner as described above to obtain the reference semantic network and the reference causal influence graph.
[0063] Based on different acquisition needs, the number of nodes and the information on connections between nodes can be mapped to various analytical indicators for different analytical items. For example, a sub-item called "Comprehensive Analysis" includes analytical indicators such as breadth of thinking, depth of thinking, focus of argumentation, and thematic relevance. In other words, the results of each analytical indicator can be calculated based on the number of nodes and the information on connections between nodes.
[0064] In this invention, dynamic features are determined based on a semantic network. These dynamic features are obtained by analyzing the dynamic processes of words corresponding to nodes on the semantic network, reflecting the local characteristics and dynamic transition properties of the response text. Analyzing the words corresponding to nodes in the semantic network yields information such as directed paths between nodes. Analyzing these directed paths allows for the identification of the following situations: staying at the current node (manifested as repetition of the same word), random walks on the semantic network (semantic continuity), or jumping to disconnected parts of the network. By reasonably calculating the number and distance of nodes on the paths in the above situations, and based on reasonable statistical model assumptions, the corresponding analytical indicators are obtained. For example, if the analytical item is adaptability, the analytical indicators include: coherence of thought, agility of thought, and the characteristics of thought jumps / coherence, etc.
[0065] In this invention, logical features are determined based on the causal influence diagram and a preset reference causal influence diagram. These logical features are obtained by analyzing the causal influence diagram and the reference causal influence diagram and can reflect the characteristics related to decision-making and coordination ability in the respondent's argument.
[0066] A comprehensive analysis of the causal impact diagram and the reference causal impact diagram yields information such as the directed connections between nodes in the causal impact diagram. The pre-defined reference causal impact diagram is constructed by analyzing a large number of standard answer texts (i.e., reference answers). Through the directed connections between nodes in the causal impact diagram, the error rate of the causal path and the effective validation rate of the causal path for the answer topic can be determined. These error rates and effective validation rates reflect the results of the analysis indicators. For example, if the analysis item is organization and coordination, the included analysis indicators include: decision-making ability, coordination ability, etc.
[0067] In this invention, the index values of analytical indicators are determined based on static, dynamic, and logical characteristics. The analytical results are then determined based on these index values, and the results include the results of multiple analytical items. Here, a preset algorithm can be used to calculate the index values of the analytical indicators based on the static, dynamic, and logical characteristics. This preset algorithm can be a function formula or a numerical rule for judgment conditions, etc.
[0068] The analytical items and analytical indicators mentioned in this invention are shown in Table 1 below.
[0069] Table 1 shows the assessment framework for answering abilities: ability name, actual performance, and measurement indicators.
[0070]
[0071]
[0072]
[0073]
[0074] In this invention, determining the analysis result based on the index value includes:
[0075] Obtain the weights of each analysis indicator for each sub-item;
[0076] The analysis results of the analysis items are determined based on the index values and weights of the analysis indicators.
[0077] Based on the analysis results, a relatively reasonable assessment of the respondent's ability can be made.
[0078] The data analysis method for answering scenarios provided by this invention determines the semantic network and causal influence diagram based on the answer text, and obtains different features based on the semantic network and causal influence diagram formed by the answer, as well as reference semantic network and causal influence diagram. Based on the different features, the corresponding analysis index values are determined, and the analysis results are determined based on the index values. This achieves specific and hierarchical identification of the respondent's key abilities and performance, and completes a comprehensive quantitative and qualitative assessment of the respondent's performance.
[0079] In a further step of the above method, the explanation of the process for determining static features based on the semantic network and a preset reference semantic network is as follows:
[0080] The number of nodes in the semantic network is counted. This involves statistically analyzing all nodes in the semantic network obtained from the responses to determine the total number of nodes. The number of nodes in the semantic network reflects the breadth of thought in the comprehensive analysis: more nodes indicate more keywords discussed and repeated, covering a wider range of topics, which can be correlated with metrics such as topic coverage.
[0081] The first similarity distance between nodes is determined based on the mutual information between words in the semantic network and the reference semantic network. This first similarity distance is the semantic similarity between words. Mutual information, a useful information metric in information theory, refers to the relevance between two sets of events; here, it is a numerical value. The semantic similarity between words is calculated based on mutual information. This similarity can be linked to the topic fit analysis metric.
[0082] The semantic network is clustered to identify node clusters, and the distance between each node in the cluster and the first central node is calculated; the first central node is the center of the node cluster. In other words, network community discovery is performed on the semantic network using the Louvain algorithm, which yields node clusters to help determine the topics discussed and, to some extent, the structural strength of the respondents' answers. This distance can be correlated with the topic modularity metric, meaning that the network's modularity can be calculated using this distance.
[0083] Calculate the PMI matrix distance between the semantic network and the reference semantic network. The PMI matrix distance between the respondent's semantic network and the reference semantic network (normalized using the number of nodes in the respondent's semantic network) can be correlated with an analytical indicator of the degree of semantic cognitive deviation.
[0084] The degree of each node in the semantic network is calculated, which is the number of edges connected to that node. Then, the skewness of the degree distribution for each node is calculated; skewness is a statistical measure of the distribution's characteristics. Skewness measures the direction and degree of skewness in the distribution of statistical data, representing the degree of asymmetry in the distribution. This skewness indicates the density of the network and can be correlated with the analytical indicator of associative ability.
[0085] In a further method described above, the process of determining dynamic features based on semantic networks is explained in detail below:
[0086] Multiple traversal paths are identified based on the semantic network, and the distance between adjacent nodes in the path is determined based on the shortest path among the multiple traversal paths.
[0087] The number of steps to travel between nodes that meet the continuity condition, the number of times to stay at the current node, and the number of steps to travel between nodes that meet the jump condition are determined based on the distance.
[0088] Based on the movement direction of nodes in the walk path, determine the number of nodes that satisfy the regressivity condition and the number of nodes that satisfy the divergence condition.
[0089] It should be noted that in this invention, a walking path is constructed by utilizing the movement process of the respondent's argument on the topic viewpoint on the semantic network. This walking path can be decomposed into features such as walking direction and speed. The walking direction can be referenced to the direction between nodes in the walking path. The walking speed can be referenced to the distance moved between adjacent nodes in the walking path. Multiple walking paths that can express a certain topic viewpoint may be identified based on the semantic network. Therefore, the shortest path must be found from these multiple walking paths. This shortest path can simply and clearly express the topic idea, and then the distance between adjacent nodes in the path is determined based on the shortest path.
[0090] In this invention, the number of steps between nodes satisfying the continuity condition, the number of stops at the current node, and the number of steps between nodes satisfying the jump condition are determined based on the movement distance. It should be noted that the continuity of the respondent's description refers to the description's ability to grasp the main point and its smooth, continuous expression. For this purpose, a distance range can be set. Two adjacent nodes within this range can be counted as nodes satisfying the continuity condition. When multiple consecutive nodes are identified, the number of steps between nodes can be counted; for example, if there are 6 nodes, the number of steps is 5.
[0091] Correspondingly, when the movement distance is 0, it indicates that the respondent paused in their description, that is, paused within the same word. The node corresponding to that word is the node that satisfies the pause condition.
[0092] Correspondingly, when the movement distance exceeds the above-mentioned distance range, it indicates that the respondent's description is a jumpy content. In this case, two adjacent nodes or the latter of two adjacent nodes are considered nodes that meet the jumpy condition. After identifying multiple jumpy nodes, the number of steps between nodes can be counted. For example, if there are 5 nodes, the number of steps is 4.
[0093] By using the above processing method, we can determine the number of steps between nodes that meet the continuity condition, the number of times the current node stays, and the number of steps between nodes that meet the jump condition.
[0094] This invention extends the search bias function of random walks in the node2vec algorithm, characterizing the directional bias of the walk path through parameterization. Specifically, the search direction can be characterized by the average distance between nodes over multiple steps during the walk: the respondent's semantic expression may tend to revolve around a central argument (returning to the same node multiple times on the network), that is, based on the movement direction of nodes in the walk path, the number of nodes satisfying the regressive condition and the number of nodes satisfying the divergent condition are determined.
[0095] A further method described above primarily explains the process of determining logical features based on the causal influence diagram and a pre-defined reference causal influence diagram, as detailed below:
[0096] Multiple causal paths were identified based on the causal impact diagram;
[0097] Identify the words and phrases corresponding to the nodes in the causal path and match them with the words and phrases in the reference causal influence map. Based on the matching words and phrases, determine the corresponding causal path in the reference causal image map.
[0098] Based on the reference causal influence diagram, determine the causal relationship information between the words and phrases in the corresponding two causal paths;
[0099] The error rate of determining causal paths based on causal relationship information, and the effective argumentation rate of causal paths for answering the topic.
[0100] It should be noted that in this invention, a comprehensive analysis of the causal influence diagram and a reference causal influence diagram is performed to determine multiple causal paths. The words corresponding to nodes in each causal path are then identified and matched with words in the reference causal influence diagram. Based on these matching words, the corresponding causal path is determined in the reference causal image diagram. This allows for the calculation of the causal relationship information between words in the corresponding two causal paths, which represents the similarity deflection angle between the two causal paths. By comparing the similarity deflection angle and its range, the error rate of the causal path and the effective verification rate of the causal path for the answer topic can be determined. These error rates and effective verification rates reflect the results of the analysis indicators. For example, if the analysis item is organization and coordination, the included analysis indicators include: decision-making ability, coordination ability, etc.
[0101] In a further step of the above method, the respondent makes their statements on the spot. Their statements on certain topics may be emotionally charged, while those on others may be overly calm. Therefore, it is also necessary to judge the positiveness or persuasiveness of the statements from the perspective of the respondent's emotions. Thus, an audio file corresponding to the text of the response is obtained, which can include changes in the timbre and volume of the respondent's statements.
[0102] The audio file is analyzed and divided into multiple audio segments;
[0103] Determine the emotional characteristics of each speech segment, and based on these characteristics, determine the positive / negative results and the impact of each speech segment.
[0104] Using Plutchik's emotion wheel as a basic framework, and based on an emotion lexicon obtained from crowdsourced data, a multiple linear regression model of emotion is established. Subsequently, based on the audio files corresponding to the original response text, the model identifies the respondent's emotional inclination and fluctuations during the response process, analyzes the persuasiveness of the candidate's language expression, and the level of tension during the response. A sentiment analysis module can also be added to provide more information for comprehensive evaluation.
[0105] In a further step of the above method, different answering scenarios require different approaches from the respondent. For example, in an interview scenario, attention might be paid to appearance. In a presentation scenario, demonstrating professionalism might be emphasized. Therefore, it is also necessary to assess the respondent's compliance with etiquette norms. This requires obtaining the corresponding scene image of the answer text, identifying the morphological features of the respondent in the scene, and determining the respondent's compliance with etiquette norms based on these features.
[0106] The data analysis device for answering scenarios provided by the present invention is described below. The data analysis device for answering scenarios described below and the data analysis method for answering scenarios described above can be referred to and correspond to each other.
[0107] Figure 2 This diagram illustrates the structure of a data analysis device for a response scenario provided by the present invention. (See attached diagram.) Figure 2 The device includes a construction module 21, a first acquisition module 22, a second acquisition module 23, a third acquisition module 24, and an analysis module 25, wherein:
[0108] Module 21 is used to obtain the response text and determine the semantic network and causal influence graph based on the response text. The semantic network is a node network with words screened out from the response text as nodes and the connection edges are determined based on the mutual information between words. The causal influence graph is a node graph with words and sentences screened out from the response text as nodes and the causal relationship between words and sentences as connection edges.
[0109] The first acquisition module 22 is used to determine static features based on the semantic network and the preset reference semantic network. The static features include information representing the number of nodes in the semantic network and the connections between nodes.
[0110] The second acquisition module 23 is used to determine dynamic features based on the semantic network. The dynamic features include information representing directed paths between nodes in the semantic network.
[0111] The third acquisition module 24 is used to determine logical features based on the causal influence diagram and the preset reference causal influence diagram. The logical features include information that characterizes the directed connections between nodes in the causal influence diagram.
[0112] Analysis module 25 is used to determine the index values of analysis indicators based on static characteristics, dynamic characteristics and logical characteristics, and to determine the analysis results based on the index values. The analysis results include the results of multiple analysis items.
[0113] In a further embodiment of the above-described apparatus, the first acquisition module is specifically used for:
[0114] Count the number of nodes in a semantic network;
[0115] The first similarity distance between nodes is determined based on the mutual information between words in the semantic network and the reference semantic network. The first similarity distance is the semantic similarity between words.
[0116] The semantic network is clustered to determine node clusters, and the modularity of the network is calculated based on the distance between each node in the node cluster and the first central node; the first central node is the center of the node cluster.
[0117] Calculate the PMI matrix distance between the semantic network and the reference semantic network;
[0118] Calculate the skewness of the node degree distribution in the semantic network; node degree is the number of edges connected to a node, and skewness is a statistic that characterizes the distribution.
[0119] In a further embodiment of the above-described apparatus, the second acquisition module is specifically used for:
[0120] Multiple traversal paths are identified based on the semantic network, and the distance between adjacent nodes in the path is determined based on the shortest path among the multiple traversal paths.
[0121] The number of steps between nodes that meet the continuity condition, the number of times the current node is stopped, and the number of steps between nodes that meet the jump condition are determined based on the movement distance.
[0122] Based on the movement direction of the nodes in the walk path, determine the number of nodes that satisfy the regressivity condition and the number of nodes that satisfy the divergence condition.
[0123] In a further embodiment of the above-described apparatus, the third acquisition module is specifically used for:
[0124] Multiple causal paths are determined based on the aforementioned causal influence diagram;
[0125] Determine the words and phrases corresponding to the nodes in the causal path that match the words and phrases in the reference causal influence map, and determine the corresponding causal path in the reference causal image map based on the matching words and phrases;
[0126] Based on the reference causal influence diagram, determine the causal relationship information between the words and phrases in the corresponding two causal paths;
[0127] The error rate of determining causal paths based on causal relationship information, and the effective argumentation rate of causal paths for answering the topic.
[0128] In a further embodiment of the above-described device, the device further includes a fourth acquisition module, used for:
[0129] Obtain the audio file corresponding to the answer text;
[0130] The audio file is analyzed and divided into multiple audio segments;
[0131] Determine the emotional characteristics of each speech segment, and determine the positive / negative results and the persuasiveness of each speech segment based on the emotional characteristics.
[0132] In a further embodiment of the above-described device, the device further includes a fifth acquisition module, used for:
[0133] Get the corresponding scene image of the answer text;
[0134] The scene of the response is identified to obtain the morphological features of the respondent in the scene;
[0135] The degree to which the respondent in the picture conforms to etiquette norms is determined based on the morphological characteristics described.
[0136] In a further embodiment of the above apparatus, the analysis module is specifically used for:
[0137] Obtain the weights of each analysis indicator for each sub-item;
[0138] The analysis results of the analysis items are determined based on the index values and weights of the analysis indicators.
[0139] Since the device described in this embodiment of the invention is based on the same principle as the method described in the above embodiments, more detailed explanations will not be repeated here.
[0140] The data analysis device for answering scenarios provided by this invention determines semantic networks and causal influence diagrams based on the answer text, and obtains different features based on the semantic networks and causal influence diagrams formed by the answer, as well as reference semantic networks and causal influence diagrams. Based on the different features, it determines the index values of corresponding analysis indicators, and determines the analysis results based on the index values. This enables the specific and hierarchical identification of the respondent's key abilities and performance, and completes a comprehensive quantitative and qualitative assessment of the respondent's performance.
[0141] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include a processor 31, a communication interface 32, a memory 33, and a communication bus 34. The processor 31, communication interface 32, and memory 33 communicate with each other via the communication bus 34. The processor 31 can call a computer program in the memory 33 to execute steps of a data analysis method in a response scenario, such as: acquiring the response text; determining a semantic network and a causal influence graph based on the response text; the semantic network is a node network with words selected from the response text as nodes and connection edges determined based on mutual information between words; the causal influence graph is a node graph with phrases selected from the response text as nodes and causal relationships between phrases as connection edges.
[0142] Static features are determined based on the semantic network and a pre-defined reference semantic network. Static features include information representing the number of nodes in the semantic network and the connections between nodes.
[0143] Dynamic features are determined based on the semantic network, and these dynamic features include information representing directed paths between nodes in the semantic network.
[0144] Logical features are determined based on the causal influence diagram and the pre-set reference causal influence diagram. The logical features include information that characterizes the directed connections between nodes in the causal influence diagram.
[0145] The index values of the analysis indicators are determined based on static, dynamic, and logical characteristics. The analysis results are then determined based on the index values, and the analysis results include the results of multiple analysis items.
[0146] Furthermore, the logical instructions in the aforementioned memory 33 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0147] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, when the program instructions are executed by the computer, the computer is able to perform the steps of a data analysis method in a response scenario, such as: acquiring response text, determining a semantic network and a causal influence graph based on the response text, the semantic network being a node network with words screened out in the response text as nodes and connection edges determined based on mutual information between words, and the causal influence graph being a node graph with words and sentences screened out in the response text as nodes and causal relationships between words and sentences as connection edges;
[0148] Static features are determined based on the semantic network and a pre-defined reference semantic network. Static features include information representing the number of nodes in the semantic network and the connections between nodes.
[0149] Dynamic features are determined based on the semantic network, and these dynamic features include information representing directed paths between nodes in the semantic network.
[0150] Logical features are determined based on the causal influence diagram and the pre-set reference causal influence diagram. The logical features include information that characterizes the directed connections between nodes in the causal influence diagram.
[0151] The index values of the analysis indicators are determined based on static, dynamic, and logical characteristics. The analysis results are then determined based on the index values, and the analysis results include the results of multiple analysis items.
[0152] On the other hand, embodiments of the present invention also provide a processor-readable storage medium storing a computer program. The computer program is used to cause the processor to execute steps of a data analysis method in a response scenario, such as: acquiring response text, determining a semantic network and a causal influence graph based on the response text, wherein the semantic network is a node network with words screened out from the response text as nodes and connection edges determined based on mutual information between words, and the causal influence graph is a node graph with words and sentences screened out from the response text as nodes and causal relationships between words and sentences as connection edges;
[0153] Static features are determined based on the semantic network and a pre-defined reference semantic network. Static features include information representing the number of nodes in the semantic network and the connections between nodes.
[0154] Dynamic features are determined based on the semantic network, and these dynamic features include information representing directed paths between nodes in the semantic network.
[0155] Logical features are determined based on the causal influence diagram and the pre-set reference causal influence diagram. The logical features include information that characterizes the directed connections between nodes in the causal influence diagram.
[0156] The index values of the analysis indicators are determined based on static, dynamic, and logical characteristics. The analysis results are then determined based on the index values, and the analysis results include the results of multiple analysis items.
[0157] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0158] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0159] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data analysis method for answering questions, characterized in that, include: Obtain the response text, and determine the semantic network and causal influence graph based on the response text. The semantic network is a node network with words screened out from the response text as nodes and connection edges determined based on mutual information between words. The causal influence graph is a node graph with words and sentences screened out from the response text as nodes and causal relationships between words and sentences as connection edges. Static features are determined based on the semantic network and a preset reference semantic network. The static features include information characterizing the number of nodes in the semantic network and the connections between nodes. Dynamic features are determined based on the semantic network, including information representing directed paths between nodes in the semantic network; a walking path is constructed by utilizing the movement process of the respondent's argument on the topic viewpoint on the semantic network, and the walking path can be decomposed into the walking direction, which refers to the direction between nodes in the walking path. Logical features are determined based on the causal influence diagram and a preset reference causal influence diagram. The logical features include information characterizing the directed connections between nodes in the causal influence diagram. The index values of the analysis indicators are determined based on the static characteristics, the dynamic characteristics, and the logical characteristics. The analysis results are determined based on the index values, and the analysis results include the results of multiple analysis items. The step of determining static features based on the semantic network and a preset reference semantic network includes: Count the number of nodes in the semantic network; The first similarity distance between nodes is determined based on the mutual information between words in the semantic network and the reference semantic network, and the first similarity distance is the semantic similarity between words. The semantic network is clustered to determine node clusters, and the modularity of the network is calculated based on the distance between each node in the node cluster and the first central node; the first central node is the center of the node cluster. Calculate the PMI matrix distance between the semantic network and the reference semantic network; Calculate the skewness of the node degree distribution in the semantic network; node degree is the number of edges connected to a node, and skewness is a statistic that characterizes the distribution. The step of determining the logical features based on the causal influence diagram and the preset reference causal influence diagram includes: Multiple causal paths are determined based on the aforementioned causal influence diagram; Determine the words and phrases corresponding to the nodes in the causal path that match the words and phrases in the reference causal influence map, and determine the corresponding causal path in the reference causal image map based on the matching words and phrases; Based on the reference causal influence diagram, determine the causal relationship information between the words and phrases in the corresponding two causal paths; The error rate of determining the causal path and the effective argumentation rate of the causal path for the answer topic are determined based on the causal relationship information.
2. The data analysis method for answering scenarios according to claim 1, characterized in that, The step of determining dynamic features based on the semantic network includes: Multiple traversal paths are identified based on the semantic network, and the distance between adjacent nodes in the path is determined based on the shortest path among the multiple traversal paths. The number of steps between nodes that satisfy the continuity condition, the number of times the current node is stopped, and the number of steps between nodes that satisfy the jump condition are determined based on the distance. Based on the movement direction of the nodes in the walk path, determine the number of nodes that satisfy the regressivity condition and the number of nodes that satisfy the divergence condition.
3. The data analysis method for answering scenarios according to claim 1, characterized in that, The method further includes: Obtain the audio file corresponding to the answer text; The audio file is analyzed and divided into multiple audio segments; Determine the emotional characteristics of each speech segment, and determine the positive / negative results and the persuasiveness of each speech segment based on the emotional characteristics.
4. The data analysis method for answering scenarios according to claim 3, characterized in that, The method further includes: Get the corresponding scene image of the answer text; The scene of the response is identified to obtain the morphological features of the respondent in the scene; The degree to which the respondent conforms to etiquette norms is determined based on the aforementioned morphological characteristics.
5. The data analysis method for answering scenarios according to claim 1 or 4, characterized in that, The analysis results are determined based on the aforementioned indicator values, including: Obtain the weights of each analytical indicator for each analytical item; The analysis results of the analysis items are determined based on the index values and weights of the analysis indicators.
6. A data analysis device for answering questions, characterized in that, include: A construction module is used to obtain the response text and determine the semantic network and causal influence graph based on the response text. The semantic network is a node network with words screened out from the response text as nodes and the connection edges are determined based on the mutual information between words. The causal influence graph is a node graph with words and sentences screened out from the response text as nodes and the causal relationship between words and sentences as connection edges. The first acquisition module is used to determine static features based on the semantic network and a preset reference semantic network. The static features include information characterizing the number of nodes in the semantic network and the connections between nodes. Specifically, it is used for: counting the number of nodes in the semantic network; determining the first similarity distance between nodes based on the mutual information between words in the semantic network and the reference semantic network, wherein the first similarity distance is the semantic similarity between words; clustering the semantic network to determine node clusters, and calculating the modularity of the network based on the distance between each node in the node cluster and the first central node; wherein the first central node is the center of the node cluster. Calculate the PMI matrix distance between the semantic network and the reference semantic network; calculate the skewness of the node degree distribution in the semantic network; node degree is the number of edges connected to a node, and skewness is a statistic that characterizes the distribution characteristics; The second acquisition module is used to determine dynamic features based on the semantic network. The dynamic features include information representing directed paths between nodes in the semantic network. The module constructs a walking path by utilizing the movement process of the respondent's argument on the topic viewpoint on the semantic network. The walking path can be decomposed into the direction of walking, and the direction of walking refers to the direction between nodes in the walking path. The third acquisition module is used to determine logical features based on the causal influence diagram and the preset reference causal influence diagram. The logical features include information characterizing the directed connections between nodes in the causal influence diagram. Specifically, it is used to: determine multiple causal paths based on the causal influence diagram; determine the words corresponding to the nodes in the causal path and the words that match in the reference causal influence diagram, and determine the corresponding causal path in the reference causal image diagram based on the matching words; Based on the reference causal influence diagram, determine the causal relationship information between the words and phrases in the corresponding two causal paths; The error rate of the causal path and the effective argumentation rate of the causal path to the answer topic are determined based on the causal relationship information. The analysis module is used to determine the index values of the analysis indicators based on the static features, the dynamic features, and the logical features, and to determine the analysis results based on the index values. The analysis results include the results of multiple analysis items.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the data analysis method for the answering scenario as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the data analysis method in the answering scenario as described in any one of claims 1 to 5.