Question and Answer Processing Method, Device, Electronic Device and Storage Medium
Through dependency syntax analysis and high-frequency dependency subgraph matching, the word segmentation type in the target question is determined and standard problems are generated, which solves the difficulty of dealing with synonyms and multi-category label problems in user input problems in the prior art, and improves the accuracy and coverage of problem standardization.
Patent Information
- Application Number
- CN202111290391.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-02
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-02
AI Technical Summary
Existing table question and answer techniques are difficult to effectively deal with attribute synonyms or attribute value alias in questions entered by users, and when the number of category labels is large, the prediction accuracy of the model is low.
By obtaining the target question, performing dependency syntax analysis, matching the dependency relationship in the high-frequency dependency subgraph, determining the type of the target word segmentation, and determining the standard question based on the type to obtain the reply answer.
It improves the coverage and accuracy of problem standardization, is suitable for problem matching in different scenarios, and alleviates the pressure of cold-start problem standardization analysis.
Smart Images

Figure CN114138951B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the fields of intelligent search, natural language processing, and knowledge graph technology, and particularly to a question and answer processing method, apparatus, electronic device, and storage medium. Background Art
[0002] In recent years, with the rapid development of artificial intelligence, table question and answer has received increasing attention. Table question and answer (TableQA) is a technology that asks questions based on existing table knowledge to obtain accurate answers. Currently, table question and answer mainly standardizes the questions input by users, and then queries the database to obtain the final answers. Summary of the Invention
[0003] The present disclosure provides a question and answer processing method, apparatus, electronic device, and storage medium.
[0004] According to one aspect of the present disclosure, there is provided a question and answer processing method, including: obtaining a target question sentence; performing dependency syntactic analysis on the target question sentence to obtain the dependency relationships between the respective participles in the target question sentence; matching the dependency relationships between the respective participles in the target question sentence with the dependency relationships between the respective participle types in a set high-frequency dependency subgraph; determining the types of the respective target participles corresponding to the target question sentence according to the participle types in the high-frequency dependency subgraph that match the dependency relationships; and determining a standard question corresponding to the target question sentence according to the types of the target participles to determine a reply answer.
[0005] According to another aspect of the present disclosure, there is provided a question and answer processing apparatus, including: a first obtaining module for obtaining a target question sentence; an analysis module for performing dependency syntactic analysis on the target question sentence to obtain the dependency relationships between the respective participles in the target question sentence; a matching module for matching the dependency relationships between the respective participles in the target question sentence with the dependency relationships between the respective participle types in a set high-frequency dependency subgraph; a first determining module for determining the types of the respective target participles corresponding to the target question sentence according to the participle types in the high-frequency dependency subgraph that match the dependency relationships; and a second determining module for determining a standard question corresponding to the target question sentence according to the types of the target participles to determine a reply answer.
[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to the first aspect embodiment of the present disclosure.
[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the embodiment of the first aspect of the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the method described in the embodiment of the first aspect of the present disclosure are implemented.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0011] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0012] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;
[0013] Figure 3 is a schematic diagram of a dependency graph according to an embodiment of the present disclosure;
[0014] Figure 4 is a schematic diagram of nodes of a high-frequency dependency sub-graph according to an embodiment of the present disclosure;
[0015] Figure 5 is a schematic diagram of the corresponding relationship between a target question sentence and a high-frequency dependency sub-graph according to an embodiment of the present disclosure;
[0016] Figure 6 is a schematic diagram according to the third embodiment of the present disclosure;
[0017] Figure 7 is a schematic diagram of tabular data according to an embodiment of the present disclosure;
[0018] Figure 8 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0019] Figure 9 is a schematic diagram according to the fifth embodiment of the present disclosure;
[0020] Figure 10 is a schematic diagram of a high-frequency dependency sub-graph and the probability of corresponding nodes according to an embodiment of the present disclosure;
[0021] Figure 11 is a schematic diagram according to the sixth embodiment of the present disclosure;
[0022] Figure 12 It is a block diagram of an electronic device for implementing the question - answering processing method of the present disclosure embodiment. Specific embodiments
[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well - known functions and structures are omitted below for clarity and conciseness.
[0024] In recent years, with the rapid development of artificial intelligence, table question - answering has received increasing attention. Table question - answering (TableQA) is a technology that asks questions based on existing table knowledge and obtains accurate answers. Currently, table question - answering mainly standardizes the questions input by users, and then queries the database to obtain the final answers.
[0025] In related technologies, dictionary matching combined with model prediction is mainly used for named - entity recognition to standardize the questions input by users. However, it is difficult for the dictionary to fully cover the attribute synonyms or attribute - value aliases in the questions input by users, and when the number of category labels corresponding to the questions input by users is large, the prediction accuracy of the model is low.
[0026] Therefore, in view of the above - mentioned problems, the present disclosure proposes a question - answering processing method, device, electronic device, and storage medium.
[0027] Figure 1 It is a schematic diagram according to the first embodiment of the present disclosure. It should be noted that the question - answering processing method of the embodiments of the present disclosure can be applied to the question - answering processing device of the embodiments of the present disclosure, and this device can be configured in an electronic device. Among them, the electronic device can be a mobile terminal, for example, a hardware device with various operating systems such as a mobile phone, a tablet computer, a personal digital assistant, etc.
[0028] As Figure 1 shown, the question - answering processing method may include the following steps:
[0029] Step 101, obtain a target question sentence.
[0030] In the embodiments of the present disclosure, the target question sentence can be collected online. For example, through web - crawler technology, sentences with an interrogative structure can be collected online as the target question sentence. Or, the target question sentence can also be a sentence with an interrogative structure collected offline, or the target question sentence can also be an artificially synthesized question sentence, etc. The embodiments of the present disclosure do not limit this.
[0031] Step 102: Perform dependency syntactic analysis on the target question sentence to obtain the dependency relationships among the individual participles in the target question sentence.
[0032] For example, performing dependency syntactic analysis on the target question sentence "What is the price of Pa Xte" can obtain the dependency relationships among "Pa Xte", "of", "price", "is", and "what". Among them, the dependency relationships can include, but are not limited to: subject-predicate relationship, verb-object relationship, attributive-middle relationship, and function word components, etc.
[0033] Step 103: Match the dependency relationships among the individual participles in the target question sentence with the dependency relationships among the types of individual participles in the set high-frequency dependency subgraph.
[0034] In the embodiments of the present disclosure, dependency syntactic analysis can be performed on each sample question sentence in the sample question sentence training set to determine the high-frequency dependency subgraph. It should be noted that the high-frequency dependency subgraph can include the dependency relationships among various nodes and the probabilities of the types to which the various nodes belong.
[0035] Furthermore, according to the dependency relationships among the individual participles in the target question sentence, a dependency relationship graph corresponding to the target question sentence can be determined, and the dependency relationships in the dependency relationship graph are matched with the dependency relationships in the high-frequency dependency subgraph to determine the types of individual participles that match the dependency relationships in the high-frequency dependency subgraph.
[0036] Step 104: Determine the types of the corresponding target participles in the target question sentence according to the types of individual participles that match the dependency relationships in the high-frequency dependency subgraph.
[0037] Furthermore, each participle in the target question sentence can be determined as the corresponding node in the high-frequency dependency subgraph. According to the probabilities recorded by each node in the high-frequency dependency subgraph, the probabilities of each participle belonging to each type can be determined. According to the probabilities of each participle belonging to each type, the types of each participle can be determined.
[0038] Step 105: Determine the standard question corresponding to the target question sentence according to the types of the target participles to determine the reply answer.
[0039] Furthermore, according to the types of each target participle, the standard data corresponding to each target participle can be determined. The target standard data is determined from the standard data, and the target question sentence is standardized according to the target standard data to generate the standard question corresponding to the target question sentence. Furthermore, the reply answer corresponding to the target question sentence can be determined.
[0040] In summary, by obtaining a target question; performing dependency syntactic analysis on the target question to obtain the dependency relationships between the respective participles in the target question; matching the dependency relationships between the respective participles in the target question with the dependency relationships between the respective participle types in a set high-frequency dependency sub-graph; determining the types of the respective target participles in the target question according to the participle types with which the dependency relationships in the high-frequency dependency sub-graph are matched; and determining the standard question corresponding to the target question according to the types of the target participles, so as to determine the reply answer. Thus, the types of the respective participles in the target question can be determined through the types of the respective nodes of the matched high-frequency dependency sub-graph, the standardization of the target participles in the target question can be effectively achieved, and the matching of questions in different scenarios can be applied, improving the coverage and accuracy of question standardization.
[0041] To determine that there is a corresponding relationship between the respective participle types in the high-frequency dependency sub-graph and the participles having the same dependency relationship in the target question, as Figure 2 shown, Figure 2 is a schematic diagram according to the second embodiment of the present disclosure. In the embodiments of the present disclosure, the dependency relationships in the dependency relationship graph corresponding to the target question can be matched with the dependency relationships in the high-frequency dependency sub-graph to determine that there is a corresponding relationship between the respective participle types in the high-frequency dependency sub-graph and the participles having the same dependency relationship in the target question. Figure 2 The embodiments shown may include the following steps:
[0042] Step 201, obtain a target question.
[0043] Step 202, perform dependency syntactic analysis on the target question to obtain the dependency relationships between the respective participles in the target question.
[0044] Step 203, determine the dependency relationship graph corresponding to the target question according to the dependency relationships between the respective participles of the target question.
[0045] For example, as Figure 3 shown, perform dependency syntactic analysis on the target question "What is the price of Pa X Te" to obtain the dependency relationships between "Pa X Te", "of", "price", "is", and "how much". According to the dependency relationships between "Pa X Te", "of", "price", "is", and "how much", the dependency relationship graph corresponding to the target question can be determined.
[0046] Step 204, query for the dependency relationships in the high-frequency dependency sub-graph in the dependency relationship graph.
[0047] Further, query whether there are the dependency relationships in the high-frequency dependency sub-graph in the dependency relationship graph. For example, if the dependency relationship in the high-frequency dependency sub-graph is a "modifier-head relationship", it can be queried in the Figure 3 shown dependency relationship graph whether this dependency relationship graph contains a "modifier-head relationship".
[0048] Step 205, in the case that there is a dependency relationship in the high-frequency dependency subgraph in the dependency graph, it is determined that there is a corresponding matching relationship between each word segmentation type in the high-frequency dependency subgraph and the word segmentations with the same dependency relationship in the target question sentence.
[0049] For example, the dependency graph is Figure 3 the right part in, and the high-frequency dependency subgraph is as Figure 4 shown. Figure 3 The dependency graph in contains Figure 4 the high-frequency dependency subgraph in, and "price" in the target question sentence can correspond to Figure 4 node 0 in, and "Passat" in the target question sentence corresponds to Figure 4 node 1 in.
[0050] Step 206, according to each word segmentation type that matches the dependency relationship in the high-frequency dependency subgraph, determine the types of the corresponding target word segmentations in the target question sentence.
[0051] It should be noted that the high-frequency dependency subgraph includes the dependency relationships between each node and the probabilities of the types to which each node belongs. After determining that there is a corresponding matching relationship between each word segmentation type in the high-frequency dependency subgraph and the word segmentations with the same dependency relationship in the target question sentence, the types of the corresponding target word segmentations in the target question sentence can be determined according to each word segmentation type that matches the dependency relationship in the high-frequency dependency subgraph.
[0052] To accurately determine the type of word segmentation, optionally, for each word segmentation in the target sentence, query the probabilities of each type recorded by the corresponding node in the matching high-frequency dependency subgraph; according to the queried probabilities of each type, determine the probabilities of the word segmentation belonging to each type; according to the probabilities of the word segmentation belonging to each type, determine the type of the word segmentation. Among them, it should be noted that the number of matching high-frequency dependency subgraphs can be one or more.
[0053] As an example, the number of matching high-frequency dependency subgraphs is one. For each word segmentation in the target sentence, query the probabilities of each type recorded by the corresponding node in the matching high-frequency dependency subgraph. The type of the word segmentation can be determined according to the highest probability of the word segmentation belonging to each type. For example, the word segmentation "price" corresponds to node 0 in the high-frequency dependency subgraph. The probability that node 0 is of type "attribute" is 96.4%, the probability that node 0 is of type "attribute value" is 1.3%, and the probability that node 0 is of type "other" is 2.2%. The type of "price" can be "attribute".
[0054] As another example, the number of matching high-frequency dependency subgraphs is at least two. For each word segment in the target sentence, query the probabilities of each type recorded by the corresponding nodes in each of the matching high-frequency dependency subgraphs. The probabilities of each type retrieved can be normalized, and based on the normalization results, determine the probabilities of the word segment belonging to each type.
[0055] For example, the word segment "price" corresponds to node 0 in the high-frequency dependency subgraph A. The probability that node 0 is of the "attribute" type is 96.4%, the probability that node 0 is of the "attribute value" type is 1.3%, and the probability that node 0 is of the "other" type is 2.2%. The word segment "price" corresponds to node 1 in the high-frequency dependency subgraph B. The probability that node 1 is of the "attribute" type is 58.4%, the probability that node 1 is of the "attribute value" type is 38.8%, and the probability that node 1 is of the "other" type is 2.6%. The probability that node 0 in the high-frequency dependency subgraph A is of the "attribute" type can be added to the probability that node 1 in the high-frequency dependency subgraph B is of the "attribute" type to obtain the probability corresponding to the "attribute" type (96.4% + 58.4%); the probability that node 0 in the high-frequency dependency subgraph A is of the "attribute value" type is added to the probability that node 1 in the high-frequency dependency subgraph B is of the "attribute value" type to obtain the probability of the "attribute value" type (1.3% + 38.8%); the probability that node 0 in the high-frequency dependency subgraph A is of the "other" type is added to the probability that node 1 in the high-frequency dependency subgraph B is of the "other" type to obtain the probability of the "other" type (2.2% + 2.6%); furthermore, perform probability normalization on the probabilities of the "attribute" type, the "attribute value" type, and the "other" type respectively. For example, the normalized probability of the "attribute" type = (96.4% + 58.4%) / ((96.4% + 58.4%) + (1.3% + 38.8%) + (2.2% + 2.6%)) = 77.5%; the normalized probability of the "attribute value" type = (1.3% + 38.8%) / ((96.4% + 58.4%) + (1.3% + 38.8%) + (2.2% + 2.6%)) = 20.08%; the normalized probability of the "other" type = (2.2% + 2.6%) / ((96.4% + 58.4%) + (1.3% + 38.8%) + (2.2% + 2.6%)) = 2.4%. Furthermore, the "attribute" type with the highest probability value can be used as the type of the target word segment "price", and the probability value that "price" is of the "attribute" type is 77.5%.
[0056] It should be noted that, in order to further improve the accuracy of determining the types of each participle in the target question, in the embodiments of the present disclosure, before querying the probabilities of each type recorded by the corresponding node in the matching high-frequency dependency subgraph for each participle in the target statement, the type probabilities of each participle in the target question can be initialized. For example, the probability of the "attribute" type of "price" in the target question can be initialized to 0%, the probability of the "attribute value" type to 0%, and the probability of the "other" type to 100%.
[0057] For example, as Figure 5 shown, the target question is "What are the brands with a price exceeding 200,000?" According to the types of each target participle corresponding in the target question based on the dependency relationships of each participle matched in the high-frequency dependency subgraph, for example, "exceeding" and "price" can correspond to the subject-predicate relationship in the high-frequency dependency subgraph, and "brand" and "exceeding" can correspond to the attributive-middle relationship in the high-frequency dependency subgraph. At the same time, according to the probability of each type recorded by each node in the high-frequency dependency subgraph, it can be determined that the type of "price" is the type "attribute", the type of "exceeding" is the type "other", and the type of "brand" is the type "attribute".
[0058] Step 207, determine the standard question corresponding to the target question according to the type of the target participle to determine the reply answer.
[0059] It should be noted that the execution processes of steps 201 to 202 and step 207 can be implemented in any one of the embodiments of the present disclosure respectively. The embodiments of the present disclosure do not make any limitations on this and will not be elaborated further.
[0060] In summary, by determining the dependency relationship graph corresponding to the target question according to the dependency relationships between each participle of the target question; querying the dependency relationships in the high-frequency dependency subgraph in the dependency relationship graph; in the case where there are dependency relationships in the high-frequency dependency subgraph in the dependency relationship graph, it is determined that there is a corresponding matching relationship between the types of each participle in the high-frequency dependency subgraph and the participles with the same dependency relationship in the target question. Thus, it can be accurately determined whether there is a high-frequency dependency subgraph in the dependency relationship graph corresponding to the target question, and in the case where there are dependency relationships in the high-frequency dependency subgraph in the dependency relationship graph, it can be accurately determined that there is a corresponding matching relationship between the types of each participle in the high-frequency dependency subgraph and the participles with the same dependency relationship in the target question.
[0061] In order to accurately determine the standard question corresponding to the target question, as Figure 6 shown,[[]]END]] Figure 6It is a schematic diagram according to the third embodiment of the present disclosure. In the embodiments of the present disclosure, standard data matching the type can be queried according to the type of the target participle, the target standard data can be determined from the standard data, and the target question sentence can be standardized according to the target standard data to generate a standard question corresponding to the target question sentence. Figure 6 The illustrated embodiment may include the following steps:
[0062] Step 601, obtain the target question sentence.
[0063] Step 602, perform dependency syntactic analysis on the target question sentence to obtain the dependency relationships between the participles in the target question sentence.
[0064] Step 603, match the dependency relationships between the participles in the target question sentence with the dependency relationships between the participle types in the set high-frequency dependency subgraphs.
[0065] Step 604, determine the types of the corresponding target participles in the target question sentence according to the participle types with which the dependencies in the high-frequency dependency subgraphs are matched.
[0066] Step 605, query the standard data matching the type according to the type of the target participle.
[0067] In the embodiments of the present disclosure, after determining the type of the target participle, the target participle can be matched with the tabular data of the same type as the participle to determine the standard data matching the type.
[0068] For example, the tabular data is as Figure 7 shown. The target question sentence is "What is the brand with a price exceeding 200,000?" By matching the dependency graph corresponding to the target question sentence with the set high-frequency dependency subgraph, it can be determined that "price" and "brand" in the target question sentence belong to the "attribute" type. The attribute slots in the tabular data, such as "vehicle model", "brand", "price", "country of origin", "number of seats", and "panoramic sunroof", can be used as the standard data matching "price" and "brand".
[0069] Step 606, for each standard data, determine the semantic similarity with the target participle.
[0070] For example, the standard data is "vehicle model", "brand", "price", "country of origin", "number of seats", and "panoramic sunroof", and the target participles are "price" and "brand". Each participle in the standard data can be respectively calculated for the semantic similarity with the target participle "price", and each participle in the standard data can be respectively calculated for the semantic similarity with the target participle "brand".
[0071] Step 607, determine the target standard data matching the type from the standard data according to the semantic similarity.
[0072] Furthermore, based on the semantic similarity between the standard data and the target word segmentation, sort each piece of standard data to determine the target standard data that matches the type from each piece of standard data. For example, sort each piece of standard data from the highest to the lowest semantic similarity, and use the standard data with the highest semantic similarity to the target word segmentation as the target standard data that matches the target word segmentation type. For example, the target standard data that matches the target word segmentation "price" is "price", and the target standard data that matches the target word segmentation "brand" is "brand".
[0073] Step 608: Standardize the target question sentence according to the target standard data to generate a standard question corresponding to the target question sentence, so as to determine the reply answer.
[0074] In the embodiments of the present disclosure, the target word segmentation in the target question sentence can be replaced with the target standard data to generate a standard question corresponding to the target question sentence. For example, if the target question sentence is "What is the brand with a price exceeding 200,000", the target word segmentation "price" can be replaced with the target standard data "price", and the target word segmentation "brand" can be replaced with the target standard data "brand", and the generated standard question is "What is the brand with a price exceeding 200,000". According to the table data, an answer can be replied to this standard question.
[0075] It should be noted that the execution processes of steps 601 to 604 can be implemented in any one of the embodiments of the present disclosure respectively. The embodiments of the present disclosure do not limit this and will not be elaborated further.
[0076] In summary, by querying each piece of standard data that matches the type according to the type of the target word segmentation; for each piece of standard data, determining the semantic similarity to the target word segmentation; according to the semantic similarity, determining the target standard data that matches the type from each piece of standard data, and standardizing the target question sentence according to the target standard data to generate a standard question corresponding to the target question sentence. Thus, the target standard data corresponding to the target word segmentation in the target question sentence can be accurately determined from the standard data, and further, a standard question corresponding to the target question sentence can be accurately generated.
[0077] In order to match the dependency relationship between the word segments in the target question sentence with the dependency relationship between the word segment types in the set high-frequency dependency subgraph, the high-frequency dependency subgraph can be obtained first, as Figure 8 shown Figure 8 is a schematic diagram according to the fourth embodiment of the present disclosure. In the embodiments of the present disclosure, dependency syntactic analysis can be performed on the sample question sentences in the sample question sentence training set to generate a sample dependency relationship graph corresponding to the sample question sentences, and then, according to the set rules, determine the high-frequency dependency subgraph from the sample dependency relationship graph, Figure 8 The embodiments shown may include the following steps:
[0078] Step 801: Obtain the target question sentence.
[0079] Step 802: Perform dependency syntax analysis on the target question sentence to obtain the dependency relationships between the segmented words in the target question sentence.
[0080] Step 803: Obtain the sample question sentence training set.
[0081] Optionally, obtain multiple sample question sentences; determine the sample question sentence templates corresponding to the sample question sentences according to the multiple sample question sentences; generate rewritten sample question sentences corresponding to the sample question sentences according to the sample question sentence templates; use the sample question sentences and the rewritten sample question sentences as the sample question sentence training set.
[0082] In the embodiments of the present disclosure, the sample question sentences can be collected online. For example, sentence structures with interrogative sentence structures can be collected online through web crawler technology as sample question sentences. Or, the sample question sentences can also be sentences with interrogative sentence structures collected offline. Or, the sample question sentences can also be artificially synthesized questions, etc. The embodiments of the present disclosure do not limit this. It should be noted that the sample question sentence training set can include various types of sample question sentences. For example, the sample question sentence can be "the X-Fu H6 with the highest price", or the sample question sentence can be "What is the brand with a price exceeding 200,000?", or the sample question sentence can be "How much more is the price of X-Sheng than that of Pa-Xte?", etc.
[0083] Further, in order to increase the number of sample question sentences in the sample question sentence training set, the sample question sentence templates corresponding to the sample question sentences can be determined, and rewritten sample question sentences can be obtained according to the sample question sentence templates in combination with different scenarios and different requirements. The sample question sentences and the rewritten sample question sentences are used as the sample question sentence training set.
[0084] For example, as shown in Table 1
[0085] Table 1: Examples of sentence templates
[0086]
[0087] In Table 1, [text_1_value] indicates that the type of the first segmented word is "attribute value (value)", and the corresponding numerical type is character type (text), representing "X-Fu H6"; [text_2_name][text_3_name] indicates that the types of the second and third segmented words are "attributes (name)", and the corresponding numerical types are both character type (text), representing "brand" and "price"; [text_finish_key] is an additional identifier, which can be extended to common question endings, such as "how much is it", "what is there", etc.
[0088] In the embodiments of the present disclosure, different types of sample questions may respectively correspond to a sample question template. The rewritten sample questions corresponding to the sample questions can be obtained according to the sample question template, and the sample questions and the rewritten sample questions are used as the sample question training set. For example, as shown in Table 2, according to the sample question template, the question is rewritten with synonyms to obtain the rewritten sample question.
[0089] Table 2 Rewritten Sample Questions
[0090]
[0091] Step 804: For each sample question in the sample question training set, perform dependency syntactic analysis on the sample question to generate a sample dependency relationship graph corresponding to the sample question.
[0092] Further, perform dependency syntactic analysis on each sample question in the sample question training set to generate a corresponding sample dependency relationship graph.
[0093] Step 805: Determine high-frequency dependency subgraphs from the sample dependency relationship graphs according to the set rules.
[0094] Optionally, according to the dependency relationships of the nodes in the sample dependency relationship graph, determine the corresponding candidate dependency subgraphs in each sample dependency graph; when the number of nodes in the candidate dependency subgraph is greater than or equal to the set node number threshold and the frequency of the candidate dependency subgraph appearing in the sample question training set is greater than or equal to the set frequency threshold, the candidate dependency subgraph is used as the high-frequency dependency subgraph.
[0095] That is to say, in order to accurately determine the high-frequency dependency subgraphs, in the embodiments of the present disclosure, for the sample dependency relationship graph corresponding to the sample question, the candidate dependency subgraphs in the sample dependency relationship graph can be determined according to the dependency relationships of the nodes in the sample dependency relationship graph. For example, the candidate dependency subgraph can be the dependency subgraph of the "attributive-middle" relationship in the sample dependency relationship graph. Further, count the number of nodes of each candidate dependency subgraph and the probability of each candidate dependency subgraph appearing in all sample dependency relationship graphs. When the number of nodes in the candidate dependency subgraph is greater than or equal to the set node number threshold and the frequency of the candidate dependency subgraph appearing in the sample question training set is greater than or equal to the set frequency threshold, the candidate dependency subgraph is used as the high-frequency dependency subgraph. For example, if the set node number threshold is 3 and the set frequency threshold is 10%, and more than 10% of the questions contain the candidate dependency subgraph and the number of nodes of the candidate dependency subgraph is greater than or equal to 3, then the candidate dependency subgraph is the high-frequency dependency subgraph.
[0096] Step 806: Match the dependency relationships between the participles in the target question with the dependency relationships between the participle types in the set high-frequency dependency subgraphs.
[0097] Step 807: Determine the types of the corresponding target participles in the target question according to the types of each participle matched by the dependency relationship in the high-frequency dependency subgraph.
[0098] Step 808: Determine the standard question corresponding to the target question according to the type of the target participle to determine the answer to the reply.
[0099] It should be noted that the execution processes of steps 801 to 802 and steps 806 to 807 can be implemented in any one of the embodiments of the present disclosure respectively. The embodiments of the present disclosure do not make any limitations thereto and will not be elaborated herein.
[0100] In summary, by obtaining the sample question training set; for each sample question in the sample question training set, performing dependency syntactic analysis on the sample question to generate a sample dependency graph corresponding to the sample question; according to the set rules, determining the high-frequency dependency subgraph from the sample dependency graph, thereby, the dependency relationship in the question can be automatically captured in the sample question, and further, the high-frequency dependency subgraph can be accurately determined from the sample dependency graph.
[0101] In order to accurately determine the probability values of the types to which each node in the high-frequency dependency subgraph belongs, as Figure 9 shown, Figure 9 is a schematic diagram according to the fifth embodiment of the present disclosure. In the embodiments of the present disclosure, after taking the candidate dependency subgraph as the high-frequency dependency subgraph, the high-frequency dependency subgraph can be used to traverse the sample question to count the types of each node in the high-frequency dependency subgraph in the sample question; and according to the types of each node in the high-frequency dependency subgraph in the sample question, determine the probability of each type recorded by each node in the high-frequency dependency subgraph. Figure 9 The embodiments shown may include the following steps:
[0102] Step 901: Obtain the target question.
[0103] Step 902: Perform dependency syntactic analysis on the target question to obtain the dependency relationship between the participles in the target question.
[0104] Step 903: Obtain the sample question training set.
[0105] Step 904: For each sample question in the sample question training set, perform dependency syntactic analysis on the sample question to generate a sample dependency graph corresponding to the sample question.
[0106] Step 905: Determine the high-frequency dependency subgraph from the sample dependency graph according to the set rules.
[0107] Step 906: Use the high-frequency dependency subgraph to traverse the sample question to count the types of each node in the high-frequency dependency subgraph in the sample question.
[0108] Step 907: Determine the probability of each type recorded by each node in the high-frequency dependency subgraph according to the types of the nodes in the high-frequency dependency subgraph in the sample question sentence.
[0109] For example, as Figure 10 shown, the high-frequency dependency subgraph includes node 0, node 1, and node 2. There is a modifier-head relationship between node 0 and node 1, and a subject-predicate relationship between node 1 and node 2. Count the quantities of node 0 being of type "attribute", type "attribute value", and type "other" in the sample question sentence. Similarly, count the quantities of node 1 being of type "attribute", type "attribute value", and type "other" in the sample question sentence, and the quantities of node 2 being of type "attribute", type "attribute value", and type "other" in the sample question sentence. Determine the probability of each type recorded by the node according to the quantities of the node being of type "attribute", type "attribute value", and type "other" in the sample question sentence. For example, if there are 3 sample question sentences containing the modifier-head relationship high-frequency dependency subgraph, and in the first sample question sentence, node 0 with the modifier-head relationship is of type "attribute", in the second sample question sentence, node 0 is of type "attribute value", and in the third sample question sentence, node 0 is of type "attribute", then the probability of node 0 being of type "attribute" is 2 / 3. Similarly, the probabilities of node 0 being of type "attribute value" and type "other" can be obtained.
[0110] Step 908: Match the dependency relationships between the word segments in the target question sentence with the dependency relationships between the types of the word segments in the set high-frequency dependency subgraph.
[0111] Step 909: Determine the types of the corresponding target word segments in the target question sentence according to the types of the word segments with matching dependency relationships in the high-frequency dependency subgraph.
[0112] Step 910: Determine the standard question corresponding to the target question sentence according to the types of the target word segments to determine the reply answer.
[0113] It should be noted that the execution processes of steps 901 to 906 and steps 908 to 910 can be implemented in any one of the embodiments of the present disclosure respectively. The embodiments of the present disclosure do not make any limitations in this regard and will not be elaborated further.
[0114] In summary, by traversing the sample question sentence using the high-frequency dependency subgraph to count the types of the nodes in the high-frequency dependency subgraph in the sample question sentence; and determining the probability of each type recorded by each node in the high-frequency dependency subgraph according to the types of the nodes in the high-frequency dependency subgraph in the sample question sentence, thus, the probability values of the types to which the nodes in the high-frequency dependency subgraph belong can be accurately determined.
[0115] The question - answering processing method of the embodiments of the present disclosure includes: obtaining a target question; performing dependency syntactic analysis on the target question to obtain the dependency relationships between the respective participles in the target question; matching the dependency relationships between the respective participles in the target question with the dependency relationships between the respective participle types in a set high - frequency dependency sub - graph; determining the types of the respective target participles in the target question according to the participle types with matching dependency relationships in the high - frequency dependency sub - graph; and determining the standard question corresponding to the target question according to the types of the target participles to determine the answer. Thus, the types of the respective participles in the target question can be determined by matching the types of each node in the high - frequency dependency sub - graph, which can effectively standardize the target participles in the target question, can be applicable to question matching in different scenarios, improves the coverage and accuracy of question standardization. At the same time, using the high - frequency dependency sub - graph can relieve the pressure of question standardization parsing in the cold start stage.
[0116] To implement the above - mentioned embodiments, the present disclosure proposes a question - answering processing device. Figure 11 It is a schematic diagram according to the sixth embodiment of the present disclosure.
[0117] As Figure 11 shown, the question - answering processing device 1100 includes: a first acquisition module 1110, an analysis module 1120, a matching module 1130, a first determination module 1140, and a second determination module 1150.
[0118] Among them, the first acquisition module 1110 is configured to obtain a target question; the analysis module 1120 is configured to perform dependency syntactic analysis on the target question to obtain the dependency relationships between the respective participles in the target question; the matching module 1130 is configured to match the dependency relationships between the respective participles in the target question with the dependency relationships between the respective participle types in a set high - frequency dependency sub - graph; the first determination module 1140 is configured to determine the types of the respective target participles in the target question according to the participle types with matching dependency relationships in the high - frequency dependency sub - graph; the second determination module 1150 is configured to determine the standard question corresponding to the target question according to the types of the target participles to determine the answer.
[0119] As a possible implementation manner of the embodiments of the present disclosure, the matching module 1130 is configured to: determine a dependency relationship graph corresponding to the target question according to the dependency relationships between the respective participles of the target question; query the dependency relationships in the high - frequency dependency sub - graph in the dependency relationship graph; and when there are dependency relationships in the high - frequency dependency sub - graph in the dependency relationship graph, determine that there is a corresponding matching relationship between the respective participle types in the high - frequency dependency sub - graph and the participles with the same dependency relationship in the target question.
[0120] As a possible implementation manner of an embodiment of the present disclosure, the matching module 1130 is further configured to: for each word segment in the target statement, query probabilities of various types recorded by the corresponding nodes in the matched high-frequency dependency subgraph; determine probabilities of the word segment belonging to various types according to the queried probabilities of various types; and determine the type of the word segment according to the probabilities of the word segment belonging to various types.
[0121] As a possible implementation manner of an embodiment of the present disclosure, the second determination module 1150 is configured to: query each standard data that matches the type according to the type of the target word segment; determine the semantic similarity with the target word segment for each standard data; determine the target standard data that matches the type from each standard data according to the semantic similarity; and standardize the target question sentence according to the target standard data to generate a standard question corresponding to the target question sentence.
[0122] As a possible implementation manner of an embodiment of the present disclosure, the question-answering processing device 1100 further includes: a second acquisition module, a generation module, and a third determination module.
[0123] Wherein, the second acquisition module is configured to acquire a sample question sentence training set; the generation module is configured to perform dependency syntactic analysis on each sample question sentence in the sample question sentence training set to generate a sample dependency relationship graph corresponding to the sample question sentence; and the third determination module determines a high-frequency dependency subgraph from the sample dependency relationship graph according to a set rule.
[0124] As a possible implementation manner of an embodiment of the present disclosure, the third determination module is configured to: determine each candidate dependency subgraph corresponding to each sample dependency relationship graph according to the dependency relationships of the respective nodes in the sample dependency relationship graph; and use the candidate dependency subgraph as the high-frequency dependency subgraph when the number of nodes in the candidate dependency subgraph is greater than or equal to a set node number threshold and the frequency of occurrence of the candidate dependency subgraph in the sample question sentence training set is greater than or equal to a set frequency threshold.
[0125] As a possible implementation manner of an embodiment of the present disclosure, the second acquisition module is configured to: acquire a plurality of sample question sentences; determine a sample question sentence template corresponding to the sample question sentences according to the plurality of sample question sentences; generate a rewritten sample question sentence corresponding to the sample question sentence according to the sample question sentence template; and use the sample question sentence and the rewritten sample question sentence as the sample question sentence training set.
[0126] As a possible implementation manner of an embodiment of the present disclosure, the question-answering processing device 1100 further includes: a statistics module and a fourth determination module.
[0127] Among them, the statistical module is used to traverse the sample question sentences using high-frequency dependency subgraphs to count the types of each node in the high-frequency dependency subgraphs in the sample question sentences; the fourth determination module is used to determine the probability of each type recorded by each node in the high-frequency dependency subgraph according to the types of each node in the high-frequency dependency subgraph in the sample question sentences.
[0128] The question-answering processing device according to the embodiments of the present disclosure obtains a target question sentence; performs dependency syntactic analysis on the target question sentence to obtain the dependency relationships between the segmented words in the target question sentence; matches the dependency relationships between the segmented words in the target question sentence with the dependency relationships between the types of segmented words in the set high-frequency dependency subgraphs; determines the types of the corresponding target segmented words in the target question sentence according to the types of segmented words with matching dependency relationships in the high-frequency dependency subgraphs; determines the standard question corresponding to the target question sentence according to the types of the target segmented words to determine the answer to be replied. Thus, the types of each segmented word in the target question sentence can be determined through the types of each node in the matching high-frequency dependency subgraph, which can effectively standardize the target segmented words in the target question and can be applicable to question matching in different scenarios, improving the coverage and accuracy of question standardization. At the same time, using the high-frequency dependency subgraph can relieve the pressure of question standardization parsing in the cold start.
[0129] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved is all carried out on the premise of obtaining the user's consent, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0130] To implement the above embodiments, the present disclosure also proposes an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the question-answering processing method as described above.
[0131] To implement the above embodiments, the present disclosure also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the question-answering processing method as described above.
[0132] To implement the above embodiments, the present disclosure also proposes a computer program product, including a computer program, and the computer program realizes the question-answering processing method as described above when executed by a processor.
[0133] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0134] Figure 12FIG. shows a schematic block diagram of an exemplary electronic device 1200 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.
[0135] As Figure 12 shown, the device 1200 includes a computing unit 1201 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0136] A plurality of components in the device 1200 are connected to the I / O interface 1205, including: an input unit 1206, such as, for example, a keyboard, a mouse, etc.; an output unit 1207, such as, for example, various types of displays, speakers, etc.; a storage unit 1208, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 1209, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows the device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0137] The computing unit 1201 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes the various methods and processes described above, such as the question-and-answer processing method. For example, in some embodiments, the question-and-answer processing method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the question-and-answer processing method described above may be executed. Alternatively, in other embodiments, the computing unit 1201 may be configured to execute the question-and-answer processing method in any other suitable manner (e.g., by means of firmware).
[0138] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0139] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0140] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0141] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0142] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0143] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can also be a server of a distributed system or a server incorporating a blockchain.
[0144] Among them, it should be noted that artificial intelligence is a discipline that studies how to make computers simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has technologies at both the hardware level and the software level. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0145] It should be understood that various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recorded in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0146] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A question-answering processing method, comprising: Obtaining a target question; Performing dependency syntactic analysis on the target question to obtain the dependency relationship between each word segment in the target question; Obtaining a sample question training set; For each sample question in the sample question training set, performing dependency syntactic analysis on the sample question to generate a sample dependency relationship graph corresponding to the sample question; According to the dependency relationship of each node in the sample dependency relationship graph, determining each corresponding candidate dependency sub-graph in each of the sample dependency relationship graphs; When the number of nodes in the candidate dependency sub-graph is greater than or equal to a set node number threshold, and the frequency of occurrence of the candidate dependency sub-graph in the sample question training set is greater than or equal to a set frequency threshold, regarding the candidate dependency sub-graph as a high-frequency dependency sub-graph; Traversing the sample question by using the high-frequency dependency sub-graph to count the types of each node in the high-frequency dependency sub-graph in the sample question; According to the types of each node in the high-frequency dependency sub-graph in the sample question, determining the probability of each type recorded by each node in the high-frequency dependency sub-graph; Matching the dependency relationship between each word segment in the target question with the dependency relationship between each word segment type in the set high-frequency dependency sub-graph; According to the word segment types corresponding to the dependency relationship matching in the high-frequency dependency sub-graph, determining the types of the corresponding target word segments in the target question; According to the types of the target word segments, determining the standard question corresponding to the target question to determine the reply answer.
2. The method according to claim 1, wherein The matching the dependency relationship between each word segment in the target question with the dependency relationship between each word segment type in the set high-frequency dependency sub-graph includes: According to the dependency relationship between each word segment in the target question, determining the dependency relationship graph corresponding to the target question; Querying the dependency relationship in the high-frequency dependency sub-graph in the dependency relationship graph; When there is the dependency relationship in the high-frequency dependency sub-graph in the dependency relationship graph, determining that there is a corresponding matching relationship between each word segment type in the high-frequency dependency sub-graph and the word segment in the target question having the same dependency relationship.
3. The method according to claim 2, wherein, The determining the types of the corresponding target word segments in the target question according to the word segment types corresponding to the dependency relationship matching in the high-frequency dependency sub-graph includes: For each word segment in the target question, querying the probabilities of each type recorded by the corresponding node in the matching high-frequency dependency sub-graph; According to the queried probabilities of each type, determining the probability of the word segment belonging to each type; According to the probability of the word segment belonging to each type, determining the type of the word segment.
4. The method according to any one of claims 1-3, wherein The determining the standard question corresponding to the target question according to the types of the target word segments includes: According to the types of the target word segments, querying the standard data matching the types; For each standard data, determining the semantic similarity with the target word segments; According to the semantic similarity, determining the target standard data matching the types from each standard data; According to the target standard data, standardizing the target question to generate the standard question corresponding to the target question.
5. The method according to claim 1, wherein The obtaining of the sample question training set includes: Obtaining a plurality of sample questions; Determining a sample question template corresponding to the sample questions according to the plurality of sample questions; Generating a rewritten sample question corresponding to the sample question according to the sample question template; Using the sample questions and the rewritten sample questions as the sample question training set.
6. A question and answer processing device, comprising: A first obtaining module, configured to obtain a target question; An analysis module, configured to perform dependency syntactic analysis on the target question to obtain the dependency relationship between each participle in the target question; A matching module, configured to match the dependency relationship between each participle in the target question with the dependency relationship between each participle type in a set high-frequency dependency subgraph; A first determining module, configured to determine the type of each target participle corresponding to the target question according to each participle type with which the dependency relationship in the high-frequency dependency subgraph is matched; A second determining module, configured to determine a standard question corresponding to the target question according to the type of the target participle to determine a reply answer; A second obtaining module, configured to obtain a sample question training set; A generating module, configured to perform dependency syntactic analysis on each sample question in the sample question training set to generate a sample dependency relationship graph corresponding to the sample question; A third determining module, configured to determine a high-frequency dependency subgraph from the sample dependency relationship graph according to a set rule; The third determining module is configured to: Determine each candidate dependency subgraph corresponding to each sample dependency relationship graph according to the dependency relationship of each node in the sample dependency relationship graph; When the number of nodes in the candidate dependency subgraph is greater than or equal to a set node number threshold, and the frequency of occurrence of the candidate dependency subgraph in the sample question training set is greater than or equal to a set frequency threshold, use the candidate dependency subgraph as the high-frequency dependency subgraph; The device further includes: A statistics module, configured to traverse the sample questions by using the high-frequency dependency subgraph to count the types of each node in the high-frequency dependency subgraph in the sample questions; A fourth determining module, configured to determine the probability of each type recorded by each node in the high-frequency dependency subgraph according to the types of each node in the high-frequency dependency subgraph in the sample questions; 7. The device according to claim 6, wherein, The matching module is configured to: Determine a dependency relationship graph corresponding to the target question according to the dependency relationship between each participle in the target question; Query the dependency relationship in the high-frequency dependency subgraph in the dependency relationship graph; When there is the dependency relationship in the high-frequency dependency subgraph in the dependency relationship graph, determine that there is a corresponding matching relationship between each participle type in the high-frequency dependency subgraph and the participle having the same dependency relationship in the target question.
8. The device according to claim 7, wherein, The matching module is further configured to: For each participle in the target question, query the probabilities of each type recorded by the corresponding node in the matched high-frequency dependency subgraph; Determine the probability of the participle belonging to each type according to the queried probabilities of each type; Determine the type of the participle according to the probability of the participle belonging to each type.
9. The device according to any one of claims 6 - 8, wherein, The second determining module is configured to: Query each standard data that matches the type according to the type of the target word segmentation; For each of the standard data, determine the semantic similarity with the target word segmentation; According to the semantic similarity, determine the target standard data that matches the type from each of the standard data; Standardize the target question according to the target standard data to generate a standard question corresponding to the target question.
10. The apparatus according to claim 6, wherein, The second acquisition module is configured to: Acquire a plurality of sample questions; Determine a sample question template corresponding to the sample questions according to the plurality of sample questions; Generate a rewritten sample question corresponding to the sample question according to the sample question template; Use the sample question and the rewritten sample question as a sample question training set.
11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.
13. A computer program product, comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Chinese natural language interrogative sentence semantization knowledge base automatic question-answering method
CN105701253A
Knowledge graph question and answer training and application service system with automatically generated template
CN111339269A