Text processing method, device, equipment and computer storage medium
By using dependency grammar analysis and relationship tree construction technology in text processing, the problem of the inability to accurately determine the semantic relationship between quantitative words and other words in the prior art is solved, and the semantic information in the text is accurately extracted without a large amount of training data.
Patent Information
- Application Number
- CN202111163785.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-09-30
AI Technical Summary
The prior art cannot accurately determine the semantic relationship between quantitative words and other words, resulting in the inability to effectively extract semantic information in sentences in data relationship-related research.
By obtaining the target text fragments related to the quantitative words in the text, performing dependency grammar analysis, building a dependency grammar relationship tree, and determining the target semantic relationship structure based on the one-to-one correspondence between the preset dependency grammar relationship structure and the semantic relationship structure.
Without a large amount of training data, the semantic relationship between quantitative words and other words can be accurately determined, improving the accuracy of text information extraction.
Smart Images

Figure CN114492320B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular relates to a text processing method, apparatus, device and computer storage medium. Background Art
[0002] As the demand for data statistics increases, users often need to extract numerical data from large amounts of text, such as quantifiers. Since this data also needs to be summarized and visualized, in addition to the quantifiers themselves, it is also necessary to extract the semantic relationship between the quantifiers and other text information. For example, if the user only extracts "4%" for "car sales increased by 4%", it will be meaningless. However, if the user can extract the relationship between "growth" and "4%", it can be determined that there is a 4% increase. If the user can extract the relationship between "car sales" and "4%", it can be determined that 4% is describing "car sales".
[0003] In existing data relationship research, algorithms for extracting relationships between words can only extract the nominative and accusative semantic information of a sentence, and cannot accurately extract the relationship between quantifiers and other words. Conventional relationship extraction algorithms, such as those based on CRF, rely heavily on application scenarios and training data. Without sufficient training data, it is almost impossible to train a usable relationship extraction model.
[0004] Therefore, the prior art lacks a method for accurately determining the semantic relationship between quantifiers and other words. Summary of the Invention
[0005] The embodiments of the present application provide a text processing method, apparatus, device, and computer storage medium, which can at least solve the problem in the prior art of being unable to accurately determine the semantic relationship between quantifiers and other words.
[0006] In a first aspect, an embodiment of the present application provides a text processing method, the method comprising:
[0007] Get the first text;
[0008] determining a target text segment including a quantifier from the first text;
[0009] performing dependency grammar analysis on the multiple words in the target text segment according to the parts of speech of the multiple words in the target text segment to obtain a first dependency grammar relationship between the words in the target text segment;
[0010] Constructing a dependency grammar relationship tree according to the first dependency grammar relationship;
[0011] Determining at least one target dependency grammar relationship structure matching a preset dependency grammar relationship structure from the dependency grammar relationship tree, wherein the target dependency grammar relationship structure is at least a portion of the dependency grammar relationship tree;
[0012] According to a one-to-one correspondence between preset dependency grammatical relationship structures and preset semantic relationship structures, a target semantic relationship structure corresponding to the at least one target dependency grammatical relationship structure is determined.
[0013] In a second aspect, an embodiment of the present application provides a text processing device, the device comprising:
[0014] An acquisition module, used for acquiring a first text;
[0015] a target text segment determining module, configured to determine a target text segment including a quantifier from the first text;
[0016] an analysis module, configured to perform dependency grammar analysis on the multiple words in the target text segment according to the parts of speech of the multiple words in the target text segment, and obtain a first dependency grammar relationship between the words in the target text segment;
[0017] A construction module, configured to construct a dependency grammar relationship tree according to the first dependency grammar relationship;
[0018] a target dependency grammatical relationship structure determination module, configured to determine at least one target dependency grammatical relationship structure that matches a preset dependency grammatical relationship structure from the dependency grammatical relationship tree, wherein the target dependency grammatical relationship structure is at least a portion of the dependency grammatical relationship tree;
[0019] The target semantic relationship structure determination module is used to determine the target semantic relationship structure corresponding to the at least one target dependency grammatical relationship structure according to the one-to-one correspondence between the preset dependency grammatical relationship structure and the preset semantic relationship structure.
[0020] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing computer program instructions;
[0021] When the processor executes the computer program instructions, the text processing method as shown in any one of the embodiments of the first aspect is implemented.
[0022] In a fourth aspect, an embodiment of the present application provides a computer storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the text processing method shown in any one of the embodiments of the first aspect is implemented.
[0023] The text processing method, apparatus, device, and computer storage medium of the embodiments of the present application determine a target text segment including a quantifier from an acquired first text, and then perform dependency grammar analysis on the multiple words in the target text segment based on the parts of speech of the multiple words in the target text segment to obtain a first dependency grammar relationship between the words in the target text segment. Then, based on the first dependency grammar relationship, a dependency grammar relationship tree is constructed, and at least one target dependency grammar relationship structure that matches a preset dependency grammar relationship structure is determined from the dependency grammar relationship tree. Then, based on the one-to-one correspondence between the preset dependency grammar relationship structure and the semantic relationship structure, a target semantic relationship structure corresponding to the at least one target dependency grammar relationship structure is determined. Since the one-to-one correspondence between the preset dependency grammar relationship structure and the preset semantic relationship structure is manually preset, the target semantic relationship structure can be determined without a large amount of data for training, thereby accurately determining the semantic relationship between the quantifier and other words. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0025] Figure 1 is a flowchart of a text processing method according to an exemplary embodiment;
[0026] Figure 2 is a schematic diagram of a dependency grammar relationship tree according to an exemplary embodiment;
[0027] Figure 3 is a schematic diagram of a target semantic relationship structure according to an exemplary embodiment;
[0028] Figure 4 is another schematic diagram of a target semantic relationship structure according to an exemplary embodiment;
[0029] Figure 5 is a schematic diagram of another target semantic relationship structure according to an exemplary embodiment;
[0030] Figure 6 is a flowchart of another text processing method according to an exemplary embodiment;
[0031] Figure 7 is a structural diagram of a text processing device according to an exemplary embodiment;
[0032] Figure 8 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0033] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0034] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0035] The text processing method, device, electronic device, and computer storage medium provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.
[0036] The text processing method provided in this application can be applied to scenarios where semantic relationships between quantifiers and other words in a text are obtained. Furthermore, the text processing method provided in the embodiments of this application can be executed by a text processing module in a text processing system. In the embodiments of this application, the text processing method provided in the embodiments of this application is illustrated by using a text processing module executing the text processing method as an example.
[0037] Figure 1 A flowchart of a text processing method provided by an embodiment of the present application is shown.
[0038] like Figure 1 As shown, the text processing method may include the following steps:
[0039] First, S110, obtaining a first text;
[0040] Second, S120 , determining a target text segment including a quantifier from the first text;
[0041] Third, S130, performing dependency grammar analysis on the multiple words in the target text segment according to the parts of speech of the multiple words in the target text segment to obtain a first dependency grammar relationship between the words in the target text segment;
[0042] Fourth, S140, constructing a dependency grammar relationship tree based on the first dependency grammar relationship;
[0043] Fifth, S150, determining at least one target dependency grammar relationship structure that matches a preset dependency grammar relationship structure from the dependency grammar relationship tree;
[0044] Sixth, S160 determines a target semantic relationship structure corresponding to at least one target dependency grammatical relationship structure according to a one-to-one correspondence between the preset dependency grammatical relationship structure and the preset semantic relationship structure.
[0045] Thus, by determining a target text segment including quantifiers from the acquired first text, and then performing dependency grammar analysis on the multiple words in the target text segment based on the parts of speech of the multiple words in the target text segment, a first dependency grammar relationship between the words in the target text segment is obtained, and then a dependency grammar relationship tree is constructed based on the first dependency grammar relationship, and at least one target dependency grammar relationship structure that matches the preset dependency grammar relationship structure is determined from the dependency grammar relationship tree, and then a target semantic relationship structure corresponding to the at least one target dependency grammar relationship structure is determined based on the one-to-one correspondence between the preset dependency grammar relationship structure and the semantic relationship structure. Since the one-to-one correspondence between the preset dependency grammar relationship structure and the preset semantic relationship structure is manually preset, the target semantic relationship structure can be determined without a large amount of data for training, thereby accurately determining the semantic relationship between the quantifier and other words.
[0046] The above steps are explained in detail below:
[0047] First, regarding S110, in an embodiment of the present application, when a user needs to obtain the semantic relationship between a quantifier and other words in a first text, the user can first obtain the first text and then process the first text to determine the semantic relationship between the quantifier and other words. The first text can be an article, such as a financial news article.
[0048] Second, regarding S120, since it is necessary to obtain the semantic relationship between the quantifiers and other words in the first text, it is necessary to first determine the target text segment that includes the quantifiers in the first text, and then process the target text segment to determine the semantic relationship between the quantifiers and other words in the target text segment. No subsequent processing is performed on the text segment that does not include the quantifiers. This can reduce the number of text segments to be processed and improve work efficiency.
[0049] Specifically, a quantifier can be a numeral or a combination of a numeral and a quantifier. For example, "300 yuan" is a quantifier, where "300" is a numeral and "yuan" is a quantifier. The target text segment can be a portion of the first text that includes a quantifier, for example, a paragraph or a sentence that includes a quantifier. Furthermore, the target text segment can also be the entire first text, for example, an article that includes a quantifier. The specific form of the target text segment is not limited herein.
[0050] Third, regarding S130, determining the semantic relationship between the quantifier and other words in the target text segment requires determining a target semantic relationship structure between the quantifier and other words in the target text segment. Determining this target semantic relationship structure requires first determining the first dependency grammatical relationship between the words in the target text segment. Therefore, dependency grammatical analysis can be performed on the multiple words in the target text segment based on their parts of speech, thereby obtaining the first dependency grammatical relationship between the words in the target text segment. Specifically, dependency grammatical analysis can be performed on the multiple words in the target text segment based on a dependency grammatical analysis algorithm.
[0051] Fourth, in step S140 , a target semantic relationship structure between quantifiers and other words in the target text segment is determined, and a dependency grammatical relationship tree needs to be constructed based on the first dependency grammatical relationship.
[0052] For example, the target text segment "This year's car sales increased by about 4% year-on-year." is subjected to dependency grammar analysis of multiple words. Based on the first dependency grammar relationship obtained, the dependency grammar relationship tree constructed can be as follows: Figure 2 As shown, the explanation of the first dependency grammatical relationship included therein can be shown in Table 1 (the explanations of other dependency grammatical relationships involved in the full text can also be shown in Table 1).
[0053] Table 1 - First dependency grammar relationship description table
[0054]
[0055] Fifth, regarding S150, the target dependency relationship structure can be at least a portion of the dependency relationship tree. While any portion of the dependency relationship tree can yield a dependency relationship structure, not all dependency relationship structures can be used to accurately determine the target semantic relationship structure. Therefore, it is necessary to match the target dependency relationship structure with a preset dependency relationship structure to determine a target dependency relationship structure that matches the preset dependency relationship structure. Of course, if the target text segment is relatively simple, the target dependency relationship structure determined can be the entire dependency relationship tree.
[0056] Sixth, regarding S160, since the grammatical rules of Chinese are universal, each sentence that conforms to the grammatical rules of Chinese has its corresponding semantic relationship structure, and the number of these semantic relationship structures is limited, so these limited semantic relationship structures can be summarized manually. The preset semantic relationship structure here can be at least one semantic relationship structure related to quantifiers summarized manually. The preset semantic relationship structure can be a semantic relationship structure that can more accurately indicate the semantic relationship between quantifiers and other words. The preset semantic relationship structure can correspond one-to-one with the preset dependency grammatical relationship structure. Therefore, after determining at least one target dependency grammatical relationship structure that matches the preset dependency grammatical relationship structure from the preset dependency grammatical relationship structure, the target semantic relationship structure corresponding to the target dependency grammatical relationship structure can be determined according to the one-to-one correspondence between the preset dependency grammatical relationship structure and the preset semantic relationship structure, thereby more accurately indicating the semantic relationship between quantifiers and other words. Since the at least one target dependency grammatical relationship structure that matches the preset dependency grammatical relationship structure can be one or more, the target semantic relationship structure can be one or more. If the number of target semantic relationship structures is 1, the target semantic relationship structure can be directly output.
[0057] Exemplarily, to determine the target semantic relationship structure between quantifiers and other words in financial news articles, the semantic relationship structures related to quantifiers in financial news articles can be manually summarized and sorted to obtain at least one preset semantic relationship structure, and a structure matching algorithm including the at least one preset semantic relationship structure can be written. The structure matching algorithm can also include a one-to-one correspondence between a preset dependency grammatical relationship structure and a preset semantic relationship structure. Through this bottom-up structure matching algorithm, at least one target dependency grammatical relationship structure that matches the preset dependency grammatical relationship structure is first determined from the dependency grammatical relationship tree, and then the target semantic relationship structure corresponding to the at least one target dependency grammatical relationship structure is determined.
[0058] Specifically, when manually summarizing the preset semantic relationship structure related to quantifiers in financial news articles, labels can be added to the words in the news articles. For example, the labels can include quantifiers (NUM), contexts (ATTRIBUTE), descriptions (DESC) and actions (ACTION). Among them, ATTRIBUTE can represent the context corresponding to the quantifier, such as time adverbials; DESC can represent the subject corresponding to the quantifier; ACTION can represent the action corresponding to the quantifier, such as increase, decrease, etc.; NUM is the quantifier. Of course, other labels can also be included in addition to these, which are not limited here. Based on this, we can first count the dependency grammatical relationship between NUM and ATTRIBUTE, DESC, and ACTION in news articles, and then sort out 16 preset semantic relationship structures that conform to financial news based on news corpus induction and linguistic rules. The preset semantic relationship structure can include labels corresponding to the words included in the target dependency grammatical relationship structure and the dependency grammatical relationship between words. For example, Figure 3 As shown, the semantic relationship structures corresponding to the text segments may be 306 , 307 , 308 , 309 and 310 .
[0059] In addition, the first dependency grammatical relationship may be directly matched with a preset semantic relationship structure to determine a target semantic relationship structure.
[0060] Specifically, when matching with the preset semantic relationship structure, each quantifier (num) in each target text segment can be matched, first observing its parent node type, then observing its child nodes, as well as the parent node's parent-child node, the child node's parent-child node, and other types.
[0061] For example, if 16 preset semantic relationship structures are summarized, among which num can have 6 preset semantic relationship structures with different types of parent nodes, and num can have two preset semantic relationship structures with different types of child nodes. Then no matter what the parent node of num is, it will not deviate from one of the 6 preset semantic relationship structures. If it is determined that the parent node of num belongs to one of the 6 preset semantic relationship structures, then continue to judge whether the child node of num belongs to one of the above two child nodes, and match each node one by one. In the end, one or more of the 16 patterns will be matched. Among them, the 6 preset semantic relationship structures of num with different types of parent nodes can be as follows:
[0062] num->parent(SBV)
[0063] num->parent(VOB)
[0064] num->parent(POB)
[0065] num->parent(COO)
[0066] num->parent(ATT)
[0067] num->parent(ADV)
[0068] The two preset semantic relationship structures of num with different types of child nodes can be as follows:
[0069] num->child(ATT)
[0070] num->child(SBV)
[0071] For example, after performing dependency grammar analysis on the target text segment, the dependency grammar relationship tree constructed is matched with the preset dependency grammar relationship structure to obtain the target dependency grammar relationship structure, and according to the one-to-one correspondence between the preset dependency grammar relationship structure and the preset semantic relationship structure, the following can be determined: Figure 4 The target semantic relationship structure shown.
[0072] Based on this, in an optional implementation, when the number of target semantic relationship structures is greater than 1, after S150, the method may further include:
[0073] According to the preset priority of each target semantic relationship structure, the target semantic relationship structures are sorted and the priority order is determined;
[0074] Output the target semantic relationship structure and the priority corresponding to the target semantic relationship structure in order of priority.
[0075] Here, the number of target semantic relationship structures can be greater than 1. For example, after performing dependency grammar analysis on "This year's car sales increased by about 4% year-on-year" and constructing a dependency grammar relationship tree, two target dependency grammar relationship structures that match the preset dependency grammar relationship structure can be determined from the dependency grammar relationship tree, thereby determining two target semantic relationship structures corresponding to the two target dependency grammar relationship structures, such as Figure 5 shown.
[0076] Based on this, when the number of target semantic relationship structures is greater than 1, the target semantic relationship structures are output in order of priority, and a target semantic structure that can more accurately indicate the semantic relationship between the quantifier and other words can be determined. Specifically, according to the preset priority of each target semantic relationship structure, multiple target semantic relationship structures can be prioritized, the priority order of the multiple target semantic relationship structures can be determined, and then the above multiple target semantic relationship structures can be output in order of priority. Of course, only one or more target semantic relationship structures with the highest priority can be output, which is not limited here. In addition, in addition to outputting the target semantic relationship structure, the priority corresponding to each target semantic relationship structure can also be output.
[0077] For example, to prioritize multiple target semantic relationship structures between quantifiers and other words in financial news articles, the priority of each target semantic relationship structure can be manually pre-set, and a priority ranking algorithm can be developed that includes each target semantic relationship structure and its corresponding priority. Using this priority ranking algorithm, the multiple target semantic relationship structures can be ranked to determine the priority order.
[0078] In this way, by outputting the target semantic relationship structure and the priority corresponding to the target semantic relationship structure in order of priority, it is possible to more accurately determine the target semantic relationship structure that most accurately indicates the semantic relationship between quantifiers and other words when there are multiple target semantic relationship structures.
[0079] Based on the above S110-S150, in a possible embodiment, as Figure 6 As shown, the above S120 may specifically include: S1201-S1204, wherein:
[0080] S1201: Divide a first text into at least one text segment.
[0081] Here, the first text can be an entire article. If the entire article is processed, the complexity will be relatively high and the speed will be relatively slow. Therefore, the first text can be divided. Specifically, it can be divided into natural paragraphs or sentences, which are not limited here. Since the text segments obtained by division need to be subsequently subjected to dependency grammar analysis, and the integrity of the analysis results obtained by performing dependency grammar analysis on sentences may be relatively poor, it can be selected to divide the text segments into natural paragraphs to maximize the integrity and rationality of the dependency grammar analysis results. Accordingly, at least one text segment can be at least one natural paragraph, and of course it can also be at least one sentence, which is not limited here.
[0082] For example, a financial news article may be cut into multiple natural segments using a cutting algorithm.
[0083] S1202, perform word segmentation on each text segment to obtain multiple words corresponding to each text segment.
[0084] Here, since it is necessary to determine the target text segment including a quantifier, and it is also necessary to perform dependency grammar analysis on multiple words in the target text segment according to the词性 of multiple words in the target text segment, therefore, it is necessary to perform word segmentation on each text segment to obtain multiple words corresponding to each text segment. A word can be the smallest language unit that can be used independently. Here, multiple words can also include punctuation marks.
[0085] Exemplarily, a word segmentation algorithm based on a long short-term memory network (LSTM) can be used to perform word segmentation on each cropped natural paragraph to obtain multiple words corresponding to each natural paragraph. Word segmentation can be performed through a word segmentation algorithm based on LSTM
[0086] S1203, label the词性 of multiple words corresponding to each text segment to obtain a labeling result.
[0087] Here, since it is necessary to determine the target text segment including a quantifier, and it is also necessary to perform dependency grammar analysis on multiple words in the target text segment according to the词性 of multiple words in the target text segment, therefore, it is necessary to label the词性 of multiple words corresponding to each text segment. The词性 of these multiple words can be used to determine the quantifier in the text segment, thereby determining the target text segment, and can also be used to perform dependency grammar analysis on multiple words in the target text segment.
[0088] Exemplarily, a labeling algorithm based on LSTM can be used to label the词性 of multiple words corresponding to each text segment. The词性 can include: time word (t), punctuation mark (w), ordinary verb (v), noun (n), auxiliary word "了" (ule), quantifier (mq), numeral (m), adverb (d), verb "是" (vshi), adjective (a), modal particle (y). Of course, other词性 can also be included, which are not limited here.
[0089] It should be noted that the quantifier in the embodiments of this application can be a numeral or a combination of a numeral and a classifier. For example, "10" is a numeral, "ton" is a classifier, and "10 tons" is a quantifier. That is to say, the quantifiers in the embodiments of this application include the words with the词性 of mq and m among the above multiple words.
[0090] S1204, determine the target text segment including a quantifier according to the labeling result.
[0091] Here, since it is necessary to determine the target semantic relationship structure between quantifiers and other words, dependency parsing can be performed only on the target text segments that include quantifiers, and the target text segments can be determined based on the annotation results. Specifically, multiple text segments can be screened, and text segments that do not include quantifiers can be filtered out, and text segments that include quantifiers can be determined as target text segments.
[0092] Specifically, the corresponding annotation results of the text fragment can be analyzed. If the text fragment includes words whose parts of speech are quantifiers, the text fragment can be subsequently subjected to dependency grammar analysis; if the text fragment does not include words whose parts of speech are quantifiers, the text fragment can be filtered out and no subsequent processing is performed on the text fragment.
[0093] For example, the annotation results may be analyzed to determine that the text segment includes "10%" with part of speech m and "10 tons" with part of speech mq. Therefore, the text segment may be subsequently subjected to dependency grammar analysis.
[0094] In this way, through the above process, text segments that do not include quantifiers can be filtered out, the number of text segments that need to be processed is reduced, and processing efficiency is improved.
[0095] Based on this, in an optional implementation, before S130, the text processing method may further include:
[0096] The parts of speech of the multiple words in the target text segment are determined according to the tagging results corresponding to the respective words in the target text segment.
[0097] Here, since dependency grammar analysis needs to be performed on multiple words in the target text segment based on their parts of speech, the parts of speech of the multiple words in the target text segment can be determined first. Specifically, the parts of speech of the multiple words in the target text segment can be determined based on the annotation results corresponding to each word in the target text segment in all the annotation results.
[0098] For example, the annotation results may include the parts of speech of a total of 100 words corresponding to the three text segments A, B, and C. After determining that the target text segment is B, the parts of speech of the 30 words included in the target text segment B can be determined from the parts of speech of these 100 words, so that the 30 words in the target text segment B can be subsequently subjected to dependency grammar analysis based on the parts of speech of these 30 words.
[0099] In this way, the parts of speech of multiple words in the target text segment can be determined through the annotation results corresponding to each word in the target text segment in all the annotation results, thereby facilitating dependency grammar analysis of multiple words in the target text segment.
[0100] Based on this, in an optional implementation, after S150, the text processing method may further include:
[0101] According to the marking results, determine the offset position of the quantifier;
[0102] Output the target semantic relationship structure and offset position.
[0103] Here, in order to visualize the quantifiers, such as making a table, a line chart, and a bar chart, in addition to determining the target semantic relationship structure corresponding to the quantifier, it is also necessary to determine the offset position of the quantifier. Therefore, the offset position of the quantifier can be determined based on the annotation result, and after determining the target semantic relationship structure, the target semantic relationship structure and the offset position of the quantifier can be output. Specifically, when the number of target semantic relationship structures is 1, the offset position of the quantifier and the target semantic relationship structure corresponding to the quantifier can be directly output; when the number of target semantic relationship structures is greater than 1, the offset position of the quantifier and multiple target semantic relationship structures corresponding to the quantifier can be directly output, wherein the multiple target semantic relationship structures can be output in order of priority, and the method for determining the order of priority will not be repeated here.
[0104] By outputting the target semantic relationship structure and the offset position of the quantifiers, it is possible to facilitate visualization of the quantifiers in the target text segment, such as by making a table, a line chart, a bar chart, etc.
[0105] Based on the same inventive concept, the present application also provides a text processing device. Figure 7 The text processing device provided in the embodiment of the present application is described in detail.
[0106] Figure 7 It is a structural block diagram of a text processing device according to an exemplary embodiment.
[0107] like Figure 7 As shown, the text processing device 9 may include:
[0108] An acquisition module 901 is configured to acquire a first text;
[0109] a target text segment determining module 902, configured to determine a target text segment including a quantifier from the first text;
[0110] An analysis module 903 is configured to perform dependency grammar analysis on the multiple words in the target text segment according to the parts of speech of the multiple words in the target text segment, and obtain a first dependency grammar relationship between the words in the target text segment;
[0111] A construction module 904 is configured to construct at least one first semantic relationship structure between the quantifier and other words based on the first dependency grammatical relationship, where the other words are words in the target text segment that have a dependency grammatical relationship with the quantifier;
[0112] A target dependency grammar relationship structure determination module 905 is configured to determine at least one target dependency grammar relationship structure that matches a preset dependency grammar relationship structure from the dependency grammar relationship tree, wherein the target dependency grammar relationship structure is at least a portion of the dependency grammar relationship tree;
[0113] The target semantic relationship structure determination module 906 is configured to determine a target semantic relationship structure corresponding to at least one target dependency grammatical relationship structure according to a one-to-one correspondence between preset dependency grammatical relationship structures and preset semantic relationship structures.
[0114] In one embodiment, the target text segment determination module 902 may include:
[0115] a division submodule, configured to divide the first text into at least one text segment;
[0116] The word segmentation submodule is used to perform word segmentation on each text segment to obtain multiple words corresponding to each text segment;
[0117] The annotation submodule is used to annotate the parts of speech of multiple words corresponding to each text segment to obtain the annotation results;
[0118] The target text segment determination submodule is used to determine the target text segment including quantifiers based on the annotation results.
[0119] In one embodiment, the text processing device 9 may further include:
[0120] The part-of-speech determination module 907 is used to perform dependency grammatical analysis on multiple words in the target text segment based on the parts of speech of multiple words in the target text segment, and before obtaining the first dependency grammatical relationship between the words in the target text segment, determine the parts of speech of multiple words in the target text segment based on the annotation results corresponding to the words in the target text segment.
[0121] In one embodiment, the text processing device 9 may further include:
[0122] A position determination module 908 is configured to determine an offset position of the quantifier based on the annotation result after determining a target semantic relationship structure corresponding to at least one target dependency grammatical relationship structure based on a one-to-one correspondence between preset dependency grammatical relationship structures and preset semantic relationship structures;
[0123] The first output module 909 is used to output the target semantic relationship structure and offset position.
[0124] In one embodiment, when the number of target semantic relationship structures is greater than 1, the text processing device 9 may further include:
[0125] A sorting module 910 is configured to, after determining a target semantic relationship structure corresponding to at least one target dependency grammatical relationship structure based on a one-to-one correspondence between preset dependency grammatical relationship structures and preset semantic relationship structures, sort the target semantic relationship structures based on a preset priority of each target semantic relationship structure to determine a priority order;
[0126] The second output module 911 is configured to output the target semantic relationship structure and the priority corresponding to the target semantic relationship structure in order of priority.
[0127] Thus, by determining a target text segment including quantifiers from the acquired first text, and then performing dependency grammar analysis on the multiple words in the target text segment based on the parts of speech of the multiple words in the target text segment, a first dependency grammar relationship between the words in the target text segment is obtained, and then a dependency grammar relationship tree is constructed based on the first dependency grammar relationship, and at least one target dependency grammar relationship structure that matches the preset dependency grammar relationship structure is determined from the dependency grammar relationship tree, and then a target semantic relationship structure corresponding to the at least one target dependency grammar relationship structure is determined based on the one-to-one correspondence between the preset dependency grammar relationship structure and the semantic relationship structure. Since the one-to-one correspondence between the preset dependency grammar relationship structure and the preset semantic relationship structure is manually preset, the target semantic relationship structure can be determined without a large amount of data for training, thereby accurately determining the semantic relationship between the quantifier and other words.
[0128] Figure 8 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment.
[0129] like Figure 8 As shown in FIG, the electronic device 10 is a block diagram of an exemplary hardware architecture of an electronic device capable of implementing the text processing method and text processing apparatus according to the embodiments of the present application. The electronic device may refer to the electronic device in the embodiments of the present application.
[0130] The electronic device 10 may include a processor 1001 and a memory 1002 storing computer program instructions.
[0131] Specifically, the processor 1001 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0132] Memory 1002 may include a large capacity memory for information or instructions. By way of example and not limitation, memory 1002 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 1002 may include removable or non-removable (or fixed) media. Where appropriate, memory 1002 may be internal or external to the integrated gateway device. In a specific embodiment, memory 1002 is a non-volatile solid-state memory. In a specific embodiment, memory 1002 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0133] The processor 1001 implements the method in the above-mentioned embodiment and achieves the corresponding technical effect by reading and executing the computer program instructions stored in the memory 1002, which will not be repeated here for the sake of brevity.
[0134] In one embodiment, the electronic device 10 may further include a transceiver 1003 and a bus 1004. Figure 8 As shown, the processor 1001 , the memory 1002 and the transceiver 1003 are connected via a bus 1004 and communicate with each other.
[0135] Bus 1004 includes hardware, software or both.For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics buses, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral control interconnect (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 1004 may include one or more buses. Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.
[0136] An embodiment of the present application further provides a computer storage medium, in which computer executable instructions are stored. The computer executable instructions are used to implement the text processing method described in the embodiment of the present application.
[0137] In some possible implementations, various aspects of the method provided in the present application may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the method according to the various exemplary implementations of the present application described above in this specification. For example, the computer device can execute the text processing method described in the embodiments of the present application.
[0138] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0139] The present application is described with reference to the flowcharts and / or block diagrams of the methods, apparatuses and computer program products according to the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable information processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable information processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0140] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable information processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0141] These computer program instructions can also be loaded onto a computer or other programmable information processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0142] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A text processing method, characterized in that, the method includes: Obtain a first text; Determine a target text segment including a quantifier from the first text; Perform dependency grammar analysis on multiple words in the target text segment according to the part-of-speech of the multiple words in the target text segment, and obtain a first dependency grammar relationship between each word in the target text segment; Construct a dependency grammar relationship tree according to the first dependency grammar relationship; Determine at least one target dependency grammar relationship structure that matches a preset dependency grammar relationship structure from the dependency grammar relationship tree, and the target dependency grammar relationship structure is at least a part of the dependency grammar relationship tree; Determine a target semantic relationship structure corresponding to the at least one target dependency grammar relationship structure according to the one-to-one correspondence between the preset dependency grammar relationship structure and the preset semantic relationship structure; Among them, the determining a target text segment including a quantifier from the first text specifically includes: Divide the first text to obtain at least one text segment; Perform word segmentation on each text segment to obtain multiple words corresponding to each text segment; Label the part-of-speech of multiple words corresponding to each text segment to obtain a labeling result; Determine the target text segment including a quantifier according to the labeling result.
2. The method according to claim 1, characterized in that, Before performing dependency grammar analysis on multiple words in the target text segment according to the part-of-speech of the multiple words in the target text segment and obtaining a first dependency grammar relationship between each word in the target text segment, the method further includes: Determine the part-of-speech of multiple words in the target text segment according to the labeling result corresponding to each word in the target text segment.
3. The method according to claim 1, characterized in that, After determining a target semantic relationship structure corresponding to the at least one target dependency grammar relationship structure according to the one-to-one correspondence between the preset dependency grammar relationship structure and the preset semantic relationship structure, the method further includes: Determine the offset position of the quantifier according to the labeling result; Output the target semantic relationship structure and the offset position.
4. The method according to claim 1, characterized in that, When the number of the target semantic relationship structures is greater than 1, after determining a target semantic relationship structure corresponding to the at least one target dependency grammar relationship structure according to the one-to-one correspondence between the preset dependency grammar relationship structure and the preset semantic relationship structure, the method further includes: Sort the target semantic relationship structures according to the priority of each preset target semantic relationship structure, and determine the priority order; Output the target semantic relationship structure and the priority corresponding to the target semantic relationship structure according to the priority order.
5. A text processing device, characterized in that, the device includes: An acquisition module, configured to acquire a first text; A target text segment determination module, configured to determine a target text segment including a quantifier from the first text; An analysis module, configured to perform dependency grammar analysis on multiple words in the target text segment according to the part-of-speech of the multiple words in the target text segment, so as to obtain a first dependency grammar relationship between each word in the target text segment; A construction module, configured to construct a dependency grammar relationship tree according to the first dependency grammar relationship; A target dependency grammar relationship structure determination module, configured to determine at least one target dependency grammar relationship structure that matches a preset dependency grammar relationship structure from the dependency grammar relationship tree, where the target dependency grammar relationship structure is at least a part of the dependency grammar relationship tree; A target semantic relationship structure determination module, configured to determine a target semantic relationship structure corresponding to the at least one target dependency grammar relationship structure according to the one-to-one correspondence between the preset dependency grammar relationship structure and the preset semantic relationship structure; Wherein, the target text segment determination module specifically includes: A division sub-module, configured to divide the first text to obtain at least one text segment; A word segmentation sub-module, configured to perform word segmentation processing on each text segment to obtain multiple words corresponding to each text segment; A tagging sub-module, configured to tag the part-of-speech of multiple words corresponding to each text segment to obtain a tagging result; A target text segment determination sub-module, configured to determine the target text segment including a quantifier according to the tagging result.
6. The apparatus according to claim 5, wherein, the apparatus further includes: A part-of-speech determination module, configured to determine the part-of-speech of multiple words in the target text segment according to the tagging results corresponding to each word in the target text segment before performing dependency grammar analysis on multiple words in the target text segment according to the part-of-speech of the multiple words in the target text segment, so as to obtain a first dependency grammar relationship between each word in the target text segment.
7. An electronic device, wherein, the device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the text processing method according to any one of claims 1-4 is implemented.
8. A computer storage medium, wherein, Computer program instructions are stored on the computer storage medium, and when the computer program instructions are executed by a processor, the text processing method according to any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Entity relationship extraction method and device, electronic equipment and storage medium
CN112232074A
Semantic relationship recognition method and device, electronic equipment and readable storage medium
CN113010642A