Multi-granularity tree representation method for text data
By constructing a multi-granularity tree representation method, the problem of difficulty in mining text data information from different granularities in existing technologies is solved, enabling the acquisition of key information from text data at different levels and improving the utilization value of text data.
Patent Information
- Application Number
- CN202211703634.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-12-29
AI Technical Summary
Existing technologies lack fast and effective means to extract information from text data at different granularities, making it difficult to accurately grasp the key information contained in the text.
By constructing a multi-granularity tree, candidate keywords in the text data are obtained, and a multi-granularity tree is constructed based on the candidate keywords. The multi-granularity tree includes Ni keywords and their weights at the i-th level. The relationship of the number of keywords is N1≥N2≥···≥NM-2≥NM-1≥NM, thereby realizing the representation of key information of text data from different levels.
It enhances the utilization value of text data, enabling the rapid acquisition of key information at different levels within the text data.
Smart Images

Figure CN116049255B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text data mining, and more specifically, to a multi-granularity tree representation method for text data. Background Technology
[0002] Text data is a crucial medium for information transmission. Various industries continuously generate vast amounts of text data, such as news reports, product reviews, and teaching comments, all of which can be represented, stored, and transmitted in text format. Extracting key information from text data is a vital requirement for many applications. Obtaining information from text data at different granularities allows for a more accurate grasp of the key information contained within the text. Current technologies primarily focus on understanding text data at a single granularity, lacking effective means to quickly and efficiently extract information from text data at different granularities. Summary of the Invention
[0003] In view of the above problems, this application proposes a multi-granularity tree representation method for text data to solve the above problems.
[0004] In a first aspect, embodiments of this application provide a multi-granularity tree representation method for text data. The method includes: acquiring text data; extracting candidate keywords from the text data; and constructing a multi-granularity tree based on the candidate keywords, wherein the multi-granularity tree is used to represent the text data, and the multi-granularity tree includes N layers at the i-th layer. i Keyword and N i The weights corresponding to each keyword, and the number of keywords included in each layer of the multi-granularity tree are N1, N2, ..., N. M-2 N M-1 N M The relationship between the number of keywords in each layer of the multi-granularity tree satisfies N1≥N2≥···≥N M-2 ≥N M-1 ≥N M .
[0005] The multi-granularity tree representation method for text data provided in this application involves acquiring text data, extracting candidate keywords from the text data, and constructing a multi-granularity tree based on the candidate keywords. The multi-granularity tree includes N elements at the i-th layer. i Keyword and N i The weights corresponding to each keyword. Constructing a multi-granularity tree based on candidate keywords extracted from text data is an effective means of representing key information in text data at different levels, which helps to improve the utilization value of text data. Attached Figure Description
[0006] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0007] Figure 1 A flowchart illustrating the multi-granularity tree representation method for text data provided in an embodiment of this application is shown.
[0008] Figure 2 This illustration shows a multi-granularity tree diagram of the multi-granularity tree representation method for text data provided in an embodiment of this application.
[0009] Figure 3 A flowchart illustrating the multi-granularity tree representation method for text data provided in an embodiment of this application is shown. Detailed Implementation
[0010] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0011] Text data is a crucial medium for information transmission. Various industries continuously generate vast amounts of text data, such as news reports, product reviews, and teaching comments, all of which can be represented, stored, and transmitted in text format. Extracting key information from text data is a vital requirement for many applications. Obtaining information from text data at different granularities allows for a more accurate grasp of the key information contained within the text. However, current text data mining techniques lack rapid and effective methods for extracting information from text data at different granularities.
[0012] To address the aforementioned issues, this application proposes a multi-granularity tree representation method for text data. The method involves acquiring text data, extracting candidate keywords from the text data, and constructing a multi-granularity tree based on the candidate keywords. The multi-granularity tree includes N elements at the i-th layer. i One keyword and 1 i The weights corresponding to each keyword, and the relationship between the number of keywords in each layer of the multi-granularity tree, satisfy N1≥N2≥···≥N M-2 ≥N M-1 ≥N M This allows for the construction of a multi-granularity tree, which can represent key information in text data at different levels, thus improving the utilization value of the text data. The specific multi-granularity tree representation method for text data will be described in detail in subsequent embodiments.
[0013] Please see Figure 1 , Figure 1A flowchart illustrating the multi-granularity tree representation method for text data provided in an embodiment of this application is shown. The following will focus on... Figure 1 The process shown will be described in detail. The multi-granularity tree representation method for text data may specifically include the following steps:
[0014] Step S110: Obtain text data.
[0015] In some implementations, the specific content of the text data is not limited here. As one approach, the text data may be a specific range of instructor comments.
[0016] Step S120: Extract candidate keywords from the text data.
[0017] In some implementations, word segmentation techniques can be used to extract candidate keywords from text data. For example, assuming the text data is teacher evaluation text data, word segmentation and stop word removal can be performed on the teacher evaluation text data to obtain candidate keywords.
[0018] Step S130: Based on the candidate keywords, construct a multi-granularity tree, wherein the multi-granularity tree is used to represent the text data, and the multi-granularity tree includes N layers at the i-th layer. i Keyword and N i The weights corresponding to each keyword, and the number of keywords included in each layer of the multi-granularity tree are N1, N2, ..., N. M-2 N M-1 N M The relationship between the number of keywords in each layer of the multi-granularity tree satisfies N1≥N2≥···≥N M-2 ≥N M-1 ≥N M .
[0019] In this embodiment, a multi-granularity tree for representing text data can be constructed based on candidate keywords. Each layer of the multi-granularity tree represents the text data at a certain granularity, and the multi-granularity tree at layer i includes N... i Keyword and N i The weights corresponding to each keyword, and the number of keywords included in each level of the multi-granularity tree are N1, N2, ..., N. M-2 N M-1 N M The relationship between the number of keywords at each level satisfies N1≥N2≥···≥N M-2 ≥N M-1 ≥N M The more keywords a tree has, the finer the granularity; the fewer keywords, the coarser the granularity. The construction process of a multi-granularity tree proceeds from fine to coarse; that is, the first level is built first, and the Mth level is built last. The number of keywords in the Mth level is N.M It can be 1. The first level of the multi-granularity tree consists of leaf nodes, represented by N1 keywords and their weights, representing the finest-grained representation of the text data; the second level uses N2 keywords and their weights, representing the next finest-grained representation of the text data; the Mth level is the root node, represented by 1 keyword and its weight, representing the coarsest-grained representation of the text data. Please refer to [link / reference]. Figure 2 , Figure 2 A schematic diagram of a multi-granularity tree representation method for text data provided in an embodiment of this application is shown. Figure 2 The circles in the text represent candidate keywords and keywords. The candidate keywords in the text data are filtered to obtain the first layer of keywords. The first layer is the finest granularity of the multi-granularity tree, the second layer is the second finest granularity of the multi-granularity tree, and the Mth layer is the coarsest granularity of the multi-granularity tree.
[0020] In some implementations, N1 keywords are extracted from the candidate keywords, and the weights corresponding to the N1 keywords are obtained to obtain the first-level representation of the multi-granularity tree of the text data. Then, the N1 keywords in the first level are clustered according to the keyword information to obtain N2 classes. The keyword with the largest weight in each of the N2 classes, and the sum of the weights corresponding to all keywords in each of the N2 classes, are used as the second-level representation of the multi-granularity tree of the text data. The process of obtaining the second level from the first level is repeated to construct the multi-granularity tree representation of the text data until the (M-1)th level representation of the text data is obtained. The keyword with the largest weight in the (M-1)th level and the sum of the weights corresponding to all keywords in the (M-1)th level are used as the representation of the Mth level. Alternatively, a specified keyword and the sum of the weights corresponding to all keywords in the (M-1)th level of the multi-granularity tree are used as the representation of the Mth level.
[0021] The multi-granularity tree representation method for text data provided in this application obtains text data, extracts candidate keywords from the text data, and constructs a multi-granularity tree representing the text data based on the candidate keywords, thereby realizing the multi-granularity tree representation of text data. This method can quickly obtain key information at different levels in the text data and helps to improve the utilization value of the text data.
[0022] Please see Figure 3 , Figure 3 A flowchart illustrating the multi-granularity tree representation method for text data provided in this application is shown. In a specific embodiment, keyword information may include keyword distance and keyword tags. The following will focus on... Figure 3 The process shown will be described in detail. The multi-granularity tree representation method for text data may specifically include the following steps:
[0023] Step S210: Obtain text data.
[0024] Step S220: Extract candidate keywords from the text data.
[0025] For a detailed description of steps S210-S220, please refer to steps S110-S120, which will not be repeated here.
[0026] Step S230: Based on the keyword dictionary, select N1 keywords from the candidate keywords and obtain the weights of the N1 keywords in the text data, which are used as the first layer of the multi-granularity tree.
[0027] In this embodiment, word segmentation technology can be used to extract candidate keywords from the text data. Then, N1 keywords are extracted from the candidate keywords based on a keyword dictionary. That is, the candidate keywords in the keyword dictionary are used as the N1 keywords of the first layer, and the word frequencies of the N1 keywords in the text data are determined as the weights corresponding to the 11 keywords. It is understood that the weights corresponding to the keywords are obtained through the text data, so different text data will yield different keyword weights.
[0028] It should be noted that a keyword dictionary can include keywords, keyword tags, and keyword distances. Understandably, using different keyword dictionaries can yield different multi-granularity tree representations of text data. Conversely, using the same keyword dictionary with different text data can produce different multi-granularity tree representations. Keyword distances and keyword tags in a keyword dictionary greatly facilitate subsequent keyword clustering. Keyword tags provide a new clustering method for constructing multi-granularity trees; the tags often include the information that users most want to extract. For example, in the keyword dictionary for flight training instructor comments, based on the nine core competency evaluation methods for pilots, partial data of keywords and keyword tags are shown in Table 1, and partial data of keyword distances are shown in Table 2.
[0029] Table 1
[0030] Keywords Label Ground procedures Procedure execution and compliance with regulations Situational awareness Situational awareness and information management Correction deviation Problem Solving and Decision Making
[0031] Table 2
[0032] airspace procedures Pine pole rod airspace procedures 0 0.923 0.916 Pine pole 0.923 0 0.014 rod 0.916 0.014 0
[0033] Step S240: Based on the N i Clustering the keyword information corresponding to each keyword yields N. i+1 Each class.
[0034] In this embodiment, it can be based on N i Clustering the keyword information corresponding to each keyword yields N. i+1There are several classes. For example, clustering based on the keyword information corresponding to N1 keywords yields N2 classes; clustering based on the keyword information corresponding to N3 keywords yields N4 classes; and so on. M-2 Clustering the keyword information corresponding to each keyword yields N. M-1 Each class.
[0035] In some implementations, when the text data is teacher comments, after clustering the keywords in the first layer, one of the N2 classes can include 6 keywords as shown in Table 3. Table 3 shows the keywords of one of the N2 classes and the word frequencies of the keywords.
[0036] Table 3
[0037] Keywords Word frequency (weight) altimeter 14 PFD 3 heading indicator 34 Attitude instrument 22 meter 56 airspeed indicator 8
[0038] In some implementations, keyword information may include keyword distance and keyword tags, which are not limited here. As one approach, it can be based on N... i Clustering based on keyword distances for each keyword yields N. i+1 There are several categories. Alternatively, it can be based on a keyword dictionary and N... i Keyword tags corresponding to each keyword, for N i Clustering by keywords yields N i+1 Each class.
[0039] In some implementations, clustering can be performed based on keyword distance, using a clustering algorithm to cluster N. i Clustering by keywords yields N i+1 There are several classes. As one implementation method, based on N... i The keyword distances corresponding to N keywords were analyzed using the K-Means Clustering Algorithm. i Clustering by keywords yields N i+1 There are several categories. Different keyword clustering methods can be used in different implementations, and no specific method is specified here.
[0040] In this embodiment, the N can be selected based on a keyword dictionary. i Among the keywords, those with the same keyword tags are clustered into one class, resulting in N. i+1 Each class.
[0041] In some implementations, users can specify a target layer for clustering based on keyword tags. For example, suppose a user specifies clustering for layer M-2 based on keyword tags. When the keyword tag is "nine core competencies for pilots," these nine core competencies include knowledge application, procedure execution and compliance, flight path management (autopilot), flight path management (manual), communication, leadership and teamwork, situational awareness and information management, workload management, and problem-solving and decision-making. Therefore, the maximum number of classes in layer M-2 is nine. Layer M-2 can be clustered into two categories: "procedure execution and compliance" and "situational awareness and information management." Information for the "procedure execution and compliance" category is shown in Table 4, and information for the "situational awareness and information management" category is shown in Table 5. As shown in Tables 4 and 5, in the M-2 level, "ground procedures" is the keyword with the highest weight in the "procedure execution and compliance" category, and "situational awareness" is the keyword with the highest weight in the "situational awareness and information management" category. Therefore, the keywords for the M-1 level are "ground procedures" and "situational awareness," with a weight of 508 (352+21+9+126=508) for "ground procedures" and 928 (32+20+32+512+108+224=928) for "situational awareness." The keyword for the root node of the M-1 level of the multi-granularity tree is designated as "core competency," and the weight of "core competency" is the sum of the frequencies of all keywords in the M-1 level, i.e., 508+928=1436.
[0042] Table 4
[0043] Keywords Word frequency (weight) Label Ground procedures 352 Procedure execution and compliance with regulations Stall procedure 21 Procedure execution and compliance with regulations Modify the program 9 Procedure execution and compliance with regulations airspace procedures 126 Procedure execution and compliance with regulations
[0044] Step S250: Place the N i+1 The keyword with the highest weight in each of the N categories, and the N i+1 The sum of the weights corresponding to all keywords in each class is used as the (i+1)th layer of the multi-granularity tree.
[0045] Table 5
[0046] Keywords Word frequency (weight) Label Flight status 32 Situational awareness and information management Sports Trends 20 Situational awareness and information management Model data 32 Situational awareness and information management Situational awareness 512 Situational awareness and information management attitude 108 Situational awareness and information management Attention allocation 224 Situational awareness and information management
[0047] In this embodiment, N can be i+1 The keyword with the highest weight in each class, and N i+1 The sum of the weights of all keywords in each class is used as the (i+1)th level of the multi-granularity tree. For example, as shown in Table 3, if the keyword with the largest weight in one of the N2 classes is "instrument", then "instrument" is used as one of the keywords in the second level of the multi-granularity tree, and the weight of "instrument" is 137 (14+3+34+22+56+8=137).
[0048] Step S260: Repeat the above steps to obtain Ni+1 The process of obtaining the (i+1)th layer of the multi-granularity tree continues until the Mth layer of the multi-granularity tree is obtained.
[0049] In this embodiment, the above process of obtaining N can be repeated. i+1 The process of obtaining the (i+1)th layer of the multi-granularity tree continues until the Mth layer of the multi-granularity tree is obtained, where the number of keywords in the Mth layer is 1.
[0050] In some implementations, it can be determined whether i+1 equals M. If i+1 is less than M, meaning the Mth level of the multi-granularity tree has not been constructed, then N is obtained repeatedly. i+1 The process of obtaining the (i+1)th layer of the multi-granularity tree continues until the Mth layer of the multi-granularity tree is obtained.
[0051] In some implementations, N can be obtained repeatedly. i+1 The process of classifying and obtaining the (i+1)th layer of the multi-granularity tree continues until the (M-1)th layer of the multi-granularity tree is obtained. The keyword with the largest weight in the (M-1)th layer of the multi-granularity tree, along with the sum of the weights of all keywords in the (M-1)th layer, is taken as the Mth layer of the multi-granularity tree. For example, the keywords and corresponding weights of the (M-1)th layer are shown in Table 6. The keyword with the largest weight in Table 6 is "ground program". Therefore, "ground program" is taken as the keyword of the Mth layer, and the weight of "ground program" in the Mth layer is 4628 (2808 + 1820 = 4628).
[0052] In some implementations, N is obtained repeatedly. i+1 The process of classifying and obtaining the (i+1)th layer of the multi-granularity tree continues until the (M-1)th layer of the multi-granularity tree is obtained. The target keyword is then acquired, and the sum of the weights corresponding to the target keyword and all keywords in the (M-1)th layer of the multi-granularity tree is used as the Mth layer of the multi-granularity tree. The keyword of the root node of the Mth layer can be the target keyword. Alternatively, the user can specify the keyword of the Mth layer as the target keyword. For example, if the user specifies "aircraft instruments" as the keyword of the Mth layer, then the keyword of the Mth layer is "aircraft instruments," and the weight corresponding to "aircraft instruments" in the Mth layer is the sum of the weights of all keywords in the (M-1)th layer.
[0053] Table 6
[0054] Keywords Word frequency (weight) Ground procedures 2808 meter 1820
[0055] One embodiment of this application provides a multi-granularity tree representation method for text data, which, compared to... Figure 1The multi-granularity tree representation method for text data shown can extract keywords at the first level based on a keyword dictionary, and then construct a multi-granularity tree up to the Mth level based on these keywords. Furthermore, by selecting different keyword dictionaries, multi-granularity trees that meet different needs can be constructed, which helps to improve the utilization value of text data. Additionally, it can be based on N... i The keyword distance pairs corresponding to each keyword are N. i Clustering of keywords, and based on the keyword dictionary and N i Each keyword corresponds to a keyword tag for N. i Keyword clustering, based on keyword distance and keyword tags, can efficiently and conveniently construct multi-granularity trees.
[0056] In summary, the multi-granularity tree representation method for text data provided in this application obtains text data, extracts candidate keywords from the text data, and constructs a multi-granularity tree representing the text data based on the candidate keywords. The multi-granularity tree includes N elements at the i-th layer. i Keyword and N i The weights corresponding to each keyword, and the relationship between the number of keywords in each layer of the multi-granularity tree, satisfy N1≥N2≥···≥N M-2 ≥N M-1 ≥N M This enables a multi-granularity tree representation of text data, which can represent key information of text data at different levels, thus helping to improve the utilization value of text data.
[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A multi-granularity tree representation method for text data, characterized in that, The method includes: Get text data; Extract candidate keywords from the text data; Based on the candidate keywords, a multi-granularity tree is constructed, wherein the multi-granularity tree is used to represent the text data, and the multi-granularity tree is in the first... Layers include Keyword and The weights corresponding to each keyword, and the number of keywords included in each layer of the multi-granularity tree. The relationship between the number of keywords in each layer of the multi-granularity tree satisfies ; The construction of a multi-granularity tree based on the candidate keywords includes: Based on the keyword dictionary, the candidate keywords are selected. One keyword, and obtain the The weights corresponding to each keyword in the text data are used as the first layer of the multi-granularity tree; Based on the above Clustering of keyword information corresponding to each keyword to obtain There are several classes, among which... It can be 1 to Positive integers; The The keyword with the highest weight in each of the categories, and the... The sum of the weights corresponding to all keywords in each of the classes is used as the th weight of the multi-granularity tree. layer; Repeat the above to obtain The class and the first granularity tree obtained The process of layering continues until the first layer of the multi-granularity tree is obtained. layer.
2. The method according to claim 1, characterized in that, The above-described method is repeated to obtain... The class and the first granularity tree obtained The process of layering continues until the first layer of the multi-granularity tree is obtained. Layers, including: Repeat the above to obtain The class and the first granularity tree obtained The process of layering continues until the first layer of the multi-granularity tree is obtained. layer; The multi-granularity tree The keyword with the highest weight in the keyword of the layer, and the first keyword of the multi-granularity tree. The sum of the weights corresponding to all keywords in the layer is used as the first weight of the multi-granularity tree. layer.
3. The method according to claim 1, characterized in that, The above-described method is repeated to obtain... The class and the first granularity tree obtained The process of layering continues until the first layer of the multi-granularity tree is obtained. Layers, including: Repeat the above to obtain The class and the first granularity tree obtained The process of layering continues until the first layer of the multi-granularity tree is obtained. layer; Obtain target keywords; The target keywords, and the multi-granularity tree The sum of the weights corresponding to all keywords in the layer is used as the first weight of the multi-granularity tree. layer.
4. The method according to claim 1, characterized in that, The keyword information includes keyword distance, which is based on the Clustering of keyword information corresponding to each keyword to obtain There are several classes, including: Based on the above The keyword distance corresponding to each keyword, for the... Clustering by keywords to obtain Each class.
5. The method according to claim 1, characterized in that, The keyword information includes keyword tags, which are based on the Clustering of keyword information corresponding to each keyword to obtain There are several classes, including: Based on the keyword dictionary and the The keyword tags corresponding to each keyword, for the... Clustering by keywords to obtain Each class.
6. The method according to claim 5, characterized in that, The keyword dictionary and the The keyword tags corresponding to each keyword, for the... Clustering by keywords to obtain There are several classes, including: The Keywords with the same keyword tags are clustered together to obtain... Each class.
Citation Information
Patent Citations
Recommendation method and system of video text labels
CN103164471A
Short text classification and intelligent analysis system for multi-granularity demand
CN114840677A