Document classification method and device

By constructing the hierarchical relationships and information gain calculations of document keywords, and automated processing of large-scale unclassified documents, solving the problems of low classification accuracy and insufficient systematicity in the existing technology, and achieving efficient document classification.

CN120492625APending Publication Date: 2025-08-15联想诺谛(北京)智能科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510370272.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively deal with the automatic classification of large-scale unclassified documents. The existing automatic classification algorithm has low classification accuracy, is difficult to meet complex or fuzzy conceptual needs, and lacks systematicity.

Method used

By extracting the keywords of each document in the document set, using a large model to build the first-level relationship, integrating the hierarchical relationships of similar keywords, calculating information gain, and iteratively dividing the document set to build a document system.

Benefits of technology

It realizes automated processing of a large number of unmarked documents, forming a document classification system that is accurate and meets the needs of complex organizations without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492625A_ABST
    Figure CN120492625A_ABST
Patent Text Reader

Abstract

The invention provides a document classification method and device, and the method comprises the steps: extracting a first keyword of each document in a document set, obtaining a first hierarchical relationship of the first keywords through a large model based on a hyponymy relationship between the first keywords, and obtaining a second hierarchical relationship of the first keywords based on the similarity between the first keywords, fusing the hierarchical relationship of the first keywords by using a large model to obtain a second hierarchical relationship, calculating a first information gain of the first keywords in the second hierarchical relationship in the document set, determining a first classification point based on the first information gain, dividing the document set into at least one document subset by using the first classification point, calculating a second information gain of the second keyword to the document subset based on the document subset, determining a second classification point of the document subset, and determining a classification mode of the document set to construct the document system when the document set is divided and iterated until a preset condition is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular to a document classification method and device. Background Art

[0002] When a business or organization accumulates a large amount of unclassified documents, such as digital assets accumulated over many years, comprehensive archiving and classification of these documents becomes crucial to improve efficiency and utilization. However, due to the sheer volume and diverse content of these documents, establishing a comprehensive and effective documentation system becomes extremely complex and challenging. Summary of the Invention

[0003] The present disclosure provides a document classification method and device to at least solve the above technical problems existing in the prior art.

[0004] According to a first aspect of the present disclosure, a document classification method is provided, the method comprising:

[0005] Extract the first keyword of each document in the document set;

[0006] Based on the hierarchical relationship between the first keywords, a first-level relationship of the first keywords is obtained using a large model;

[0007] Based on the similarity between the first keywords, the hierarchical relationships of the first keywords are fused using a large model to obtain a second hierarchical relationship;

[0008] calculating a first information gain of a first keyword in the second hierarchical relationship in the document set, and determining a first classification point based on the first information gain;

[0009] Dividing the document set into at least one document subset using the first classification point, calculating a second information gain of a second keyword on the document subset based on the document subset, and determining a second classification point for the document subset;

[0010] When the document set is divided and iterated until a preset condition is met, a classification method of the document set is determined to construct a document system.

[0011] In one possible implementation, obtaining the first hierarchical relationship of the first keywords using a large model based on the hierarchical relationship between the first keywords includes:

[0012] Input the first keyword into the macro model;

[0013] Based on the first keyword, the large model performs hyponym and hyponym analysis on the first keyword, and determines the first hierarchical relationship of the first keyword according to the analysis result.

[0014] In one possible implementation, the step of fusing the hierarchical relationships of the first keywords using a large model based on the similarity between the first keywords to obtain a second hierarchical relationship includes:

[0015] Calculating similarities between first keywords with different hierarchical relationships, and extracting similar keywords whose similarities exceed a similarity threshold;

[0016] Using the large model to fuse the hierarchical relationships corresponding to the similar keywords to obtain a target hierarchical relationship;

[0017] Determine whether similar keywords in the target hierarchical relationship need to be supplemented with upper-level concepts;

[0018] If supplementation is required, the second-level relationship is obtained after supplementation;

[0019] Otherwise, the target hierarchical relationship is taken as the second hierarchical relationship.

[0020] In one possible implementation, calculating the first information gain of the first keyword in the second hierarchical relationship in the document set includes:

[0021] Calculate the probability of each first keyword in the document set;

[0022] Calculating an initial entropy of the first keyword based on the probability;

[0023] determining a second keyword in the second hierarchical relationship, and calculating the information entropy of the second keyword;

[0024] The information gain of the second keyword in the document set is calculated based on the initial entropy and the information entropy.

[0025] In one embodiment, the classification point corresponding to the second keyword includes multiple lower-level classification points, each lower-level classification point has a corresponding sub-keyword, and calculating the information entropy of the second keyword includes:

[0026] Calculating the probability of each sub-keyword value;

[0027] Calculating the conditional entropy of each of the sub-keywords;

[0028] The information entropy of the second keyword is determined based on the value probability and conditional entropy of the sub-keyword.

[0029] In one possible implementation manner, determining the information entropy of the second keyword based on the value probability and conditional entropy of the sub-keyword includes:

[0030] Calculate the information entropy of each sub-keyword in the lower classification point based on the value probability and conditional entropy of the sub-keyword;

[0031] The information entropy of all sub-keywords is summed to obtain the information entropy of the second keyword.

[0032] In one possible implementation manner, calculating the information gain of the second keyword in the document set based on the initial entropy and the information entropy includes:

[0033] Based on the difference between the initial entropy and the information entropy, the information gain of the second keyword in the document set is determined.

[0034] In one embodiment, determining the classification point based on the information gain includes:

[0035] sorting the information gains of all the first keywords on the document set to obtain a sorting result;

[0036] Determine a first classification point according to the sorting result; and use the first classification point as a classification point in the document system.

[0037] In one embodiment, the preset condition includes at least one of the following:

[0038] The keywords are not related, the depth of the current classification point reaches the maximum depth, the number of documents in the category is less than the preset document number threshold, and the gain rate of all classification points is less than the preset gain rate threshold.

[0039] According to a second aspect of the present disclosure, a document classification device is provided, the device comprising:

[0040] A keyword extraction module, used to extract the first keyword of each document in the document set;

[0041] a first calculation module, configured to obtain a first hierarchical relationship of the first keywords using a large model based on the hierarchical relationship between the first keywords;

[0042] a second calculation module, configured to fuse the hierarchical relationships of the first keywords using a large model based on the similarities between the first keywords to obtain a second hierarchical relationship;

[0043] a third calculation module, configured to calculate a first information gain of a first keyword in the second hierarchical relationship in the document set, and determine a first classification point based on the first information gain;

[0044] a classification point determination module, configured to divide the document set into at least one document subset using the first classification point, calculate a second information gain of a second keyword for the document subset based on the document subset, and determine a second classification point for the document subset;

[0045] The system output module is used to determine the classification method of the document set to construct a document system when it iterates the division of the document set until a preset condition is met.

[0046] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0047] at least one processor; and

[0048] a memory communicatively connected to the at least one processor; wherein,

[0049] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.

[0050] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described in the present disclosure.

[0051] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:

[0053] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.

[0054] Figure 1 A schematic diagram of the implementation process of the document classification method according to an embodiment of the present disclosure is shown;

[0055] Figure 2 A schematic diagram of the first level relationship of an embodiment of the present disclosure is shown;

[0056] Figure 3 shows a schematic diagram of the second hierarchical relationship of an embodiment of the present disclosure;

[0057] Figure 4 A schematic diagram of a document system according to an embodiment of the present disclosure is shown;

[0058] Figure 5 A schematic diagram of the structure of a document classification device according to an embodiment of the present disclosure is shown;

[0059] Figure 6 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0060] To make the purposes, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative work shall fall within the scope of protection of the present disclosure.

[0061] In related technologies, experts can define categories and categorize documents, but this method is time-consuming and labor-intensive, making it difficult to handle large document collections. Alternatively, automated classification algorithms, such as clustering and keyword-based classification methods, can be used. However, current automated classification algorithms have low classification accuracy and struggle to handle complex or ambiguous concepts; classification systems can also be overly simplistic and fail to meet complex organizational needs. Currently, retrieval methods are also used to find documents, which do not require document classification. However, this method lacks systematicity and cannot fully understand the overall structure of the document collection.

[0062] The document classification method provided in this application can confirm the initial first-level classification based on the hierarchical relationship of keywords in the document set; introduce information gain calculation based on different classification systems for the document library, calculate the contribution of each classification point to the document classification, and automatically mine the best document classification system from a large number of documents. This application does not require human intervention and can automatically process a large number of unlabeled documents to form a document classification system.

[0063] The following describes a document classification method and device provided by the present application in conjunction with the accompanying drawings.

[0064] like Figure 1 As shown, the present application provides a document classification method, the method comprising:

[0065] S101, extracting the first keyword of each document in the document set;

[0066] The document set is composed of multiple documents, each of which may include different contents, that is, each document includes different first keywords. In this application, as many first keywords as possible are extracted from each document, and each document may include multiple first keywords.

[0067] For example, Figure 2As shown, a document set may include multiple documents, such as Document A, Document B, and Document C. Each document may contain multiple different first keywords, and the different first keywords have a hierarchical relationship. The first keyword in each document may be of different types. For example, the first keyword in Document A may be a small device, the first keyword in Document B may be a connector, and the first keyword in Document C may be an air separation device. Other documents may also be included, and the first keywords in other documents may be the same as or different from the first keywords in Document A, Document B, and Document C.

[0068] S102, based on the hierarchical relationship between the first keywords, using a large model to obtain a first hierarchical relationship between the first keywords;

[0069] Continue as Figure 2 As shown, there is a hierarchical relationship between the first keywords in Document A, Document B, and Document C, or similar first keywords can be ranked higher using a large model. Similar first keywords are first keywords whose similarity between the first keywords exceeds a similarity threshold. In this application, a pre-trained large model is used to classify the first keywords based on the hierarchical relationship between the first keywords, thereby obtaining a first-level relationship of the first keywords.

[0070] The large model used in this application is a large language model (LLM). The large language model can stratify the upper and lower relationships of the input first keyword to obtain the first hierarchical relationship of the first keyword.

[0071] For example, you can add a hyperlink of device type to the small device in document A, and a hyperlink of material to the insulation layer. Similarly, you can add hyperlinks to similar first keywords in documents B and C, thereby obtaining the first-level relationship of the first keywords in each document.

[0072] S103, based on the similarity between the first keywords, using a large model to fuse the hierarchical relationships of the first keywords to obtain a second hierarchical relationship;

[0073] After obtaining the first-level relationships of the first keywords in each document, the large model is used again to determine whether the first keywords between each first-level relationship are similar. It is then determined whether the first-level relationships of similar first keywords can be fused. Those first-level relationships that can be fused are fused to obtain the second-level relationships. Fusion of the first keywords in the first-level relationships can make the resulting document system more accurate.

[0074] For example, Figure 3As shown, the first-level relationship in document A includes two levels: device type and small device, and the first-level relationship in document C includes three levels: device type, large system, and air separation equipment. Therefore, the first-level relationships in document A and document C both include device types. The device types can be fused to obtain device types, the lower layers of device types are large systems and small devices, and the lower layers of large systems are air separation equipment, thereby realizing the fusion of the first-level relationship of device types in document A and the first-level relationship of device types in document C.

[0075] It should be noted that the large model can also supplement the upper-level concepts of the fused second-level relationships. For example, medium-sized equipment can be added to the lower layer of equipment types, so that the second-level relationship becomes equipment types, with the lower layers of equipment types being medium-sized equipment, large systems, and small equipment, where the lower layer of large systems is air separation equipment. This application uses the large model to fuse and supplement the scattered first-level relationships to obtain the second-level relationships. The fusion of the first keyword in the first-level shutdown can make the final document system simple and accurate.

[0076] S104, calculating a first information gain of a first keyword in the second hierarchical relationship in the document set, and determining a first classification point based on the first information gain;

[0077] The first information gain is an indicator that measures the amount of information loss caused by the first keyword when dividing the document set. Therefore, the first information gain is used to select the first classification point to divide the document set into subsets with similar first keywords.

[0078] For example, Figure 3 As shown, medium-sized equipment is added to the lower layer of the equipment type, resulting in the second-level relationship becoming equipment type. The lower layers of the equipment type are medium-sized equipment, large systems, and small equipment, with the lower layer of the large system being air separation equipment. At this point, the first information gain of equipment type, medium-sized equipment, large systems, small equipment, and air separation equipment in the document set is calculated, and the first keyword with the highest first information gain is determined as the first classification point. It is understandable that although the second-level relationship is obtained after fusion, the first information gain of each first keyword in the document set still needs to be calculated. The first keyword with the highest first information gain is determined as the first classification point, and classification is performed based on this first classification point.

[0079] S105, dividing the document set into at least one document subset using the first classification point, calculating a second information gain of a second keyword for the document subset based on the document subset, and determining a second classification point for the document subset;

[0080] After determining the first classification point, the document set is divided into at least one document subset based on the first classification point. It is understood that after determining the first keyword as the first classification point, the remaining keywords are used as second keywords, and the second information gain of the second keywords for the document subset is calculated, and then the second classification point of the document subset is determined, thereby performing a step-by-step division.

[0081] For example, the first keyword with the largest first information gain is the equipment type, and then the second information gains of medium-sized equipment, large systems, small equipment, and air separation equipment in the document set are calculated respectively. The second keyword corresponding to the second information gain ranked first is used as the second classification point, and then the document subsets are divided.

[0082] S106 , when the document set is divided and iterated until a preset condition is met, a classification method of the document set is determined to construct a document system.

[0083] This application repeats the steps of dividing the remaining document sets until the preset conditions are met, and uses all the obtained classification points to divide the document sets to obtain a document system.

[0084] For example, Figure 4 As shown, the document system finally obtained by gain sorting can be the equipment type, and the lower layer of the equipment type includes medium-sized equipment, large systems and small equipment. The lower layer of medium-sized equipment is xx equipment, the lower layer of large systems is air separation equipment, and the lower layer of small equipment is zz equipment and yy equipment.

[0085] Among them, the preset conditions can be that the labels are unrelated, the current depth reaches the maximum depth, the number of documents in the category is less than the minimum document number threshold, the gain rate of all classification points is less than the minimum gain rate threshold, etc.

[0086] The document classification method provided in the present application extracts the first keyword of each document in the document set, inputs the first keyword into the large model, obtains the first hierarchical relationship of the first keyword in each document, and then fuses the first hierarchical relationship of the first keyword in each document, and calculates the first information gain of each first keyword in the second hierarchical relationship in the document set after the fusion, so as to determine the first classification point, divides the document set into at least one document subset based on the first classification point, and calculates the second information gain of the second keyword in the document subset for all document subsets, so as to determine the second classification point of the document subset, and iterates the division until the preset conditions are met, so as to construct a document system according to the classification method of the above-mentioned document set.

[0087] This application confirms the initial first-level classification for each document by analyzing the hierarchical relationship of keywords in the document set; introduces information gain calculation based on different classification systems for the document library, calculates the contribution of each classification point to the document classification, and automatically mines the best document classification system from a large number of documents. The document classification method provided by this application does not require human intervention and can automatically process a large number of unlabeled documents.

[0088] In some embodiments, obtaining the first hierarchical relationship of the first keywords using a large model based on the hyponymous and hyponymous relationships between the first keywords includes:

[0089] Input the first keyword into the macro model;

[0090] Based on the first keyword, the large model performs hyponym and hyponym analysis on the first keyword, and determines the first hierarchical relationship of the first keyword according to the analysis result.

[0091] The present application pre-trains the large model so that the trained large model can perform hyponym and hyponym analysis on the input first keyword and construct an initial first-level relationship based on the identified hyponyms and hyponyms. After obtaining the first-level relationship, the large model can also add hypernyms to similar first keywords.

[0092] For example, Figure 2 As shown, the document set may include Document A, Document B, Document C, etc. The first keywords in Document A may include: fixed bracket, reducer, compensator, reinforced pipe, bellows expansion joint, thermal insulation layer, external thread, reinforced pipe joint, small equipment, pipeline terminology, etc.; the first keywords in Document B may include: safety valve, carbon steel, corrugated pipe, set pressure, outlet valve, balancing bellows, etc.; the first keywords in Document C may include: air separation equipment, main cooling, explosion-proof measures, etc. After performing a hyponym analysis on the first keywords contained in Documents A, B, and C, the macro model constructs an initial first-level relationship. This first-level relationship also includes hypernyms added by the macro model for similar first keywords.

[0093] The first-level relationships in Document A include: equipment types and piping system components. Under equipment types are small equipment, under small equipment are materials, and under materials are insulation layers. Under piping system components are reducers, reinforced pipes, connection and fixing components, and expansion and compensation components. Under connection and fixing components are reinforced pipe joints, and under expansion and compensation components are compensators and bellows expansion joints. The third level also includes article types, under which are terminology explanations. Equipment, piping system components, materials, and article types are all upper-level concepts supplemented by the larger model.

[0094] The first-level relationships in Document B include: expansion and compensation components, materials, and components. The lower layers of the expansion and compensation components are the bellows and balancing bellows; the lower layer of materials is carbon steel; and the lower layers of components are the safety valve, outlet valve, bellows, and balancing bellows. The expansion and compensation components, materials, and components are all upper-level concepts supplemented by the larger model.

[0095] The first-level relationship of document C includes: equipment type and article type. The lower level of equipment type is large system, and the lower level of large system is air separation equipment; the lower level of article type is measure discussion; among them, equipment type and article type are both upper-level concepts supplemented by the large model.

[0096] In some embodiments, the step of fusing the hierarchical relationships of the first keywords using a large model based on the similarity between the first keywords to obtain a second hierarchical relationship includes:

[0097] Calculating similarities between first keywords with different hierarchical relationships, and extracting similar keywords whose similarities exceed a similarity threshold;

[0098] Using the large model to fuse the hierarchical relationships corresponding to the similar keywords to obtain a target hierarchical relationship;

[0099] Determine whether similar keywords in the target hierarchical relationship need to be supplemented with upper-level concepts;

[0100] If supplementation is required, the second-level relationship is obtained after supplementation;

[0101] Otherwise, the target hierarchical relationship is taken as the second hierarchical relationship.

[0102] Specifically, on the basis of the first-level relationship, the similarity between all the first keywords in the first-level relationship is determined, similar keywords whose similarity exceeds the similarity threshold are extracted, the hierarchical relationships corresponding to the similar keywords are fused to obtain the target hierarchical relationship, and the large model is used to judge whether the similar keywords in the target hierarchical relationship need to be supplemented with upper-level concepts. In response to the need for supplementation, the second-level relationship is obtained after the supplementation. If no supplementation is required, the target hierarchical relationship is used as the second-level relationship.

[0103] Exemplarily, similar keywords in the initial hierarchical relationship among document A, document B, and document C are fused, for example, the equipment types are fused, and the components and pipeline system components are fused to obtain the target hierarchical relationship, and then similar keywords that need to be elevated are elevated, for example, the pipeline system components and liquid storage system components are elevated to the component type, thereby obtaining the second hierarchical relationship.

[0104] like Figure 3As shown, the second-level relationship after the fusion of similar keywords in the initial hierarchical relationship in Document A, Document B, and Document C includes: equipment type and component type; among which, the lower layer of equipment type is medium-sized equipment, large system and small equipment, and the lower layer of large system is air separation equipment; the lower layer of component type is pipeline system components and liquid storage system components, and the lower layer of pipeline system components is reducers, reinforced pipes, connecting and fixing components, expansion and compensation components, among which the lower layer of connecting and fixing components is reinforced pipe joints; the lower layer of expansion and compensation components is compensators, bellows expansion joints, corrugated pipes, and balanced bellows.

[0105] In some embodiments, calculating the first information gain of the first keyword in the second hierarchical relationship in the document set includes:

[0106] Calculate the probability of each first keyword in the document set;

[0107] Calculating an initial entropy of the first keyword based on the probability;

[0108] determining a second keyword in the second hierarchical relationship, and calculating the information entropy of the second keyword;

[0109] The information gain of the second keyword in the document set is calculated based on the initial entropy and the information entropy.

[0110] After obtaining the second-level relationship, this application calculates the probability of each first keyword in the second-level relationship in the document set, and then calculates the initial entropy of the first keyword based on the probability. The remaining keywords after removing the first keyword are used as the second keyword, and the information gain of the second keyword is calculated. Therefore, the information gain of the second keyword in the document set can be calculated based on the initial entropy and information entropy.

[0111] Calculating the information entropy for the second keyword at the highest level in the second-level relationship includes:

[0112] Calculating the probability of each sub-keyword value;

[0113] Calculating the conditional entropy of each of the sub-keywords;

[0114] The information entropy of the second keyword is determined based on the value probability and conditional entropy of the sub-keyword.

[0115] Specifically, determining the information entropy of the second keyword based on the value probability and conditional entropy of the sub-keyword includes:

[0116] Calculate the information entropy of each sub-keyword in the lower classification point based on the value probability and conditional entropy of the sub-keyword;

[0117] The information entropy of all sub-keywords is summed to obtain the information entropy of the second keyword.

[0118] Specifically, such as Figure 3 As shown, the information gain of the highest level classification point F, namely the device type or component type, is calculated. In this application, the information gain of each classification to the document set is calculated in order to find the classification point that can distinguish the document set to the greatest extent. After calculating the probability of the first keyword in the document set, the initial entropy of the first keyword is calculated in the following way:

[0119] H(S)=-∑(p(i)*log(p(i)))

[0120] Here, H(S) is the initial entropy, and p(i) is the probability of the first keyword appearing in the document set. Calculate the conditional entropy H(S|F=v) of classification point F. Assume that classification point F has k different values {v1, v2, …, vk}. For example, if device types include medium-sized devices, large systems, and small devices, v can be 1, 2, or 3. Based on classification point F, the document set S under certain conditions can be divided into k subsets {S1, …, Sk}. For example, if device type is the second keyword, the classification points for device type include three: 1, 2, and 3. Calculate the conditional entropy for 1, 2, and 3. Then, calculate the information entropy by summing the product of the value probabilities and the conditional entropy.

[0121] Therefore, the information entropy of the document set S can be calculated as follows:

[0122] H(S|F)=∑(p(v)*H(S|F=v))

[0123] Among them, p(v) is the probability that the feature F takes the value v, and H(S|F=v) is the entropy under the condition of F=v.

[0124] In some embodiments, calculating the information gain of the second keyword in the document set based on the initial entropy and the information entropy includes:

[0125] Based on the difference between the initial entropy and the information entropy, the information gain of the second keyword in the document set is determined.

[0126] Specifically, the information gain IG(S,F) is calculated in the following way:

[0127] IG(S,F)=H(S)-H(S|F)

[0128] In some embodiments, determining the classification point based on the information gain includes:

[0129] sorting the information gains of all the first keywords on the document set to obtain a sorting result;

[0130] Determine a first classification point according to the sorting result; and use the first classification point as a classification point in the document system.

[0131] According to different values of classification point F, such as 1, 2, and 3, the document set is divided into several subsets, and the above steps are repeated to recalculate the gains of the remaining classification points (excluding the used classification points F); when the stopping conditions are met, the cumulative gain brought by the classification is confirmed, for example: the labels are unrelated, the current depth reaches the maximum depth, the number of documents in the category is less than the minimum document number threshold, and the gain rate of all classification points is less than the minimum gain rate threshold.

[0132] Finally, the document system obtained by gain sorting is as follows Figure 4 As shown. Viewing by equipment type includes: Equipment Type. The lower level of Equipment Type includes Medium Equipment, Large Systems, and Small Equipment. Medium Equipment includes xx equipment, Large Systems includes Air Separation Equipment, and Small Equipment includes zz equipment and yy equipment. Viewing by associated components: Component Type. The lower level of Component Type includes Piping System Components and Liquid Storage System Components. Piping System Components include Linking and Fixing Components and Expansion and Compensation Components. Linking and Fixing Components include Reinforced Pipe Joints, and Expansion and Compensation Components include Bellows Expansion Joints, Corrugated Pipes, and Balancing Bellows. The lower level of Liquid Storage System Components includes AA and BB.

[0133] The document classification method provided in this application determines the initial hierarchical classification for each document by analyzing the hyponym and hyponym relationships of the document's keywords; introduces information gain calculation based on different classification systems for the document library, calculates the contribution of each classification point to the document classification, and automatically mines the best document classification system from a large number of documents.

[0134] like Figure 5 As shown, the present application provides a document classification device, the device comprising:

[0135] Keyword extraction module 501, used to extract the first keyword of each document in the document set;

[0136] A first calculation module 502 is configured to obtain a first hierarchical relationship of the first keywords using a large model based on the hierarchical relationship between the first keywords;

[0137] A second calculation module 503 is configured to fuse the hierarchical relationships of the first keywords using a large model based on the similarities between the first keywords to obtain a second hierarchical relationship;

[0138] A third calculation module 504 is configured to calculate a first information gain of the first keyword in the second hierarchical relationship in the document set, and determine a first classification point based on the first information gain;

[0139] a classification point determination module 505, configured to divide the document set into at least one document subset using the first classification point, calculate a second information gain of a second keyword for the document subset based on the document subset, and determine a second classification point for the document subset;

[0140] The system output module 506 is configured to determine a classification method for the document set to construct a document system when the document set is divided and iterated until a preset condition is met.

[0141] The document classification device provided in the present application extracts the first keyword of each document in the document set through the keyword extraction module 501; the first calculation module 502 obtains the first hierarchical relationship of the first keywords based on the hierarchical relationship between the first keywords using a large model; the second calculation module 503 fuses the hierarchical relationship of the first keywords based on the similarity between the first keywords using a large model to obtain a second hierarchical relationship; the third calculation module 504 calculates the first information gain of the first keyword in the second hierarchical relationship in the document set, and determines the first classification point based on the first information gain; the classification point determination module 505 divides the document set into at least one document subset using the first classification point, calculates the second information gain of the second keyword for the document subset based on the document subset, and determines the second classification point of the document subset; the system output module 506 determines the classification method of the document set to construct a document system when it iterates the division of the document set until the preset conditions are met.

[0142] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0143] Figure 6 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0144] like Figure 6As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0145] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0146] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the document classification method. For example, in some embodiments, the document classification method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the document classification method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the document classification method by any other suitable means (e.g., via firmware).

[0147] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0148] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0149] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0150] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0151] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0152] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0153] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0154] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0155] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A document classification method, comprising: Extract the first keyword of each document in the document set; Based on the hierarchical relationship between the first keywords, a first-level relationship of the first keywords is obtained using a large model; Based on the similarity between the first keywords, the hierarchical relationships of the first keywords are fused using a large model to obtain a second hierarchical relationship; calculating a first information gain of a first keyword in the second hierarchical relationship in the document set, and determining a first classification point based on the first information gain; Dividing the document set into at least one document subset using the first classification point, calculating a second information gain of a second keyword on the document subset based on the document subset, and determining a second classification point for the document subset; When the document set is divided and iterated until a preset condition is met, a classification method of the document set is determined to construct a document system.

2. The method according to claim 1, wherein obtaining the first hierarchical relationship of the first keywords using a large model based on the hierarchical relationship between the first keywords comprises: Input the first keyword into the macro model; Based on the first keyword, the large model performs hyponym and hyponym analysis on the first keyword, and determines the first hierarchical relationship of the first keyword according to the analysis result.

3. The method according to claim 2, wherein the step of fusing the hierarchical relationships of the first keywords using a large model based on the similarity between the first keywords to obtain a second hierarchical relationship comprises: Calculating similarities between first keywords with different hierarchical relationships, and extracting similar keywords whose similarities exceed a similarity threshold; Using the large model to fuse the hierarchical relationships corresponding to the similar keywords to obtain a target hierarchical relationship; Determine whether similar keywords in the target hierarchical relationship need to be supplemented with upper-level concepts; If supplementation is required, the second-level relationship is obtained after supplementation; Otherwise, the target hierarchical relationship is taken as the second hierarchical relationship.

4. The method according to claim 1, wherein calculating the first information gain of the first keyword in the second hierarchical relationship in the document set comprises: Calculate the probability of each first keyword in the document set; Calculating an initial entropy of the first keyword based on the probability; determining a second keyword in the second hierarchical relationship, and calculating the information entropy of the second keyword; The information gain of the second keyword in the document set is calculated based on the initial entropy and the information entropy.

5. The method according to claim 4, wherein the classification point corresponding to the second keyword includes multiple lower-level classification points, each lower-level classification point has a corresponding sub-keyword, and the calculating the information entropy of the second keyword includes: Calculating the probability of each sub-keyword value; Calculating the conditional entropy of each of the sub-keywords; The information entropy of the second keyword is determined based on the value probability and conditional entropy of the sub-keyword.

6. The method according to claim 5, wherein determining the information entropy of the second keyword based on the value probability and conditional entropy of the sub-keyword comprises: Calculate the information entropy of each sub-keyword in the lower classification point based on the value probability and conditional entropy of the sub-keyword; The information entropy of all sub-keywords is summed to obtain the information entropy of the second keyword.

7. The method according to claim 4, wherein the calculating the information gain of the second keyword in the document set based on the initial entropy and the information entropy comprises: Based on the difference between the initial entropy and the information entropy, the information gain of the second keyword in the document set is determined.

8. The method according to claim 4, wherein determining the classification point based on the information gain comprises: sorting the information gains of all the first keywords on the document set to obtain a sorting result; determining a first classification point according to the sorting result; The first classification point is used as the classification point in the document system.

9. The method according to claim 1, wherein the preset condition comprises at least one of the following: The keywords are not related, the depth of the current classification point reaches the maximum depth, the number of documents in the category is less than the preset document number threshold, and the gain rate of all classification points is less than the preset gain rate threshold.

10. A document classification device, comprising: A keyword extraction module, used to extract the first keyword of each document in the document set; a first calculation module, configured to obtain a first hierarchical relationship of the first keywords using a large model based on the hierarchical relationship between the first keywords; a second calculation module, configured to fuse the hierarchical relationships of the first keywords using a large model based on the similarities between the first keywords to obtain a second hierarchical relationship; a third calculation module, configured to calculate a first information gain of a first keyword in the second hierarchical relationship in the document set, and determine a first classification point based on the first information gain; a classification point determination module, configured to divide the document set into at least one document subset using the first classification point, calculate a second information gain of a second keyword for the document subset based on the document subset, and determine a second classification point for the document subset; The system output module is used to determine the classification method of the document set to construct a document system when it iterates the division of the document set until a preset condition is met.