Text classification methods, systems, devices, and media
The text classification method improves accuracy by quantifying keyword interactions and semantic relationships, addressing the limitations of existing methods through directional traction strength corrections and non-linear competition modulation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2026-02-17
- Publication Date
- 2026-04-06
AI Technical Summary
Existing text classification methods fail to accurately depict the complex semantic relationships between keywords, leading to low classification accuracy in scenarios with similar categories or multiple themes due to insufficient description of semantic diffusion and non-linear competition relationships.
A text classification method that calculates directional traction strength, tension diffusion coefficients, and interaction correction coefficients to correct and modulate keyword interactions, incorporating non-linear competitive modulation based on semantic distance to establish a classification objective function.
Enhances classification accuracy by quantitatively characterizing keyword interactions, dynamically adjusting category attribution based on semantic relevance, and stabilizing classification decisions through comprehensive potential energy modeling.
Smart Images

Figure 0007840628000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text classification, and specifically to text classification methods, systems, devices, and media.
Background Art
[0002] With the development of natural language processing technology, text classification has been widely applied in scenarios such as information retrieval, public opinion analysis, and content recommendation. In the prior art, common text classification methods generally map text to a feature space based on keyword statistical features or vectorized representations, and then calculate the similarity or probability distribution between the text and various categories using a classification model. Such methods generally consider the contributions of different types of keywords to be independent of each other, or only perform simple weighting processing, making it difficult to depict the semantic relationship between keywords and their mutual influence in different types of contexts.
[0003] However, in actual applications, there are often complex semantic connections among the keywords in the text. The similarity, co-occurrence relationship, and pulling direction for different categories among different keywords jointly affect the classification result. When the prior art deals with multiple keyword interactions, semantic competition between categories, and the propagation effect of semantic influence, it generally uses a linear superposition or static weight mechanism, leading to insufficient description of the semantic diffusion of keywords and the non-linear competition relationship between categories. As a result, in scenarios with similar categories or multiple theme text scenarios, problems such as ambiguity of the classification boundary and instability of the attribution degree easily appear, and there is a problem of low classification accuracy.
Summary of the Invention
Problems to be Solved by the Invention
[0004] In response to the above problems in the prior art, the text classification method, system, device, and medium according to the present invention solve the problem of low classification accuracy existing in the prior art.
Means for Solving the Problems
[0005] To achieve the above-mentioned objective of the invention, the technical solution used in this invention is a text classification method, Step S1 involves obtaining the directional traction strength for each type of keyword in the text, Step S2 involves obtaining the tension diffusion coefficient of a keyword based on the semantic similarity between keywords in the text and the direction of traction. Step S3 involves multiplying the tension diffusion coefficient of the keyword by the corresponding directional tensile strength to obtain a first corrected directional tensile strength for the keyword type, Step S4 involves correcting the corresponding directional traction strength based on the interaction correction coefficient to obtain a second corrected directional traction strength for the keyword type. Step S5 involves fusing the first corrected directional traction strength and the second corrected directional traction strength, and further introducing nonlinear competitive modulation based on the semantic distance between the types to obtain the final degree of belonging to each type of text. Step S6 includes obtaining the first semantic potential energy value and the second semantic potential energy value, combining them with the final degree of belonging to establish a classification objective function and obtain the type of text.
[0006] Furthermore, S1 is, Substep S11 calculates the semantic similarity between keywords in the text and reference words in various reference word sets, Substep S12 involves multiplying the semantic similarity by the TF-ICF weight of the corresponding reference word to obtain a traction value. The process includes a substep S13 in which the maximum traction value is selected from the traction values corresponding to each reference word of each type to obtain the directional traction strength of this keyword for this type.
[0007] Furthermore, S2 is Substep S21 calculates the semantic similarity of two keywords in the text, Substep S22 obtains the traction direction based on the directional traction strength of two keywords of the same type, The process includes a substep S23 which calculates the tension diffusion coefficient of a keyword based on semantic similarity and traction direction.
[0008] Furthermore, the formula for calculating the tension diffusion coefficient of the keyword in S23 is as follows: JPEG0007840628000002.jpg11170JPEG0007840628000003.jpg31170
[0009] Furthermore, S4 is, Substep S41 calculates an interaction correction coefficient based on the co-occurrence weights, semantic similarity, and TF-IDF weights of the two keywords, Substep S42 involves multiplying the interaction correction coefficient by the interaction influence coefficient and adding 1 to obtain the global interaction gain coefficient for the keyword. The process includes a substep S43 in which the global interaction gain coefficient of the keyword is multiplied by the corresponding directional traction strength to obtain a second corrected directional traction strength for the keyword type.
[0010] Furthermore, the formula for calculating the interaction correction coefficient in S41 is as follows: JPEG0007840628000004.jpg10170JPEG0007840628000005.jpg21170JPEG0007840628000006.jpg24170
[0011] Furthermore, S5 is, Substep S51 involves multiplying the normalized TF-IDF weight of the keyword by the first corrected directional traction strength to obtain a first enhancement value. Substep S52 involves multiplying the normalized TF-IDF weight of the keyword by a second corrected directional traction strength to obtain a second enhancement value. Substep S53 involves adding the second enhancement value and the first enhancement value, taking the average, and obtaining the total enhancement value. Substep S54 involves summing the overall reinforcement values of all keywords in the type dimension to obtain the initial degree of belonging to each type of text, Substep S55 calculates the semantic distance between any two types based on the reference word sets of each type, The process includes a substep S56 in which nonlinear competitive modulation is performed on the initial degree of belonging for each type of text based on semantic distance to obtain the final degree of belonging for each type of text.
[0012] Furthermore, the formula for calculating the final degree of belonging in S56 is as follows: JPEG0007840628000007.jpg12170JPEG0007840628000008.jpg34170
[0013] Furthermore, the S6 is Substep S61 calculates a first semantic potential energy value based on a first enhancement value corresponding to each keyword in the text, Substep S62 calculates a second semantic potential energy value based on a second enhancement value corresponding to each keyword in the text, Substep S63 involves weighting the final degree of belonging, the first semantic potential energy value, and the second semantic potential energy value to obtain a classification objective function. The process includes a substep S64 in which a determination value is calculated for all categories based on a classification objective function, and the category with the highest determination value is determined to be the category to which the text belongs.
[0014] Furthermore, the formula for calculating the first semantic potential energy value in S61 is as follows: The formula for calculating the second semantic potential energy value in S62 is as follows: JPEG0007840628000011.jpg13170JPEG0007840628000012.jpg22170
Advantages of the Invention
[0015] The beneficial effects of the present invention are as follows.
[0016] 1. The present invention realizes the quantitative characterization of the semantic relationship of keywords by obtaining the directional traction strength for each type of keyword and calculating the tension diffusion coefficient by combining the semantic similarity and traction direction between keywords, breaking through the limitations of the prior art that regards keyword contributions as mutually independent or only simple weighting. At the same time, the directional traction strength is corrected by the tension diffusion coefficient, accurately capturing the propagation effect of the semantic influence of keywords, enabling the type traction effect of each keyword to be reasonably diffused according to semantic relevance, conforming to the semantic law of keyword interaction in the actual text, and solving the problem of insufficient description of keyword semantic diffusion in the prior art.
[0017] 2. The present invention introduces an interaction correction coefficient to perform secondary correction on the directional traction strength, realizing quantitative correction for the complex interaction relationship between multiple keywords, and compensating for the defects in the prior art of processing multiple keyword interactions using linear superposition and static weight mechanisms. By conforming to the interaction characteristics of different keyword combinations with the interaction correction coefficient, the traction strength for the type of keyword is made more conformable to the actual semantic expression, and the distortion of the traction strength caused by ignoring keyword interactions is avoided.
[0018] 3. After fusing the directional traction strength after two corrections, the present invention introduces non-linear competition modulation based on the semantic distance between categories, breaking through the limitations of the linear modeling of semantic competition between categories in the prior art. For categories with similar meanings and multiple theme text scenarios, non-linear competition modulation can dynamically adjust the category attribution degree of the text based on the degree of semantic relevance between categories, weaken the semantic overlap interference between categories, and strengthen the dedicated feature expressions of different categories, effectively solving the problems in the prior art such as the ambiguity of classification boundaries and the instability of attribution degrees in such scenarios.
[0019] 4. The present invention constructs a classification objective function based on the first semantic potential energy value, the second semantic potential energy value, and the final attribution degree, and models by fusing the potential features at the semantic level and the category attribution degree. The classification decision not only depends on the traction effect of keywords on categories, but also in combination with the semantic potential distribution of the entire text, which can further enhance the rationality and reliability of the classification result. Compared with the classification logic of the prior art that only depends on similarity or probability distribution, the basis of the classification decision of the present invention is more comprehensive, effectively avoiding classification deviations due to one-dimensional judgment, and improving the classification accuracy.
Brief Description of the Drawings
[0020] [Figure 1] It is a flowchart of a text classification method.
Embodiments for Implementing the Invention
[0021] In the following, specific embodiments of the present invention will be described to facilitate those skilled in the art to understand the present invention. It should be understood that the present invention is not limited to the scope of the embodiments for implementing the invention. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention determined by the appended claims. All invention creations using the concept of the present invention are within the scope of protection.
[0022] Example 1, as shown in Figure 1, is a text classification method, Step S1 involves obtaining the directional traction strength for each type of keyword in the text, Step S2 involves obtaining the tension diffusion coefficient of a keyword based on the semantic similarity between keywords in the text and the direction of traction. Step S3 involves multiplying the tension diffusion coefficient of the keyword by the corresponding directional tensile strength to obtain a first corrected directional tensile strength for the keyword type, Step S4 involves correcting the corresponding directional traction strength based on the interaction correction coefficient to obtain a second corrected directional traction strength for the keyword type. Step S5 involves fusing the first corrected directional traction strength and the second corrected directional traction strength, and further introducing nonlinear competitive modulation based on the semantic distance between the types to obtain the final degree of belonging to each type of text. Step S6 includes obtaining the first semantic potential energy value and the second semantic potential energy value, combining them with the final degree of belonging to establish a classification objective function and obtain the type of text.
[0023] JPEG0007840628000013.jpg25170
[0024] JPEG0007840628000014.jpg31170
[0025] JPEG0007840628000015.jpg46170
[0026] In this embodiment, the various types of samples refer to samples classified into their corresponding categories.
[0027] This invention constructs a reference word set for each type and assigns TF-ICF weights to accurately capture type-specific core semantic features, providing clear type-specific semantic anchors for subsequent calculations and avoiding semantic ambiguity problems caused by generalized keyword statistical features.
[0028] In this embodiment, S1 is Substep S11 calculates the semantic similarity between keywords in the text and reference words in various reference word sets, Substep S12 involves multiplying the semantic similarity by the TF-ICF weight of the corresponding reference word to obtain a traction value. The process includes a substep S13 in which the maximum traction value is selected from the traction values corresponding to each reference word of each type to obtain the directional traction strength of this keyword for this type.
[0029] In this example, the formula for calculating the taste similarity in S11 is as follows: JPEG0007840628000016.jpg12170JPEG0007840628000017.jpg37170
[0030] In this embodiment, the expression for obtaining the directional traction strength for a given keyword type is as follows: JPEG0007840628000018.jpg9170JPEG0007840628000019.jpg29170
[0031] The calculation of traction values in this invention simultaneously fuses semantic similarity and TF-ICF weights to ensure semantic relevance between keywords and reference words, while also introducing a degree of type distinction for reference words, so that traction values not only reflect semantic relationships but also directly embody the actual value of this relationship for text classification. In order for this invention to calculate traction values for each reference word and keyword in each category, the maximum traction value is selected as the directional traction strength from the traction values corresponding to one category.
[0032] In this embodiment, S2 is Substep S21 calculates the semantic similarity of two keywords in the text, Substep S22 obtains the traction direction based on the directional traction strength of two keywords of the same type, The process includes a substep S23 which calculates the tension diffusion coefficient of a keyword based on semantic similarity and traction direction.
[0033] In this embodiment, the formula for obtaining the traction direction in S22 is as follows: JPEG0007840628000020.jpg9170JPEG0007840628000021.jpg33170
[0034] In this embodiment, the formula for calculating the tension diffusion coefficient of the keyword in S23 is as follows: JPEG0007840628000022.jpg11170JPEG0007840628000023.jpg36170
[0035] JPEG0007840628000024.jpg12170
[0036] In this embodiment, semantic similarity between keywords can be calculated using cosine similarity, or it can be calculated using conventional techniques. For example, similarity can be calculated after extracting the contextual semantic vectors of keywords using context-sensing pre-trained models such as selective Sentence-BERT or the general text embedding model GTE (General Text Embedding). Alternatively, similarity can be calculated based on semantic-original concept relationships by relying on conventional semantic knowledge bases such as CNKI or WordNet. Furthermore, semantic relationships can be quantified using the neighboring features of keywords in a large corpus using the Secondary Co-origin Mutual Information Method (SOC-PMI).
[0037] JPEG0007840628000025.jpg28170
[0038] In this embodiment, the multiplication formula in S3 is as follows: JPEG0007840628000026.jpg8170JPEG0007840628000027.jpg20170
[0039] In this embodiment, S4 is Substep S41 calculates an interaction correction coefficient based on the co-occurrence weights, semantic similarity, and TF-IDF weights of the two keywords, Substep S42 involves multiplying the interaction correction coefficient by the interaction influence coefficient and adding 1 to obtain the global interaction gain coefficient for the keyword. The process includes a substep S43 in which the global interaction gain coefficient of the keyword is multiplied by the corresponding directional traction strength to obtain a second corrected directional traction strength for the keyword type.
[0040] In this embodiment, the formula for calculating the interaction correction coefficient in S41 is as follows: JPEG0007840628000028.jpg11170JPEG0007840628000029.jpg44170
[0041] JPEG0007840628000030.jpg23170
[0042] The TF-IDF weight of a keyword is the TF-IDF weight of the keyword in the text to be classified.
[0043] In this embodiment, the formula for calculating the global interaction gain coefficient of the keyword in S42 is as follows: JPEG0007840628000031.jpg8170JPEG0007840628000032.jpg18170
[0044] In this embodiment, the multiplication formula in S43 is as follows: JPEG0007840628000033.jpg8170JPEG0007840628000034.jpg13170
[0045] JPEG0007840628000035.jpg29170
[0046] This invention obtains an interaction correction coefficient by multiplying co-occurrence weights, semantic similarity, and the normalized TF-IDF weight of another keyword, representing the interaction contribution of the j-th keyword in the text to be classified to the m-th keyword of the m-th type, further converting it into a global interaction gain coefficient through control of the interaction influence coefficient, and finally completing a quadratic correction for directional traction strength, thereby effectively quantifying the multidimensional interaction relationships between keywords in the text to be classified.
[0047] In this embodiment, S5 is Substep S51 involves multiplying the normalized TF-IDF weight of the keyword by the first corrected directional traction strength to obtain a first enhancement value. Substep S52 involves multiplying the normalized TF-IDF weight of the keyword by a second corrected directional traction strength to obtain a second enhancement value. Substep S53 involves adding the second enhancement value and the first enhancement value, taking the average, and obtaining the total enhancement value. Substep S54 involves summing the overall reinforcement values of all keywords in the type dimension to obtain the initial degree of belonging to each type of text, Substep S55 calculates the semantic distance between any two types based on the reference word sets of each type, The process includes a substep S56 in which nonlinear competitive modulation is performed on the initial degree of belonging for each type of text based on semantic distance to obtain the final degree of belonging for each type of text.
[0048] In this embodiment, the formula for obtaining the first enhancement value in S51 is as follows: JPEG0007840628000036.jpg10170JPEG0007840628000037.jpg20170
[0049] In this example, the formula for obtaining the second enhancement value in S52 is as follows: JPEG0007840628000038.jpg10170JPEG0007840628000039.jpg15170
[0050] In this embodiment, the formula for calculating the overall reinforcement value in S53 is as follows: JPEG0007840628000040.jpg12170JPEG0007840628000041.jpg8170
[0051] In this embodiment, S53 may be obtained by weighting the second enhancement value and the first enhancement value using a weighting coefficient to obtain an overall enhancement value.
[0052] In this embodiment, the formula for calculating the initial degree of belonging to each type of text in S54 is as follows: JPEG0007840628000042.jpg9170JPEG0007840628000043.jpg8170
[0053] In this embodiment, the formula for calculating the semantic distance between any two types in S55 is as follows: JPEG0007840628000044.jpg15170JPEG0007840628000045.jpg28170
[0054] JPEG0007840628000046.jpg14170
[0055] In this embodiment, the formula for calculating the final degree of belonging in S56 is as follows: JPEG0007840628000047.jpg12170JPEG0007840628000048.jpg35170
[0056] This invention first weights and enhances the directional traction strength, which has been corrected twice by normalized TF-IDF weights, to allow more important keywords in the text to be classified to contribute a stronger type-pulling effect. Furthermore, by mean value fusion, it takes into account the feature values of two correction logics, semantic tension diffusion and multidimensional interaction of multiple keywords, to more comprehensively reflect the type-pulling ability of keywords in the overall enhancement value. Subsequently, by summing the type dimensions, it generates an initial degree of belonging to each type in the text. Based on this, it calculates the semantic distance by the intersection of type-reference words with IDF weighting, quantifies the core semantic relationships between types, avoids interference from generic words, and finally generates stronger degree of belonging suppression for semantically closer types through exponential nonlinear competitive modulation.
[0057] In this embodiment, the competition intensity coefficient β is used to control the "amplitude valve" of the competitive suppression intensity between different species. A relatively small value of β indicates weak mutual suppression between both species, and the final degree of belonging retains more initial features. In this embodiment, β ∈ [0.5, 1.5].
[0058] In this embodiment, S6 is Substep S61 calculates a first semantic potential energy value based on a first enhancement value corresponding to each keyword in the text, Substep S62 calculates a second semantic potential energy value based on a second enhancement value corresponding to each keyword in the text, Substep S63 involves weighting the final degree of belonging, the first semantic potential energy value, and the second semantic potential energy value to obtain a classification objective function. The process includes a substep S64 in which a determination value is calculated for all categories based on a classification objective function, and the category with the highest determination value is determined to be the category to which the text belongs.
[0059] In this embodiment, the formula for calculating the first semantic potential energy value in S61 is as follows: JPEG0007840628000049.jpg13170JPEG0007840628000050.jpg28170In this embodiment, the formula for calculating the second semantic potential energy value in S62 is as follows: JPEG0007840628000051.jpg13170JPEG0007840628000052.jpg22170
[0060] In this embodiment, the classification objective function in S64 is as follows: JPEG0007840628000053.jpg8170JPEG0007840628000054.jpg33170
[0061] JPEG0007840628000055.jpg7170
[0062] In this embodiment, the first and second weighting coefficients are selected by scanning different combinations of η1 and η2 within a predetermined range and calculating metrics such as accuracy and F1 value in the validation set of the classification objective function to select the optimal combination of η1 and η2.
[0063] The calculation of semantic potential energy values combines the sum and variance of reinforcement values to maintain the overall strength of the pull on keyword types, while the variance term reflects the distributional stability of reinforcement values among keywords. A larger variance indicates an unbalanced pull on this type of keyword, dynamically weakening the potential energy value, while a smaller variance indicates a more stable pull, preserving the potential energy value. This allows classification decisions to consider not only the degree of belonging but also the internal stability of keyword features, thus avoiding misjudgments driven by a few strong keywords.
[0064] Example 2, a text classification system comprising a traction strength acquisition unit, a tension diffusion coefficient acquisition unit, a first correction unit, a second correction unit, a fusion attribute calculation unit, and a classification unit, The traction strength acquisition unit is used to acquire the directional traction strength for each type of keyword in the text. The tension diffusion coefficient acquisition unit is used to acquire the tension diffusion coefficient of keywords based on the semantic similarity between keywords in the text and the direction of traction. The first correction unit is used to obtain the first corrected directional tensile strength for the keyword type by multiplying the tension diffusion coefficient of the keyword by the corresponding directional tensile strength. The second correction unit is used to correct the corresponding directional traction strength based on the interaction correction coefficient to obtain a second corrected directional traction strength for the keyword type. The fused belongingness calculation unit is used to obtain the final belongingness for each type of text by fusing the first corrected directional traction intensity and the second corrected directional traction intensity, and further introducing nonlinear competitive modulation by semantic distance between types. The classification unit obtains the first and second semantic potential energy values, respectively, and combines them with the final degree of belonging to establish a classification objective function and obtain the text type.
[0065] The specific implementation process for Example 2 is the same as that for Example 1.
[0066] Example 3 is a text classification device comprising a processor and a computer program, the computer program being executed by the processor to implement the text classification method described in Example 1.
[0067] Example 4 is a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the text classification method described in Example 1 is realized.
[0068] This invention overcomes the assumption of independent keyword contributions, describing semantic relationships and traction direction propagation between keywords using tension diffusion coefficients, capturing the complex interaction effects of multiple keywords using interaction correction coefficients, and achieving dual refinement correction for directional traction strength. Next, it introduces a nonlinear competitive modulation mechanism based on type semantic distance, dynamically describing semantic relationships and competitive constraints between types, avoiding the ambiguity problem of classification boundaries due to linear superposition. Finally, by fusing the final degree of belonging with semantic potential energy values that reflect keyword traction stability, it constructs a classification objective function, further improving the stability and accuracy of the classification results, thereby effectively solving the problem of low classification accuracy in semantically similar types or multiple thematic text scenarios.
[0069] The foregoing are merely preferred embodiments of the present invention and are not intended to limit it. To those skilled in the art, the present invention is subject to various modifications and changes. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should all fall within the scope of protection of the present invention.
Claims
1. A text classification method performed by a processor, Step S1 involves obtaining the directional traction strength for each type of keyword in the text, Step S2 involves obtaining the tension diffusion coefficient of a keyword based on the semantic similarity between keywords in the text and the direction of traction. Step S3 involves multiplying the tension diffusion coefficient of the keyword by the corresponding directional tensile strength to obtain a first corrected directional tensile strength for the keyword type, S4 corrects the corresponding directional traction strength based on the interaction correction coefficient to obtain a second corrected directional traction strength for the keyword type. Step S5 involves fusing the first corrected directional traction strength and the second corrected directional traction strength, and further introducing nonlinear competitive modulation based on the semantic distance between the types to obtain the final degree of belonging to each type of text. Step S6 involves obtaining the first semantic potential energy value and the second semantic potential energy value, combining them with the final degree of belonging to establish a classification objective function and obtain the type of text. A text classification method characterized by including
2. The aforementioned S1 is, Substep S11 calculates the semantic similarity between keywords in the text and reference words in various reference word sets, Substep S12 involves multiplying the semantic similarity by the TF-ICF weight of the corresponding reference word to obtain a traction value, The text classification method according to claim 1, characterized by including a substep S13 of selecting the maximum traction value from the traction values corresponding to each reference word of each type to obtain the directional traction strength of this keyword for this type.
3. The aforementioned S2 is, Substep S21 calculates the semantic similarity of two keywords in the text, Substep S22 obtains the traction direction based on the directional traction strength of two keywords of the same type, The text classification method according to claim 1, further comprising a substep S23 for calculating the tension diffusion coefficient of a keyword based on semantic similarity and traction direction.
4. The formula for calculating the tension diffusion coefficient of the keyword in S23 is as follows:
5. The aforementioned S4 is, Substep S41 calculates an interaction correction coefficient based on the co-occurrence weights, semantic similarity, and TF-IDF weights of two keywords, Substep S42 involves multiplying the interaction correction coefficient by the interaction influence coefficient and adding 1 to obtain the global interaction gain coefficient for the keyword. The text classification method according to claim 1, further comprising substep S43, which involves multiplying the global interaction gain coefficient of a keyword by the corresponding directional traction strength to obtain a second corrected directional traction strength for the keyword type.
6. The formula for calculating the interaction correction coefficient in S41 is as follows:
7. The aforementioned S5 is, Substep S51 involves multiplying the normalized TF-IDF weight of the keyword by the first corrected directional traction strength to obtain a first enhancement value. Substep S52 involves multiplying the normalized TF-IDF weight of the keyword by a second corrected directional traction strength to obtain a second enhancement value. Substep S53 involves adding the second enhancement value and the first enhancement value, taking the average, and obtaining the total enhancement value. Substep S54 involves summing the overall reinforcement values of all keywords in the type dimension to obtain the initial degree of belonging to each type of text, Substep S55 calculates the semantic distance between any two types based on the set of reference words for each type, The text classification method according to claim 1, characterized by including a substep S56 in which nonlinear competitive modulation is performed on the initial degree of belonging for each type of text based on semantic distance to obtain a final degree of belonging for each type of text.
8. The formula for calculating the final degree of affiliation in S56 is as follows:
9. The aforementioned S6 is, Substep S61 calculates a first semantic potential energy value based on a first enhancement value corresponding to each keyword in the text, Substep S62 calculates a second semantic potential energy value based on a second enhancement value corresponding to each keyword in the text, Substep S63 involves weighting the final degree of belonging, the first semantic potential energy value, and the second semantic potential energy value to obtain a classification objective function. The text classification method according to claim 7, further comprising substep S64, which calculates a determination value for all types based on a classification objective function, and determines the type with the largest determination value as the text's type of classification.
10. The formula for calculating the first semantic potential energy value in S61 is as follows: The formula for calculating the second semantic potential energy value in S62 is as follows:
11. A text classification system, which is implemented based on the text classification method described in any one of claims 1 to 10, It includes a traction strength acquisition unit, a tension diffusion coefficient acquisition unit, a first correction unit, a second correction unit, a fusion belongingness calculation unit, and a classification unit. The aforementioned traction strength acquisition unit is used to acquire the directional traction strength for each type of keyword in the text. The tension diffusion coefficient acquisition unit is used to acquire the tension diffusion coefficient of a keyword based on the semantic similarity between keywords in the text and the direction of traction. The first correction unit is used to obtain a first corrected directional tensile strength for the keyword type by multiplying the tension diffusion coefficient of the keyword by the corresponding directional tensile strength. The second correction unit is used to correct the corresponding directional traction strength based on the interaction correction coefficient to obtain a second corrected directional traction strength for the keyword type. The aforementioned fused belongingness calculation unit is used to obtain the final belongingness for each type of text by fusing the first corrected directional traction intensity and the second corrected directional traction intensity, and further introducing nonlinear competitive modulation by semantic distance between types. The classification unit obtains a first semantic potential energy value and a second semantic potential energy value, and combines them with the final degree of belonging to establish a classification objective function and obtain the type of text. A text classification system characterized by the following features.
12. A text classification device comprising a processor and a computer program, wherein the processor executes the computer program to realize the text classification method described in any one of claims 1 to 10. A text classification device characterized by the following features.
13. A computer-readable storage medium in which a computer program is stored, wherein the processor stores the computer program for realizing the text classification method described in any one of claims 1 to 10. A computer-readable storage medium characterized by the following features.
Citation Information
Patent Citations
Text processing method and apparatus, computer readable storage medium and computer device
CN109446525A
Text classification model training methods, text classification methods and related devices
CN111382269B
Document processing method, device and equipment
CN118657138A
Information processor, information processing method and computer program
JP2007193380A