A text classification method, system, device, and medium

By obtaining the directional traction strength and tension diffusion coefficient of keywords to categories, and combining the interaction correction factor and nonlinear competitive modulation, the problem of insufficient characterization of semantic relationships in text classification is solved, thereby improving classification accuracy and stability.

CN121681836BActive Publication Date: 2026-04-21SICHUAN LIANGSHANSHUILUOHE ELECTRICITY DEV CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN LIANGSHANSHUILUOHE ELECTRICITY DEV CO LTD
Filing Date
2026-02-10
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively characterize the semantic relationships between keywords and their mutual influence in different category contexts during text classification, resulting in low classification accuracy. This is especially true in semantically similar categories or multi-topic text scenarios where classification boundaries are blurred and attribution is unstable.

Method used

By obtaining the directional traction strength of keywords for each category, and combining the semantic similarity between keywords and the traction direction to calculate the tension diffusion coefficient, an interaction correction factor is introduced to correct the directional traction strength, and nonlinear competitive modulation is introduced through the semantic distance between categories to construct a classification objective function.

Benefits of technology

It achieves quantitative representation of semantic associations of keywords, accurately captures the propagation effect of semantic influence, solves the problem of low classification accuracy in existing technologies, and improves the rationality and reliability of classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121681836B_ABST
    Figure CN121681836B_ABST
Patent Text Reader

Abstract

This invention discloses a text classification method, system, device, and medium, belonging to the field of text classification technology. The invention obtains the directional pull strength of each keyword in the text to different categories; then, combining the semantic similarity and pull direction between keywords, it calculates the tension diffusion coefficient of the keywords, multiplies it by the directional pull strength to obtain a first corrected directional pull strength; simultaneously, it introduces an interaction correction factor to further correct the directional pull strength, obtaining a second corrected directional pull strength. Based on this, the two types of corrected directional pull strengths are fused, and nonlinear competitive modulation is introduced using the semantic distance between categories to determine the final assignment degree of the text to each category. Finally, by obtaining the first and second semantic potential values ​​and combining them with the final assignment degree, a classification objective function is established to accurately determine the text category. This invention effectively strengthens the correlation matching degree between keywords and categories, improving the accuracy and reliability of text classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text classification technology, specifically to a text classification method, system, device, and medium. Background Technology

[0002] With the development of natural language processing technology, text classification has been widely applied in scenarios such as information retrieval, public opinion analysis, and content recommendation. In existing technologies, common text classification methods are usually based on keyword statistical features or vectorized representations. After mapping the text to a feature space, a classification model is used to calculate the similarity or probability distribution between the text and each category. These methods generally treat the contributions of keywords to different categories as independent or perform only simple weighting, making it difficult to characterize the semantic relationships between keywords and their mutual influence in different category contexts.

[0003] However, in practical applications, keywords in text often have complex semantic relationships. The similarity, co-occurrence relationship, and directional influence on categories among different keywords all affect the classification results. Existing technologies typically employ linear superposition or static weighting mechanisms when dealing with multi-keyword interactions, semantic competition between categories, and the propagation effect of semantic influence. This leads to insufficient characterization of keyword semantic diffusion and non-linear competition between categories. Consequently, in semantically similar categories or multi-topic text scenarios, problems such as blurred classification boundaries, unstable attribution, and low classification accuracy easily arise. Summary of the Invention

[0004] In view of the above-mentioned shortcomings in the prior art, the present invention provides a text classification method, system, device and medium that solves the problem of low classification accuracy in the prior art.

[0005] To achieve the above-mentioned objectives, the technical solution adopted by this invention is: a text classification method, comprising the following steps:

[0006] S1. Obtain the directional pull strength of keywords in the text for each category;

[0007] S2. Based on the semantic similarity between keywords in the text and the direction of attraction, obtain the tension diffusion coefficient of the keywords;

[0008] S3. Multiply the tension diffusion coefficient of the keyword by the corresponding directional traction strength to obtain the first corrected directional traction strength of the keyword for the category;

[0009] S4. Based on the interaction correction factor, the corresponding directional traction strength is corrected to obtain the second corrected directional traction strength of the keyword to the category;

[0010] S5. The first modified directional traction strength and the second modified directional traction strength are fused together, and then nonlinear competitive modulation is introduced through the semantic distance between categories to obtain the final belonging degree of the text to each category.

[0011] S6. Obtain the first semantic potential value and the second semantic potential value respectively, and combine them with the final attribution degree to establish a classification objective function to obtain the category of the text.

[0012] Furthermore, S1 includes the following sub-steps:

[0013] S11. Calculate the semantic similarity between keywords in the text and reference words in the reference word set for each category;

[0014] S12. Multiply the semantic similarity by the TF-ICF weight of the corresponding reference word to obtain the traction value;

[0015] S13. Select the maximum traction value from the traction values ​​corresponding to each reference word in each category to obtain the directional traction strength of the keyword for that category.

[0016] Furthermore, S2 includes the following sub-steps:

[0017] S21. Calculate the semantic similarity between two keywords in the text;

[0018] S22. Based on the directional traction strength of two keywords in the same category, obtain the traction direction;

[0019] S23. Calculate the tension diffusion coefficient of keywords based on semantic similarity and traction direction.

[0020] Furthermore, the formula for calculating the tension diffusion coefficient of the keyword in S23 is:

[0021] ,

[0022] in, For the first The keyword in the first The tension diffusion coefficient of each category, For the first The first keyword and the first The semantic similarity of keywords. For the first The first keyword and the first The keyword in the first The traction direction of each category, and For keyword numbering, To select the maximum value, This represents the number of keywords in the text.

[0023] Furthermore, S4 includes the following sub-steps:

[0024] S41. Calculate the interaction correction factor based on the co-occurrence weights and semantic similarity of the two keywords, as well as the TF-IDF weights.

[0025] S42. Multiply the interaction correction factor by the interaction influence coefficient and add 1 to obtain the global interaction gain coefficient of the keyword.

[0026] S43. Multiply the global interaction gain coefficient of the keyword by the corresponding directional traction strength to obtain the second modified directional traction strength of the keyword to the category.

[0027] Furthermore, the formula for calculating the interaction correction factor in S41 is:

[0028] ,

[0029] in, For the first The keyword for the first Interaction correction factors for each category, For the first The first keyword and the first The keyword in the first Co-occurrence weights of each category, For the first The first keyword and the first The semantic similarity of keywords. For the first The TF-IDF weights of each keyword. The TF-IDF weight of the first keyword. For the first The TF-IDF weights of each keyword, with max representing the maximum value. The number of keywords in the text. and This is the keyword number.

[0030] Furthermore, S5 includes the following sub-steps:

[0031] S51. Multiply the normalized TF-IDF weight of the keyword by the first modified directional traction strength to obtain the first enhancement value;

[0032] S52. Multiply the normalized TF-IDF weight of the keyword by the second modified directional traction strength to obtain the second reinforcement value;

[0033] S53. Add the second enhancement value and the first enhancement value and take the average value to obtain the comprehensive enhancement value;

[0034] S54. Sum the comprehensive reinforcement values ​​of all keywords along the category dimension to obtain the initial affiliation degree of the text to each category;

[0035] S55. Calculate the semantic distance between any two categories based on the reference word set for each category;

[0036] S56. Based on semantic distance, perform nonlinear competitive modulation on the initial classification degree of each category of the text to obtain the final classification degree of the text to each category.

[0037] Furthermore, the formula for calculating the final affiliation degree in S56 is as follows:

[0038] ,

[0039] in, For the text pair The final affiliation of each category, It is an exponential function. For the first The category and the first The semantic distance between each category For the text pair The initial affiliation of each category, For the text pair The initial affiliation of each category, To prevent parameters with a denominator of 0, This is the competition intensity coefficient. For the number of categories, and This is the category number.

[0040] Furthermore, S6 includes the following sub-steps:

[0041] S61. Calculate the first semantic potential value based on the first reinforcement value corresponding to each keyword in the text;

[0042] S62. Calculate the second semantic potential value based on the second reinforcement value corresponding to each keyword in the text;

[0043] S63. Weight the final attribution score, the first semantic potential value, and the second semantic potential value to obtain the classification objective function;

[0044] S64. Based on the classification objective function, by calculating the decision values ​​of all categories, the category with the largest decision value is determined as the category to which the text belongs.

[0045] Furthermore, the formula for calculating the first semantic potential value in S61 is as follows:

[0046] ,

[0047] in, For the text pair The first semantic potential value of each category The number of keywords in the text. For the first The keyword for the first The first enhancement value for each category, for indivual variance For keyword numbering, For the category number;

[0048] The formula for calculating the second semantic potential energy value in S62 is:

[0049] ,

[0050] in, For the text pair The second semantic potential value of each category For the first The keyword for the first The second enhancement value for each category, for indivual The variance.

[0051] The beneficial effects of this invention are as follows:

[0052] 1. This invention achieves a quantitative representation of the semantic association between keywords by obtaining the directional traction strength of keywords to each category and calculating the tension diffusion coefficient by combining the semantic similarity between keywords with the traction direction. This breaks through the limitations of existing technologies that treat keyword contributions as mutually independent or simply weighted. Simultaneously, by correcting the directional traction strength through the tension diffusion coefficient, it accurately captures the propagation effect of keyword semantic influence, enabling the category traction effect of each keyword to diffuse reasonably according to semantic association. This aligns with the semantic rules of keyword mutual influence in actual texts and solves the problem of insufficient characterization of keyword semantic diffusion in existing technologies.

[0053] 2. This invention introduces an interaction correction factor to perform a secondary correction on the directional traction strength, achieving quantitative correction of complex interaction relationships between multiple keywords. This overcomes the shortcomings of existing technologies that use linear superposition and static weighting mechanisms to handle multi-keyword interactions. By adapting the interaction characteristics of different keyword combinations to the interaction correction factor, the traction strength of keywords to categories is made more consistent with the actual semantic expression, avoiding distortion of traction strength caused by ignoring keyword interactions.

[0054] 3. Based on the directional traction strength after two corrections, this invention introduces nonlinear competitive modulation through semantic distance between categories, overcoming the limitations of existing technologies in linear modeling of semantic competition between categories. For semantically similar categories and multi-topic text scenarios, nonlinear competitive modulation can dynamically adjust the category affiliation of the text according to the degree of semantic association between categories, weakening semantic overlap interference between categories, strengthening the expression of unique features of different categories, and effectively solving the problems of ambiguous classification boundaries and unstable affiliation in existing technologies in such scenarios.

[0055] 4. This invention constructs a classification objective function using a first semantic potential value, a second semantic potential value, and a final attribution degree. It integrates semantic potential features with category attribution degree into a fusion model, enabling classification decisions to not only rely on the keyword's influence on the category but also consider the overall semantic potential distribution of the text, further improving the rationality and reliability of the classification results. Compared to existing technologies that rely solely on similarity or probability distribution, this invention provides a more comprehensive classification decision-making basis, effectively avoiding classification bias caused by single-dimensional judgments and improving classification accuracy. Attached Figure Description

[0056] Figure 1 This is a flowchart of a text classification method. Detailed Implementation

[0057] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0058] Example 1, as Figure 1 As shown, a text classification method includes the following steps:

[0059] S1. Obtain the directional pull strength of keywords in the text for each category;

[0060] S2. Based on the semantic similarity between keywords in the text and the direction of attraction, obtain the tension diffusion coefficient of the keywords;

[0061] S3. Multiply the tension diffusion coefficient of the keyword by the corresponding directional traction strength to obtain the first corrected directional traction strength of the keyword for the category;

[0062] S4. Based on the interaction correction factor, the corresponding directional traction strength is corrected to obtain the second corrected directional traction strength of the keyword to the category;

[0063] S5. The first modified directional traction strength and the second modified directional traction strength are fused together, and then nonlinear competitive modulation is introduced through the semantic distance between categories to obtain the final belonging degree of the text to each category.

[0064] S6. Obtain the first semantic potential value and the second semantic potential value respectively, and combine them with the final attribution degree to establish a classification objective function to obtain the category of the text.

[0065] This invention first constructs a reference word set for each category in the category set. Category set: ,in, For a set of categories, As the first category, For the first Categories For the first Categories For the number of categories, This is the category number.

[0066] Reference vocabulary list: ,in, For the first A collection of reference words for each category. For the first The first reference word in each category, For the first The first category One reference word, For the first The first category One reference word, For the reference word in a category, The number of reference words in a set of reference words.

[0067] No. The first category one reference word TF-ICF weights (term frequency - inverse class frequency) It equals the product of the average term frequency within the category and the inverse document frequency, i.e. ,in, For the first The first category one reference word In the Frequency of occurrence of all samples in each category For the first The total number of words in all samples of each category For the number of categories, For the first The first category one reference word The number of different categories that appeared in the samples corresponding to all classification categories.

[0068] In this embodiment, each category of sample refers to a sample that has been classified into the corresponding category.

[0069] This invention accurately captures the core semantic features specific to each category by constructing a reference word set for each category and assigning TF-ICF weights, thus providing clear category semantic anchors for subsequent calculations and avoiding the semantic ambiguity caused by generalized keyword statistical features.

[0070] In this embodiment, S1 includes the following sub-steps:

[0071] S11. Calculate the semantic similarity between keywords in the text and reference words in the reference word set for each category;

[0072] S12. Multiply the semantic similarity by the TF-ICF weight of the corresponding reference word to obtain the traction value;

[0073] S13. Select the maximum traction value from the traction values ​​corresponding to each reference word in each category to obtain the directional traction strength of the keyword for that category.

[0074] In this embodiment, the formula for calculating semantic similarity in S11 is:

[0075] ,

[0076] in, For the first The first keyword and the first The first category Semantic similarity of reference words, For the first The first category One reference word, It is a norm 2. For vector dot product, For the reference word in a category, For the category number, This represents the keyword's identifier. When calculating similarity, if the vector dimensions are inconsistent, zeros can be used for padding. For the first One keyword.

[0077] In this embodiment, the expression for the directional traction strength of the keyword to the category is obtained as follows:

[0078] ,

[0079] in, For the first The keyword for the first The directional traction strength of each category, For the first The first category one reference word TF-ICF weights, For from the first Filter the maximum traction value from all traction values ​​corresponding to each category. This is the traction value.

[0080] The traction value calculation of this invention integrates semantic similarity and TF-ICF weights, ensuring both the semantic fit between keywords and reference words and introducing the category discrimination of reference words. This allows the traction value to not only reflect semantic association but also directly embody the practical value of that association for text classification. Since this invention calculates the traction value between each reference word and keyword in each category, the maximum traction value is selected from the traction values ​​corresponding to a category as the directional traction strength.

[0081] In this embodiment, S2 includes the following sub-steps:

[0082] S21. Calculate the semantic similarity between two keywords in the text;

[0083] S22. Based on the directional traction strength of two keywords in the same category, obtain the traction direction;

[0084] S23. Calculate the tension diffusion coefficient of keywords based on semantic similarity and traction direction.

[0085] In this embodiment, the formula for obtaining the traction direction in S22 is:

[0086] ,

[0087] in, For the first The first keyword and the first The keyword in the first The traction direction of each category, For the first The keyword for the first The directional traction strength of each category, For the first The keyword for the first The directional traction strength of each category, For sign functions, sign functions When the value inside the parentheses is greater than or equal to 0, for The value is assigned as 1, and if the value inside the parentheses is less than 0, then... The value is assigned to -1.

[0088] In this embodiment, the formula for calculating the tension diffusion coefficient of the keyword in S23 is:

[0089] ,

[0090] in, For the first The keyword in the first The tension diffusion coefficient of each category, For the first The first keyword and the first The semantic similarity of keywords. For the first The first keyword and the first The keyword in the first The traction direction of each category, and For keyword numbering, To select the maximum value, The number of keywords in the text. For the first One keyword, For the first One keyword.

[0091] In this embodiment, and The calculation methods are the same, both using cosine similarity.

[0092] In this embodiment, in addition to using cosine similarity to calculate the semantic similarity between keywords, existing technologies can also be used: context-aware pre-trained models such as Sentence-BERT and General Text Embedding (GTE) can be selected to extract the contextual semantic vectors of keywords and then calculate the similarity; or existing semantic knowledge bases such as CNKI and WordNet can be used to calculate the similarity based on semantic primitive concept relations; or the second-order co-occurrence point mutual information method (SOC-PMI) can be used to quantify semantic association through the neighborhood features of keywords in a large-scale corpus.

[0093] This invention calculates the semantic similarity between text keywords and determines the traction direction by combining the difference in directional traction strength under the same category, and finally obtains the tension diffusion coefficient. It can quantify the directional diffusion effect of semantic association between keywords on the traction strength of the category. It not only reflects the synergistic and competitive relationship between semantically similar keywords, but also avoids the diffusion failure problem caused by the coefficient being too small through the constraint of max[0.1,...].

[0094] In this embodiment, the multiplication formula in S3 is:

[0095] ,

[0096] in, For the first The keyword for the first The first modified directional traction strength of each category For the first The keyword in the first The tension diffusion coefficient of each category, For the first The keyword for the first The directional traction strength of each category.

[0097] In this embodiment, S4 includes the following sub-steps:

[0098] S41. Calculate the interaction correction factor based on the co-occurrence weights and semantic similarity of the two keywords, as well as the TF-IDF weights.

[0099] S42. Multiply the interaction correction factor by the interaction influence coefficient and add 1 to obtain the global interaction gain coefficient of the keyword.

[0100] S43. Multiply the global interaction gain coefficient of the keyword by the corresponding directional traction strength to obtain the second modified directional traction strength of the keyword to the category.

[0101] In this embodiment, the formula for calculating the interaction correction factor in S41 is:

[0102] ,

[0103] in, For the first The keyword for the first Interaction correction factors for each category, For the first The first keyword and the first The keyword in the first Co-occurrence weights of each category, For the first The first keyword and the first The semantic similarity of keywords. For the first The TF-IDF weights of each keyword. The TF-IDF weight of the first keyword. For the first The TF-IDF weights of each keyword, with max representing the maximum value. The number of keywords in the text. and This is the keyword number.

[0104] No. The first keyword and the first The keyword in the first Co-occurrence weights of each category The methods for obtaining this include: first, counting the first... The number of samples containing both keywords in all samples of each category is then divided by the number of samples in the first category. The total number of samples in each category is the co-occurrence weight.

[0105] The TF-IDF weight of a keyword is the TF-IDF weight of the keywords in the text to be classified.

[0106] In this embodiment, the formula for calculating the global interaction gain coefficient of the keyword in S42 is as follows:

[0107] ,

[0108] in, For the first The keyword for the first Global interaction gain coefficients for each category, In this embodiment, the interaction coefficient is used. Take 0.6.

[0109] In this embodiment, the multiplication formula in S43 is:

[0110] ,

[0111] in, For the first The keyword for the first The second modified directional traction strength for each category.

[0112] In this embodiment, an interaction influence coefficient is set. The core purpose is to modify the interaction factor. The effect of the global interaction gain coefficient can be flexibly and controllably adjusted so that it not only conforms to the actual keyword interaction rules of the text, but also avoids distortion of the directional traction strength caused by excessive or insufficient interaction correction.

[0113] This invention multiplies the co-occurrence weight, semantic similarity, and the normalized TF-IDF weight of another keyword to obtain an interaction correction factor, which reflects the interaction correction factor in the text to be classified. The keyword for the first The keyword in the first The interaction contribution of each category is then converted into a global interaction gain coefficient through interaction influence coefficient adjustment, which finally completes the secondary correction of the directional traction strength and effectively quantifies the multidimensional interaction relationship between keywords in the text to be classified.

[0114] In this embodiment, S5 includes the following sub-steps:

[0115] S51. Multiply the normalized TF-IDF weight of the keyword by the first modified directional traction strength to obtain the first enhancement value;

[0116] S52. Multiply the normalized TF-IDF weight of the keyword by the second modified directional traction strength to obtain the second reinforcement value;

[0117] S53. Add the second enhancement value and the first enhancement value and take the average value to obtain the comprehensive enhancement value;

[0118] S54. Sum the comprehensive reinforcement values ​​of all keywords along the category dimension to obtain the initial affiliation degree of the text to each category;

[0119] S55. Calculate the semantic distance between any two categories based on the reference word set for each category;

[0120] S56. Based on semantic distance, perform nonlinear competitive modulation on the initial classification degree of each category of the text to obtain the final classification degree of the text to each category.

[0121] In this embodiment, the formula for obtaining the first enhancement value in S51 is:

[0122] ,

[0123] in, For the first The keyword for the first The first enhancement value for each category, For the first The TF-IDF weights of each keyword. For the first The keyword for the first The first modified directional traction strength for each category.

[0124] In this embodiment, the formula for obtaining the second strengthening value in S52 is:

[0125] ,

[0126] in, For the first The keyword for the first The second enhancement value for each category, For the first The keyword for the first The second modified directional traction strength for each category.

[0127] In this embodiment, the formula for calculating the overall enhancement value in S53 is as follows:

[0128] ,

[0129] in, For the first The keyword for the first The overall enhancement value for each category.

[0130] In this embodiment, S53 can also use a weighting coefficient to weight the second enhancement value and the first enhancement value to obtain a comprehensive enhancement value.

[0131] In this embodiment, the formula for calculating the initial affiliation degree of the text to each category in S54 is as follows:

[0132] ,

[0133] in, For the text pair The initial affiliation of each category.

[0134] In this embodiment, the formula for calculating the semantic distance between any two categories in S55 is:

[0135] ,

[0136] in, For the first The category and the first The semantic distance between each category For the first A collection of reference words for each category. For the first A collection of reference words for each category. For the first The first category One reference word, Reference word Inverse document frequency.

[0137] Reference words Inverse document frequency This refers to using all categorized samples as a global document library and calculating reference terms. Inverse document frequency.

[0138] In this embodiment, the formula for calculating the final affiliation degree in S56 is:

[0139] ,

[0140] in, For the text pair The final affiliation of each category, It is an exponential function. For the first The category and the first The semantic distance between each category For the text pair The initial affiliation of each category, For the text pair The initial affiliation of each category, To prevent parameters with a denominator of 0, This is the competition intensity coefficient. For the number of categories, and This is the category number.

[0141] This invention first strengthens the directional traction intensity of the two corrections by normalizing TF-IDF weights, allowing more important keywords in the text to be classified to contribute a stronger category traction effect. Then, through mean fusion, it takes into account the feature value of both semantic tension diffusion and multi-keyword multi-dimensional interaction correction logics, so that the comprehensive enhancement value more comprehensively reflects the keyword's ability to traction the category. Subsequently, by summing and aggregating the category dimensions, the initial affiliation degree of the text to each category is generated. On this basis, the semantic distance is calculated by the intersection of category reference words weighted by IDF, which quantifies the core semantic association between categories and avoids interference from common words. Finally, through exponential nonlinear competitive modulation, the affiliation degree of semantically similar categories is suppressed more strongly.

[0142] In this embodiment, the competition intensity coefficient An "amplitude valve" is used to control the intensity of competition suppression between different categories. Smaller values ​​result in weaker mutual inhibition between the two factors, leading to the final assignment degree retaining more of the initial features. In this embodiment, .

[0143] In this embodiment, S6 includes the following sub-steps:

[0144] S61. Calculate the first semantic potential value based on the first reinforcement value corresponding to each keyword in the text;

[0145] S62. Calculate the second semantic potential value based on the second reinforcement value corresponding to each keyword in the text;

[0146] S63. Weight the final attribution score, the first semantic potential value, and the second semantic potential value to obtain the classification objective function;

[0147] S64. Based on the classification objective function, by calculating the decision values ​​of all categories, the category with the largest decision value is determined as the category to which the text belongs.

[0148] In this embodiment, the formula for calculating the first semantic potential energy value in S61 is:

[0149] ,

[0150] in, For the text pair The first semantic potential value of each category The number of keywords in the text. For the first The keyword for the first The first enhancement value for each category, for indivual variance For keyword numbering, For the category number;

[0151] In this embodiment, the formula for calculating the second semantic potential energy value in S62 is:

[0152] ,

[0153] in, For the text pair The second semantic potential value of each category For the first The keyword for the first The second enhancement value for each category, for indivual The variance.

[0154] In this embodiment, the classification objective function in S64 is:

[0155] ,

[0156] in, For the text pair The category decision value for each category, For the text pair The final affiliation of each category, For the text pair The first semantic potential value of each category For the text pair The second semantic potential value of each category The first weighting coefficient, This is the second weighting coefficient.

[0157] exist Judgment values ​​for each category In the selection process, the category corresponding to the maximum value is chosen as the category of the text.

[0158] In this embodiment, the first weighting coefficient and the second weighting coefficient traverse different values ​​within a preset range. and Combining these metrics, we calculate the accuracy, F1 score, and other indicators of the classification objective function on the validation set, thereby selecting the optimal one. and combination.

[0159] The calculation of semantic potential value combines the sum of reinforcement values ​​and variance, preserving the overall strength of the keyword's pull on the category while reflecting the distribution stability of reinforcement values ​​among keywords through the variance term. A larger variance indicates a more unbalanced pull on the category by the keyword, and the potential value will be dynamically weakened; a smaller variance indicates a more stable pull, and the potential value will be preserved. This allows classification decisions to not only depend on the degree of belonging but also consider the internal stability of keyword features, avoiding misjudgments caused by a few strong keywords dominating the classification.

[0160] Example 2: A text classification system, comprising: a traction strength acquisition unit, a tension diffusion coefficient acquisition unit, a first correction unit, a second correction unit, a fusion affiliation calculation unit, and a classification unit;

[0161] The traction strength acquisition unit is used to acquire the directional traction strength of keywords in the text for each category;

[0162] The tension diffusion coefficient acquisition unit is used to obtain the tension diffusion coefficient of keywords based on the semantic similarity between keywords in the text and the traction direction;

[0163] The first correction unit is used to multiply the tension diffusion coefficient of the keyword by the corresponding directional traction strength to obtain the first corrected directional traction strength of the keyword for the category;

[0164] The second correction unit is used to correct the corresponding directional traction strength according to the interaction correction factor, so as to obtain the second corrected directional traction strength of the keyword to the category.

[0165] The fusion attribution calculation unit is used to fuse the first modified orientation traction strength and the second modified orientation traction strength, and then introduce nonlinear competitive modulation through the semantic distance between categories to obtain the final attribution of the text to each category;

[0166] The classification unit is used to obtain the first semantic potential value and the second semantic potential value respectively, and combined with the final attribution degree, to establish a classification objective function and obtain the category of the text.

[0167] The specific implementation process of Example 2 is the same as that of Example 1.

[0168] Example 3: A text classification device includes a processor and a computer program, the computer program being executed by the processor to implement the text classification method as described in Example 1.

[0169] Example 4: A computer-readable storage medium storing a computer program that, when executed by a processor, implements the text classification method as described in Example 1.

[0170] This invention breaks through the assumption of independent keyword contributions. It characterizes the semantic association and directional propagation between keywords through the tension diffusion coefficient, and captures the complex interaction effects of multiple keywords using an interaction correction factor, achieving a dual fine-grained correction of the directional traction strength. Secondly, it introduces a nonlinear competitive modulation mechanism based on category semantic distance to dynamically characterize the semantic association and competitive constraints between categories, avoiding the problem of blurred classification boundaries caused by linear superposition. Finally, by fusing the final attribution degree with the semantic potential value reflecting the stability of keyword traction to construct a classification objective function, it further improves the stability and accuracy of the classification results, thereby effectively solving the problem of low classification accuracy in semantically similar categories or multi-topic text scenarios.

[0171] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A text classification method, characterized in that, Includes the following steps: S1. Obtain the directional pull strength of keywords in the text for each category: , in, For the first The keyword for the first The directional traction strength of each category, For the first The first category one reference word TF-ICF weights, For from the first Filter the maximum traction value from all traction values ​​corresponding to each category. This is the traction value; S2. Based on the semantic similarity between keywords in the text and the direction of attraction, obtain the tension diffusion coefficient of the keywords; S2 includes the following steps: S21. Calculate the semantic similarity between two keywords in the text; S22. Based on the directional traction strength of two keywords in the same category, obtain the traction direction; S23. Calculate the tension diffusion coefficient of keywords based on semantic similarity and traction direction; The formula for obtaining the traction direction in S22 is: , in, For the first The first keyword and the first The keyword in the first The traction direction of each category, For the first The keyword for the first The directional traction strength of each category, For sign functions, sign functions When the value inside the parentheses is greater than or equal to 0, for The value is assigned as 1, and if the value inside the parentheses is less than 0, then... The value is assigned to -1; The formula for calculating the tension diffusion coefficient of the keyword in S23 is: , in, For the first The keyword in the first The tension diffusion coefficient of each category, For the first The first keyword and the first The semantic similarity of keywords. and For keyword numbering, To select the maximum value, The number of keywords in the text. For the first One keyword, For the first One keyword; S3. Multiply the tension diffusion coefficient of the keyword by the corresponding directional traction strength to obtain the first corrected directional traction strength of the keyword for the category; S4. Based on the interaction correction factor, the corresponding directional traction strength is corrected to obtain the second corrected directional traction strength of the keyword to the category; S4 includes the following steps: S41. Calculate the interaction correction factor based on the co-occurrence weights and semantic similarity of the two keywords, as well as the TF-IDF weights. S42. Multiply the interaction correction factor by the interaction influence coefficient and add 1 to obtain the global interaction gain coefficient of the keyword. S43. Multiply the global interaction gain coefficient of the keyword by the corresponding directional traction strength to obtain the second modified directional traction strength of the keyword to the category; The formula for calculating the interaction correction factor in S41 is: , in, For the first The keyword for the first Interaction correction factors for each category, For the first The first keyword and the first The keyword in the first Co-occurrence weights of each category, For the first The TF-IDF weights of each keyword. The TF-IDF weight of the first keyword. For the first The TF-IDF weights of each keyword, with max being the maximum value; S5. The first modified directional traction strength and the second modified directional traction strength are fused together, and then nonlinear competitive modulation is introduced through the semantic distance between categories to obtain the final belonging degree of the text to each category. S5 includes the following steps: S51. Multiply the normalized TF-IDF weight of the keyword by the first modified directional traction strength to obtain the first enhancement value; S52. Multiply the normalized TF-IDF weight of the keyword by the second modified directional traction strength to obtain the second reinforcement value; S53. Add the second enhancement value and the first enhancement value and take the average value to obtain the comprehensive enhancement value; S54. Sum the comprehensive reinforcement values ​​of all keywords along the category dimension to obtain the initial affiliation degree of the text to each category; S55. Calculate the semantic distance between any two categories based on the reference word set for each category; S56. Based on semantic distance, perform nonlinear competitive modulation on the initial classification degree of each category of the text to obtain the final classification degree of the text to each category; The formula for calculating the final affiliation in S56 is: , in, For the text pair The final affiliation of each category, It is an exponential function. For the first The category and the first The semantic distance between each category For the text pair The initial affiliation of each category, For the text pair The initial affiliation of each category, To prevent parameters with a denominator of 0, This is the competition intensity coefficient. For the number of categories, and For the category number; S6. Obtain the first semantic potential value and the second semantic potential value respectively, and combine them with the final attribution degree to establish a classification objective function to obtain the category of the text; S6 includes the following steps: S61. Calculate the first semantic potential value based on the first reinforcement value corresponding to each keyword in the text; S62. Calculate the second semantic potential value based on the second reinforcement value corresponding to each keyword in the text; S63. Weight the final attribution score, the first semantic potential value, and the second semantic potential value to obtain the classification objective function; S64. Based on the classification objective function, by calculating the decision values ​​of all categories, the category with the largest decision value is determined as the category to which the text belongs; The formula for calculating the first semantic potential energy value in S61 is: , in, For the text pair The first semantic potential value of each category The number of keywords in the text. For the first The keyword for the first The first enhancement value for each category, for indivual The variance; The formula for calculating the second semantic potential energy value in S62 is: , in, For the text pair The second semantic potential value of each category For the first The keyword for the first The second enhancement value for each category, for indivual The variance.

2. The text classification method according to claim 1, characterized in that, S1 includes the following steps: S11. Calculate the semantic similarity between keywords in the text and reference words in the reference word set for each category; S12. Multiply the semantic similarity by the TF-ICF weight of the corresponding reference word to obtain the traction value; S13. Select the maximum traction value from the traction values ​​corresponding to each reference word in each category to obtain the directional traction strength of the keyword for that category.

3. A text classification system, characterized in that, The text classification method based on any one of claims 1 to 2 includes: a traction strength acquisition unit, a tension diffusion coefficient acquisition unit, a first correction unit, a second correction unit, a fusion affiliation calculation unit, and a classification unit; The traction strength acquisition unit is used to acquire the directional traction strength of keywords in the text for each category; The tension diffusion coefficient acquisition unit is used to acquire the tension diffusion coefficient of keywords based on the semantic similarity between keywords in the text and the traction direction. The first correction unit is used to multiply the tension diffusion coefficient of the keyword by the corresponding directional traction strength to obtain the first corrected directional traction strength of the keyword for the category; The second correction unit is used to correct the corresponding directional traction strength according to the interaction correction factor to obtain the second corrected directional traction strength of the keyword to the category; The fusion affiliation calculation unit is used to fuse the first modified directional traction strength and the second modified directional traction strength, and then introduce nonlinear competitive modulation through the semantic distance between categories to obtain the final affiliation degree of the text to each category; The classification unit is used to obtain the first semantic potential value and the second semantic potential value respectively, and combine them with the final attribution degree to establish a classification objective function to obtain the category of the text.

4. A text classification device, comprising: A processor and a computer program, characterized in that the computer program is executed by the processor to implement the text classification method as described in any one of claims 1 to 2.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the text classification method as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Metro industry-based knowledge file reading method

    CN121144356A

  • Equipment operation and maintenance intelligent management method and system based on artificial intelligence

    CN121412397A