Text classification method based on machine learning and hyperparticle method

By constructing and segmenting super-grain squares, the overlap and instability problems in traditional particle sphere calculations are solved, efficient and accurate text classification is achieved, and high-dimensional large-scale text data processing is adapted to high-dimensional large-scale text data processing, improving the performance and efficiency of text analysis technology.

CN120492630AActive Publication Date: 2025-08-15CHONGQING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510609564.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-15
Estimated Expiration
2045-05-13

Smart Images

  • Figure CN120492630A_ABST
    Figure CN120492630A_ABST
Patent Text Reader

Abstract

The invention provides a text classification method based on machine learning and a hyperparticle method, which comprises the following steps: acquiring a preprocessed text data set, and extracting a feature vector of text data by utilizing a feature extraction model; according to the feature vectors of all the text data in the text data set, constructing an initial hyper-particle square; the purity of the initial super-particle square is calculated, and if the purity of the initial super-particle square is lower than a set threshold value, the initial super-particle square is divided into a plurality of super-particle squares which are not overlapped with one another; repeating the operations of purity calculation and segmentation on the newly generated super-particle squares until the purities of all the super-particle squares meet the conditions; according to the label of the text data in each hyper-particle party, determining the label of each hyper-particle party by utilizing a majority principle; and outputting the label of the hyper-particle party to which the to-be-classified text data belongs as a classification result of the to-be-classified text data. The method can adapt to the unique property of text data, improve the performance and efficiency of text classification, and promote the development of a text analysis technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a text classification method based on machine learning and super-granularity. Background Art

[0002] In the era of information explosion, text data is growing exponentially. Traditional data processing and analysis technologies face challenges such as inefficiency and lack of robustness when faced with massive amounts of text data. Text data is characterized by high dimensionality, sparseness, and complex semantics, making it extremely challenging to extract effective information and achieve accurate classification from large amounts of text. Granular computing (GrC), a method that simulates human cognitive processes, offers a new approach to solving complex tasks. It improves problem-solving efficiency and flexibility by processing data at multiple granularity levels. The introduction of information granularity further facilitates data simplification and efficient analysis. Granular sphere computing (GBC), proposed by Xia et al., holds a prominent position in the field of granular computing due to its efficiency, robustness, and adaptability. It replaces data points with spheres, reducing data size and improving algorithm efficiency. It can also adaptively generate spheres based on data characteristics, achieving significant results in data dimensionality reduction and multi-domain applications.

[0003] However, in text applications, GBC's limitations are becoming increasingly prominent. During text data processing, the geometric structure of the spheres can have serious consequences: overlapping issues can cause the same text data to be misclassified, undermining the accuracy of text classification; incomplete coverage can cause key text information to be overlooked, affecting the integrity of text analysis. Furthermore, the instability of sphere generation can lead to significant differences in the classification results of text data processed in different batches, making it impossible to guarantee the reliability and repeatability of text classification models. These shortcomings make GBC difficult to meet the text field's demand for stable processing and accurate classification of high-dimensional, large-scale text data. An innovative granular computing method is urgently needed to adapt to the unique properties of text data, improve the performance and efficiency of text classification, and promote the development of text analysis technology. Summary of the Invention

[0004] In order to solve the problems existing in the background technology, adapt to the unique properties of text data, improve the performance and efficiency of text classification, and promote the development of text analysis technology, one aspect of the present invention provides a text classification method based on machine learning and super-granularity, including:

[0005] S1: Obtain the preprocessed text dataset and use the feature extraction model to extract the feature vector of the text data;

[0006] S2: Construct an initial hypergranular cube based on the feature vectors of all text data in the text dataset;

[0007] S3: Calculate the purity of the initial super-granular cube. If the purity of the initial super-granular cube is lower than the set threshold, split the initial super-granular cube into multiple non-overlapping super-granular cubes. Repeat the above purity calculation and segmentation operations for the newly generated super-granular cubes until the purity of all super-granular cubes meets the requirements.

[0008] S4: Determine the label of each super-granular cube using the majority principle based on the label of the text data in each super-granular cube;

[0009] S5: Outputting the label of the super-granularity to which the text data to be classified belongs as the classification result of the text data to be classified.

[0010] Another aspect of the present invention provides a text classification method based on machine learning and super-granularity, wherein the system includes a memory and a processor; the memory is used to store an application; the processor is used to run the application and execute the text classification method based on machine learning and super-granularity.

[0011] Another aspect of the present invention provides a computer storage medium having a program stored thereon, which, when executed by a processor, implements the text classification method based on machine learning and super-granularity.

[0012] The present invention has at least the following beneficial effects

[0013] The present invention replaces traditional spheres with hyper-granular cubes, utilizing the highly symmetrical geometric properties of n-dimensional hypercubes to avoid data overlap and effectively prevent the same text data from being misclassified. This fundamentally ensures the accuracy of text classification and addresses the drawback of GBC methods that can cause text classification errors due to overlap. Hyper-granular cubes can more evenly represent data, achieving complete coverage of the text data space, ensuring that key text information is not missed. This addresses the shortcomings of GBC methods, where incomplete coverage affects the integrity of text analysis, and enables the text classification process to fully capture all types of text information. The hyper-granular cube generation strategy adopted in the present invention avoids the problem of randomly selecting center points, effectively eliminates the instability of sphere generation, ensures consistency in the classification results of text data processed from different batches, significantly improves the reliability and repeatability of the text classification model, and overcomes the problem of GBC methods' unstable results in practical applications. The partitioning and generation algorithm based on hyper-granular cubes reduces time overhead while maintaining comprehensive data coverage. Compared with traditional GBC methods, it can more efficiently process high-dimensional, large-scale text data, meeting the demand for efficient processing in the text field and effectively promoting the development of text analysis technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION

[0015] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0016] See also Figure 1 One aspect of the present invention provides a text classification method based on machine learning and super-granularity, comprising:

[0017] S1: Obtain the preprocessed text dataset and use the feature extraction model to extract the feature vector of the text data;

[0018] Preferably, the feature extraction model includes: a bag-of-words model, a TF-IDF model, a Word2Vec model or a BERT model.

[0019] In this embodiment, the bag-of-words model, TF-IDF model, Word2Vec model, and BERT model are all classic models for text feature extraction, and each has its own advantages, providing a variety of options. The bag-of-words model converts text into vectors by counting the frequency of occurrence of words in the text. It is simple and intuitive and can quickly quantify the text. The TF-IDF model, based on the bag-of-words model, further considers the importance of words in the document, highlights the key features of the text, and helps to distinguish the topic differences between different texts. The Word2Vec model is based on a neural network and can learn the semantic vectors of words, capture the semantic similarity between words, and make text features more semantically representative. The BERT model, as a pre-trained deep bidirectional Transformer model, can dynamically understand the semantics of each word based on the context and extract more accurate and richer text semantic features. Through the feature extraction model, text data can be converted into structured feature vectors, effectively reducing the high-dimensionality and sparsity of text data, and providing a high-quality data foundation for subsequent super-granular construction and classification tasks. It not only improves the efficiency of text data processing, but also enhances the model's understanding and expression of text semantic information, thereby improving the accuracy and reliability of text classification.

[0020] S2: Construct an initial hypergranular cube based on the feature vectors of all text data in the text dataset;

[0021] Preferably, the constructing of the initial super-granular cube includes: using the feature vector of the text data processed by the feature extraction model as the coordinates of the text data in the n-dimensional space, and calculating the center and side length of the initial super-granular cube according to the coordinates of all the text data in the n-dimensional space, and constructing the initial super-granular cube in the n-dimensional space, wherein the center and side length of the initial super-granular cube include:

[0022]

[0023] d0=2·max{d(x ij ,C 0j )}

[0024] Where D represents the feature vector set corresponding to the text dataset, m represents the number of text data in the text dataset D; x ij ∈x i C represents the eigenvalue of the feature vector of the i-th text data in the j-th dimension; 0j represents the eigenvalue of the central eigenvector C0 of the initial hypergranular cube GH0 in the jth dimension; d0 represents the side length of the initial hypergranular cube GH0; x i Represents the feature vector of the i-th text data; d(x ij ,C 0j ) represents x ij and C 0j The Euclidean distance function between .

[0025] In step S2 of this embodiment, constructing the initial hyper-cube lays the foundation for subsequent text classification. First, the feature vectors of the text data processed by the feature extraction model are used as the coordinates of the text in n-dimensional space, where n depends on the dimensionality of the feature vectors. The center of the initial hyper-cube is calculated by averaging the eigenvalues of all text data in each dimension. The eigenvalues of the j-th dimension of all text data in the text dataset are summed and divided by the number of text data m to obtain the eigenvalue of the center of the initial hyper-cube in the j-th dimension. To calculate the side length, the maximum Euclidean distance between the eigenvectors of all text data and the center is first found, and then multiplied by 2 to obtain the side length. Based on the calculated center and side length, the position and size of the initial hyper-cube can be determined in n-dimensional space, completing the construction of the initial hyper-cube. Constructing the initial hyper-cube in this way allows text data to be represented in n-dimensional space as a geometric structure. This provides a basic framework for subsequent text data classification and processing, giving the previously abstract text data a more intuitive spatial distribution. Reasonably determined center and side lengths can reflect the distribution characteristics of text data to a certain extent, which helps in subsequent calculation of the purity of the hypergranular cube, hypergranular cube segmentation and other operations, thereby improving the efficiency and accuracy of text classification and making the text classification process more systematic and logical.

[0026] S3: Calculate the purity of the initial super-granular cube. If the purity of the initial super-granular cube is lower than the set threshold, split the initial super-granular cube into multiple non-overlapping super-granular cubes. Repeat the above purity calculation and segmentation operations for the newly generated super-granular cubes until the purity of all super-granular cubes meets the requirements.

[0027] Preferably, the purity of the supergranules includes:

[0028]

[0029] Where P represents the purity of the super-granular GH, |GH| represents the number of samples in the super-granular GH; n k It represents the number of samples belonging to the kth class in the hypergranular cube.

[0030] In this embodiment, calculating the purity of the super-granular cube is the key basis for determining whether the super-granular cube needs to be further divided. In the purity formula, |GH| is the total number of samples in the super-granular cube, and n k is the number of samples belonging to the kth class within the super-granular cube. The purity of the super-granular cube is calculated by calculating the ratio of the number of samples of the largest class within the super-granular cube to the total number of samples within the super-granular cube. If the purity of the initial super-granular cube falls below a pre-set threshold, it indicates that the sample categories within the super-granular cube are mixed, making accurate classification difficult. In this case, it is necessary to segment the super-granular cube into multiple non-overlapping super-granular cubes. After segmentation, the purity of the newly generated super-granular cube is calculated again. If it still does not meet the requirements, the segmentation process continues until the purity of all super-granular cubes meets the set requirements. Calculating the purity of the super-granular cube and performing segmentation based on the results can make the sample categories within the super-granular cube more uniform and concentrated. This allows for more accurate classification of text data based on the labels of the super-granular cube during subsequent text classification. This avoids classification errors caused by mixed sample categories within the super-granular cube, improving text classification accuracy. At the same time, by continuously optimizing the hypergranular structure, the entire text classification system becomes more reasonable and efficient, and the model's ability to process complex text data is enhanced, which helps to cope with the challenges brought by the high-dimensional, sparse, and semantically complex characteristics of text data.

[0031] Preferably, when the purity calculation and segmentation operations are repeated on the newly generated super-granular cube, if the newly generated super-granular cube does not contain any data points, it is called a meaningless super-granular cube and is deleted.

[0032] In this embodiment, as super-granules are continuously segmented to meet purity requirements, newly generated super-granules may contain no data points. These super-granules, lacking actual data, are ineffective in text classification tasks and are therefore considered meaningless. Deleting these super-granules simplifies the super-granule set and prevents invalid structures from occupying computing resources and storage space.

[0033] Removing meaningless hypergranules optimizes the overall structure of the hypergranule, making subsequent computation and classification more efficient. This reduces unnecessary computation, avoids wasting time and resources on meaningless structures, and improves the efficiency of the text classification algorithm. It also helps maintain the simplicity and effectiveness of the hypergranule system, allowing hypergranule-based text classification to focus more on areas supported by real data, further improving classification accuracy and reliability.

[0034] Preferably, dividing the super-granular cube into a plurality of non-overlapping super-granular cubes comprises:

[0035] S31: Take the center C of the super-granular square GH as the reference center and the side length Construction of reference supergranular GH re ;

[0036] S32: Reference super granular GH re Each vertex is the center point, with the side length as Construct a new hyper-granular cube and obtain a hyper-granular cube with each vertex as the center point and without overlapping.

[0037] In this embodiment, the center C of the supercube is used as the reference center, and a reference supercube is constructed using half the side length of the original supercube, d / 2. This step establishes an intermediate transition structure for the subsequent construction of a new supercube. By reducing the side length and using the original center as the reference, a relatively small supercube is obtained as a reference.

[0038] Using the vertices of the reference supercube as new centers, a new supercube is constructed with a side length of d / 2. Because the construction is centered around the vertices and the side lengths are fixed, the newly generated supercubes are guaranteed to not overlap, thus splitting the original supercube into multiple sub-supercubes that meet the requirements.

[0039] This segmentation method has clear geometric logic and regularity. From a spatial layout perspective, it ensures that newly generated hypergranules are rationally distributed within the original hypergranule space and do not interfere with each other, effectively avoiding data overlap and laying the foundation for more accurate subsequent text data classification. Through this orderly segmentation, each new hypergranule can more purely contain a certain category or categories of similar text data, helping to improve the purity of the hypergranule and, in turn, enhance the accuracy and efficiency of text classification. Furthermore, this regular segmentation method facilitates algorithm implementation and the efficient use of computing resources, reducing algorithm complexity and operating costs.

[0040] S4: Determine the label of each super-granular cube using the majority principle based on the label of the text data in each super-granular cube;

[0041] In this embodiment, each super-granule contains a number of text data items, each of which is labeled (e.g., belonging to different text category labels). The "majority rule" involves counting the number of text data items with different labels within a super-granule and determining the label of the super-granule with the label that appears the most. For example, if a super-granule contains 10 text data items, 9 of which are labeled "technology" and 1 is labeled "entertainment," then according to the majority rule, the label of this super-granule will be determined to be "technology."

[0042] Determining supergranule labels using the majority principle ensures that the labels are representative, reflecting the dominant text category within the supergranule. This provides a clear basis for subsequent text classification. When text data to be classified is determined to belong to a specific supergranule, it can be directly assigned the label of that supergranule, simplifying the classification process. Furthermore, this approach can, to a certain extent, offset the interference of small amounts of outlier text data within the supergranule, enhancing the stability and accuracy of classification and ensuring that the classification results are more consistent with the overall categorical tendencies of the text data within the supergranule.

[0043] S5: Outputting the label of the super-granularity to which the text data to be classified belongs as the classification result of the text data to be classified.

[0044] In this embodiment, after the previous steps, the super-granules have been constructed and the labels of each super-granule have been determined. At this point, the text data to be classified is first processed using a feature extraction model to extract its feature vectors. These feature vectors serve as the coordinates of the text data to be classified in n-dimensional space, and the super-granule to which the text data belongs is then determined. If a sample point is located on the boundary between two super-granules, the distance from the sample point to the centers of the super-granules on both sides of the boundary is calculated, and the sample point is assigned to the nearest super-granule. The label of this super-granule is then used as the classification result.

[0045] Another aspect of the present invention provides a text classification method based on machine learning and super-granularity, wherein the system includes a memory and a processor; the memory is used to store an application; the processor is used to run the application and execute the text classification method based on machine learning and super-granularity.

[0046] Another aspect of the present invention provides a computer storage medium having a program stored thereon, which, when executed by a processor, implements the text classification method based on machine learning and super-granularity.

[0047] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus Direct 10RAM (RDRAM), Direct Memory Bus Dynamic RAM (DRDRAM), and Rambus Dynamic RAM (RDRAM), etc.

[0048] In summary, the present invention replaces traditional spheres with hyper-granular cubes, utilizing the highly symmetrical geometric properties of n-dimensional hypercubes to avoid data overlap and effectively prevent the same text data from being misclassified. This fundamentally ensures the accuracy of text classification and addresses the drawback of GBC methods that can lead to text classification errors due to overlap. Hyper-granular cubes can more evenly represent data, achieving complete coverage of the text data space, ensuring that key text information is not missed. This addresses the shortcomings of GBC methods, where incomplete coverage affects the integrity of text analysis, and enables the text classification process to fully capture all types of text information. The hyper-granular cube generation strategy adopted in the present invention avoids the problem of randomly selecting center points, effectively eliminates the instability of sphere generation, ensures consistency in the classification results of text data processed from different batches, significantly improves the reliability and repeatability of the text classification model, and overcomes the problem of GBC methods' unstable results in practical applications. The partitioning and generation algorithm based on hyper-granular cubes reduces time overhead while maintaining comprehensive data coverage. Compared with traditional GBC methods, it can more efficiently process high-dimensional, large-scale text data, meeting the demand for efficient processing in the text field and effectively promoting the development of text analysis technology.

[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A text classification method based on machine learning and super-granularity, characterized in that: include: S1: Obtain the preprocessed text dataset and use the feature extraction model to extract the feature vector of the text data; S2: Construct an initial hypergranular cube based on the feature vectors of all text data in the text dataset; S3: Calculate the purity of the initial super-granular cube. If the purity of the initial super-granular cube is lower than the set threshold, split the initial super-granular cube into multiple non-overlapping super-granular cubes. Repeat the above purity calculation and segmentation operations for the newly generated super-granular cubes until the purity of all super-granular cubes meets the requirements. S4: Determine the label of each super-granular cube using the majority principle based on the label of the text data in each super-granular cube; S5: Outputting the label of the super-granularity to which the text data to be classified belongs as the classification result of the text data to be classified.

2. A text classification method based on machine learning and super-granularity according to claim 1, characterized in that: The feature extraction model includes: a bag-of-words model, a TF-IDF model, a Word2Vec model or a BERT model.

3. The text classification method based on machine learning and super-granularity according to claim 1, characterized in that: The constructing of the initial super-granular cube includes: using the feature vector of the text data processed by the feature extraction model as the coordinate of the text data in the n-dimensional space, and calculating the center and side length of the initial super-granular cube according to the coordinates of all the text data in the n-dimensional space, and constructing the initial super-granular cube in the n-dimensional space, wherein the center and side length of the initial super-granular cube include: d0=2·max{d(x ij ,C 0j )} Where D represents the feature vector set corresponding to the text dataset, m represents the number of text data in the text dataset D; x ij ∈x i C represents the eigenvalue of the feature vector of the i-th text data in the j-th dimension; 0j represents the eigenvalue of the central eigenvector C0 of the initial hypergranular cube GH0 in the jth dimension; d0 represents the side length of the initial hypergranular cube GH0; x i Represents the feature vector of the i-th text data; d(x ij ,C 0j ) represents x ij and C 0j The Euclidean distance function between .

4. The text classification method based on machine learning and super-granularity according to claim 1, characterized in that: The purity of the supergranular square includes: Where P represents the purity of the super-granular GH, |GH| represents the number of samples in the super-granular GH; n k It represents the number of samples belonging to the kth class in the hypergranular cube.

5. The text classification method based on machine learning and super-granularity according to claim 1, characterized in that: Splitting a supergranule into multiple non-overlapping supergranules includes: S31: Take the center C of the super-granular square GH as the reference center and the side length Construction of reference supergranular GH re ; S32: Reference super granular GH re Each vertex is the center point, with the side length as Construct a new hyper-granular cube and obtain a hyper-granular cube with each vertex as the center point and without overlapping.

6. The text classification method based on machine learning and super-granularity according to claim 1, characterized in that: When the purity and segmentation operations are repeated on the newly generated super-granular cube, if the newly generated super-granular cube does not contain any data points, it is called a meaningless super-granular cube and is deleted.

7. The text classification method based on machine learning and super-granularity according to claim 1, characterized in that: Determining the label of each super-granular entity by using the majority principle according to the label of the text data in each super-granular entity includes: Among them, n k represents the number of samples belonging to the kth class in the hypergranular cube; L represents the category label of the hypergranular cube.

8. A text classification method based on machine learning and super-granularity, characterized in that: The system includes a memory and a processor; the memory is used to store an application; the processor is used to run the application and execute the text classification method based on machine learning and super-granularity as described in any one of claims 1 to 7.

9. A computer storage medium, characterized in that The computer storage medium stores a program, and when the program is executed by the processor, the text classification method based on machine learning and super-granularity according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Granular ball and metric learning-based adversarial attack text classification method

    CN119621980A