Unsupervised automatic news classification method
Through the unsupervised automatic news classification method, unsupervised classification processing and identification of the news data sets is solved, and the complexity of text and picture content in news classification is achieved, and accurate and efficient automatic news classification is achieved.
Patent Information
- Application Number
- CN202510627729.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the field of news classification, because news content includes both text and pictures, traditional natural language processing is difficult to achieve accurate classification.
An unsupervised automatic classification method for news is proposed. By unsupervised classification processing of the news data set, a classified unmarked news cluster is generated, and then these clusters are identified, classified identification is marked, and a basic learning model is trained using classification mark news. Finally, the news to be classified is input to the model to obtain its type.
This method can effectively solve the complexity of text and picture content in news classification, realize accurate and automatic classification of news, and improve classification efficiency and accuracy.
Smart Images

Figure CN120144765A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an unsupervised news automatic classification method. Background Art
[0002] Text classification (Text Classification or Text Categorization, TC), also known as automatic text classification (Automatic Text Categorization), refers to the process of a computer mapping a text carrying information to a pre-given category or several category topics. Text classification is mainly applied to fields such as sentiment analysis, topic tagging, news classification, question answering systems, dialogue behavior classification, natural language inference, relationship classification, event prediction, etc. Currently, in the field of news classification, due to the existence of not only text content but also picture content, it is difficult to accurately classify news using traditional natural language processing. Summary of the Invention
[0003] In view of the above technical problems, an embodiment of the present application proposes an unsupervised news automatic classification method, which can solve the problem that it is difficult to accurately classify news using traditional natural language processing in the current news classification field due to the existence of not only text content but also picture content.
[0004] In a first aspect, an embodiment of the present application provides an unsupervised news automatic classification method, including: Performing unsupervised classification processing on a news data set to obtain classified unlabeled news clusters; Performing identification recognition on the classified unlabeled news clusters to obtain corresponding news classification identifications; Marking the news classification identifications to the classified unlabeled news in the classified unlabeled news clusters to obtain classified labeled news; Training a basic learning model using the classified labeled news to obtain a news classification model; Inputting the news to be classified into the news classification model to obtain the news type of the news to be classified.
[0005] In some embodiments, the performing unsupervised classification processing on the news data set to obtain classified unlabeled news clusters includes: Using a large language model to perform summary processing on the news data set to obtain a news summary data set; Performing unsupervised classification processing on the news summary data set to obtain the classified unlabeled news clusters.
[0006] In some embodiments, the news summary data set includes multiple news summaries; Performing unsupervised classification on the news summary dataset to obtain the classified unlabeled news clusters includes: Performing semantic segmentation and extraction on the news summary to obtain at least one feature semantics; Performing unsupervised classification on the feature semantics to obtain the classified unlabeled news clusters.
[0007] In some embodiments, performing unsupervised classification on the feature semantics to obtain the classified unlabeled news clusters includes: Setting the sample distance and sample cosine similarity for classification; Clustering multiple feature semantics based on the sample distance and the sample cosine similarity to obtain the classified unlabeled news clusters.
[0008] In some embodiments, identifying the classified unlabeled news clusters and generating corresponding news classification identifiers includes: Obtaining news classification identifiers and a classification identifier feature set corresponding to the news classification identifiers; Obtaining a representative vector of the classified unlabeled news clusters based on the classified unlabeled news clusters; Obtaining the news classification identifier of the classified unlabeled news clusters based on the representative vector and the classification identifier feature set.
[0009] In some embodiments, obtaining the news classification identifier of the classified unlabeled news clusters based on the representative vector and the classification identifier feature set is specifically: Calculating the similarity between the representative vector and the classification identifier features in the classification identifier feature set to obtain a similarity score; If the similarity score is greater than the similarity score threshold, recording the similarity score and the candidate news classification identifier corresponding to the similarity score; Determining the news classification identifier from the candidate news classification identifiers based on the similarity score.
[0010] In some embodiments, determining the news classification identifier from the candidate news classification identifiers based on the similarity score further includes: Obtaining a semantic feature weight corresponding to the feature semantics based on the structure and content of the news summary; Obtaining a news classification identifier score of the candidate news data based on the semantic feature weight and the similarity score; Obtaining the news classification identifier based on the news classification identifier score, where the news classification identifier includes a news classification main identifier and a news classification secondary identifier arranged in sequence.
[0011] In some embodiments, the step of marking the news classification identifier to the unclassified news in the unclassified news cluster to obtain classified news is specifically as follows: Obtain the news classification identifier scores of the main news classification identifier and the secondary news classification identifier; Mark the main news classification identifier, the secondary news classification identifier, and the corresponding news classification identifier scores to the unclassified news in the unclassified news cluster to obtain the classified news.
[0012] Second, an unsupervised news automatic classification system provided by an embodiment of the present application includes: A processing module, configured to perform unsupervised classification processing on a news data set to obtain an unclassified news cluster; An identification module, configured to perform identification on the unclassified news cluster to obtain corresponding news classification identifiers; mark the news classification identifiers to the unclassified news in the unclassified news cluster to obtain classified news; A training module, configured to train a basic learning model using the classified news to obtain a news classification model; A classification module, configured to input the news to be classified into the news classification model to obtain the news type of the news to be classified.
[0013] Third, an electronic device provided by an embodiment of the present application includes the unsupervised news automatic classification system described in the second aspect.
[0014] The present application provides an unsupervised news automatic classification method, including performing unsupervised classification processing on a news data set to obtain an unclassified news cluster; performing identification on the unclassified news cluster to obtain corresponding news classification identifiers; marking the news classification identifiers to the unclassified news in the unclassified news cluster to obtain classified news; training a basic learning model using the classified news to obtain a news classification model; inputting the news to be classified into the news classification model to obtain the news type of the news to be classified, which can classify news using a machine learning model and can solve the problem that in the current news classification field, due to the existence of not only text content but also picture content, it is difficult to accurately classify news using traditional natural language processing. Description of the Drawings
[0015] Hereinafter, the present invention will be described in more detail based on embodiments and with reference to the drawings.
[0016] Figure 1 is a flowchart of an unsupervised news automatic classification method provided by an embodiment of the present invention; Figure 2Schematic diagram of an unsupervised news automatic classification system provided by an embodiment of the present invention. Detailed implementation manners
[0017] The present invention will be further described below with reference to the accompanying drawings.
[0018] Text classification was initially carried out through expert rules (Patterns), and an expert system was established using knowledge engineering. The advantage of this approach is that it can solve problems more intuitively, but it is time-consuming and laborious, and both the coverage and accuracy are limited.
[0019] Today's text classification refers to automatically classifying and tagging texts (or other entities) by a computer according to a certain classification system or standard. With the explosive growth of information, manually annotating data has become time-consuming, of low quality, and affected by the subjective awareness of the annotator. Therefore, it is of practical significance to use machine automation to implement text annotation. Handing over the repetitive and boring text annotation tasks to a computer can effectively overcome the above problems, and at the same time, the annotated data has characteristics such as consistency and high quality. However, currently in the field of news classification, since it not only has text content but also picture content, it is difficult to accurately classify news using traditional natural language processing.
[0020] In a first aspect, as Figure 1 shown, in view of the above technical problems, the embodiments of the present application provide an unsupervised news automatic classification method, including: S101: Perform unsupervised classification processing on a news data set to obtain classified unlabeled news clusters; In some embodiments, the performing unsupervised classification processing on the news data set to obtain classified unlabeled news clusters includes: Using a large language model to perform summary processing on the news data set to obtain a news summary data set; Performing unsupervised classification processing on the news summary data set to obtain the classified unlabeled news clusters.
[0021] It should be noted that a large language model (Large Language Model, abbreviated as LLM) can generate natural language texts or understand the meaning of language texts. A large language model can not only perform simple language tasks such as spelling check and grammar correction, but also handle complex tasks such as text summarization, machine translation, sentiment analysis, dialogue generation, and content recommendation. By performing summary processing on the news data set, the computational amount for calculating and classifying the news data can be reduced, and the calculation speed can be improved. Among them, the news data set includes a large amount of news texts (or news data).
[0022] In some embodiments, the news summary data set includes multiple news summaries; Performing unsupervised classification processing on the news abstract dataset to obtain the classified unlabeled news clusters, including: Performing semantic segmentation and extraction on the news abstract to obtain at least one feature semantic; Performing unsupervised classification processing on the feature semantic to obtain the classified unlabeled news clusters.
[0023] It should be noted that the semantic segmentation and extraction can be implemented through a large language model. By performing semantic segmentation and processing on the news abstract and obtaining the feature semantic, the news data can be classified based on the feature semantic, further reducing the computational complexity during classification. Among them, the at least one feature semantic is usually multiple feature semantics.
[0024] In some embodiments, performing unsupervised classification processing on the feature semantic to obtain the classified unlabeled news clusters includes: Setting the sample distance and sample cosine similarity for classification; Clustering multiple feature semantics based on the sample distance and sample cosine similarity to obtain the classified unlabeled news clusters.
[0025] It should be noted that after obtaining the feature semantic, the feature semantic is usually vectorized to obtain the vectorized feature semantic. By simultaneously defining the sample distance and sample cosine similarity, the similarity of the clustering samples (or feature semantics) can be made higher and the clustering more accurate. Among them, the classified unlabeled news clusters include multiple feature semantics (or vectorized feature semantics). The specific values of the sample distance and sample cosine similarity can be set according to actual needs, and the present application does not make specific limitations in this regard.
[0026] S102: Identifying the classification of the classified unlabeled news clusters to obtain the corresponding news classification identifier; It should be noted that although there are some news datasets available for use, these datasets are usually manually labeled, which has a relatively large subjectivity, that is, the classification criteria are unstable and the classification results are inaccurate. By setting the sample distance and sample cosine similarity in the embodiments of the present application and obtaining the corresponding news classification identifier, the news classification identifier corresponding to the classified unlabeled news clusters can be obtained quickly and accurately, and generating the news classification identifier for the classified unlabeled news clusters can improve the identification speed.
[0027] In some embodiments, identifying the classification of the classified unlabeled news clusters and generating the corresponding news classification identifier includes: Obtaining the news classification identifier and the classification identifier feature set corresponding to the news classification identifier; Obtain the representative vector of the classified unlabeled news cluster based on the classified unlabeled news cluster; Obtain the news classification identifier of the classified unlabeled news cluster based on the representative vector and the classification identifier feature set.
[0028] It should be noted that the news classification identifier is an existing news classification identifier, such as finance, entertainment, politics, health, etc. The classification identifier feature set is usually the representative vocabulary and / or common vocabulary in this news field. For example, the classification identifier feature set of financial news may include vocabulary such as import volume, export volume, total trade volume, stock market, real estate market, etc. Among them, due to the different importance levels of each representative vocabulary and / or the occurrence frequencies of common vocabulary, the vocabulary weights of the representative vocabulary (and / or common vocabulary) in the classification identifier feature set can be set, so as to classify the news samples (or news data) more comprehensively. Among them, the representative vector can be the median vector or average vector of multiple vectorized feature semantics in the classified unlabeled news cluster, so as to be able to represent the classified unlabeled news cluster.
[0029] In some embodiments, the obtaining of the news classification identifier of the classified unlabeled news cluster based on the representative vector and the classification identifier feature set is specifically as follows: Calculate the similarity between the representative vector and the classification identifier features in the classification identifier feature set to obtain a similarity score; If the similarity score is greater than the similarity score threshold, record the similarity score and the candidate news classification identifier corresponding to the similarity score; Determine the news classification identifier from the candidate news classification identifiers based on the similarity score.
[0030] It should be noted that when calculating the similarity between the representative vector and the classification identifier feature set, the vector distance and / or vector cosine similarity can be calculated, and then the similarity score is obtained based on the vector distance and / or the connected cosine similarity. Among them, the similarity score threshold is used to determine whether the representative vector is the same as or similar to the representative vocabulary and / or common vocabulary in the classification identifier feature set. If the similarity score is less than or equal to the similarity score threshold, it means that the representative vector is not the same as or not relevant to the representative vocabulary or common vocabulary in the classification identifier feature set. The specific value of the similarity score threshold can be set according to actual needs, and this application does not make specific limitations on this.
[0031] It should be noted that there can be multiple candidate news classification identifiers, that is, the same representative vector can be the same as or similar to the representative vocabulary (and / or common vocabulary) in multiple classification identifier feature sets, so there are multiple candidate news classification identifiers.
[0032] In some embodiments, determining the news classification identifier based on the similarity score by the candidate news classification identifier further includes: Obtaining semantic feature weights corresponding to the feature semantics based on the structure and content of the news summary; Obtaining a news classification identifier score for the candidate news data based on the semantic feature weights and the similarity score; Obtaining the news classification identifier based on the news classification identifier score, where the news classification identifier includes a news classification main identifier and a news classification secondary identifier arranged in sequence.
[0033] It should be noted that since the importance of each feature semantics in each news text (or news data) is different, it is necessary to treat them separately. When obtaining the semantic feature weights, usually the first sentence and the last sentence of the news summary are key positions, and their corresponding semantic feature weights are relatively large; the semantic feature weights can also be set based on the occurrence times of the corresponding feature semantics; the semantic feature weights can also be set based on both the structure of the news summary and the occurrence times of the feature semantics. Among them, obtaining the semantic feature weights can be achieved through a large language model or obtained from the original news text (or original news data) based on a large language model, so as to ensure the accuracy of the obtained semantic feature weights. It should be noted that after obtaining the semantic feature weights and the similarity score, the news classification identifier score can be obtained by performing operations such as multiplication or addition on the semantic feature weights and the similarity score. Among them, each news text (or news data, news summary) usually includes multiple feature semantics, and each feature semantics has a semantic feature weight, and each feature semantics can obtain one or more news classification identifier scores for the candidate news data. After obtaining all the candidate news classifications and the corresponding news classification identifier scores, the total score can be calculated for each candidate news classification identifier, and the total identifier score of the news text (or news data, news summary) for the candidate news classification identifier can be obtained. Based on the numerical size of the total identifier score, the news classification main identifier and the news classification secondary identifier can be determined in sequence.
[0034] S103: Marking the news classification identifier to the unclassified news in the unclassified news cluster to obtain classified news; In some embodiments, the marking the news classification identifier to the unclassified news in the unclassified news cluster to obtain classified news is specifically: Obtaining the news classification identifier scores of the news classification main identifier and the news classification secondary identifier; Mark the main news classification identifier, the secondary news classification identifier, and the corresponding news classification identifier score to the unclassified news in the unclassified news cluster to obtain the classified news.
[0035] It should be noted that by marking both the main news classification identifier and the secondary news classification identifier on multiple news data (or news texts) in the unclassified news cluster, the classification (or field) to which the news text (or news data) belongs can be more comprehensively reflected. By jointly marking the corresponding news classification identifier score with the main news classification identifier and the secondary news classification identifier on the unclassified news cluster, the probability that the news text (or news data) belongs to each classification (or field) can be calculated simultaneously, facilitating multi-level and comprehensive classification of the news text (or news data). When predicting the classification of the news text, all classifications of the news text and their corresponding probabilities (or scores) can be output simultaneously.
[0036] S104: Train a basic learning model using the classified news to obtain a news classification model; It should be noted that by marking the main news classification identifier, the secondary news classification identifier, and the corresponding news classification identifier score on the unclassified news in the unclassified news cluster, the classified news can be obtained.
[0037] S105: Input the news to be classified into the news classification model to obtain the news type of the news to be classified.
[0038] In summary, the embodiment of the present application provides an unsupervised news automatic classification method, including performing unsupervised classification processing on a news data set to obtain an unclassified news cluster; performing identifier recognition on the unclassified news cluster to obtain corresponding news classification identifiers; marking the news classification identifiers to the unclassified news in the unclassified news cluster to obtain classified news; training a basic learning model using the classified news to obtain a news classification model; inputting the news to be classified into the news classification model to obtain the news type of the news to be classified, which can classify news using a machine learning model and can solve the problem that in the current news classification field, due to the existence of not only text content but also picture content, it is difficult to accurately classify news using traditional natural language processing.
[0039] In a second aspect, as Figure 2 shown, the embodiment of the present application provides an unsupervised news automatic classification system, including: A processing module 210, configured to perform unsupervised classification processing on a news dataset to obtain classified unlabeled news clusters; An identification module 220, configured to perform identification recognition on the classified unlabeled news clusters to obtain corresponding news classification identifications; label the news classification identifications to the classified unlabeled news in the classified unlabeled news clusters to obtain classified labeled news; A training module 230, configured to use the classified labeled news to train a basic learning model to obtain a news classification model; A classification module 240, configured to input news to be classified into the news classification model to obtain the news type of the news to be classified.
[0040] In some embodiments, the news summary dataset includes multiple news summaries; The performing unsupervised classification processing on the news summary dataset to obtain the classified unlabeled news clusters includes: Performing semantic segmentation and extraction on the news summary to obtain at least one feature semantic; Performing unsupervised classification processing on the feature semantic to obtain the classified unlabeled news clusters.
[0041] In some embodiments, the performing unsupervised classification processing on the feature semantic to obtain the classified unlabeled news clusters includes: Setting a sample distance and a sample cosine similarity for classification; Clustering a plurality of the feature semantics based on the sample distance and the sample cosine similarity to obtain the classified unlabeled news clusters.
[0042] In some embodiments, the performing identification recognition on the classified unlabeled news clusters and generating corresponding news classification identifications includes: Obtaining news classification identifications and a classification identification feature set corresponding to the news classification identifications; Obtaining a representative vector of the classified unlabeled news clusters based on the classified unlabeled news clusters; Obtaining the news classification identifications of the classified unlabeled news clusters based on the representative vector and the classification identification feature set.
[0043] In some embodiments, the obtaining the news classification identifications of the classified unlabeled news clusters based on the representative vector and the classification identification feature set is specifically: Calculating a similarity between the representative vector and classification identification features in the classification identification feature set to obtain a similarity score; If the similarity score is greater than a similarity score threshold, recording the similarity score and a candidate news classification identification corresponding to the similarity score; Determine the news classification identifier based on the similarity score and the candidate news classification identifier.
[0044] In some embodiments, determining the news classification identifier based on the similarity score and the candidate news classification identifier further includes: Obtain a semantic feature weight corresponding to the feature semantics based on the structure and content of the news summary; Obtain a news classification identifier score for the candidate news data based on the semantic feature weight and the similarity score; Obtain the news classification identifier based on the news classification identifier score, where the news classification identifier includes a news classification main identifier and a news classification secondary identifier arranged in sequence.
[0045] In some embodiments, marking the news classification identifier to the unclassified news in the unclassified news cluster to obtain classified news specifically includes: Obtain the news classification identifier scores of the news classification main identifier and the news classification secondary identifier; Mark the news classification main identifier, the news classification secondary identifier, and the corresponding news classification identifier scores to the unclassified news in the unclassified news cluster to obtain the classified news.
[0046] In a third aspect, an embodiment of the present application provides an electronic device, including an unsupervised news automatic classification system according to any one of the embodiments in the second aspect.
[0047] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatuses, systems), and / or computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0048] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes.
[0049] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in Figure 1 one process or more processes and / or boxes Figure 1 the functions specified in one box or more boxes.
[0050] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not used to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. An unsupervised news automatic classification method, characterized in that: include: Perform unsupervised classification on the news dataset to obtain classified unlabeled news clusters; Identify the classified unlabeled news cluster to obtain a corresponding news classification identifier; Marking the news classification identifier to the classified unmarked news in the classified unmarked news cluster to obtain classified marked news; Using the classified and labeled news to train the basic learning model, to obtain a news classification model; The news to be classified is input into the news classification model to obtain the news type of the news to be classified.
2. The unsupervised news automatic classification method according to claim 1, characterized in that: The unsupervised classification process is performed on the news data set to obtain classified unlabeled news clusters, including: Using a large language model to perform summary processing on the news dataset to obtain a news summary dataset; An unsupervised classification process is performed on the news summary data set to obtain the classified unlabeled news clusters.
3. The unsupervised news automatic classification method according to claim 2 is characterized in that: The news summary dataset includes multiple news summaries; The performing unsupervised classification processing on the news summary data set to obtain the classified unlabeled news clusters includes: Performing semantic segmentation and extraction on the news summary to obtain at least one characteristic semantic; The feature semantics are subjected to unsupervised classification processing to obtain the classified unlabeled news cluster.
4. The unsupervised news automatic classification method according to claim 3 is characterized in that: The performing unsupervised classification processing on the feature semantics to obtain the classified unlabeled news cluster includes: Set the sample distance and sample cosine similarity for classification; The plurality of feature semantics are clustered based on the sample distance and the sample cosine similarity to obtain the classified unlabeled news cluster.
5. The unsupervised news automatic classification method according to claim 3 is characterized in that: The step of identifying the classified unlabeled news cluster and generating a corresponding news classification identifier includes: Acquire a news classification identifier and a classification identifier feature set corresponding to the news classification identifier; Obtaining a representative vector of the classified unlabeled news cluster based on the classified unlabeled news cluster; A news classification identification of the classified unlabeled news cluster is obtained based on the representative vector and the classification identification feature set.
6. The unsupervised news automatic classification method according to claim 5, characterized in that: The news classification identification of the classified unlabeled news cluster is obtained based on the representative vector and the classification identification feature set, specifically: Calculating similarity between the representative vector and the classification identification features in the classification identification feature set to obtain a similarity score; If the similarity score is greater than the similarity score threshold, the similarity score and the to-be-selected news category identifier corresponding to the similarity score are recorded; The news category identifier is determined from the candidate news category identifiers based on the similarity score.
7. The unsupervised news automatic classification method according to claim 6, characterized in that: The determining the news classification identifier from the candidate news classification identifiers based on the similarity score also includes: Obtaining a semantic feature weight corresponding to the feature semantics based on the structure and content of the news summary; Obtaining a news classification identification score of the news data to be selected based on the semantic feature weight and the similarity score; The news classification identifier is obtained based on the news classification identifier score, wherein the news classification identifier includes a news classification primary identifier and a news classification secondary identifier arranged in sequence.
8. The unsupervised news automatic classification method according to claim 7, characterized in that: The step of marking the news classification identifier to the classified untagged news in the classified untagged news cluster to obtain the classified marked news is specifically as follows: Obtaining news category identification scores of the news category primary identification and the news category secondary identification; The news classification main identifier, the news classification auxiliary identifier and the corresponding news classification identifier score are marked to the classified unmarked news in the classified unmarked news cluster to obtain the classified marked news.
9. An unsupervised news automatic classification system, characterized in that: include: A processing module is used to perform unsupervised classification processing on the news data set to obtain classified unlabeled news clusters; An identification module is used to identify the classified untagged news cluster to obtain a corresponding news classification identification; mark the news classification identification to the classified untagged news in the classified untagged news cluster to obtain classified marked news; A training module, used to train a basic learning model using the classified and labeled news to obtain a news classification model; The classification module is used to input the news to be classified into the news classification model to obtain the news type of the news to be classified.
10. An electronic device, characterized in that: Comprising an unsupervised automatic news classification system as described in claim 9.
Citation Information
Patent Citations
News processing method, device, equipment and medium
CN110990705A
Scientific literature knowledge entity-oriented unsupervised identification method and system
CN116050419A
Streaming news clustering method and device and computer equipment
CN116304755A
Internet news analysis system and method based on big data
CN118093979A
Method and system to classify news snippets into categories using an ensemble of machine learning models
US20240330780A1