Industry classification method and system
Patent Information
- Application Number
- CN202311112879.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-08-30
AI Technical Summary
然而,在冷启动阶段,没有相关的行业类别打标数据时,无法使用该类有监督分类模型进行行业分类,并且有监督分类模型的表现强依赖于行业类别打标数据的数量和质量,当打标数据比较少或者打标质量比较差时,均会导致热点事件分类准确率较低
[0018]由以上技术方案可知,本说明书提供的行业分类方法和系统,该方案在获取待分类的多个样本数据和多个行业类别对应的关键词库之后,基于关键词库对多个样本数据进行匹配,并将与关键词库相匹配的样本数据作为第一样本数据,以及将与关键词库不匹配的样本数据作为第二样本数据,之后基于关键词库确定第一样本数据的行业类别,采用第一样本数据对第二样本数据进行匹配,以及基于与第二样本数据相匹配的第一样本数据的行业类别确定与其对应的第二样本数据的行业类别。该方案中,通过利用已分类的第一样本数据对第二样本数据进行行业分类,从而将未实现行业分类的样本数据归属到与其相似度较高的样本数据的行业分类中,能够提高样本数据的匹配覆盖度,达到对更多的样本数据进行行业分类的效果,进而实现及时准确地发掘出更多热点事件的效果。
Smart Images

Figure CN117194659B_ABST
Abstract
Description
Technical Field
[0001] This manual relates to the Internet field, and in particular to an industry classification method and system. Background Technology
[0002] The internet is a crucial channel for users to obtain information. Understanding user needs and promptly capturing trending events are key to effective information delivery. Currently, many methods employ supervised algorithm-based or rule-based event industry classification and topic generation models. However, in the cold start phase, without relevant industry category labeling data, these supervised classification models cannot be used. Furthermore, the performance of supervised classification models heavily depends on the quantity and quality of industry category labeling data; insufficient or poor-quality labeling data leads to low accuracy in trending event classification.
[0003] In summary, there is a need to provide a new industry classification method and system that can improve the accuracy of classifying trending events.
[0004] The information in the background section is merely information known only to the inventor and does not imply that such information had entered the public domain before the date of this application, nor does it imply that it can be considered prior art in this disclosure. Summary of the Invention
[0005] This manual provides an industry classification method and system with higher accuracy.
[0006] Firstly, this specification provides an industry classification method for classifying trending events by industry. The method includes: acquiring multiple sample data to be classified, each sample data including information on at least one trending event; acquiring a keyword library corresponding to multiple industry categories; matching the multiple sample data with the keyword library to determine the industry category of a first sample data that matches the keyword library; and matching a second sample data that does not match the keyword library with the first sample data, and determining the industry category of the corresponding second sample data based on the industry category of the successfully matched first sample data.
[0007] In some embodiments, the step of matching the second sample data that does not match the keyword database with the first sample data and determining the industry category of the corresponding second sample data based on the industry category of the successfully matched first sample data includes, for each second sample data: determining the similarity between the current second sample data and the first sample data pairwise; and taking the industry category of at least one first sample data with a similarity greater than a preset first similarity threshold or the largest similarity as the industry category of the current second sample data.
[0008] In some embodiments, matching the plurality of sample data with the keyword library to determine the industry category of the first sample data among the plurality of sample data includes, for each sample data in the plurality of sample data: extracting sample keywords of the current sample data; determining the similarity between the sample keywords and a plurality of keywords in the keyword library; and when it is determined that the current sample data matches the keyword library based on the similarity, determining its corresponding industry category based on the industry category corresponding to at least one keyword that matches the current sample data, or, when it is determined that the current sample data does not match the keyword library based on the similarity, using the current sample data as the second sample data.
[0009] In some embodiments, determining the corresponding industry category based on the industry category corresponding to at least one keyword that matches the current sample data includes: determining the industry category corresponding to at least one keyword with a similarity greater than a preset second similarity threshold, or with a similarity greater than the preset second similarity threshold and the largest similarity, as the initial industry category of the current sample data; and determining the industry category of the current sample data based on the initial industry category of the current sample data.
[0010] In some embodiments, determining the industry category of the current sample data based on the initial industry category of the current sample data includes: obtaining at least one historical sample data whose industry category was incorrectly classified and corrected, and its corresponding corrected industry category; determining the similarity between the current sample data and the at least one historical sample data; and obtaining the industry category of the current sample data based on the similarity between the current sample data and the at least one historical sample data, and the initial industry category of the current sample data.
[0011] In some embodiments, obtaining the industry category of the current sample data based on the similarity between the current sample data and the at least one historical sample data, and the initial industry category of the current sample data, includes: determining that when the maximum similarity among the similarities between the current sample data and the at least one historical sample data is greater than a preset third similarity threshold, using the corrected industry category of the historical sample data corresponding to the maximum similarity as the industry category of the current sample data; or, determining that when the maximum similarity among the similarities between the current sample data and the at least one historical sample data is less than a preset third similarity threshold, using the initial industry category as the industry category of the current sample data.
[0012] In some embodiments, obtaining a keyword library corresponding to multiple industry categories includes: obtaining a preset original keyword library; determining expanded keywords corresponding to the multiple industry categories, wherein the expanded keywords are obtained based on feature descriptions of their corresponding industry categories; and expanding the original keyword library based on the expanded keywords to obtain the keyword library.
[0013] In some embodiments, expanding the original keyword library based on the expanded keywords corresponding to the plurality of industry categories to obtain the keyword library includes: determining the similarity between the expanded keywords and the plurality of original keywords in the original keyword library; and adding the expanded keywords with a similarity greater than a preset fourth similarity threshold to the industry category corresponding to their original keywords to obtain the keyword library.
[0014] In some embodiments, after matching the second sample data that does not match the keyword database with the first sample data and determining the industry category of the second sample data based on the industry category of the first sample data, the method further includes: determining the topic of the sample data corresponding to at least one of the multiple industry categories.
[0015] In some embodiments, determining the topic of sample data corresponding to at least one of the plurality of industry categories includes, for each of the at least one industry category: clustering the sample data in the current industry category based on the topic similarity between sample data in the current industry category to obtain at least one sample data set; for at least a portion of the sample data sets in the at least one sample data set, sequentially using each sample data set in the at least a portion of the sample data sets as the current sample data set, normalizing the popularity values of sample data from different channels in the current sample data set, and obtaining the popularity value of the topic corresponding to the current sample data set based on the weighted sum of the normalized popularity values; determining the event name of the sample data with the highest normalized popularity value in the current sample data set as the name of the topic; and outputting the popularity value and name of the topic corresponding to the at least a portion of the sample data sets.
[0016] In some embodiments, the topic similarity includes string similarity and geographical similarity of the sample data.
[0017] Secondly, this specification also provides an industry classification system, comprising: at least one storage medium storing at least one instruction set for performing industry classification; and at least one processor communicatively connected to the at least one storage medium, wherein, when the industry classification system is running, the at least one processor reads the at least one instruction set and executes the industry classification method described in the first aspect of this specification according to the instructions of the at least one instruction set.
[0018] As can be seen from the above technical solutions, the industry classification method and system provided in this specification, after acquiring multiple sample data to be classified and a keyword database corresponding to multiple industry categories, matches the multiple sample data based on the keyword database. Sample data matching the keyword database is used as the first sample data, and sample data not matching the keyword database is used as the second sample data. Then, the industry category of the first sample data is determined based on the keyword database. The first sample data is then used to match the second sample data, and the industry category of the second sample data is determined based on the industry category of the first sample data that matches the second sample data. In this solution, by using the already classified first sample data to classify the second sample data by industry, sample data that has not yet been classified by industry is assigned to the industry category of sample data with high similarity. This improves the matching coverage of sample data, achieving the effect of classifying more sample data by industry, and thus enabling the timely and accurate discovery of more trending events.
[0019] Other functions of the industry classification methods and systems provided in this specification will be partially listed in the following description. The figures and examples described below will be readily apparent to those skilled in the art. The inventive aspects of the industry classification methods and systems provided in this specification can be fully understood through practice or use of the methods, apparatus, and combinations described in the detailed examples below. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A hardware structure diagram of a computing device provided according to an embodiment of this specification is shown;
[0022] Figure 2 A flowchart of an industry classification method P100 provided according to an embodiment of this specification is shown;
[0023] Figure 3 A flowchart illustrating the creation of a keyword library provided in an embodiment of this specification is shown;
[0024] Figure 4 A flowchart illustrating the matching process between multiple sample data and a keyword database according to embodiments of this specification is shown;
[0025] Figure 5 A schematic diagram of negative feedback correction provided according to embodiments of this specification is shown; and
[0026] Figure 6 A schematic diagram illustrating topic generation provided according to embodiments of this specification is shown. Detailed Implementation
[0027] The following description provides specific application scenarios and requirements for this specification, intended to enable those skilled in the art to make and use the contents of this specification. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but rather to the widest scope consistent with the claims.
[0028] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not restrictive. For example, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. When used in this specification, the terms “comprising,” “including,” and / or “containing” mean that the associated integers, steps, operations, elements, and / or components are present, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups, or that other features, integers, steps, operations, elements, components, and / or groups may be added to the system / method.
[0029] Considering the following description, these and other features of this specification, as well as the operation and function of the related components of the structure, and the economy of assembly and manufacture of the parts, can be significantly improved. All of these form part of this specification with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0030] The flowcharts used in this specification illustrate operations implemented according to some embodiments of this specification. It should be clearly understood that the operations in the flowcharts may not be implemented in a sequential order. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0031] In this specification, "X includes at least one of A, B, or C" means that X includes at least A, or X includes at least B, or X includes at least C. That is, X may include only one of A, B, and C, or any combination of A, B, and C, as well as other possible content / elements. The arbitrary combination of A, B, and C can be A, B, C, AB, AC, BC, or ABC.
[0032] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.
[0033] For ease of description, the terms that will appear in the following descriptions will be explained as follows:
[0034] Hot topics refer to events that attract widespread attention, spark discussion, incite public sentiment, and generate strong reactions in society. They are generally characterized by their timeliness, challenging nature, universality, sensitivity, and variability.
[0035] Text classification is a natural language processing technique used to categorize large amounts of text data into different categories according to specific criteria. Text classification helps users analyze and classify massive amounts of text data, thereby better understanding the data and extracting valuable information. Text classification has been widely applied in areas such as internet search engines, spam filtering, product review analysis, news categorization, and sentiment analysis.
[0036] Machine learning: Machine learning is a multidisciplinary field that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance.
[0037] Unsupervised learning: Unsupervised learning is a training method for machine learning. It is essentially a statistical method that can discover potential structures in unlabeled data.
[0038] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Its main applications include machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, and Chinese OCR (Optical Character Recognition).
[0039] String similarity: String similarity is a metric that measures the distance between two text strings to perform approximate string matching or comparison.
[0040] TF-IDF stands for Term Frequency-Inverse Document Frequency. It is a statistical method used to assess the importance of a word to a document within a set of documents or a corpus. The importance of a word increases proportionally to the frequency of its occurrence in the document, but decreases inversely proportionally to its frequency of occurrence in the corpus.
[0041] Hot topics refer to events that attract widespread attention, spark discussion, incite public sentiment, and generate strong reactions in society. They generally possess characteristics such as timeliness, challenge, universality, sensitivity, and variability. For example, the occurrence of disease outbreaks can lead to heated discussions about certain medical industry keywords online, resulting in a surge in search volume on search engines. For operations personnel, the ability to quickly grasp industry hot topics, understand user needs, and implement corresponding operational strategies is crucial for industry operations. Therefore, a method that can automatically categorize hot topics by industry is needed to improve operational efficiency. This method can also summarize hot topics into subtopics and push them to operations personnel, enabling them to promptly identify hot topic operational opportunities and further improve operational efficiency.
[0042] When classifying trending events by industry, a text classification model can be built based on a large amount of labeled data on the industry categories of the events, using traditional machine learning algorithms such as Support Vector Machine (SVM) and LightGBM, or deep learning algorithms such as Long Short-Term Memory (LSTM) and Bidirectional Encoder Representations (BERT). This method requires preparing a large amount of labeled data on the industry categories of the events in advance, consuming significant manpower for data labeling. It is not suitable for the cold start phase when there is no labeled data, and the model's performance is highly dependent on the quantity and quality of the labeled samples. When the amount of labeled data is small, or the coverage or types are limited, the classification accuracy of the model trained on labeled data is low. Alternatively, a keyword library can be built based on experience, and industry classification and topic generation can be performed based on manually generated rules. This method requires significant manpower to maintain the keyword library and human judgment logic, is highly dependent on human experience, and due to the diversity of natural language expressions, the accuracy of manual industry classification is also low.
[0043] The industry classification method provided in this manual, after classifying trending events by industry, can be further applied to topic generation. This allows for the generation of popular industry topics based on the industry classification results, which can then be pushed to relevant personnel, such as operations staff, to facilitate their decision-making regarding operational opportunities and product feature optimization suggestions. Alternatively, related services and content can be pushed to users based on these popular topics. The industry classification method also allows for sample labeling during the training phase of classification models, efficiently acquiring labeled data, reducing labor costs, and improving the accuracy of data labeling. Finally, the industry classification method can be applied in information recommendation scenarios. For example, by classifying trending events by industry and generating popular industry topics based on these classifications, related services and content can be pushed to users based on these popular topics. The industry classification method described in this manual can also be applied to public opinion monitoring scenarios. For example, hot topics can be categorized by industry using this method, and trending topics can be generated based on these categorized events. These trending topics can then be pushed to operations personnel, enabling them to monitor public opinion information on social media, forums, news channels, etc., thereby understanding public attitudes and concerns about the industry and formulating corresponding operational strategies. For instance, in the public utility payment industry, based on the captured trending topic of "electricity price adjustment," operations personnel can publish related content to increase the concurrent article views. Furthermore, in the travel industry, the captured trending topic of "noise is the reason for missing the subway stop in location A" can be pushed to operations personnel, allowing them to decide to provide users with a "subway alighting reminder" function, and so on.
[0044] Those skilled in the art should understand that the industry classification methods and systems described in this specification, when applied to other use cases, are also within the scope of protection of this specification.
[0045] Figure 1 A hardware structure diagram of a computing device 600 provided according to an embodiment of this specification is shown. The computing device 600 can be a device for classifying trending events by industry. In this case, the computing device 600 can store data or instructions for executing the industry classification method described in this specification, and can execute or be used to execute said data or instructions. In some embodiments, the computing device 600 may include hardware devices with data processing capabilities and necessary programs for driving the hardware devices to operate. In some embodiments, the computing device 600 may include mobile devices, tablet computers, laptop computers, personal computers, servers, server clusters, distributed servers, cloud servers, etc., or any combination thereof.
[0046] like Figure 1As shown, the computing device 600 may include at least one storage medium 630 and at least one processor 620. In some embodiments, the computing device 600 may also include a communication port 650 and an internal communication bus 610. Additionally, the computing device 600 may include I / O components 660.
[0047] The internal communication bus 610 can connect different system components, including storage medium 630, processor 620 and communication port 650.
[0048] I / O component 660 supports input / output between computing device 600 and other components.
[0049] Communication port 650 is used for data communication between computing device 600 and external sources. For example, communication port 650 can be used for data communication between computing device 600 and network 400. Communication port 650 can be a wired communication port or a wireless communication port.
[0050] Storage medium 630 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 632, a read-only storage medium (ROM) 634, or a random access storage medium (RAM) 636. Storage medium 630 also includes at least one instruction set stored in the data storage device. The instructions are computer program code, which may include programs, routines, objects, components, data structures, procedures, modules, etc., that execute the industry classification methods provided in this specification.
[0051] At least one processor 620 can be communicatively connected to at least one storage medium 630 and a communication port 650 via an internal communication bus 610. The at least one processor 620 is used to execute the at least one instruction set described above. When the computing device 600 is running, the at least one processor 620 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the industry classification method provided in this specification. The processor 620 can execute all the steps included in the industry classification method. The processor 620 can be in the form of one or more processors. In some embodiments, the processor 620 may include one or more hardware processors, such as a microcontroller, microprocessor, reduced instruction set computer (RISC), application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), central processing unit (CPU), graphics processing unit (GPU), physical processing unit (PPU), microcontroller unit, digital signal processor (DSP), field-programmable gate array (FPGA), advanced RISC machine (ARM), programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof. For illustrative purposes only, only one processor 620 is described in this specification for the computing device 600. However, it should be noted that the computing device 600 in this specification may also include multiple processors. Therefore, the operation and / or method steps disclosed in this specification may be executed by one processor as described in this specification, or they may be executed jointly by multiple processors. For example, if the processor 620 of the computing device 600 in this specification executes steps A and B, it should be understood that steps A and B may also be executed jointly or separately by two different processors 620 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0052] Figure 2 A flowchart of an industry classification method P100 according to an embodiment of this specification is shown. As previously illustrated, the computing device 600 can execute the industry classification method P100 of this specification. Specifically, the processor 620 can read an instruction set stored in its local storage medium and then execute the industry classification method P100 of this specification according to the specifications of the instruction set. Figure 2 As shown, method P100 may include:
[0053] S120: Obtain multiple sample data to be classified.
[0054] Each sample data set in the multiple sample data sets may include information about at least one hot topic event. The hot topic event information is descriptive information about the hot topic event. The sample data can be in any form, such as text data, video data, image data, audio data, web page data, etc. For ease of description, we will use text data as an example. When the sample data is video data, image data, audio data, or web page data, the computing device 600 can extract the text content from the sample data, such as the title, as text-based sample data. Alternatively, the computing device 600 can also extract text data after converting video data, image data, audio data, or web page data into text format as text-based sample data. A sample data set can contain information about one hot topic event or multiple hot topic events. When a sample data set contains information about multiple hot topic events, these multiple hot topic events can belong to the same industry category or different industry categories.
[0055] The computing device 600 can automatically crawl information on trending events from various industries (such as people's livelihood, medical care, travel, etc.) from multiple information sources (such as news websites, social media, etc.) using web crawlers. The information on trending events can be in text or non-text format; non-text format can include images, audio, video, etc. Text-based information on trending events can be directly used as sample data. For non-text-based trending events, the computing device 600 needs to convert it to text or extract the text content (such as titles) from the non-text-based trending events, and then use the converted text-based information or the text content from the non-text-based trending events as sample data.
[0056] S140: Obtain keyword libraries corresponding to multiple industry categories.
[0057] A keyword database is a list of keywords. This list contains several keywords used to identify industry categories. Keywords can be terms, nouns, phrases, or expressions related to the industry category. For example, various drug names, medical device names, medical service names, and medical service institution names used in the medical industry category can all be keywords for that category. Similarly, nouns such as school names, educational institution names, teachers, and students in the education industry category can all be keywords for that category. The keyword database in this manual is a database of keywords related to various service names or industry terms within each industry category; therefore, it can also be called an industry keyword database. Keywords within keywords can also be called industry keywords.
[0058] The keyword database corresponds to multiple industry categories. These categories can be defined by various industry classification standards, such as the National Industrial Classification of Economic Activities (NISA) or the International Standard Industrial Classification of All Economic Activities (ISIC). Examples include categories like food, travel, hotels, transportation, lifestyle, and maternal and infant products. Keywords in the database are categorized according to these industry groups. Each keyword corresponds to an specific industry category. The database can also store the mapping between keywords and their corresponding industry categories. Keywords across different industry categories may overlap; therefore, each keyword may correspond to one or more industry categories.
[0059] The computing device 600 can build a keyword library based on how it interacts with users. Figure 3 A flowchart illustrating the creation of a keyword library provided in an embodiment of this specification is shown. For example... Figure 3 As shown, users can manually add industry keywords for various industry categories to the database to establish a mapping relationship between keywords and industry categories, thereby obtaining the original keyword library. The aforementioned keyword library can include the original keyword library.
[0060] Natural language has diverse forms of expression, and different people can express the same meaning in very different ways. To enrich the expression of keywords under industry categories in the original keyword database, the computing device 600 can also determine expanded keywords for each industry category based on the feature descriptions of each industry category, thereby determining expanded keywords corresponding to multiple industry categories, and expanding the obtained original keyword database based on the expanded keywords to obtain the keyword database.
[0061] Some service platforms contain service data sources with industry-specific tagging data, such as service names for various industry categories. These service names describe the characteristics of each industry category. The computing device 600 acquires these service names from various industry categories on these platforms and then uses natural language processing methods such as word segmentation, stop word removal, and TF-IDF to automatically extract expanded keywords from them. Furthermore, for the same industry category, the computing device 600 can determine the pairwise similarity between the expanded keywords in that industry category and the keywords in the original keyword library for that industry category. It also adds expanded keywords with similarity greater than a preset fourth similarity threshold to the industry category corresponding to their original keywords, thus obtaining a keyword library. The similarity here can be achieved using string similarity algorithms, such as the Longest Common Subsequence algorithm, the Jaccard Similarity Coefficient algorithm, the Cosine Similarity algorithm, the SimHash algorithm, the Hamming Distance algorithm, the Levenshtein Distance algorithm, and so on. This allows for the selection of keywords with high similarity to those in the original keyword database, thereby expanding the original keyword database and making the expression of keywords with the same meaning under various industry categories more diverse.
[0062] It should be noted that for expanded keywords with a similarity equal to the preset fourth similarity threshold, the computing device 600 may use them as expanded keywords or may not use them as expanded keywords. Those skilled in the art can choose according to actual needs, and this embodiment does not impose any restrictions on this.
[0063] Furthermore, the preset fourth similarity threshold can be determined based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations. Taking the computing device 600 setting the fourth similarity threshold based on historical records as an example, historical records can reflect the mapping relationship between the similarity between the expanded keywords and the keywords in the original keyword library. When the similarity between the expanded keywords and the keywords in the original keyword library is low, it will lead to low matching accuracy for the industry category of the sample data in the future. The computing device 600 can set the similarity corresponding to the required matching accuracy as the fourth similarity threshold.
[0064] S160: Match multiple sample data with the keyword database to obtain the industry category of the first sample data that matches the keyword database.
[0065] After multiple sample data sets are processed into text format, the computing device 600 can use text processing methods, such as keyword matching, to match the multiple sample data sets with a keyword database to determine whether the sample data sets contain keywords from the keyword database. For example, the computing device 600 can compare each sample data set with the keyword list in the keyword database to determine if there are any matches, and determine the industry category of the successfully matched sample data based on the matching results. The specific implementation process of S160 is described in detail below with reference to the accompanying drawings:
[0066] Figure 4 A flowchart illustrating the matching process between multiple sample data sets and a keyword database, as described in embodiments of this specification, is presented. Figure 4 As shown, for each sample data in the multiple sample data sets, S160 may include the following steps:
[0067] S160-2: Extract sample keywords from the current sample data.
[0068] Sample keywords are words or phrases with specific meanings and importance within the information of trending events corresponding to the current sample data. They can be used to summarize and describe the industry category of the trending events, helping users or search engines quickly understand their industry category. Sample keywords can be terms, nouns, verbs, or adjectives related to the industry category of the trending events.
[0069] The number of sample keywords can be less than the number of words extracted from the current sample data. Hot topic text includes a title and body, with the title typically indicating the industry category. The computing device 600 can extract sample keywords from the title of the hot topic event, or from the text itself, or from both the title and body. Keyword extraction can employ natural language processing methods such as word segmentation, stop word removal, and retention of industry keywords.
[0070] S160-4: Determine the similarity between the sample keywords and multiple keywords in the keyword library.
[0071] The current sample data contains at least one sample keyword. When there is only one sample keyword, the computing device 600 can calculate the similarity between this sample keyword and every keyword in the keyword library, thus obtaining the similarity between the sample keyword in the current sample data and multiple keywords in the keyword library. When there are multiple sample keywords, the computing device 600 can calculate the similarity between each of these multiple sample keywords and each keyword in the keyword library, thus obtaining the similarity between the multiple sample keywords in the current sample data and multiple keywords in the keyword library. The similarity calculation can be implemented using string similarity algorithms such as the longest common subsequence algorithm, Jaccard similarity coefficient algorithm, cosine similarity algorithm, SimHash algorithm, Hamming distance algorithm, and edit distance algorithm.
[0072] After determining the similarity between the current sample data and the keyword database based on step S160-4, the computing device 600 can perform a preliminary match between the current sample data and the keyword database based on the similarity. This can also be understood as performing a preliminary search in the keyword database based on the current sample data. For example, the computing device 600 can compare the similarity with a preset second similarity threshold and determine the preliminary matching result based on the comparison result. The preliminary matching result indicates whether the current sample data matches the keyword database. For example, if the similarity is greater than the preset second similarity threshold, the computing device 600 can determine that the current sample data matches the keyword database. Conversely, if the similarity is less than the preset second similarity threshold, the computing device 600 determines that the current sample data does not match the keyword database. A similarity greater than the preset second similarity threshold can mean that at least one keyword in the current sample data has a similarity greater than the preset second similarity threshold with at least one keyword in the keyword database. A similarity less than the preset second similarity threshold can mean that the similarity between all keywords in the current sample data and all keywords in the keyword database is less than the preset second similarity threshold.
[0073] The preset second similarity threshold can be determined by the computing device 600 based on requirements, historical records, experiments, etc., and this embodiment does not limit this. Similarly, the preset second similarity threshold can be set by the computing device 600 based on requirements, historical records, experiments, etc., and this embodiment does not limit this. The principle of the computing device 600 setting the second similarity threshold can be found in the description of the computing device 600 setting the fourth similarity threshold, and will not be repeated here.
[0074] It should be noted that, for cases where the similarity is equal to the second preset similarity threshold, the computing device 600 can identify whether the current sample data matches or does not match the keyword database. Those skilled in the art can choose according to actual needs, and this specification does not impose any restrictions on this.
[0075] S160-6: When determining the match between the current sample data and the keyword library based on similarity, the corresponding industry category is determined based on the industry category corresponding to at least one keyword that matches the current sample data.
[0076] If the current sample data matches the keyword library, it means that the computing device 600 can initially search the keyword library and find at least one keyword that matches the current sample data. The at least one keyword that matches the current sample data can be a keyword in the keyword library whose similarity to any sample keyword in the current sample data is greater than a preset first similarity threshold.
[0077] Sometimes, keywords matching the current sample data may include multiple keywords. These multiple keywords may belong to the same industry category or to multiple different industry categories. Therefore, the computing device 600 determines the industry category corresponding to the current sample data based on the industry category corresponding to at least one keyword matching the current sample data. Step S160-6 can be implemented in several ways, specifically as follows:
[0078] In some embodiments, the computing device 600 may determine the industry category corresponding to at least one keyword with a similarity greater than a preset second similarity threshold as the initial industry category of the current sample data. The computing device 600 may choose the industry category corresponding to the keyword with the highest similarity as the initial industry category of the current sample data, or it may choose the industry categories corresponding to multiple keywords with the highest similarity as the initial industry category of the current sample data. For example, the computing device 600 may sort multiple keywords matching the current sample data in descending order of similarity and select the industry categories corresponding to the top K keywords as the initial industry categories of the current sample data, thereby improving the accuracy of industry classification. Here, K is a positive integer. The initial industry category of the current sample data may be one or multiple.
[0079] In some embodiments, the computing device 600 may further determine the industry category corresponding to at least one keyword with a similarity greater than a preset second similarity threshold as the initial industry category of the current sample data. The computing device 600 may select the industry categories corresponding to all keywords among the at least one keyword with a similarity greater than the preset second similarity threshold as the initial industry category of the current sample data, or it may select the industry categories corresponding to some keywords among the at least one keyword with a similarity greater than the preset second similarity threshold as the initial industry category of the current sample data; this embodiment does not limit this selection.
[0080] Subsequently, the computing device 600 can also determine the industry category of the current sample data based on the initial industry category. For example, the computing device 600 can directly use the initial industry category as the industry category corresponding to the current sample data. Alternatively, the computing device 600 can also determine the classification accuracy of the initial industry category and correct it to obtain the final industry category of the current sample data.
[0081] Sometimes, the initial industry category obtained in the preliminary search results may be incorrect. To further improve the accuracy of industry classification, the computing device 600 can determine the industry category of the current sample data based on the historical sample data of the negative feedback and the initial industry category. The historical sample data of the negative feedback refers to information on historical hot events that were incorrectly classified and have been corrected. The industry category of the corrected historical sample data is the corrected industry category. The negative feedback correction process is described below with reference to the accompanying diagram. Figure 5 A schematic diagram of negative feedback correction provided according to embodiments of this specification is shown. Figure 5 As shown, the computing device 600 can correct the initial industry category of the current sample data based on historical sample data with negative feedback that is similar to the current sample data. For example, the computing device 600 can determine the similarity between the current sample data and at least one historical sample data by acquiring at least one historical sample data whose industry category was incorrectly classified and corrected, and its corresponding corrected industry category. Then, based on the similarity between the current sample data and at least one historical sample data, and the initial industry category of the current sample data, the computing device 600 can obtain the industry category of the current sample data.
[0082] The accuracy of industry classification results for historical sample data can be determined manually or through model recognition. If errors are identified in the industry classification results of historical sample data through manual verification or model identification, manual correction can be used to correct the classification results. The historical sample data and its corresponding corrected industry category are then stored in a historical sample database to correct the industry classification results of subsequent sample data. The historical sample database includes at least one historical sample data point and its corresponding corrected industry category.
[0083] After obtaining the initial industry category of the current sample data, the computing device 600 can determine whether the initial industry category of the current sample data needs to be corrected based on the similarity between the current sample data and historical sample data. For example, the computing device 600 can calculate the similarity between the current sample data and each of at least one historical sample data, select the historical sample data with the highest similarity to the current sample data, and determine whether the initial industry category of the current sample data needs to be corrected based on the corrected industry category of the historical sample data with the highest similarity. When historical sample data similar to the current sample data is identified, the computing device 600 can use some of the string similarity algorithms described above to determine the similarity between the current sample data and each historical sample data. In some embodiments, when performing similarity calculations, the computing device 600 can determine the pairwise similarity between the current sample data and each historical sample data based on keywords. For example, the computing device 600 extracts keywords from the current historical sample data and the current sample data respectively, and obtains the similarity between the current sample data and the current historical sample data based on the similarity between the keywords of the current historical sample data and the keywords of the current sample data. The similarity between current sample data and current historical sample data can be a statistical measure of the pairwise similarity between keywords, such as a weighted sum, a weighted average, etc. In some embodiments, the computing device 600 can also determine the similarity between them based on the words in the entire text. For example, after preprocessing the current historical sample data and the current sample data by word segmentation and stop word removal, the computing device 600 uses a natural language processing algorithm to determine the vector representations of the current historical sample data and the current sample data based on the words in the preprocessed current historical sample data and the words in the preprocessed current sample data, thereby calculating the similarity between the two.
[0084] The computing device 600 can also determine whether the initial industry category needs to be corrected based on a comparison between the maximum similarity among at least one similarity between the current sample data and at least one historical sample data and a preset third similarity threshold. The historical sample data corresponding to the maximum similarity is the historical sample data in the historical sample database that is most similar to the current sample data in content. The maximum similarity value can vary, and its magnitude will affect the correction result of the initial industry category. To improve the accuracy of the correction of the initial industry category, the computing device 600 can compare the maximum similarity with a preset third similarity threshold and determine whether the initial industry category needs to be corrected based on the comparison result.
[0085] When the maximum similarity exceeds a preset third similarity threshold, it indicates a high similarity between the current sample data and the historical sample data with the highest similarity. This implies a high similarity in their respective industry categories, suggesting a higher probability of an initial industry category misclassification for the current sample data. Therefore, the computing device 600 needs to correct the initial industry category of the current sample data. In this case, the computing device 600 can use the corrected industry category corresponding to the historical sample data with the highest similarity to the current sample data, along with the initial industry category, to jointly determine the industry category of the current sample data. For example, the computing device 600 can use the corrected industry category of the historical sample data corresponding to the highest similarity as the industry category of the current sample data. By using the corrected industry category from historical sample data to correct the industry category of the current sample data, the computing device 600 can improve the accuracy of its industry classification and reduce the incidence of similarity errors.
[0086] When the maximum similarity is less than a preset third similarity threshold, it indicates a low similarity between the current sample data and the historical sample data with the maximum similarity. Therefore, the initial industry category of the current sample data is less likely to be misclassified. Consequently, the computing device 600 does not need to correct the initial industry category of the current sample data. In this case, the computing device 600 can directly use the initial industry category of the current sample data as its industry category.
[0087] It should be noted that when the maximum similarity is equal to the preset third similarity threshold, the initial industry category of the current sample data may or may not be corrected, and this embodiment does not limit this.
[0088] Furthermore, the preset third similarity threshold can be determined based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations. Similarly, the preset third similarity threshold can be set by the computing device 600 based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations. Moreover, the principle of the computing device 600 setting the third similarity threshold can be found in the description of the computing device 600 setting the fourth similarity threshold, and will not be repeated here.
[0089] After correcting misclassified sample data based on historical sample data with negative feedback, the corrected sample data can be added to the historical sample database to continuously enrich the historical sample database, thereby correcting subsequent sample data and further improving the accuracy of industry classification.
[0090] Continue reading Figure 4 After step S160-6, step S160 may also include step S160-8.
[0091] S160-8: When it is determined that the current sample data does not match the keyword database based on similarity, the current sample data shall be used as the second sample data.
[0092] If the current sample data does not match the keyword database, it means that an initial match in the keyword database cannot yield keywords that match the current sample data. In this case, the computing device 600 can use the current data as second sample data for subsequent matching.
[0093] In the embodiments shown in steps S160-2 to S160-8, the computing device 600 extracts keywords for each sample data from multiple sample data sets individually, and then matches them with a keyword database. To improve the convenience of keyword extraction, in some embodiments, the computing device 600 can also extract sample keywords uniformly for multiple sample data sets, and then match the sample keywords of each sample data set with a keyword database. The specific matching process can be found in the descriptions of steps S160-2, S160-4, S160-6, and S160-8 in the above embodiments, and will not be repeated here.
[0094] Continue reading Figure 3 After step S160, the method P100 may further include step S180.
[0095] S180: Match the second sample data that does not match the keyword database with the first sample data from multiple sample data, and determine the industry category of the corresponding second sample data based on the industry category of the first sample data that is successfully matched.
[0096] After the initial search, the sample data that successfully matches the keyword database is designated as the first sample data, while the sample data that fails to match or does not match the keyword database is designated as the second sample data. As mentioned earlier, a failure to match or a mismatch with the keyword database can be defined as the similarity to all keywords in the keyword database being less than a preset first similarity threshold. The second sample data, because it does not match the keyword database, therefore does not find a corresponding industry category in the keyword database. Due to the diversity of natural language expression, it is impossible to exhaustively list all keywords under each industry category in the keyword database; therefore, some sample data will fail to match the keyword database. However, there may be similarity between the successfully matched first sample data and the unmatched second sample data. This specification can classify the second sample data into industry categories based on the classified first sample data that has a high similarity to the second sample data. The higher the similarity between the first and second sample data, the more similar the content or information they describe, and the higher the probability that they belong to the same category. In some embodiments, for each second sample data in at least one second sample data set, the computing device 600 can further determine the similarity between the current second sample data and the first sample data pairwise, and determine the industry category of the at least one first sample data whose similarity is greater than a preset first similarity threshold as the industry category of the current second sample data. The computing device 600 can determine the industry category of the current second sample data by identifying the industry categories corresponding to all first sample data in at least one first sample data set whose similarity is greater than the preset first similarity threshold, or it can determine the industry categories corresponding to some first sample data in at least one first sample data set whose similarity is greater than the preset first similarity threshold as the industry category of the current second sample data. This embodiment does not limit this approach. In this embodiment, industry classification can be achieved for second sample data with a similarity greater than the preset first similarity threshold; however, second sample data with a similarity less than the preset first similarity threshold can be classified into other categories, i.e., unclassified data.
[0097] In some embodiments, to improve the matching coverage of sample data, for each second sample data in at least one second sample data set, the computing device 600 can further determine the industry category of the at least one first sample data set with the highest similarity by determining the similarity between each pair of the current second sample data and the first sample data. For example, the computing device 600 can sort the industry categories corresponding to keywords in the first sample data in descending order of similarity, and select the industry categories of the first sample data corresponding to the top P keywords as the industry categories of the current second sample data, thereby improving the accuracy of industry classification. Here, P is a positive integer. The current second sample data may have one or more industry categories. This allows for matching coverage of all sample data.
[0098] It should be noted that for sample data with a similarity equal to a preset first similarity threshold, the computing device 600 may classify it or not. Those skilled in the art can choose according to actual needs, and this embodiment does not impose any restrictions on this.
[0099] Furthermore, the preset first similarity threshold can be determined based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations. Similarly, the preset first similarity threshold can be set by the computing device 600 based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations. Moreover, the principle of the computing device 600 setting the first similarity threshold can be found in the description of the computing device 600 setting the fourth similarity threshold, and will not be repeated here.
[0100] In summary, the method P100 provided in this specification uses an unsupervised algorithm to classify hot topics by industry. It does not require a large amount of manually labeled data in advance; only industry keywords are needed. Based on keyword expansion, keyword matching, and string similarity calculation, it can classify hot topics by industry, reducing reliance on data labeling and significantly lowering manual maintenance costs. Therefore, method P100 is suitable for the cold start phase when there is no labeled data. Compared to deep learning models, method P100 has advantages such as low complexity, fast running speed, and low memory consumption. Furthermore, method P100 can expand the original keyword library based on other data sources to automatically expand keywords for each industry, alleviating the problem of the diversity of natural language expressions and the inability to exhaustively list keywords, thus reducing manual maintenance costs. On the other hand, in step S160 of method P100, a second similarity threshold is set to determine whether the keyword library matches the sample data, and sample data with similarity higher than the second similarity threshold are classified by industry based on the magnitude of the similarity. For sample data whose similarity is below a preset second similarity threshold, they are considered mismatched with the keyword database. In step S180, industry classification is re-performed. The similarity between the first sample data that matches an industry category and the second sample data that does not match an industry category is used to propagate the industry classification of the second sample data. This further alleviates the problems of inaccurate classification caused by the diversity of natural language expressions and the inability to exhaustively list all cases when manually maintaining keywords and classification logic. Moreover, method P100 has negative feedback reinforcement learning capabilities, preventing similar classification errors from recurring and continuously improving classification accuracy. In summary, method P100 provided in this specification can improve the industry classification accuracy of sample data, preventing sample data from being assigned to an incompatible industry category when similarity is low. Simultaneously, method P100 can also improve the sample coverage of industry classification, enabling accurate industry classification for all sample data.
[0101] For new sample data, the computing device 600 can first match the newly generated sample data with a keyword database. If the newly generated sample data matches the keyword database, the computing device 600 can use the industry category corresponding to the matching keyword as its corresponding industry category. If the newly generated sample data does not match the keyword database, the computing device 600 can match it with pre-classified sample data and use the industry category of the pre-classified sample data that matches it as its corresponding industry category.
[0102] After acquiring industry categories from multiple sample data sets, the computing device 600 can further apply the classification results of these industry categories to other scenarios. For example, the computing device 600 can apply the industry category classification results to topic generation scenarios to generate topics from sample data under different industry categories. Another example is that the computing device 600 can use the industry category classification results as labeled data for a text classification model and apply it to the training process of the text classification model. Yet another example is that the computing device 600 can apply the industry category classification results to data mining scenarios to extract valuable information, such as suggestions and opinions, from trending events under different industry categories. For ease of description, we will further illustrate this by taking the application of industry category classification results to a topic generation scenario as an example.
[0103] After determining the industry categories of multiple sample data, the computing device 600 can also generate topics for the sample data under different industry categories, thereby pushing these topics to operations personnel so that they can perform relevant operational actions based on the topics, thus improving operational efficiency. Therefore, after step S180, the method P100 may further include: the computing device 600 determining topics for the sample data corresponding to at least one of the multiple industry categories. For example, at least one industry category could be all industry categories. Alternatively, at least one industry category could be an industry category that meets preset conditions. The preset conditions could be that the number of sample data under an industry category is greater than a preset quantity threshold or that the overall popularity value of the sample data under an industry category is greater than a preset popularity threshold. If the number of sample data under certain industry categories is small or the overall popularity value of the sample data under an industry category is low, it indicates that people have low attention to these sample data, and generating topics for them may not attract user attention or promote information dissemination. Therefore, the computing device 600 can only generate topics for the sample data corresponding to at least one industry category that meets the preset conditions among the multiple industry categories.
[0104] Figure 6 A schematic diagram illustrating topic generation according to embodiments of this specification is shown. (e.g.) Figure 6As shown, when generating topics for sample data corresponding to at least one industry category, the computing device 600 can perform the following for each industry category: clustering the sample data in the current industry category based on the topic similarity between sample data in the current industry category to obtain at least one sample data set; and for at least a portion of the sample data sets in the at least one sample data set, taking each sample data set in the at least a portion of the sample data sets as the current sample data set in turn; normalizing the popularity values of sample data from different channels in the current sample data set; and obtaining the popularity value of the topic corresponding to the current sample data set based on the weighted sum of the normalized popularity values; and determining the event name of the sample data with the highest normalized popularity value in the current sample data set as the name of the topic, and outputting the popularity value and name of the topic corresponding to at least a portion of the sample data sets.
[0105] In some embodiments, the topic similarity between sample data in the current industry category may include string similarity. In some embodiments, the topic similarity between sample data in the current industry category may include string similarity and geographic similarity. String similarity can be obtained using a string similarity algorithm. Geographic similarity can be determined based on whether the geographic information contained in the sample data is consistent. Then, the computing device 600 can cluster the sample data with string similarity greater than a preset fifth similarity threshold and consistent geographic information to obtain at least one sample data set. Consistent geographic information can mean that the distance between geographic coordinates is less than a preset distance, or that the regional information is consistent, such as consistent province information, consistent city information, etc.
[0106] It should be noted that for sample data with a similarity equal to the preset fifth similarity threshold, the computing device 600 may or may not cluster them. Those skilled in the art can choose according to actual needs, and this embodiment does not impose any restrictions on this.
[0107] Furthermore, the preset fifth similarity threshold can be determined based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations. Similarly, the preset fifth similarity threshold can be set by the computing device 600 based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations. Moreover, the principle of the computing device 600 setting the fifth similarity threshold can be found in the description of the computing device 600 setting the fourth similarity threshold, and will not be repeated here.
[0108] After obtaining at least one sample dataset, the computing device 600 can also determine the topic popularity value and topic name for each sample dataset. Since these sample datasets may come from different channels, and the range of popularity values for sample data varies greatly from channel to channel, the computing device 600 can normalize the popularity values of the sample data from different channels, unify the range of popularity values for these sample datasets, and then perform a weighted sum of the normalized popularity values of the sample data in each sample dataset to obtain the topic popularity value for each sample dataset.
[0109] Furthermore, the computing device 600 can also determine the topic name for the current sample data set. For example, the computing device 600 can determine the event name of the sample data corresponding to the highest normalized popularity value in the current sample data set as the topic name of the current sample data set. Another example is that the computing device 600 can determine the event name of the sample data with the most extracted sample keywords in the current sample data set as the topic name of the current sample data set. Yet another example is that the computing device 600 can count the occurrences of the same event name in the current sample data set and determine the event name with the highest occurrence count as the topic name of the current sample data set. Here, when the similarity between the event names of different sample data is greater than a preset sixth similarity threshold, the two event names can be identified as the same event name.
[0110] It should be noted that, for cases where the similarity is equal to the preset sixth similarity threshold, the computing device 600 may or may not recognize the event names of the two events as the same event name. Those skilled in the art can choose according to actual needs, and this embodiment does not impose any restrictions on this.
[0111] Furthermore, the preset sixth similarity threshold can be determined based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations. Similarly, the preset sixth similarity threshold can be set by the computing device 600 based on requirements, historical records, experiments, etc., and this embodiment does not impose any limitations. Moreover, the principle of the computing device 600 setting the sixth similarity threshold can be found in the description of the computing device 600 setting the fourth similarity threshold, and will not be repeated here.
[0112] Furthermore, the computing device 600 can output at least one topic. For example, the computing device 600 can visualize the at least one topic. There are various ways to visualize the topic, such as the computing device 600 displaying the at least one topic on a monitor, or issuing prompts about the at least one topic through sound and light, etc.
[0113] In summary, the industry classification method P100 and system 001 provided in this specification, after acquiring multiple sample data to be classified and keyword libraries corresponding to multiple industry categories, match multiple sample data based on the keyword libraries. Sample data matching the keyword libraries is used as the first sample data, and sample data not matching the keyword libraries is used as the second sample data. Then, the industry category of the first sample data is determined based on the keyword libraries. The first sample data is then used to match the second sample data, and the industry category of the second sample data is determined based on the industry category of the first sample data that matches the second sample data. In this scheme, by using the already classified first sample data to classify the second sample data by industry, sample data that has not yet been classified by industry is assigned to the industry category of sample data with high similarity. This improves the matching coverage of sample data, achieving the effect of classifying more sample data by industry, and thus enabling the timely and accurate discovery of more trending events.
[0114] This specification, in another aspect, provides a non-transitory storage medium storing at least one set of executable instructions for performing industry classification. When the executable instructions are executed by a processor, they instruct the processor to implement the steps of the industry classification method P100 described in this specification. In some possible embodiments, various aspects of this specification can also be implemented as a program product comprising program code. When the program product is run on a computing device 600, the program code causes the computing device 600 to perform the steps of the industry classification method P100 described in this specification. The program product for implementing the above method may employ a portable compact disk read-only memory (CD-ROM) containing program code and may run on the computing device 600. However, the program product of this specification is not limited thereto. In this specification, a readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system. The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof. Program code for performing the operations described herein can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on computing device 600, partially on computing device 600, as a standalone software package, partially on computing device 600 and partially on a remote computing device, or entirely on a remote computing device.
[0115] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0116] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure is presented by way of example only and is not restrictive. Although not explicitly stated herein, those skilled in the art will understand that this specification requires various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this specification and are within the spirit and scope of the exemplary embodiments described herein.
[0117] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.
[0118] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and aiding in the understanding of a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art may readily identify some of the devices as separate embodiments when reading this specification. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. It is also valid when each secondary embodiment contains fewer than all the features of a single foregoing disclosed embodiment.
[0119] Each patent, patent application, publication of the patent application, and other materials such as articles, books, specifications, publications, documents, articles, etc., cited herein may be incorporated by reference. All contents used for all purposes, except for any history of prosecution documents relating to it, that may be inconsistent with or conflict with this document, or any such history of prosecution documents that may have a limiting effect on the widest extent of the claims, are now or hereafter associated with this document. For example, in the event of any inconsistency or conflict between the description, definition, and / or use of terms associated with any of the included materials and the terms, description, definition, and / or used in connection with this document, the terms used herein shall prevail.
[0120] Finally, it should be understood that the embodiments disclosed herein are illustrative of the principles of the embodiments described in this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can implement the applications described in this specification using alternative configurations based on the embodiments in this specification. Therefore, the embodiments in this specification are not limited to the embodiments precisely described in the applications.
Claims
1. An industry classification method for classifying trending events by industry, including: Acquire multiple sample data to be classified, wherein the multiple sample data includes information on at least one hotspot event; Obtain keyword libraries corresponding to multiple industry categories; The multiple sample data includes first sample data that can be successfully matched with the keyword database and second sample data that cannot be successfully matched with the keyword database; The multiple sample data are matched with the keyword database to determine at least one industry category of the first sample data; as well as The second sample data is matched with the first sample data, and the industry category of the corresponding second sample data is determined based on at least one industry category of the successfully matched first sample data. Then, for each of the at least one industry category: Cluster the sample data in the current industry category based on the topic similarity between the sample data in the current industry category to obtain at least one sample data set; For at least a portion of the sample data sets in the at least one sample data set, each sample data set in the at least a portion of the sample data sets is sequentially taken as the current sample data set. The popularity values of the sample data from different channels in the current sample data set are normalized. Based on the weighted sum of the normalized popularity values, the popularity value of the topic corresponding to the current sample data set is obtained. The event name of the sample data with the highest normalized popularity value in the current sample data set is determined as the name of the topic. as well as Output the popularity value and name of the topic corresponding to at least a portion of the sample data set.
2. The method according to claim 1, wherein, The step of matching the second sample data with the first sample data and determining the industry category of the corresponding second sample data based on at least one industry category of the successfully matched first sample data includes, for each second sample data: Determine the pairwise similarity between the current second sample data and the first sample data; as well as The industry category of the current second sample data is defined as at least one industry category of the first sample data whose similarity is greater than a preset first similarity threshold or whose similarity is the largest.
3. The method according to claim 1, wherein, The step of matching the plurality of sample data with the keyword database to determine at least one industry category of the first sample data includes matching each sample data in the plurality of sample data: Extract sample keywords from the current sample data; Determine the similarity between the sample keywords and multiple keywords in the keyword library; as well as When the current sample data is determined to match the keyword database based on the similarity, at least one industry category is determined based on the industry category corresponding to at least one keyword that matches the current sample data; or, when the current sample data is determined not to match the keyword database based on the similarity, the current sample data is used as the second sample data.
4. The method according to claim 3, wherein, The process of determining the corresponding industry category based on at least one industry category corresponding to at least one keyword matching the current sample data includes: The industry category corresponding to at least one keyword with a similarity greater than a preset second similarity threshold, or with the highest similarity greater than the preset second similarity threshold, is determined as the initial industry category of the current sample data; and The industry category of the current sample data is determined based on the initial industry category of the current sample data.
5. The method according to claim 4, wherein, The step of determining at least one industry category of the current sample data based on at least one initial industry category of the current sample data includes: Obtain at least one historical sample data of an industry category that was incorrectly classified and has been corrected, along with its corresponding corrected industry category; Determine the similarity between the current sample data and the at least one historical sample data; and Based on the similarity between the current sample data and the at least one historical sample data, and at least one initial industry category of the current sample data, at least one industry category of the current sample data is obtained.
6. The method according to claim 5, wherein, The process of obtaining the industry category of the current sample data based on the similarity between the current sample data and the at least one historical sample data, and the initial industry category of the current sample data, includes: When the maximum similarity among the similarities between the current sample data and the at least one historical sample data is greater than a preset third similarity threshold, the corrected industry category of the historical sample data corresponding to the maximum similarity is used as the industry category of the current sample data; or... When the maximum similarity among the similarities between the current sample data and the at least one historical sample data is less than a preset third similarity threshold, the initial industry category is taken as the industry category of the current sample data.
7. The method according to claim 1, wherein, The acquisition of keyword libraries corresponding to multiple industry categories includes: Obtain the preset original keyword library; Determine the expanded keywords corresponding to the multiple industry categories, wherein the expanded keywords are obtained based on the feature descriptions of their corresponding industry categories; and The original keyword library is expanded based on the expanded keywords to obtain the keyword library.
8. The method according to claim 7, wherein, The process of expanding the original keyword library based on the expanded keywords to obtain the keyword library includes: Determine the similarity between the expanded keywords and multiple original keywords in the original keyword library; and Expanded keywords with a similarity greater than a preset fourth similarity threshold are added to the industry category corresponding to their original keywords to obtain the keyword library.
9. The method according to claim 1, wherein, The topic similarity includes the string similarity and geographical similarity of the sample data.
10. An industry classification system, comprising: At least one storage medium storing at least one instruction set for industry classification; as well as At least one processor is communicatively connected to the at least one storage medium. When the industry classification system is running, the at least one processor reads the at least one instruction set and executes the industry classification method according to any one of claims 1-9 according to the instructions of the at least one instruction set.
Citation Information
Patent Citations
Policy key information extraction method and device, storage medium and electronic equipment
CN112035653A
E-commerce commodity classification method and system based on hierarchical combination model
CN112463971A
Industrial category determination method and device, storage medium and electronic equipment
CN114297347A