Enterprise product word sorting method and device, storage medium and electronic equipment
By acquiring a set of enterprise product terms, determining product term features, and constructing comparative features, and using a product term importance differentiation model, the problem of difficulty in distinguishing the importance of enterprise product terms is solved, achieving accurate ranking of enterprise product terms and improving the accuracy of enterprise competitiveness analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QIZHI TECH CO LTD
- Filing Date
- 2023-07-14
- Publication Date
- 2026-04-17
AI Technical Summary
The importance of product keywords for businesses is difficult to accurately distinguish, making it impossible to accurately rank the importance of each product keyword.
By acquiring a set of enterprise product terms, determining product term features, constructing comparative features, and inputting them into a trained product term importance differentiation model, the importance of enterprise product terms is analyzed to achieve accurate ranking.
It enables accurate ranking of the importance of enterprise product keywords, improving the accuracy of enterprise industry layout and competitiveness analysis.
Smart Images

Figure CN116933778B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technology, specifically to a method, apparatus, storage medium, and electronic device for sorting enterprise product terms. Background Technology
[0002] Enterprise product keywords are the key terms corresponding to a company's current business operations. By analyzing enterprise product keywords, one can gain a better understanding of the company's business information. In the course of business operations, analyzing enterprise product keywords helps determine a company's main business, thereby understanding the main business development of its competitors in the same industry, analyzing its own competitiveness, and focusing on improving its industrial layout and key areas to enhance its core competitiveness. This is of great significance to the healthy development of a company. Therefore, enterprise product keywords help in understanding the main businesses of other companies and provide a basis for analyzing one's own competitiveness.
[0003] However, since a company's business operations are often not singular, there are multiple product keywords for the company. It is difficult to distinguish the importance of these product keywords, which makes it impossible to accurately rank the importance of each product keyword. Summary of the Invention
[0004] In order to distinguish the importance of enterprise product terms and accurately rank the importance of each enterprise product term, this application provides an enterprise product term ranking method, apparatus, storage medium, and electronic device.
[0005] The first aspect of this application provides a method for ranking enterprise product terms, specifically including:
[0006] Obtain the set of enterprise product terms for the target enterprise, wherein the set of enterprise product terms includes at least one enterprise product term;
[0007] Determine the product term features corresponding to each of the enterprise product terms in the enterprise product term set, and use the product term features to evaluate the importance of the enterprise product terms;
[0008] Based on the product word characteristics of each of the enterprise product words, determine the comparison features between each of the enterprise product words in the enterprise product word set and the remaining enterprise product words;
[0009] The comparative features are input into the trained product word importance discrimination model to determine the importance priority of each enterprise product word, and the enterprise product words are sorted according to the importance priority.
[0010] By adopting the above technical solution, since product term features can evaluate the importance of enterprise product terms, the product term features of each enterprise product term are determined sequentially. Then, based on the product term features of each enterprise product term and the product term features of each remaining enterprise product term in the enterprise product term set, a comparison feature is constructed between each enterprise product term and each remaining enterprise product term, so that each enterprise product term corresponds to multiple comparison features. Finally, the comparison features corresponding to each enterprise product term are sequentially input into the trained product term importance discrimination model. The product term importance discrimination model analyzes the importance between each enterprise product term and the remaining enterprise product terms based on the comparison features, and then determines the importance priority of each enterprise product term, thereby enabling accurate ranking of each enterprise product term according to its importance priority.
[0011] Optionally, obtain relevant information sources corresponding to the target company;
[0012] By using a pre-defined NER model, enterprise product terms are extracted from the enterprise introduction text in each of the relevant information sources to obtain a set of enterprise product terms.
[0013] By adopting the above technical solution, relevant information sources containing information about the target company are identified. Then, the NER model is used to extract company product terms from the company introduction texts containing information about the target company in the relevant information sources, thus obtaining a set of company product terms. This enables accurate extraction of company product terms and facilitates subsequent importance ranking.
[0014] Optionally, a target product term may be selected one by one from the set of enterprise product terms;
[0015] Based on Tencent word vectors, the cos similarity between the target product word in the enterprise product word set and other enterprise product words is calculated, and each cos similarity corresponds to one other enterprise product word;
[0016] Count the frequency of each of the other companies' product terms in all the relevant information sources.
[0017] The frequency of each of the other enterprise product terms is used as a weight, and a weighted average is calculated with the corresponding cosine similarity to obtain the product term features corresponding to the target product term, and / or
[0018] Based on Tencent word vectors, the cosine similarity between each enterprise product word in the enterprise product word set and other enterprise product words is calculated.
[0019] Calculate the average value of each cos similarity, and determine the average value as the product word feature of the corresponding enterprise product word.
[0020] By employing the above technical solution, after calculating the cosine similarity between each other company's product term and the target product term, the frequency of each other company's product term in all relevant information sources is then counted. Finally, the cosine similarity and the corresponding word frequency for each other company's product term are multiplied and averaged to obtain the product term features corresponding to the target product term, thereby determining the product term features of each company's product term. Alternatively, the average cosine similarity between other company's product terms can be used to characterize the similarity between product terms, and this can be determined as the product term feature of the company's product term. This allows for the reasonable determination of indicators for evaluating the importance of each company's product term.
[0021] Optionally, determining the product term features corresponding to each enterprise product term in the enterprise product term set specifically includes:
[0022] Select one target product term from the set of enterprise product terms one by one;
[0023] Filter related product terms from the enterprise product term set that are associated with the target product term;
[0024] The first number of related product words that are consistent with the relevant information sources corresponding to the target product word is counted, and the first number is determined as the product word feature corresponding to the target product word.
[0025] By adopting the above technical solution, after screening the related product terms that are associated with the target product term, the number of related product terms that are consistent with the relevant information sources of the target product term is counted, i.e., the first number. The larger the first number, the more words are associated with the target product term and have the same source, which indicates that the target product term is more important. Therefore, determining the first number as the product term feature of the target product term can more accurately evaluate the importance of each company's product term.
[0026] Optionally, determining the comparison features between each enterprise product term and the remaining enterprise product terms in the enterprise product term set based on the product term features of each enterprise product term specifically includes:
[0027] By combining the product terms of each enterprise in pairs, the target combination is obtained;
[0028] The product word features of the same dimension of the two enterprise product words in each target combination are respectively used as the minuend and the subtrahend. The minuend and the subtrahend are subtracted to determine the comparison features between each enterprise product word in the enterprise product word set and the remaining enterprise product words.
[0029] The step of inputting each of the comparative features into the trained product word importance discrimination model to determine the importance priority of each of the enterprise product words, and ranking each of the enterprise product words according to the importance priority, specifically includes:
[0030] Each of the aforementioned comparative features is input into the trained product word importance distinction model. If the output result is 1, it is determined that the importance priority of the enterprise product word corresponding to the minuend is higher than the importance priority of the enterprise product word corresponding to the subtrahend.
[0031] If the output result is 0, it is determined that the importance priority of the enterprise product word corresponding to the minuend is lower than the importance priority of the enterprise product word corresponding to the subtrahend.
[0032] Based on their respective importance priorities, the product terms of each enterprise are sorted, with higher importance priorities resulting in the product terms appearing earlier in the list.
[0033] By adopting the above technical solution, product terms of various enterprises are combined in pairs to obtain multiple target combinations. Then, the product term features of the two enterprise product terms in each target combination are subtracted to obtain multiple comparison features. These comparison features are input into the trained product term importance discrimination model. If the output result is 1, the importance priority of the enterprise product term corresponding to the minuend is higher than that of the product term corresponding to the subtrahend. Conversely, if the output result is 0, it means that the importance priority of the enterprise product term corresponding to the minuend is lower than that of the product term corresponding to the subtrahend. This distinguishes the importance between pairs of enterprise product terms and ultimately accurately sorts each enterprise product term according to its importance.
[0034] Optionally, determining the product term features corresponding to each enterprise product term in the enterprise product term set specifically includes:
[0035] Statistically analyze the frequency of each enterprise product term in the enterprise product term set across all relevant information sources;
[0036] If there are identical frequencies of occurrence among the aforementioned frequencies, then the target information source containing all the aforementioned enterprise product terms is selected from all the aforementioned relevant information sources;
[0037] The second number of times each of the enterprise product terms appears in the enterprise introduction text in the target information source and the number of times each of the enterprise product terms appears in the target information source are counted.
[0038] The product term features of each enterprise product term are obtained by weighted summing of the second quantity and the frequency of occurrence corresponding to each enterprise product term.
[0039] By adopting the above technical solution, if the frequency of occurrence of each company's product terms is the same across all relevant information sources, frequency will not be used as a product term feature to avoid affecting the evaluation of the product term's importance. Next, a target information source containing all company product terms is identified from all relevant information sources; this is a more authoritative and reliable information source. Finally, the second quantity and frequency of occurrence of each company's product term in the company introduction text of the target information source are weighted and summed to obtain the product term feature for each company's product term, thus providing a more reasonable evaluation of the importance of the company's product terms.
[0040] Optionally, determining the product term features corresponding to each enterprise product term in the enterprise product term set specifically includes:
[0041] Select one target product term from the set of enterprise product terms one by one;
[0042] Count the first number of the product words in the enterprise product word set that have a hierarchical relationship with the target product word;
[0043] Determine the second number of the product words in each of the preceding and following product words that are in the same product field as the target product word, and calculate the ratio of the number of the second number to the number of the first number. Use the ratio of the number of the first number to determine the product word feature of the target product word.
[0044] By adopting the above technical solution, the more product terms that have a hierarchical relationship with the target product term, the stronger the representativeness of the target product term. Then, the second number of product terms that are in the same product field as the target product term is counted, and the ratio of the number of the second number to the number of the first number is calculated. The larger the ratio, the more enterprise product terms are related to the target product term and are in the same product field. This further shows that the target product term is highly representative in the product field. Identifying this as a product term feature can more accurately evaluate the importance of enterprise product terms.
[0045] A second aspect of this application provides a product keyword sorting device, specifically comprising:
[0046] The product keyword acquisition module is used to acquire the set of enterprise product keywords of the target enterprise, wherein the set of enterprise product keywords includes at least one enterprise product keyword;
[0047] The feature determination module is used to determine the product word features corresponding to each of the enterprise product words in the enterprise product word set;
[0048] The feature comparison module is used to determine the comparison features between each of the enterprise product words in the enterprise product word set and the remaining enterprise product words based on the product word features of each of the enterprise product words;
[0049] The product term ranking module is used to input the comparative features into the trained product term importance discrimination model, determine the importance priority of each enterprise product term, and rank each enterprise product term according to the importance priority.
[0050] By adopting the above technical solution, after the product term acquisition module obtains the set of enterprise product terms, the feature determination module determines the product term features corresponding to each enterprise product term. Then, the feature comparison module determines the comparison features between each enterprise product term and the remaining enterprise product terms. Finally, the product term ranking module inputs each comparison feature into the product term importance differentiation model to distinguish the importance of each enterprise product term, and ultimately achieves accurate ranking of enterprise product terms.
[0051] In summary, this application includes at least one of the following beneficial technical effects:
[0052] The product term features of each enterprise's product term are determined. Then, based on the product term features of each enterprise's product term and the product term features of each remaining enterprise's product term in the enterprise's product term set, a comparison feature is constructed between each enterprise's product term and each remaining enterprise's product term, so that each enterprise's product term corresponds to multiple comparison features. Finally, the comparison features corresponding to each enterprise's product term are sequentially input into the trained product term importance discrimination model. The product term importance discrimination model analyzes the importance between each enterprise's product term and the remaining enterprise's product terms based on the comparison features, and then determines the importance priority of each enterprise's product term, so as to accurately rank each enterprise's product term according to its importance priority. Attached Figure Description
[0053] Figure 1 This is a schematic diagram of the architecture of an enterprise product keyword ranking system provided in an embodiment of this application;
[0054] Figure 2 This is a flowchart illustrating a method for sorting enterprise product terms according to an embodiment of this application;
[0055] Figure 3 This is a schematic diagram of a method for determining product term features provided in an embodiment of this application;
[0056] Figure 4 This is a flowchart illustrating another method for sorting enterprise product terms provided in an embodiment of this application;
[0057] Figure 5 This is a schematic diagram of the structure of a product keyword sorting device provided in an embodiment of this application;
[0058] Figure labeling: 11. Product term acquisition module; 12. Feature determination module; 13. Feature comparison module; 14. Product term sorting module. Detailed Implementation
[0059] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0060] In the description of the embodiments in this application, words such as "illustrative," "for example," or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "illustrative," "for example," or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "illustrative," "for example," or "for example" is intended to present the relevant concepts in a specific manner.
[0061] See Figure 1 This application discloses an architectural diagram of an enterprise product keyword ranking system, specifically including: a terminal and a server. The terminal can be directly or indirectly connected to the server via wired or wireless network. The terminal can be an electronic device such as a mobile phone, tablet computer, e-book reader, multimedia playback device, wearable device, or PC (Personal Computer). The server 20 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers.
[0062] Specifically, the user inputs the name of the target company through the terminal, and the terminal uploads the name of the target company to the server. The server then determines the corresponding product terms of the target company, mines the product term features of each product term, and finally constructs comparative features between product terms of each company based on the product term features. These features are then input into the model to obtain the importance priority of each product term, and the importance of each product term is ranked according to the importance priority.
[0063] Additionally, it should be noted that the execution entity of the enterprise product word ranking method disclosed in this application is a server.
[0064] See Figure 2 This application discloses a flowchart illustrating a method for ranking enterprise product terms, which can be implemented using a computer program or run on an enterprise product term ranking device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application, specifically including:
[0065] S101: Obtain the set of product keywords for the target company.
[0066] In one feasible implementation, relevant information sources corresponding to the target enterprise are obtained;
[0067] By using a pre-built NER model, enterprise product terms are extracted from the enterprise introduction text in various relevant information sources to obtain a set of enterprise product terms.
[0068] Specifically, the enterprise product term set includes at least one enterprise product term, which is a keyword reflecting the target enterprise's business operations or products. The enterprise product term set can reflect all of the target enterprise's business operations or products. In this embodiment, multiple enterprise product terms reflect a single business operation or product of the target enterprise; in other embodiments, a single enterprise product term reflects a single business operation or product of the target enterprise.
[0069] The target company's name is input into a business query platform, which includes, but is not limited to, search engines, online platforms like Shunqi.com, and brand websites. If the platform retrieves the target company's description text, it is identified as a relevant information source. The description texts from each source are then imported into a pre-built NER model to extract a set of product terms representing the target company's products. The Named Entity Recognition (NER) model is used to identify named entities, such as names of people, places, and organizations, from the text. Named Entity Recognition is a fundamental task in Natural Language Processing (NLP), essentially a sequence labeling task where each character in a text is labeled. NLP refers to the technology of processing natural language text using computer algorithms. Specifically, the extraction process involves using the NER model to extract industry entities, domain entities, and relationship categories from the company description text, ultimately outputting the product terms. In other embodiments, a BiLSTM-CRF model can also be used to extract the target company's product terms.
[0070] S102: Determine the product term features corresponding to each enterprise's product term in the enterprise product term set.
[0071] In one feasible implementation, a target product term is selected one by one from the enterprise product term set;
[0072] Based on Tencent word vectors, the cos similarity between the target product word and other enterprise product words in the enterprise product word set is calculated, and each cos similarity corresponds to one other enterprise product word;
[0073] Count the frequency of product terms for each other company across all relevant information sources;
[0074] The frequency of product terms from other companies is used as a weight, and a weighted average is calculated with the corresponding cosine similarity to obtain the product term features corresponding to the target product term, and / or
[0075] Based on Tencent word vector calculation, the cosine similarity between each enterprise product word in the enterprise product word set and other enterprise product words is calculated.
[0076] Calculate the average value of each cosine similarity and determine the average value as the product word feature of the corresponding enterprise product word.
[0077] Specifically, product term features are feature information that can evaluate the importance of enterprise product terms. A single enterprise product term can have one product term feature; in other embodiments, a single enterprise product term can have multiple product term features. Tencent Word Vectors is a tool developed by Tencent AI Lab for Chinese word vector research. It provides pre-trained word vectors (200-dimensional) of 8 million Chinese words and can be applied to many downstream NLP tasks. Essentially, it is a pre-trained model that maps input Chinese words to a vector space. In this embodiment, enterprise product terms can be represented as vectors.
[0078] Cosine similarity, also known as cosine similarity, is a mathematical method used to evaluate the similarity between texts. Each text is represented as a vector, and the cosine of the angle between them is calculated.
[0079] One target product term is selected from the set of enterprise product terms one by one. This target product term is represented as a vector using Tencent Word Vectors. Simultaneously, all other enterprise product terms, excluding the target product term, are also represented as vectors using Tencent Word Vectors. The cosine similarity between the target product term and each of the other enterprise product terms is calculated using a pre-defined cosine similarity formula. That is, the cosine of the angle between the vector corresponding to the target product term and the vector corresponding to each of the other enterprise product terms. It should be noted that each other enterprise product term corresponds to one cosine similarity. In other embodiments, the target product term and other enterprise product terms can also be represented as vectors using a vector space model.
[0080] The formula for calculating cosine similarity is: cos(θ)=a*b / (||a||*||b||), where a represents the vector corresponding to the target product term, b represents the vector corresponding to the product terms of other companies, ||a|| and ||b|| are the magnitudes of vectors a and b, respectively, and cos(θ) represents the cosine similarity.
[0081] Next, the frequency of each other company product term in the set of company product terms, excluding the target product term, is counted across all relevant information sources; that is, the total number of times it appears across all relevant information sources. One feasible method is to use each other company product term as a search keyword, iterate through all relevant information sources, and count each occurrence to obtain the frequency of each other company product term. In other embodiments, the frequency of each other company product term can also be counted within the same relevant information source.
[0082] Finally, the frequency of each other company's product term is used as a weight, and the weighted average is calculated with the cos similarity of the other company's product term to obtain the product term feature corresponding to the target product term. This product term feature can be understood as the sum of the similarity frequency weights of the target product term. The larger the value of the product term feature, the more important the target product term is.
[0083] For example, such as Figure 3 As shown, there are four product terms A, B, C, and D. To determine the product term features of A, the cosine similarities between A and B, A and C, and A and D are calculated as cosine similarity 1, cosine similarity 2, and cosine similarity 3, respectively. Next, the word frequencies corresponding to B, C, and D are counted as P1, P2, and P3, respectively. Finally, the following calculation is performed: (P1 * cosine similarity 1 + P2 * cosine similarity 2 + P3 * cosine similarity 3) / 3, yielding the product term features of A. Similarly, to determine the product term features of B, the cosine similarities between B and A, B and C, and B and D are calculated. Then, the word frequencies corresponding to A, C, and D are counted, and finally, a weighted average is taken.
[0084] It should be noted that after the product term features of the target product term are determined, the next enterprise product term is selected from the enterprise product term set as the target product term to determine the same product term features, until all enterprise product terms in the enterprise product term set have been determined. Furthermore, all enterprise product terms in the enterprise product term set can have a single identical product term feature determined, or multiple product term features can be determined.
[0085] And / or, according to the same cos similarity calculation formula mentioned above, calculate the cos similarity between each enterprise product word in the enterprise product word set and other enterprise product words. Then calculate the average value between each cos similarity value. The average value is determined as the product word feature of the corresponding enterprise product word. This product word feature can be understood as the similarity between product words. The larger the average value, the higher the similarity between product words, and the more important the corresponding enterprise product word.
[0086] S103: Based on the product term features of each enterprise's product terms, determine the comparative features between each enterprise's product term and the remaining enterprise product terms in the enterprise product term set.
[0087] In one feasible implementation, the product terms of each enterprise are combined in pairs to obtain the target combination;
[0088] The product word features of the same dimension of the two enterprise product words in each target combination are taken as the minuend and the subtrahend, respectively. The minuend and the subtrahend are subtracted to determine the comparative features between each enterprise product word in the enterprise product word set and the remaining enterprise product words.
[0089] Specifically, after the product term features of each enterprise's product terms are determined, the product terms are paired to obtain multiple target combinations. One feasible implementation method is to use the combine function to perform pairwise combinations. For example, if there are four enterprise product terms A, B, C, and D in the set of enterprise product terms, the target combinations obtained by pairwise combinations are as follows: AB, AC, AD, BC, CD.
[0090] Next, the product word features of the same dimension corresponding to the two enterprise product words in each target combination are subtracted. That is, the product word feature of one enterprise product word is used as the minuend, and the product word feature of the other enterprise product word is used as the subtrahend, and the two are subtracted. Finally, the comparison features between each enterprise product word and the remaining enterprise product words in the set of enterprise product words are obtained. It should be noted that in this embodiment, the enterprise product words are combined in pairs without repetition; in other embodiments, they can also be combined in pairs with repetition.
[0091] For example, product feature of company product term A is 'a', product feature of company product term B is 'b', product feature of company product term C is 'c', and product feature of company product term D is 'd'. Then the comparison features of company product term A with the remaining company product terms are: ab, ac, ad. In other embodiments, it can also be: the comparison features of company product term A with the remaining company product terms are: ba, ca, da.
[0092] S104: Input each comparative feature into the trained product word importance discrimination model, determine the importance priority of each company's product words, and rank each company's product words according to each importance priority.
[0093] In one feasible implementation, each comparative feature is input into the trained product word importance discrimination model. If the output result is 1, it is determined that the importance priority of the enterprise product word corresponding to the minuend is higher than the importance priority of the enterprise product word corresponding to the subtrahend.
[0094] If the output result is 0, it is determined that the importance priority of the enterprise product term corresponding to the minuend is lower than the importance priority of the enterprise product term corresponding to the subtrahend.
[0095] The product keywords of each enterprise are sorted according to their importance and priority. The higher the importance and priority, the higher the corresponding product keyword is ranked.
[0096] Specifically, the trained product term importance discrimination model is a trained Gradient Boosting Decision Tree (GBDT) machine learning model. In other embodiments, it can also be a trained logistic regression model. Further, the process of obtaining the trained product term importance discrimination model can be summarized as follows: collect product terms from multiple companies, construct comparative features for each company's product terms according to steps S102-S103, and then input these features into the GBDT machine learning model for training. This allows the GBDT machine learning model to continuously learn how to distinguish the importance between pairs of company product terms.
[0097] GBDT (Boolean Decision Tree) is a binary classification machine learning model that performs regression or classification on data by continuously reducing the residuals generated during training. GBDT uses decision trees as its basic learner, with each new tree fitting the negative gradient direction of the current loss function. GBDT can handle various types of data and exhibits high accuracy and robustness. The comparative features corresponding to each company's product terms are input into the trained product term importance discrimination model for binary classification. In machine learning, binary classification means assigning the comparative features between two input company product terms to one of two categories.
[0098] If the output of the trained product word importance discrimination model is 1, it means that the importance priority of the enterprise product word corresponding to the product feature word of the minuend is higher than that of the enterprise product word corresponding to the minuend. For example, if the comparison feature ab of enterprise product word A and the remaining enterprise product word B is input into the model, and the output is 1, it means that enterprise product word A is more important than B.
[0099] If the output of the trained product term importance discrimination model is 0, it means that the importance priority of the product term corresponding to the minuend is lower than that of the product term corresponding to the subtrahend. For example, if the comparison feature ab between product term A and the remaining product term B is input into the model, and the output is 0, it means that product term B is more important than A.
[0100] Finally, using this method, the importance priority of all enterprise product terms is determined. That is, the importance of each pair of enterprise product terms is distinguished. Ultimately, based on the importance priority, the importance of each enterprise product term in the set can be accurately ranked. The higher the importance priority, the higher the ranking.
[0101] See Figure 4This application discloses a flowchart illustrating another method for ranking enterprise product terms, which can be implemented using a computer program or run on an enterprise product term ranking device based on the von Neumann architecture. This computer program can be integrated into an application or run as a standalone utility application, specifically including:
[0102] S201: Obtain relevant information sources corresponding to the target company.
[0103] S202: Extract enterprise product terms from the enterprise introduction text in various relevant information sources using a pre-built NER model to obtain a set of enterprise product terms.
[0104] For details, please refer to step S101, which will not be repeated here.
[0105] S203: Select one target product term from the enterprise's product term set one by one.
[0106] S204: Filter related product terms in the enterprise's product term set that are associated with the target product term.
[0107] Specifically, in this embodiment, the association relationship is a coexistence relationship, which means that the two enterprise product terms often appear in adjacent positions or appear frequently at the same time. In other embodiments, the association relationship can also be a hierarchical relationship, in which a specific word becomes a hierarchical word, and a more general keyword becomes a hypernym.
[0108] After selecting a target product term from the set of enterprise product terms, input the target product term along with every other enterprise product term in the set into a search engine. Count the number of pages in the search results where the target product term and the other enterprise product term co-occur. If the number of pages is greater than the preset value, it indicates that the two co-occur frequently and are related. In this case, the other enterprise product term is identified as a related product term. Continue in this manner to determine whether the next other enterprise product term is a related product term.
[0109] S205: Count the first number of related product words in each related product word that are consistent with the relevant information source corresponding to the target product word, and determine the first number as the product word feature corresponding to the target product word.
[0110] Specifically, after determining the related product terms of the target product term, the relevant information sources corresponding to each related product term are identified by traversing all relevant information sources. Then, each relevant information source corresponding to the target product term is compared one by one, and the first number of related product terms that match the relevant information sources corresponding to the target product term is selected and counted. This first number is determined as the product term feature corresponding to the target product term. The larger the first number, the more product terms that are related to and share the same source as the target product term, thus indicating that the target product term may be more important. This process is repeated, selecting the next target product term to determine the product term feature, until all enterprise product terms in the enterprise product term set have been selected, thereby determining the product term feature corresponding to each enterprise product term.
[0111] In one feasible implementation, the frequency of each enterprise's product terms in all relevant information sources is counted.
[0112] If there are identical occurrence frequencies among the various occurrence frequencies, then filter out the target information source containing all enterprise product terms from all relevant information sources;
[0113] The second number of times each company's product term appears in the company introduction text in the target information source, and the number of times each company's product term appears in the target information source;
[0114] The product term features of each enterprise's product term are obtained by weighted summing of the second quantity and frequency of occurrence corresponding to each enterprise's product term.
[0115] Specifically, the frequency of each company's product terms in all relevant information sources is counted. The specific statistical method can be found in step S102, and will not be repeated here. If the frequencies of occurrence for different company product terms are not the same, then the frequency can be used to determine the product term characteristics of each company's product terms. The higher the frequency, the more important the corresponding company product term.
[0116] If there are instances of identical frequency, to more accurately differentiate the importance of different company product terms, a target information source containing all company product terms is selected from all information sources; that is, a more authoritative and referential information source is chosen. Next, the second number of company introduction texts containing each company product term in the target information source is counted. A higher second number indicates a larger number of company introduction texts related to the corresponding company product term within the same information source. Simultaneously, the frequency of each company product term in the target information source is counted. Finally, the second number and frequency of occurrence for each company product term are weighted and summed. A larger weighted sum indicates a more important company product term, and this sum is used to determine the product term characteristic for each company product term. In other embodiments, if the number of second numbers less than a preset threshold is greater than a preset number, it indicates that many company product terms correspond to fewer company introduction texts; in this case, setting the weight of the frequency of occurrence greater than the weight of the second number is more reasonable.
[0117] In another feasible implementation, a target product term is selected one by one from the enterprise product term set;
[0118] Count the first item of each product term in the enterprise's product term set that has a hierarchical relationship with the target product term.
[0119] Determine the second number of product words in each product word that are in the same product field as the target product word, and calculate the ratio of the number of the second number to the number of the first number. Use this ratio as the product word feature of the target product word.
[0120] Specifically, hyponymy is a linguistic relationship referring to the connection between keywords with strong specificity and keywords with strong generality. One feasible method for obtaining hyponymy-hypernymy relationships between a set of enterprise product terms and a target product term is as follows: Using the K-Means clustering algorithm, other enterprise product terms in the set are assigned to the category containing the most similar cluster centers. Finally, the target product term is assigned to its corresponding category. This identifies the hyponymy-hypernymy relationships between the target product term and its corresponding product terms, and the first number of each relationship is counted. This is existing technology and will not be elaborated further. In other embodiments, hierarchical clustering algorithms can also be used. Clustering algorithms are unsupervised learning algorithms used to group similar data points together, thereby reducing the dimensionality of the data.
[0121] Next, using Natural Language Processing (NLP) techniques, the product domains of the target product term and its constituent product terms are determined. The second number of constituent product terms sharing the same product domain as the target product term is then counted. Finally, the second number is divided by the first number to obtain the ratio. A larger ratio indicates a greater number of constituent product terms sharing the same product domain as the target product term, signifying stronger universality and representativeness of the target product term within that product domain, and thus higher importance. Therefore, the ratio is determined as the product term feature of the target product term, where the product domain can be understood as the industry category to which the company's product term belongs. In other embodiments, the first number can also be directly determined as the product term feature of the target product term.
[0122] S206: Based on the product term characteristics of each enterprise's product terms, determine the comparative features between each enterprise's product term and the remaining enterprise product terms in the enterprise product term set.
[0123] S207: Input each comparative feature into the trained product word importance discrimination model, determine the importance priority of each company's product words, and rank each company's product words according to each importance priority.
[0124] For details, please refer to steps S103-S104, which will not be repeated here.
[0125] The implementation principle of the enterprise product term ranking method in this application embodiment is as follows: determine the product term features of each enterprise product term; then, based on the product term features of each enterprise product term and the product term features of each remaining enterprise product term in the enterprise product term set, construct comparison features between each enterprise product term and each remaining enterprise product term, so that each enterprise product term corresponds to multiple comparison features; finally, input the comparison features corresponding to each enterprise product term into the trained product term importance discrimination model in sequence; the product term importance discrimination model analyzes the importance between each enterprise product term and the remaining enterprise product terms based on the comparison features, and then determines the importance priority of each enterprise product term, thereby enabling accurate ranking of each enterprise product term according to the importance priority.
[0126] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0127] Please see Figure 5 This is a schematic diagram of the enterprise product term sorting device provided in this application embodiment. This enterprise product term sorting device can be implemented as all or part of a device through software, hardware, or a combination of both. The device 1 includes a product term acquisition module 11, a feature determination module 12, a feature comparison module 13, and a product term sorting module 14.
[0128] The product keyword acquisition module is used to acquire the set of enterprise product keywords of the target enterprise, and the set of enterprise product keywords includes at least one enterprise product keyword;
[0129] The feature determination module is used to determine the product word features corresponding to each enterprise's product word in the enterprise product word set;
[0130] The feature comparison module is used to determine the comparison features between each enterprise product word in the enterprise product word set and the remaining enterprise product words based on the product word features of each enterprise's product words;
[0131] The product term ranking module is used to input the various comparative features into the trained product term importance discrimination model, determine the importance priority of each company's product terms, and rank the product terms of each company according to the importance priority.
[0132] Optional, the product keyword acquisition module 11 specifically includes:
[0133] Obtain relevant information sources corresponding to the target company;
[0134] By using a pre-built NER model, enterprise product terms are extracted from the enterprise introduction text in various relevant information sources to obtain a set of enterprise product terms.
[0135] Optional, feature determination module 12, specifically includes:
[0136] Select one target product keyword from the enterprise's product keyword set one by one;
[0137] Based on Tencent word vectors, the cos similarity between the target product word and other enterprise product words in the enterprise product word set is calculated, and each cos similarity corresponds to one other enterprise product word;
[0138] Count the frequency of product terms for each other company across all relevant information sources;
[0139] The frequency of product terms from other companies is used as a weight, and a weighted average is calculated with the corresponding cosine similarity to obtain the product term features corresponding to the target product term, and / or
[0140] Based on Tencent word vector calculation, the cosine similarity between each enterprise product word in the enterprise product word set and other enterprise product words is calculated.
[0141] Calculate the average value of each cosine similarity and determine the average value as the product word feature of the corresponding enterprise product word.
[0142] Optionally, the feature determination module 12 is further used for:
[0143] Select one target product keyword from the enterprise's product keyword set one by one;
[0144] Filter related product terms from the enterprise's product terminology set that are associated with the target product term;
[0145] The first number of related product terms that are consistent with the relevant information sources corresponding to the target product term is counted, and the first number is determined as the product term feature corresponding to the target product term.
[0146] Optional, feature comparison module 13, specifically used for:
[0147] By pairing the product terms of each company, the target combination can be obtained;
[0148] The product word features of the same dimension of the two enterprise product words in each target combination are taken as the minuend and the subtrahend, respectively. The minuend and the subtrahend are subtracted to determine the comparative features between each enterprise product word in the enterprise product word set and the remaining enterprise product words.
[0149] Optional, product keyword sorting module 14, specifically used for:
[0150] Each comparative feature is input into the trained product word importance discrimination model. If the output result is 1, it is determined that the importance priority of the enterprise product word corresponding to the minuend is higher than that of the enterprise product word corresponding to the subtrahend.
[0151] If the output result is 0, it is determined that the importance priority of the enterprise product term corresponding to the minuend is lower than the importance priority of the enterprise product term corresponding to the subtrahend.
[0152] The product keywords of each enterprise are sorted according to their importance and priority. The higher the importance and priority, the higher the corresponding product keyword is ranked.
[0153] Optionally, the feature determination module 12 is further used for:
[0154] The frequency of each company's product terms in all relevant information sources was analyzed.
[0155] If there are identical occurrence frequencies among the various occurrence frequencies, then filter out the target information source containing all enterprise product terms from all relevant information sources;
[0156] The second number of times each company's product term appears in the company introduction text in the target information source, and the number of times each company's product term appears in the target information source;
[0157] The product term features of each enterprise's product term are obtained by weighted summing of the second quantity and frequency of occurrence corresponding to each enterprise's product term.
[0158] Optionally, the feature determination module 12 is further used for:
[0159] Select one target product keyword from the enterprise's product keyword set one by one;
[0160] Count the first item of each product term in the enterprise's product term set that has a hierarchical relationship with the target product term.
[0161] Determine the second number of product words in each product word that are in the same product field as the target product word, and calculate the ratio of the number of the second number to the number of the first number. Use this ratio as the product word feature of the target product word.
[0162] It should be noted that the enterprise product keyword sorting device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the enterprise product keyword sorting method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the enterprise product keyword sorting device and the enterprise product keyword sorting method embodiment provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiment, which will not be repeated here.
[0163] This application also discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it employs a method for sorting enterprise product terms as described above.
[0164] The computer program can be stored in a computer-readable medium. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or certain middleware. The computer-readable medium includes any entity or device capable of carrying computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the computer-readable medium includes, but is not limited to, the above-mentioned components.
[0165] The above-described enterprise product word sorting method is stored in the computer-readable storage medium and loaded and executed on the processor to facilitate the storage and application of the method.
[0166] This application also discloses an electronic device in which a computer program is stored in a computer-readable storage medium. When the computer program is loaded and executed by a processor, it employs the above-mentioned enterprise product word sorting method.
[0167] The electronic device can be a desktop computer, a laptop computer, or a cloud server, and the electronic device includes, but is not limited to, a processor and a memory. For example, the electronic device may also include input / output devices, network access devices, and buses.
[0168] The processor can be a central processing unit (CPU). Of course, depending on the actual use, it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc., and this application does not limit it.
[0169] The memory can be an internal storage unit of an electronic device, such as a hard disk or RAM, or an external storage device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD), or flash memory card (FC) equipped on the electronic device. Furthermore, the memory can be a combination of an internal storage unit and an external storage device. The memory is used to store computer programs and other programs and data required by the electronic device. The memory can also be used to temporarily store data that has been output or will be output. This application does not limit this.
[0170] In this electronic device, the enterprise product word sorting method of the above embodiment is stored in the memory of the electronic device and loaded and executed on the processor of the electronic device for convenient use.
[0171] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A method for ranking enterprise product terms, characterized in that, The method includes: S101. Obtain the set of enterprise product terms for the target enterprise. Step S101 includes: S201. Obtain relevant information sources corresponding to the target enterprise; S202. Extract enterprise product terms from the enterprise introduction text in each of the relevant information sources using a pre-set NER model to obtain a set of enterprise product terms. The set of enterprise product terms includes at least one enterprise product term, which is a keyword reflecting the business operations or products of the target enterprise. S102. Determine the product word features corresponding to each of the enterprise product words in the enterprise product word set. Step S102 includes: counting the frequency of each of the enterprise product words in the enterprise product word set in all the relevant information sources; if there are identical frequencies among the frequencies of occurrence, then select a target information source containing all the enterprise product words from all the relevant information sources; count the second number of enterprise introduction texts containing each enterprise product word in the target information source and the number of times each enterprise product word appears in the target information source; perform a weighted summation of the second number and the number of times each enterprise product word appears to obtain the product word features of each enterprise product word, wherein when the number of second numbers less than a preset threshold is greater than a preset number, then the weight of the number of times appears is set to be greater than the weight of the second number, and the product word features are used to evaluate the importance of enterprise product words; S103. Based on the product word features of each of the enterprise product words, determine the comparison features between each of the enterprise product words in the enterprise product word set and the remaining enterprise product words. Step S103 includes: combining each of the enterprise product words in pairs to obtain target combinations; taking the product word features of the same dimension of the two enterprise product words in each target combination as the minuend and the subtrahend respectively, and performing feature subtraction between the minuend and the subtrahend to determine the comparison features between each of the enterprise product words in the enterprise product word set and the remaining enterprise product words. S104. Input the aforementioned comparative features into the trained product word importance discrimination model to determine the importance priority of each of the aforementioned enterprise product words, and rank the aforementioned enterprise product words according to their importance priorities. Step S104 includes: Each of the aforementioned comparative features is input into the trained product term importance discrimination model. If the output result is 1, it is determined that the importance priority of the enterprise product term corresponding to the minuend is higher than that of the enterprise product term corresponding to the subtrahend. If the output result is 0, it is determined that the importance priority of the enterprise product term corresponding to the minuend is lower than that of the enterprise product term corresponding to the subtrahend. Based on each importance priority, each enterprise product term is sorted, with higher importance priority resulting in a higher ranking of the corresponding enterprise product term.
2. The enterprise product keyword sorting method according to claim 1, characterized in that, Step S102 includes: Select one target product term from the set of enterprise product terms one by one; Based on Tencent word vectors, the cos similarity between the target product word in the enterprise product word set and other enterprise product words is calculated, with each cos similarity corresponding to one other enterprise product word; Count the frequency of each of the other companies' product terms in all the relevant information sources. The frequency of each of the other enterprise product terms is used as a weight, and a weighted average is calculated with the corresponding cosine similarity to obtain the product term features corresponding to the target product term; and / or Based on Tencent word vectors, calculate the cosine similarity between each enterprise product word in the enterprise product word set and other enterprise product words; Calculate the average value of each cos similarity, and determine the average value as the product word feature of the corresponding enterprise product word.
3. The enterprise product keyword sorting method according to claim 1, characterized in that, Step S102 also includes: S203. Select one target product term from the set of enterprise product terms one by one; S204. Filter the related product terms in the enterprise product term set that are associated with the target product term; S205. Count the first number of related product words that are consistent with the relevant information sources corresponding to the target product word in each of the related product words, and determine the first number as the product word feature corresponding to the target product word.
4. The enterprise product keyword sorting method according to claim 1, characterized in that, Step S102 includes: Select one target product term from the set of enterprise product terms one by one; Count the first number of the product words in the enterprise product word set that have a hierarchical relationship with the target product word; Determine the second number of the product words in each of the preceding and following product words that are in the same product field as the target product word, and calculate the ratio of the number of the second number to the number of the first number. Use the ratio of the number of the first number to determine the product word feature of the target product word.
5. A device for sorting enterprise product terms, used to implement the enterprise product term sorting method as described in any one of claims 1 to 4, characterized in that, The enterprise product keyword sorting device includes: Product term acquisition module (11) is used to acquire the enterprise product term set of the target enterprise, wherein the enterprise product term set includes at least one enterprise product term; The feature determination module (12) is used to determine the product word features corresponding to each of the enterprise product words in the enterprise product word set; The feature comparison module (13) is used to determine the comparison features between each enterprise product word in the enterprise product word set and the remaining enterprise product words based on the product word features of each enterprise product word; The product word ranking module (14) is used to input the comparative features into the trained product word importance distinction model, determine the importance priority of each enterprise product word, and rank each enterprise product word according to the importance priority.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is loaded and executed by the processor, it employs the method described in any one of claims 1-4.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, When the processor loads and executes the computer program, it employs the method described in any one of claims 1-4.
Citation Information
Patent Citations
A method and apparatus for outputting information
CN109190123A