Entity competitor data mining method and system based on correlation between indexes
By acquiring entity similarity metrics and user review data, and combining principal component analysis, sentiment analysis, and expert evaluation, the most competitive core competitors are identified. This solves the problem of existing methods ignoring consumer needs and the correlation of metrics, and achieves more accurate competitor identification and business guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-03-31
AI Technical Summary
Existing competitor identification methods are mainly based on internal corporate data, ignoring consumer needs and failing to consider the correlation between various indicators, resulting in unsatisfactory identification results.
By acquiring entity similarity indicators, filling in missing values, using principal component analysis to remove correlations between indicators, combining user review data to extract feature words and sentiment analysis, constructing competitive indicators, using the TF-IDF algorithm and expert evaluation to determine weights, and finally using Choquet scores to generate a competitive index to identify core competitors.
It improves the accuracy and credibility of identifying core competitors, and can more comprehensively reflect consumers' subjective perceptions and objective entity characteristics, providing effective business guidance.
Smart Images

Figure CN121765409A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data mining technology, and in particular relates to a method and system for data mining of entity competitors based on the correlation between indicators. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Competitive analysis is a way for business managers to understand their market position. Through competitive analysis, they can not only clearly identify their own strengths and weaknesses, but also dynamically track the gap with industry benchmarks, providing data support for formulating differentiated development strategies. In the market landscape of 2025, the profitability of physical businesses is no longer a simple game of position; through cross-industry integration, it has evolved into a comprehensive space integrating public welfare communication, cultural experience, and community service, and the factors influencing its sustainable development are becoming increasingly diversified. Therefore, accurately identifying core competitors and systematically analyzing the gap between oneself and competitors has become a necessary prerequisite for physical businesses to gain a foothold in the market and achieve sustainable development.
[0004] Competitive analysis requires data from both the company itself and its competitors. Existing competitor identification methods mainly fall into three categories: those based on managerial perception, those based on industry structure, and those based on text mining. However, most of these methods rely on internal company data and neglect consumer needs. Therefore, some scholars have proposed identifying competitors through user reviews.
[0005] Although user reviews can more intuitively reflect consumer perceptions and represent more authentic customer needs, research on using text mining methods for competitor identification is still in its early stages. Specifically, it mainly suffers from the following shortcomings: Existing research on competitiveness analysis methods based on user reviews mostly considers only review data, neglecting other factors that reflect store competitiveness, resulting in a relatively simplistic data type. Furthermore, current technologies do not consider the correlations between various indicators during data fusion, ignoring their interrelationships, thus leading to less than ideal identification results. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, this invention provides a method and system for data mining of entity competitors based on the correlation between indicators. By considering the correlation between various indicators of an entity, the most competitive core competitors can be accurately identified, thereby providing effective guidance for the entity's subsequent operations.
[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a method for mining entity competitor data based on the correlation between indicators.
[0008] Data mining methods for entity competitors based on the correlation between indicators include: Obtain the similarity index of entities and fill in the missing values; Principal component analysis is used to remove the correlation between the various similarity indicators to obtain the natural attributes of the entities and their corresponding attribute values; based on the obtained natural attributes, cluster analysis is performed on the entities to generate a set of candidate competitors. We acquire user reviews of entities, extract feature words as customer concerns, and construct competitive indicators by combining customer concerns with natural attributes; we analyze customer concerns based on sentiment analysis methods and use the obtained sentiment scores as customer satisfaction. By combining the TF-IDF algorithm and expert evaluation, the weights of competitive indicators are determined; Choquet integrals are used to fuse the competitive indicators and their weights, generate and rank the competitive index of the entity, and identify the core competitors of the target entity.
[0009] Furthermore, the extraction of feature words includes: first, deleting invalid comments, deduplicating text, and removing special symbols from the user's comment data; then, performing word segmentation and part-of-speech tagging on the comment data to remove stop words; finally, using the LDA topic model to extract feature words, and determining the number of feature words based on the topic consistency score.
[0010] Furthermore, the customer satisfaction is represented by an emotion dictionary and a probabilistic language terminology set, specifically including: constructing an emotion dictionary, determining the emotion score corresponding to the comment data, and performing weighted calculations on customer concerns.
[0011] Furthermore, the missing value imputation is implemented based on the principle of collaborative filtering, which is to fill in the missing values by calculating the similarity between entities, and to measure the similarity based on cosine similarity and Pearson correlation coefficient.
[0012] Furthermore, the principal component analysis method includes: standardizing the original data into a standardized matrix and calculating the covariance matrix of the standardized matrix; extracting the eigenvalues and eigenvectors of the covariance matrix; selecting principal components based on the eigenvalues and constructing a dimensionality-reduced natural attribute matrix.
[0013] Furthermore, when performing cluster analysis on entities based on the obtained natural attributes, multiple clustering algorithms are compared, the clustering effect is evaluated by the Davidson-Bolding index, and the optimal clustering result is selected to generate a candidate competitor set; among them, the multiple clustering algorithms include hierarchical clustering, k-means clustering, and self-organizing map network clustering.
[0014] Furthermore, the determination of the competitive indicator weights includes: calculating the initial weights of customer concerns using the TF-IDF algorithm; generating expert weights based on expert evaluations of the competitive indicators; and integrating the initial weights and expert weights using a weighted average method to obtain the final weights of the competitive indicators. The generation of the expert weights includes constructing an expert evaluation matrix, where the matrix uses intuitive fuzzy numbers to represent evaluation information and combines expert opinions using an aggregation operator.
[0015] A second aspect of the present invention provides a data mining system for entity competitors based on the correlation between indicators.
[0016] A data mining system for entity competitors based on the correlation between indicators includes: The data preprocessing module is configured to: obtain the similarity index of entities and fill in missing values; The preliminary analysis module is configured to: use principal component analysis to remove the correlation between the various indicators in the similarity index to obtain the natural attributes of the entity and their corresponding attribute values; and perform cluster analysis on the entity based on the obtained natural attributes to generate a set of candidate competitors. The in-depth indicator analysis module is configured to: acquire user comment data on entities, extract feature words as customer concerns, and construct competitive indicators by combining customer concerns with natural attributes; analyze customer concerns based on sentiment analysis methods, and use the obtained sentiment score as customer satisfaction. The competitor identification module is configured to: combine the TF-IDF algorithm and expert evaluation to determine the weights of competitive indicators; use Choquet integrals to fuse competitive indicators and their weights to generate and rank the competitive index of entities in order to identify the core competitors of the target entity. A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps in the entity competitor data mining method based on the correlation between indicators as described in the first aspect of the present invention.
[0017] The fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the entity competitor data mining method based on the correlation between indicators as described in the first aspect of the present invention.
[0018] The above one or more technical solutions have the following beneficial effects: This invention first obtains similarity indicators for entities, then uses principal component analysis to remove correlations between the indicators. Next, it acquires user reviews of the entities, extracts feature words as customer focus points, and combines these focus points with natural attributes to construct competitive indicators. This not only reflects consumers' subjective consumption perceptions but also integrates "natural attributes" such as store environment, service attitude, and location to characterize objective entity features. By combining these two different but complementary types of indicators into competitive indicators, the limitations of a single data source are overcome, making the assessment of entity competitiveness more comprehensive and multi-dimensional, thus significantly improving the accuracy and reliability of core competitor identification results.
[0019] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0020] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0021] Figure 1 This is a flowchart of the entity competitor data mining method based on the correlation between indicators in Embodiment 1 of the present invention. Detailed Implementation
[0022] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0023] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0024] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0025] Example 1 This embodiment discloses a method for mining entity competitor data based on the correlation between indicators.
[0026] like Figure 1 As shown, the data mining method for entity competitors based on the correlation between indicators includes: Step S1: Obtain the similarity index of entities and fill in the missing values; Step S2: Use principal component analysis to remove the correlation between the various indicators in the similarity index to obtain the natural attributes of the entity and their corresponding attribute values; perform cluster analysis on the entity based on the obtained natural attributes to generate a set of candidate competitors. Step S3: Obtain user review data for entities, extract feature words as customer concerns, and construct competitive indicators by combining customer concerns with natural attributes; analyze customer concerns based on sentiment analysis methods, and use the obtained sentiment score as customer satisfaction. Step S4: Combine the TF-IDF algorithm and expert evaluation to determine the weights of competitive indicators; use Choquet integrals to fuse competitive indicators and their weights, generate and rank the competitive index of the entity to identify the core competitors of the target entity.
[0027] Based on the above process, this invention, by considering the correlations between various indicators of an entity, can accurately identify the most competitive core competitors, thereby providing effective guidance for the entity's subsequent operations. To facilitate understanding of the technical solution of this invention, the specific implementation methods of this invention will be further explained and described below.
[0028] In step S1, the similarity index of the entities is obtained and missing values are filled.
[0029] First, similarity indicators for entities are determined through research, and then the corresponding indicator values are collected. In actual implementation, these indicator values represent the actual data of the physical stores; for example, the indicator value for sales volume might be 100,000 yuan. The similarity indicators include multiple factors such as store environment, service attitude, location, repeat purchase products, and sales volume. While obtaining the similarity indicators for entities can, techniques such as web scraping can also be used; this embodiment does not impose specific limitations on this, as long as the required similarity indicators can be obtained.
[0030] In the implementation process, questionnaires can be distributed to experienced professionals in the target entity's city, asking them to select indicators that are highly relevant to physical sales and customer satisfaction, such as store environment, service attitude, traffic location, secondary sales products, and sales volume. Then, similarity indicators can be selected through data statistics. It is used for the initial identification of competitors.
[0031] Subsequently, missing values were imputed in the acquired data using similarity measurement methods. The similarity measurement method is crucial; since the similarity indicators in this invention are quantitative data, cosine similarity and Pearson correlation coefficient are used to measure the similarity between entities. ; ; in, and Representing stores and stores Cosine similarity and Pearson correlation coefficient between them; Representing stores and stores In indicators The parameter values below; Representing stores and stores The average parameter value across all indicators. Representing stores and stores The following is a set of indicators with parameter values. Indicates store and stores A set of common indicators.
[0032] ; ; ; in, Indicates store For missing indicators The predicted value, Representing stores and The nearest neighbor set of the indicator.
[0033] In step S2, principal component analysis is used to remove the correlation between the various similarity indicators to obtain the natural attributes of the entity. and corresponding attribute values Cluster analysis is performed on the entities based on the obtained natural attributes to generate a set of candidate competitors.
[0034] In this process, more clustering indicators are not necessarily better; as the number of indicators increases, the complexity of competitor identification also increases. Furthermore, clustering algorithms require independence between indicators; correlated indicators can negatively impact algorithm performance. Therefore, this invention aims to replace a larger number of indicators with relatively fewer indicators while minimizing information loss. Principal component analysis (PCA) is a statistical analysis method that seeks principal components to replace original variables based on their correlations. The resulting principal components are independent of each other and retain most of the information from the initial variables, ensuring minimal information loss during the transformation process. To make the identification results more accurate and reliable, this invention uses PCA to reduce the dimensionality of similarity indicators, obtaining the natural attributes of entities. and corresponding attribute values This includes: standardizing the original data into a standardized matrix and calculating the covariance matrix of the standardized matrix; extracting the eigenvalues and eigenvectors of the covariance matrix; selecting principal components based on the eigenvalues; and constructing a dimensionality-reduced natural attribute matrix. The specific process is as follows: First, the original similarity index data is transformed into OK Initial index matrix of columns For the initial matrix The index values are standardized. Considering that Z-Score standardization can ensure the comparability of data, Z-Score standardization is chosen to standardize the original matrix. ; in, The mean of the original data. The standard deviation of the original data; This represents the standardized index value. When... hour, That is, a positive standardization result is obtained; when hour, This results in a negative standardized result.
[0035] Subsequently, the covariance matrix is calculated. Then, calculate the eigenvalues and corresponding eigenvectors of the covariance matrix. Next, construct a matrix from the eigenvectors row-wise according to the magnitude of the eigenvalues. Take the first row... Rows form a matrix ,in The number of principal components.
[0036] Finally, the calculation reduces the original moments to dimensionality. The new matrix after dimension : ; in, .
[0037] In practice, SPSS software can be used to perform KMO values and Bartlett's test. If the conditions for principal component analysis are met, the number of principal components among various similarity indicators can be determined based on eigenvalues and variance contribution rates, serving as the natural attributes of the entity. Secondly, based on the indicators and the component matrix of the principal components, the attribute performance of each entity under the principal components can be calculated. For example, the first principal component The linear equation can be expressed as follows: ; in, This is a similarity indicator.
[0038] When performing clustering analysis on entities based on the obtained natural attributes, multiple clustering algorithms are compared, and the clustering effect is evaluated using the Davidson-Bolding index. The optimal clustering result is then selected to generate a candidate competitor set. These clustering algorithms include hierarchical clustering, k-means clustering, and self-organizing map network clustering. Then, by analyzing the scree plot, the optimal number of clusters for each method can be obtained. Based on this, the DBI of each clustering algorithm is calculated according to the following formula, and the optimal clustering method is selected for preliminary identification of competitors based on the magnitude of the clustering index DBI: ; in, For the first The cluster center of each cluster class For the first The average distance of each element in each cluster to the cluster center. Therefore, a smaller DBI indicates smaller intra-cluster distances and larger inter-cluster distances. Based on the clustering results of the algorithm corresponding to the optimal DBI, a competitor candidate set (CCs) is constructed, where... , This represents the number of candidate competitors.
[0039] In step S3, user reviews of entities are acquired, feature words are extracted as customer concerns, and these customer concerns, along with natural attributes, are used to construct a competitive indicator. Customer concerns are then analyzed using sentiment analysis methods, and the resulting sentiment score is used as customer satisfaction. This can be achieved through the following methods: Step S3-1: Set up a questionnaire, invite experienced customers of the physical store to evaluate the store, and extract characteristic words from user reviews. As a customer focus, and in conjunction with the natural attributes of the entity. Together they constitute competitive indicators .
[0040] When extracting feature words, firstly, invalid comments, text deduplication, and special character removal are performed on the user comment data. Then, word segmentation and part-of-speech tagging are performed on the comment data to remove stop words. Finally, the LDA topic model is used to extract feature words, and the number of feature words is determined based on topic consistency scores. Furthermore, to improve the accuracy of keyword extraction using the LDA method, the number of feature words is determined based on topic consistency coherence, resulting in the final customer focus. In this embodiment, Python tools are used to extract candidate feature words for entities using the LDA model, and hyperparameters are set. It is 0.1. The initial score is 0.01; then, the number of topics is determined using the topic consistency score; finally, the final customer focus is obtained based on the number of topics. .
[0041] ; in, Indicates the first Candidate feature words, Characteristic words in representative comments The probability of them occurring simultaneously Indicating characteristic words The probability of it appearing alone This represents the number of topic-candidate feature word pairs. A higher consistency score indicates better topic coherence. Finally, it links customer focus with the entity's natural attributes. Construct a set of competitive indicators, namely: .
[0042] Step S3-2: Identify customer focus points based on sentiment analysis methods Sentiment analysis was conducted, and the resulting sentiment score was used as a measure of customer satisfaction. .
[0043] Customer satisfaction is represented using a sentiment lexicon and a probabilistic linguistic term set (PLTS). Specifically, this involves: constructing a sentiment lexicon, determining the sentiment scores corresponding to the review data, and weighting the customer's concerns. First, a sentiment lexicon is constructed, and the sentiment score of each text is analyzed using Leximancer software. Then, the texts containing customer concerns are weighted to calculate the sentiment score for each concern, represented by a probabilistic linguistic term set (PLTS). For example, the sentiment score for the customer concern "service," which corresponds to a customer satisfaction score of [missing value], is [missing value]. .
[0044] In step S4, the weights of competitive indicators are determined by combining the TF-IDF algorithm and expert evaluation. The competitive indicators and their weights are then fused using Choquet integrals to generate and rank the entities' competitive indices, thereby identifying the target entity's core competitors. This can be achieved through the following methods: Step S4-1: Determine the weights of competitive indicators.
[0045] Calculate customer focus using the TF-IDF algorithm initial weights Based on expert evaluations of competitive indicators, expert weights are generated. The initial weights and expert weights are then integrated using a weighted average method to obtain the final weights of the competitive indicators. The generation of expert weights includes constructing an expert evaluation matrix, where the matrix uses intuitionistic fuzzy numbers to represent evaluation information, and combines expert opinions using an aggregation operator. Specifically, this can be achieved through the following methods: Using Python, a TF-IDF algorithm was developed to obtain the TF-IDF values of words related to customer concerns. For each customer concern... The TF-IDF values of related words are aggregated to obtain each TF-IDF value : ; In the above formula, A set of related words used to describe customer concerns. For the TF-IDF set of keywords related to the focus; where, . and These represent the maximum and minimum TF-IDF values among all customer focus words, respectively. The higher the TF-IDF of a relevant word, the higher its TF-IDF value for that focus.
[0046] To make the weights more convincing, the probability of each point of interest is obtained based on the LDA model, and combined with the TF-TDF value to form the influence weight of the point of interest under the frequency of customer mention (this influence weight is a set of intuitive language): ; in, Indicates the focus on usage To assess the degree of certainty of an entity's competitiveness. Indicates the focus on usage To assess the degree of negativity in competitiveness, Indicate focus The probability of occurrence.
[0047] Comments Every word in the comments The probability can be expressed as: ; in, Indicates the first One topic, The number of topics.
[0048] Invite experts to focus on the physical aspects and natural attributes Influence weight To evaluate, also using intuitive language sets Expert evaluation information is collected in a specific format. Based on this, a weighted average method is used to obtain the evaluation results for each focus point. And the final weight of natural attributes The expert decision-making process is as follows: 1) Construct the evaluation matrix. Let... Represents a set of competitive indicators, consisting of Experts Intuitive fuzzy numbers are used to provide... Competitive indicators Influence weight The expert's weight is... Then the expert evaluation matrix It can be represented as: ; in, , Experts think The role it plays in obtaining the competition index. The numbers represent, in order, very unimportant, unimportant, average, important, and very important. This indicates the degree of certainty of the decision-making information provided. This indicates the degree of negativity of the decision-making information provided. Furthermore, It can be represented as: .
[0049] 2) Gather expert evaluation information. Based on the expectations of the intuitive language set. and aggregation operators The opinions of experts were compiled and calculated. Influence weight .
[0050] ; ; At this point, the aggregated decision matrix for: ; Weighting customer focus in relation to frequency of customer mention Compared to the subset of the influence weight decision matrix under expert evaluation By pooling information, the ultimate impact weight of customer focus points can be obtained. This leads to the acquisition of each competitive indicator. The final set of influence weights .
[0051] Step S4-2: Use Choquet integrals to fuse multiple types of competitive indicators, namely the natural attributes of entities. and customer focus ; Calculate the competition index of the target entity and its candidate competitors, and obtain the core competitors of the target entity by comparing the competition index.
[0052] First, based on And the operational rules of the intuitive language set, calculating parameters The interaction weights between the values and competitive indicators : ; ; ; ; Subsequently, based on the interaction weights among the competitive indicators, the competition index for each entity is calculated. : ; ; in, Indicates a competitive indicator. ,and . is the scoring function for the intuitive language set.
[0053] Finally, by comparing the competition index of each entity in the candidate competitor set, and ranking them according to the competition index, the entities with higher rankings are the core competitors of the target entity. That is: first calculate the parameters... and the interaction weights between each pair of competitive indicators Then, based on customer satisfaction and natural attribute scores, the competitiveness index of each entity is calculated. Finally, according to the competition index Sorting from largest to smallest, the top-ranked stores represent the target entity's core competitors. The higher the ranking, the stronger the core competitiveness, and the more reliable the metrics.
[0054] Based on the above methods, this invention, by considering the various competitive indicators affecting entities and the correlations between these indicators, proposes a method for identifying core competitors of an entity, building upon the initial identification of competitors using clustering algorithms. This provides methodological support for managers to understand the competitive position of a store in the market and determine improvement directions. First, considering the similarities in store layout among competitors, principal component analysis is performed on similarity indicators to obtain the natural attributes of the entities as clustering indicators. Further clustering algorithms are used to perform cluster analysis on the entities, resulting in a candidate set of competitors. Then, customer evaluations of the target entity and its candidate competitors are obtained through surveys. Natural language processing (NLP) technology is used to identify customer concerns, and these concerns, along with the natural attributes of the entities, are combined to form competitive indicators. Correlation relationships between these indicators are then mined based on expert evaluations. Finally, Choquet integrals are used to fuse different types of competitive indicator data, thereby identifying the core competitors of the target entity. By comparing and analyzing with competitors, managers can more quickly and accurately determine directions for improving store quality and service.
[0055] Because this invention starts from the actual needs of customers, combines multiple types of indicators to analyze the competitiveness of entities, and does not ignore the correlation between indicators, it has higher reliability and accuracy than existing methods. Furthermore, this invention is embedded in a program, requiring only the import of survey text for automatic analysis, greatly saving labor costs.
[0056] Example 2 This embodiment discloses a data mining system for entity competitors based on the correlation between indicators.
[0057] A data mining system for entity competitors based on the correlation between indicators includes: The data preprocessing module is configured to: obtain the similarity index of entities and fill in missing values; The preliminary analysis module is configured to: use principal component analysis to remove the correlation between the various indicators in the similarity index to obtain the natural attributes of the entity and their corresponding attribute values; and perform cluster analysis on the entity based on the obtained natural attributes to generate a set of candidate competitors. The in-depth indicator analysis module is configured to: acquire user comment data on entities, extract feature words as customer concerns, and construct competitive indicators by combining customer concerns with natural attributes; analyze customer concerns based on sentiment analysis methods, and use the obtained sentiment score as customer satisfaction. The competitor identification module is configured to: combine the TF-IDF algorithm and expert evaluation to determine the weights of competitive indicators; use Choquet integrals to fuse competitive indicators and their weights to generate and rank the competitive index of entities in order to identify the core competitors of the target entity. Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.
[0058] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps in the entity competitor data mining method based on the correlation between indicators as described in Embodiment 1 of this disclosure.
[0059] Example 4 The purpose of this embodiment is to provide an electronic device.
[0060] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the entity competitor data mining method based on the correlation between indicators as described in Embodiment 1 of this disclosure.
[0061] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0062] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0063] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. An entity competitor data mining method based on correlation between indicators, characterized in that, The method comprises the following steps: Obtain the similarity indicators of the entities and fill in the missing values; remove the correlation between each indicator in the similarity indicators by using principal component analysis to obtain the natural attributes of the entities and the corresponding attribute values; Perform clustering analysis on the entities based on the obtained natural attributes to generate a candidate competitor set; Obtain the comment data of the entities by users, extract feature words as customer concerns, and construct the customer concerns and the natural attributes into competitive indicators; analyze the customer concerns by using a sentiment analysis method, and use the obtained sentiment scores as customer satisfaction; Determine the weights of the competitive indicators by combining the TF-IDF algorithm and expert evaluation; Fuse the competitive indicators and their weights by using Choquet integral to generate a competition index of the entities and sort the competition index to identify the core competitors of the target entity.
2. The method of claim 1, wherein the method further comprises: The feature words are extracted by the following steps: firstly, delete invalid comments, remove duplicate texts, and remove special symbols from the comment data of the users; then, perform word segmentation and part-of-speech tagging on the comment data to remove stop words; finally, extract feature words by using an LDA topic model, and determine the number of feature words according to the theme consistency score.
3. The method of claim 1, wherein the method further comprises: determining a correlation between the first and second indicators; and determining a correlation between the second and third indicators. The customer satisfaction is represented by using a sentiment dictionary and a probability language term set, and specifically comprises the following steps: construct a sentiment dictionary, determine the sentiment scores corresponding to the comment data, and perform weighted calculation on the customer concerns.
4. The method of claim 1, wherein the method further comprises: determining a correlation between the first and second indicators; and determining a correlation between the second and third indicators. The filling of the missing values is realized based on the principle of collaborative filtering, and the missing values are filled by calculating the similarity between the entities and measuring the similarity based on cosine similarity and Pearson correlation coefficient.
5. The method of claim 1, wherein the method further comprises: determining a correlation between the first and second indicators; and determining a correlation between the second and third indicators. The principal component analysis method comprises the following steps: standardize the original data into a standardized matrix, calculate the covariance matrix of the standardized matrix, extract the eigenvalues and eigenvectors of the covariance matrix, select principal components according to the eigenvalues, and construct a natural attribute matrix after dimension reduction.
6. The method of claim 1, wherein the method further comprises: determining a correlation between the first and second indicators; and determining a correlation between the second and third indicators. When performing clustering analysis on the entities based on the obtained natural attributes, a plurality of clustering algorithms are compared, the clustering effect is evaluated by using the Davies-Bouldin index, and the optimal clustering result is selected to generate the candidate competitor set; wherein the plurality of clustering algorithms comprise hierarchical clustering, k-means clustering, and self-organizing mapping network clustering.
7. The method of claim 1, wherein the method further comprises: determining a correlation between the first and second indicators; and determining a correlation between the second and third indicators. The determination of the weights of the competitive indicators comprises the following steps: calculate the initial weights of the customer concerns by using the TF-IDF algorithm; generate expert weights according to the evaluation of the competitive indicators by experts; integrate the initial weights and the expert weights by using the weighted average method to obtain the final weights of the competitive indicators; wherein the generation of the expert weights comprises constructing an expert evaluation matrix, the matrix uses intuitionistic fuzzy numbers to represent the evaluation information, and combines the expert opinions by using an aggregation operator.
8. An entity competitor data mining system based on inter-indicator correlations, characterized in that, The method comprises the following steps: The data preprocessing module is configured to obtain the similarity indicators of the entities and fill in the missing values; The indicator preliminary analysis module is configured to remove the correlation between each indicator in the similarity indicators by using principal component analysis to obtain the natural attributes of the entities and the corresponding attribute values; perform clustering analysis on the entities based on the obtained natural attributes to generate a candidate competitor set; The index depth analysis module is configured to: acquire user comment data on the entity, extract feature words as customer attention points, and construct the customer attention points and natural attributes together as competitive indexes; analyze the customer attention points based on a sentiment analysis method, and use the obtained sentiment scores as customer satisfaction degrees; The competitor identification module is configured to: combine a TF-IDF algorithm and expert evaluation to determine the weights of the competitive indexes; The competitive indexes and the weights are fused by using Choquet integral to generate competitive indexes of the entity and sort the competitive indexes, so as to identify core competitors of the target entity.
9. A computer-readable storage medium having stored thereon a program, characterized in that, The program is executed by the processor to implement the steps in the entity competitor data mining method based on the correlation between indexes according to any one of claims 1-7.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and capable of running on the processor, characterized by The processor executes the program to implement the steps in the entity competitor data mining method based on the correlation between indexes according to any one of claims 1-7.