Company similarity calculation method
Through multi-source data integration and advanced feature engineering, combined with entropy weight method and feedback mechanism, the problem of insufficient data utilization in the company's similarity analysis is solved, accurate similarity calculation and continuous optimization are achieved, and the scientificity and reliability of the analysis are improved.
Patent Information
- Application Number
- CN202510444255.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-18
AI Technical Summary
It is difficult for the existing technology to fully utilize multi-source heterogeneous data in company similarity analysis, and the lack of advanced feature engineering and objective weight allocation mechanisms, resulting in a lack of depth and breadth of evaluation results, and there are artificial biases, which affects the scientificity and fairness of the evaluation system.
By obtaining structured and unstructured data from multiple data sources, key attributes are extracted using natural language processing and image recognition, feature weights are determined using entropy weight method, and similarity among companies is calculated by cosine similarity, and a feedback mechanism is introduced to optimize features and weight allocation.
Accurate quantitative evaluation of company similarity has been achieved, the accuracy and efficiency of feature extraction has been improved, the scientificity and fairness of the evaluation system has been ensured, and the calculation accuracy and applicability have been continuously improved.
Smart Images

Figure CN120336876A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing and analysis, and specifically refers to a method for calculating company similarity. Background Art
[0002] In the context of the current rapid development of economic globalization and digitalization, the accurate analysis of the similarity between companies has become an indispensable part of key business activities such as investment decision-making and market strategy formulation. Traditional company similarity analysis mainly relies on single-dimensional data sources, such as financial statements or industry classification information. Although this method can provide certain reference value, due to its limitation to specific types of data, it is often difficult to comprehensively reflect the multi-faceted characteristics and dynamic changes of companies. In addition, existing technologies have significant deficiencies in processing unstructured data (such as news reports and social media content), and are unable to effectively extract key information that helps understand the essential characteristics of companies, resulting in evaluation results lacking depth and breadth. At the same time, traditional methods mostly use subjective weighting methods when determining the importance weights of various features, which not only increases the risk of human bias, but also reduces the scientificity and fairness of the evaluation system. In response to these challenges, existing technical solutions fail to fully integrate multi-source heterogeneous data, lack advanced feature engineering technologies, and objective and systematic weight allocation mechanisms, thus limiting the accuracy and practicality of company similarity calculation. Therefore, there is an urgent need for a method for calculating company similarity that can comprehensively utilize various data resources, adopt cutting-edge data analysis technologies, and have self-optimization capabilities to meet the needs of the increasingly complex business environment. Summary of the Invention
[0003] I. Technical Problems to be Solved
[0004] The technical problem to be solved by the present invention is the various problems mentioned in the above background art, and a method for calculating company similarity is provided.
[0005] II. Technical Solution
[0006] To solve the above technical problems, the technical solution provided by the present invention is: A method for calculating company similarity, comprising the following steps:
[0007] S1. Data collection: Obtain structured and unstructured data on multiple target companies from multiple data sources;
[0008] S2. Feature engineering: Process the collected data, and extract a key attribute set that can reflect the essential characteristics of the target company through natural language processing and image recognition;
[0009] S3. Weight determination: According to the influence degree and importance of each feature on the company's business, use the entropy weight method to determine the weight value corresponding to each feature;
[0010] S4, Similarity calculation: Based on the extracted features and their corresponding weight values, apply cosine similarity to calculate the similarity scores between any two companies;
[0011] S5, Result optimization: Introduce a feedback mechanism to continuously adjust the feature selection criteria and weight allocation scheme to continuously improve the accuracy and reliability of the similarity calculation results.
[0012] As an improvement, the data sources in step S1 specifically include public databases, business reports, news and social media, the structured data specifically includes financial data, operating indicators, organizational structure information, and industry classification codes, and the unstructured data specifically includes news reports and announcements, social media posts, business reports, and product pictures.
[0013] As an improvement, the data processing in step S2 specifically includes cleaning the irrelevant or incorrect information in the original data. At the same time, for the missing data items, the methods of mean filling, median filling, etc. can be used to complete them.
[0014] As an improvement, the steps of determining the weight values corresponding to each feature by the entropy weight method are as follows:
[0015] A. Construct the original data matrix. Suppose there are n companies as evaluation objects and m features are used to describe these companies, then an n×m original data matrix can be constructed
[0016] X = [x ij n×m , where Xij represents the value of the i-th company on the j-th feature.
[0017] B. Data standardization. The standardization formula is
[0018]
[0019] where max(xj) and min(xj) are the maximum and minimum values of the j-th feature respectively.
[0020] C. Calculate the entropy value of each feature. First, calculate the proportion pij of the j-th feature in the i-th sample:
[0021]
[0022] Then, calculate the entropy value Ej of the j-th feature according to the definition of information entropy:
[0023]
[0024] D. Calculate the redundancy. The redundancy dj reflects the importance of the j-th feature, and the calculation formula is
[0025] d j = 1 - E j
[0026] C. Calculate the weights, calculate the weight wj of each feature
[0027]
[0028] As an improvement, in the result optimization step, the feedback mechanism includes the collection and analysis of internal algorithm performance evaluation and external user evaluation information to guide the iterative update of the algorithm.
[0029] III. Beneficial effects
[0030] The advantages of the present invention compared with the prior art are as follows:
[0031] A method for calculating company similarity proposed by the present invention realizes the accurate quantitative evaluation of the similarity between companies through multi-source data integration, advanced feature engineering technology, and a scientific weight allocation mechanism. First, in the data collection stage, this method not only relies on traditional structured data (such as financial data, business indicators, etc.), but also makes full use of unstructured data (such as news reports, social media posts, etc.), thus providing a more comprehensive and in-depth perspective to understand the essential characteristics of companies. Second, through natural language processing and image recognition technology for feature engineering, key attribute sets are effectively extracted from massive information, greatly improving the accuracy and efficiency of feature extraction. Furthermore, the entropy weight method is used to objectively determine the importance weights of each feature, ensuring the scientificity and fairness of the evaluation system. The similarity calculation step based on cosine similarity further guarantees the reliability and interpretability of the results. Finally, a feedback mechanism including internal algorithm performance evaluation and external user feedback is introduced, enabling this method to continuously self-optimize and improve the accuracy and applicability of similarity calculation. In summary, while improving the accuracy of company similarity analysis, the present invention also provides strong data support and technical guarantee for fields such as investment decision-making and market analysis, and has significant practical value and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of a method for calculating company similarity of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0034] As Figure 1 shown, the present invention provides a method for calculating company similarity, specifically including the following detailed implementation steps:
[0035] S1 Data collection
[0036] Obtain structured and unstructured data about multiple target companies from multiple data sources. The data sources specifically include but are not limited to public databases, business reports, news and social media. Among them, structured data includes financial data (such as revenue, profit), operating metrics (such as market share, growth rate), organizational structure information (such as number of employees, management structure), and industry classification codes; unstructured data covers news reports and announcements, social media posts, business reports, and product pictures, etc.
[0037] In specific implementation, web crawler tools can be used to scrape relevant data from financial websites and social media platforms, and combined with API interfaces to obtain information such as financial statements from public databases.
[0038] S2 Feature engineering
[0039] Process the collected data, and extract a key attribute set that can reflect the essential characteristics of the target company through natural language processing (NLP) technology and image recognition technology. First, clean the irrelevant or incorrect information in the original data, and supplement missing values by means of mean filling, median filling, etc.
[0040] Text data processing: Use a tokenizer to split the text into words or phrases and remove stop words. Apply topic modeling algorithms to identify potential topics in the text. Use sentiment analysis to evaluate public sentiment or customer satisfaction.
[0041] Image data processing: Preprocess the pictures (resize, crop, etc.), and then apply a convolutional neural network (CNN) to automatically identify key elements and label them.
[0042] S3 Weight determination
[0043] According to the degree of influence and importance of each feature on the company's business, the entropy weight method is used to determine the weight value corresponding to each feature.
[0044] Construct the original data matrix: Assume that n companies are used as evaluation objects, and m features are used to describe these companies, then an n×m original data matrix X can be constructed.
[0045] Data standardization: For each feature j, use the formula for standardization.
[0046] Calculate the entropy value Ej: Based on the standardized data, calculate the entropy value of each feature according to the given formula.
[0047] Calculate redundancy degree dj: Calculate the redundancy degree using dj = 1 - Ej.
[0048] Calculate weight wj: Finally, through Obtain the weights of each feature.
[0049] S4 Similarity calculation
[0050] Based on the extracted features and their corresponding weight values, apply cosine similarity to calculate the similarity score between any two companies.
[0051] Dot product calculation: Calculate the dot product of two feature vectors A and B
[0052] Norm calculation: Calculate the norms ||A|| and ||B|| of the two vectors respectively.
[0053] Cosine similarity: According to the formula Obtain the similarity score.
[0054] S5 Result optimization
[0055] Introduce a feedback mechanism to continuously adjust the feature selection criteria and weight allocation scheme to continuously improve the accuracy and reliability of the similarity calculation results. The feedback mechanism includes not only the internal algorithm performance evaluation but also the collection and analysis of external user evaluation information to guide the iterative update of the algorithm.
[0056] Internal evaluation: Regularly test the model performance and check the prediction accuracy.
[0057] External feedback: Set up a user feedback channel to collect the opinions and suggestions of users, and adjust the feature selection criteria and weight settings accordingly.
[0058] Through the above specific implementation manners, the present invention can effectively integrate multi-source heterogeneous data, extract key features using advanced data analysis techniques, and determine their weights using scientific methods, so as to achieve accurate quantitative evaluation of the similarity between companies, providing strong support for fields such as investment decision-making and market analysis.
[0059] According to the influence degree and importance of each feature on the company's business, use the entropy weight method to determine the weight values corresponding to each feature..
[0060] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0061] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
[0062] The above description of the present invention and its embodiments is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of the present invention, design similar structural modes and embodiments to this technical solution without creative efforts, they should all fall within the protection scope of the present invention.
Claims
1. A method for calculating company similarity, characterized in that It includes the following steps: S1. Data collection: Obtain structured and unstructured data on multiple target companies from multiple data sources; S2. Feature engineering: Process the collected data, and extract a key attribute set that can reflect the essential characteristics of the target company through natural language processing and image recognition; S3. Weight determination: According to the influence degree and importance of each feature on the company's business, use the entropy weight method to determine the weight value corresponding to each feature; S4. Similarity calculation: Based on the extracted features and their corresponding weight values, apply cosine similarity to calculate the similarity score between any two companies; S5. Result optimization: Introduce a feedback mechanism to continuously adjust the feature selection criteria and weight allocation scheme to continuously improve the accuracy and reliability of the similarity calculation results.
2. The method for calculating the similarity of companies according to claim 1, wherein In step S1, the data sources specifically include public databases, commercial reports, news and social media. The structured data specifically includes financial data, operating indicators, organizational structure information, and industry classification codes. The unstructured data specifically includes news reports and announcements, social media posts, commercial reports, and product pictures.
3. A method for calculating company similarity according to claim 1, characterized in that, In step S2, the processing of the data specifically includes cleaning irrelevant or incorrect information in the original data. At the same time, for data items with vacancies, methods such as mean filling, median filling, or the like can be used to supplement them completely.
4. A method for calculating company similarity according to claim 1, characterized in that, The steps of using the entropy weight method to determine the weight value corresponding to each feature are as follows: A. Construct an original data matrix. Suppose there are n companies as evaluation objects and m features are used to describe these companies, then an n×m original data matrix can be constructed X = [x ij n×m , where Xij represents the value of the i-th company on the j-th feature. B. Data standardization. The standardization formula is where max(xj) and min(xj) are the maximum and minimum values of the j-th feature respectively. C. Calculate the entropy value of each feature. First, calculate the proportion pij of the j-th feature in the i-th sample: Then calculate the entropy value Ej of the j-th feature according to the definition of information entropy: D. Calculate the redundancy. The redundancy dj reflects the importance of the j-th feature, and the calculation formula is d j = 1 - E j C. Calculate the weight. Calculate the weight wj of each feature 5. A method for calculating company similarity according to claim 1, characterized in that, In the result optimization step, the feedback mechanism includes internal algorithm performance evaluation and collection and analysis of external user evaluation information to guide the iterative update of the algorithm.