Short-period anonymized enterprise similarity calculation method and system
By employing a short-cycle anonymization method for enterprise similarity calculation, the problems of data security risks and insufficient adaptability in existing technologies are addressed, thereby achieving accuracy and compliance in enterprise similarity calculation and improving the system's practicality and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-03-27
AI Technical Summary
Existing methods for calculating enterprise similarity pose data security risks, fail to fully consider various enterprise characteristics, lack adaptability, struggle to find a balance between short-term recognition accuracy and long-term privacy protection, and are not compliant enough.
A short-cycle anonymization method for calculating enterprise similarity is adopted. By acquiring enterprise data, feature extraction and feature vector construction are performed. For identifiable features, the duration is determined and anonymization or reversible encryption is performed. A similarity calculation model is used to calculate enterprise similarity, combined with an adaptive learning mechanism and a differentiated processing strategy.
This approach achieves the goal of improving the accuracy and adaptability of enterprise similarity calculations while protecting data privacy, ensuring compliant operation, and enhancing the system's usability and maintenance efficiency.
Smart Images

Figure CN121743884A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a data processing method, more particularly to a short-period anonymized enterprise similarity calculation method and system. BACKGROUND
[0002] Enterprise similarity calculation helps to identify market competitors and potential partners, providing data support for market strategy formulation, risk assessment and business cooperation. At the same time, it is also of great significance for intellectual property management, personalized service provision and resource optimization.
[0003] Existing systems tend to permanently save enterprise identification information or only perform one-time data desensitization processing. These practices pose data security risks, and once attacked, they can lead to the leakage of a large amount of sensitive enterprise data, especially the leakage of core business information such as transaction relationships and sales categories, which will cause significant losses to enterprises. Traditional similarity calculation methods usually only focus on single-dimensional features such as enterprise name, industry classification or business scope text, while ignoring the comprehensive consideration of enterprise multi-dimensional characteristics. Such one-sided analysis cannot accurately reflect the actual similarity between enterprises. Current technologies mostly use uniform standards for similarity calculation, without personalized adjustments for different types of companies such as entity and service companies, leading to similarity evaluation results that do not match the actual situation. Existing systems have deficiencies in distinguishing between authorized and unauthorized data usage, either over-relying on limited authorized data such as invoice data, resulting in narrow data coverage, or ignoring necessary authorization restrictions, causing compliance issues. Traditional rule-based similarity calculation methods require manual intervention to design and update rules, making them difficult to adapt to new data sets and industry changes, lacking the ability to automatically learn and optimize, and increasing maintenance costs. There is a lack of fine-grained management of enterprise data lifecycle, making it difficult to balance short-term identification accuracy and long-term privacy protection, often falling into the dilemma of complete identifiability or complete anonymity.
[0004] Therefore, it is necessary to design a new method to effectively protect data privacy, provide comprehensive enterprise feature analysis, implement differentiated processing for different types of enterprises, clearly define data authorization boundaries, have adaptive ability, and reasonably manage the timeliness of data. SUMMARY
[0005] The present application aims to overcome the shortcomings of the prior art and provide a short-period anonymized enterprise similarity calculation method and system.
[0006] To achieve the above-mentioned purpose, the present application adopts the following technical solution: a short-period anonymized enterprise similarity calculation method, comprising:
[0007] obtaining a plurality of enterprise data;
[0008] feature extraction and feature vector construction are performed on each of the enterprise data to obtain a feature vector;
[0009] a time length judgment is performed on identifiable features in the feature vector, and the corresponding feature vector is processed according to a judgment result to obtain a processing result;
[0010] an anonymous feature in the feature vector and the processing result are input into a similarity calculation model for similarity calculation to obtain a calculation result;
[0011] the calculation result is output.
[0012] The application further provides a short-period anonymized enterprise similarity calculation device, which comprises:
[0013] a data acquisition unit configured to acquire a plurality of enterprise data;
[0014] a feature vector construction unit configured to perform feature extraction and feature vector construction on each of the enterprise data to obtain a feature vector;
[0015] a processing unit configured to perform a time length judgment on identifiable features in the feature vector, and process the corresponding feature vector according to a judgment result to obtain a processing result;
[0016] a calculation unit configured to input an anonymous feature in the feature vector and the processing result into a similarity calculation model for similarity calculation to obtain a calculation result;
[0017] an output unit configured to output the calculation result.
[0018] The application further provides a computer device, which comprises a memory and a processor, the memory has a computer program stored thereon, and the processor realizes the above method when executing the computer program.
[0019] Compared with the prior art, the present application has the beneficial effects that: the present application ensures comprehensive capture of enterprise characteristics by acquiring multiple enterprise data and performing feature extraction and feature vector construction; the identifiable features in the feature vector are subjected to duration judgment, and are processed according to a set short life cycle anonymization mechanism to protect data privacy, while retaining sufficient feature information for subsequent analysis, so as to meet business needs while protecting privacy; different types of enterprises are subjected to differentiated feature engineering and similarity calculation strategies, improving the accuracy and adaptability of analysis; the data use rules of authorized and non-authorized enterprises are clearly distinguished to ensure compliance operation; an adaptive learning mechanism is used to reduce manual intervention, so that the system can automatically adapt to new data distribution and industry characteristics; and the anonymized features and processing results are input into a similarity calculation model for calculation, and finally the output results are output, realizing efficient and safe evaluation of the similarity between enterprises, reasonably managing data timeliness, and enhancing the practicability and maintenance efficiency of the system.
[0020] The present application will be further described below in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0022] Figure 1 The application scenario diagram of the short cycle anonymized enterprise similarity calculation method provided by the embodiment of the present application;
[0023] Figure 2 The flowchart of the short cycle anonymized enterprise similarity calculation method provided by the embodiment of the present application;
[0024] Figure 3 The schematic block diagram of the short cycle anonymized enterprise similarity calculation device provided by the embodiment of the present application;
[0025] Figure 4 The schematic block diagram of the computer device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0027] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0028] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0029] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0030] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating an application scenario of the short-cycle anonymization enterprise similarity calculation method provided in this embodiment of the invention. Figure 2 This is a schematic flowchart illustrating the short-cycle anonymization method for calculating enterprise similarity provided in this invention. This method is applied to a server that interacts with terminals. It acquires and processes data from multiple enterprises, extracts features and constructs feature vectors for each enterprise, and determines whether to anonymize or reversibly encrypt identifiable features based on their duration, thus effectively protecting data privacy. Differentiated processing is implemented for different types of enterprises to ensure comprehensive enterprise feature analysis while protecting individual privacy. By using a similarity calculation model including a feature encoder and a similarity calculation module to perform multi-level, multi-type similarity calculations on the processed features, and utilizing a comprehensive loss function to optimize the model's accuracy and consistency, effective assessment of inter-enterprise similarity is achieved. Furthermore, the model employs dual-track training, knowledge distillation techniques, and continuous learning and sliding window strategies, enabling it to adapt dynamically to changes in data distribution, reasonably manage data timeliness, and clearly define data authorization boundaries to ensure compliance in data use. This method not only improves the security and efficiency of data processing but also provides a scientific basis for inter-enterprise comparisons.
[0031] Figure 2 This is a flowchart illustrating the short-cycle anonymization method for calculating enterprise similarity provided in an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S150.
[0032] S110: Obtain data from multiple enterprises.
[0033] In this embodiment, enterprise data refers to a collection of various information about different enterprises gathered from multiple sources. This information is used for subsequent feature extraction, similarity calculation, and model training. Specifically, enterprise data includes, but is not limited to, the following aspects:
[0034] Basic Information: This refers to the basic information provided during company registration, including company name, registered capital, date of establishment, company type (such as limited liability company, joint-stock company, etc.), legal representative, and registered address. This information provides the company with basic identification.
[0035] Scope of business: This describes the various business activities that the company is authorized to conduct. It can be a standardized text description, or it can be further transformed into structured data or vector representations to facilitate the calculation of similarity.
[0036] Industry characteristics: This covers the primary industry, secondary industry, and their sub-sectors to which the company belongs. This helps identify the company's position in the industry chain and its relationships with other companies.
[0037] Product and service characteristics: Based on data from invoices or other sales records, this section showcases the types of goods and services sold by the company. This may include sales distribution, major product categories, and product concentration indicators, used to measure the company's operating characteristics and market positioning.
[0038] Transaction network characteristics: These reflect the transaction relationships between a company and its suppliers and customers. For example, the number of core suppliers, the number of core customers, the number of provinces with which transactions take place, the proportion of suppliers with similar characteristics, and the proportion of customers with similar characteristics, all contribute to constructing a network diagram that reflects the interaction patterns between companies.
[0039] Characteristics of qualification certificates: These refer to various qualification certificates held by the company, such as ISO certification and industry-specific licenses. These certificates not only prove the company's capabilities but also serve as an important dimension for judging the similarity between companies.
[0040] Financial indicators include those related to scale, growth, and stability, such as operating revenue, net profit, and debt-to-equity ratio, which provide a basis for assessing the company's economic condition.
[0041] Anonymized features: For enterprise data exceeding a set lifespan (e.g., 15 days), directly identifiable information is anonymized and replaced with a randomly generated UUID, while the feature information is retained for model iteration and optimization. This approach protects enterprise privacy while ensuring the effective use of data.
[0042] After acquiring this enterprise data, the next steps involve preprocessing, feature engineering, and training deep learning models to accurately calculate enterprise similarity. This process not only facilitates intelligent matching between enterprises but also emphasizes the importance of data privacy, ensuring that business needs are met without infringing on user privacy.
[0043] S120. Perform feature extraction and feature vector construction on each of the enterprise data to obtain a feature vector.
[0044] In this embodiment, a feature vector refers to a numerical vector representation obtained after processing, fusing, selecting, and reducing the dimensionality of various features extracted from enterprise data. These feature vectors can capture the core characteristics of an enterprise and are used for subsequent enterprise similarity calculations, model training, and other analytical tasks.
[0045] In one embodiment, step S120 described above may include steps S121 to S122.
[0046] S121. Extract different categories of features from each of the enterprise data.
[0047] In this embodiment, multiple types of features are extracted from the raw enterprise data. These features can reflect the characteristics of the enterprise in different dimensions. According to the provided technical solution overview, the enterprise data is divided into multiple categories, each containing different features:
[0048] Basic information features: such as registered capital, years of establishment, province of registration, and type of enterprise.
[0049] Business scope characteristics: The business scope is transformed into a vector representation through text processing techniques (such as TF-IDF or Word2Vec).
[0050] Industry characteristics: including industry classifications and information on sub-sectors.
[0051] Product / service characteristics: Indicators such as product category distribution and sales concentration are constructed based on sales data.
[0052] Transaction network characteristics: reflecting the number of core suppliers and customers in the upstream and downstream relationship, the geographical distribution of transactions, etc.
[0053] Characteristics of qualification certificates: the various qualification certificates held and their certification status.
[0054] Financial indicators: such as revenue size, profit level, etc.
[0055] S122. Process, fuse, select, and perform principal component analysis on features of different categories to obtain feature vectors.
[0056] In this embodiment, the extracted features of each category are first processed to make them suitable for subsequent analysis and model input. Then, through a series of technical means, these features are fused into a unified feature vector.
[0057] Feature processing: Apply appropriate processing methods to different types of data. For example, for text data, word segmentation and stop word removal may be necessary; for numerical data, standardization or normalization may be required.
[0058] Feature fusion: Combining features from different sources to form a comprehensive feature representation. This can be achieved by directly concatenating the feature vectors or by using more complex fusion strategies.
[0059] Feature selection aims to identify and retain features that contribute most to the final objective (such as enterprise similarity calculation) while removing redundant or irrelevant features. Common methods include selection based on statistical tests and recursive feature elimination.
[0060] Principal Component Analysis (PCA): As a dimensionality reduction technique, PCA can reduce the dimensionality of feature vectors while preserving as much information as possible from the original data. This is crucial for improving computational efficiency and avoiding overfitting.
[0061] Through the above process, the resulting feature vectors not only contain rich information about the companies but also have a well-structured form, facilitating further use by machine learning algorithms. These feature vectors form the basis of the input layer of the deep neural network, supporting accurate calculation of company similarity.
[0062] S130. The duration of the identifiable features in the feature vector is determined, and the corresponding feature vector is processed according to the determination result to obtain the processing result.
[0063] In this embodiment, the processing result refers to anonymizing or reversibly encrypting these features based on whether the lifecycle of the identifiable features in the feature vector exceeds a set number of days, in order to generate a new feature representation that protects privacy while maintaining data validity.
[0064] Specifically, if the lifecycle exceeds the set number of days, the identifiable features will be replaced with a random UUID for anonymization; if it does not exceed the set number of days, the enterprise identifier will be stored using reversible encryption.
[0065] In one embodiment, please refer to Figure 4 The above-mentioned step S130 may include steps S131 to S132.
[0066] S131. Determine whether the lifecycle of the identifiable features in the feature vector exceeds a set number of days.
[0067] In this embodiment, this step mainly checks whether the duration of the identifiable features (such as enterprise tax number, name, and other identification information that can be directly associated with a specific enterprise) contained in each feature vector exceeds a preset number of days, which in this embodiment is 15 days. This process may involve database queries or recording timestamps to determine how many days have passed since the identifiable features in a certain feature vector were created or last updated.
[0068] S132. If the lifecycle of an identifiable feature in the feature vector exceeds a set number of days, the identifiable feature is anonymized to generate a random identifier, thus forming a processing result.
[0069] In this embodiment, when an identifiable feature in a feature vector is found to have exceeded a set lifespan (e.g., 15 days), the feature needs to be anonymized. This typically includes the following steps:
[0070] Anonymization: The original identifiable features are replaced with a completely randomly generated UUID (Universally Unique Identifier), so that even if someone possesses the UUID, they cannot trace back to the original corporate identity.
[0071] Preserve feature information but remove traceable fields: Although anonymization is performed, other features related to the enterprise (such as business scope, product sales distribution, etc.) are still retained to facilitate model training and optimization.
[0072] S133. If the lifecycle of the identifiable features in the feature vector does not exceed a set number of days, then the enterprise identifier of the identifiable features is reversibly encrypted and stored to obtain the processing result.
[0073] In this embodiment, if the identifiable features in the feature vector are still within the set lifespan (i.e., not exceeding 15 days), then these features will be encrypted and stored in a reversible manner. This means:
[0074] Reversible encryption: This method uses a specific encryption algorithm to encrypt and store a company's identification information, allowing the original identification to be recovered through decryption when necessary. This approach ensures short-term data availability and meets business needs while enhancing data security through encryption.
[0075] Maintain traceability: Within the validity period, the system can trace back to the specific enterprise based on the encrypted identifier, meeting immediate business needs such as precise matching or similarity calculation.
[0076] Through the steps described above, this embodiment effectively protects enterprise data privacy while also considering short-term business operational needs. This dynamic adjustment mechanism not only improves data security but also ensures that the data used during model training is of high quality and relevance, thereby enhancing the accuracy and efficiency of enterprise similarity calculation.
[0077] S140. The anonymized features in the feature vector and the processing result are input into the similarity calculation model to perform similarity calculation to obtain the calculation result.
[0078] In this embodiment, the calculation result refers to the quantitative index reflecting the degree of similarity between data items obtained after processing the anonymized features and the processing results through the similarity calculation model.
[0079] The similarity calculation model includes a feature encoder and a similarity calculation module;
[0080] The feature encoder includes a neural network structure, which includes a fully connected layer, a batch normalization layer, a LeakyReLU activation function, and a Dropout layer.
[0081] The similarity calculation module includes a similarity feature calculation layer, a fully connected layer, a ReLU activation function, and a Dropout layer.
[0082] In one embodiment, step S140 described above may include steps S141 to S147.
[0083] S141. Input the anonymized features in the feature vector and the processing result into the similarity calculation model.
[0084] In this embodiment, the anonymized features and processing results in the feature vector are used as inputs and fed into the similarity calculation model for subsequent processing.
[0085] S142. The anonymized features in the feature vector and the processing result are preprocessed by the corresponding encoder to obtain the encoded features.
[0086] In this embodiment, a feature encoder is used to preprocess the anonymized features of the input and the processing results. The encoded feature representation is generated through a neural network structure that includes a fully connected layer, a batch normalization layer, a LeakyReLU activation function, and a Dropout layer.
[0087] S143. Integrate all encoded features in the feature fusion layer to form a unified feature representation to obtain fused features.
[0088] In this embodiment, all encoded features are integrated in the feature fusion layer to form a unified feature representation, namely the fused feature, which prepares for subsequent similarity calculation.
[0089] S144. The fused features are sequentially processed through a fully connected layer containing 512 neurons and a batch normalization layer. The LeakyReLU activation function is used to introduce non-linear processing, and Dropout technology is applied to randomly discard neurons. The features are then extracted through a fully connected layer containing 256 neurons, batch normalization is performed again, and the LeakyReLU activation function is used again. Dropout technology is applied to randomly discard neurons to obtain the representation vector.
[0090] In this embodiment, the fused features are sequentially passed through two fully connected layers (containing 512 and 256 neurons respectively), a batch normalization layer, and a LeakyReLU activation function for nonlinear transformation, and the Dropout technique is applied to randomly discard neurons, finally obtaining the representation vector.
[0091] S145. Perform multi-type similarity calculations on the representation vector to obtain multiple similarity features.
[0092] In this embodiment, multiple similarity features refer to quantitative indicators that reflect the multi-dimensional similarity between enterprise representation vectors, obtained through calculation methods such as absolute difference, cosine similarity, and Hadamard product.
[0093] In one embodiment, step S145 described above may include steps S1451 to S1453.
[0094] S1451. Calculate the absolute difference element by element for the representation vectors of multiple enterprises to obtain the absolute difference.
[0095] S1452. Calculate the cosine similarity between the representation vectors of multiple enterprises to obtain the cosine similarity.
[0096] S1453. Perform the Hadamard product of the representation vectors of the multiple enterprises to obtain the Hadamard product;
[0097] The various similarity features include absolute difference, cosine similarity, and Hadamard product.
[0098] In this embodiment, multiple similarity calculation methods are applied to the representation vectors of multiple enterprises, including:
[0099] Absolute difference calculation: Calculate the absolute difference between the representation vectors element by element.
[0100] Cosine similarity calculation: measures the angular difference between representation vectors to assess similarity.
[0101] Hadamard product: Multiply corresponding positions and then sum them up to provide another perspective on similarity measurement.
[0102] These operations together constitute a variety of similarity features, namely absolute difference, cosine similarity, and Hadamard product.
[0103] S146. The various similarity features are concatenated to form a comprehensive similarity feature vector.
[0104] In this embodiment, the various similarity features obtained above are concatenated to form a comprehensive similarity feature vector for easier overall analysis.
[0105] S147. After inputting the similarity feature vector into a fully connected layer containing 128 neurons, the ReLU activation function is used to perform a non-linear transformation on the output of the fully connected layer. The Dropout technique is applied to randomly discard neurons. Then, the feature is extracted and compressed through a fully connected layer containing 64 neurons. The ReLU activation function is used again to process the feature to obtain the calculation result.
[0106] In this embodiment, the comprehensive similarity feature vector is finally input into a module containing two fully connected layers (128 and 64 neurons respectively), and the ReLU activation function is used for nonlinear transformation and Dropout technique is applied to obtain the final calculation result, a quantitative index reflecting the degree of similarity between data items.
[0107] This process ensures that even after corporate information is anonymized, the system can still effectively calculate the similarity between different companies, while guaranteeing both privacy protection and computational efficiency.
[0108] S150, Output the calculation results.
[0109] In this embodiment, the calculation results are output to the terminal for display.
[0110] In this embodiment, the definition of enterprise similarity is first clarified, and enterprises are classified into entity class and service class according to their business characteristics. Then, a series of feature engineering and deep learning models are used to accurately calculate enterprise similarity. This includes technical steps such as data collection and preprocessing, sample labeling, feature engineering, deep neural network structure design, and multi-objective optimization.
[0111] Enterprise similarity is mainly used to measure the similarity of enterprises in their business operations, covering the degree of matching in multiple dimensions such as industry characteristics, business scope, products or services, and trading partners.
[0112] Enterprises are divided into two main categories:
[0113] Physical enterprises: such as agriculture, forestry, animal husbandry and fishery, mining, manufacturing, etc.
[0114] Service-oriented enterprises: such as transportation, warehousing and postal services, information transmission, software and information technology services, etc. If the service fee invoice amount accounts for the highest proportion of a physical enterprise's total invoices, it is also considered a service-oriented enterprise.
[0115] Enterprise similarity tagging adopts a multi-level rule system:
[0116] For enterprises that have been authorized to access data (mainly those with invoice data):
[0117] Based on the similarity of product categories sold:
[0118] If a single product category accounts for ≥50% of sales, then companies with the same product category and a sales share of ≥50% will be identified as similar companies.
[0119] If the top 3 product categories account for ≥70% of sales, then search for companies that contain the company's top 3 product categories and whose total sales account for ≥70% as similar companies.
[0120] If the sales share of the top 10 product categories is ≥80%, then search for companies that contain the top 10 product categories of this company and whose total sales share is ≥80% as similar companies.
[0121] Based on the similarity of the transaction objects:
[0122] At least one supplier / customer accounts for ≥10% of the total valid invoice amount.
[0123] There are at least two identical suppliers / customers whose valid invoice amount accounts for ≥5% of the total;
[0124] There are at least 3 consistent suppliers / customers.
[0125] For companies that have not authorized data access:
[0126] Based on the similarity of business scope: enterprises whose first 3 business scopes have at least 1 identical item and whose total business scopes have at least 50% identical items;
[0127] Based on the similarity of qualification certificates: the secondary industry is consistent and includes any qualification of the user company;
[0128] Based on the similarity of bidding projects: participating in the same bidding project and being in the same secondary industry.
[0129] Starting with raw enterprise data, labeling rules are applied to identify and extract enterprise groups with similar characteristics. Positive sample pairs (label = 1) and negative sample pairs (label = 0) are constructed, and features are extracted from these sample pairs. Finally, short-lifecycle processing is applied to the features to ensure the freshness and validity of the data, forming the final training sample set.
[0130] The original enterprise identification information is reversibly encrypted and stored, with a validity period of 15 days. After the expiration period, it is replaced with a completely random UUID, while retaining the feature information for model training.
[0131] Based on different dimensions of information about an enterprise, characteristics can be categorized into basic information characteristics, business scope characteristics, industry characteristics, product / service characteristics, transaction network characteristics, qualification certificate characteristics, and financial indicator characteristics.
[0132] Text feature processing involves word segmentation and standardization, constructing TF-IDF or Word2Vec vectors, and calculating text similarity; commodity sales distribution features are constructed based on invoice data to include TOPN commodity sales distribution features, commodity category diversity indicators, and sales concentration indicators; transaction network features include the proportion of core transaction objects in the construction of inter-enterprise transaction relationship networks, transaction network diversity, and transaction geographic distribution, etc.
[0133] Based on different dimensions of information about the enterprise, the characteristics can be divided into the following categories:
[0134] Basic information features include basic information such as company name, registered capital, company type, and geographical location.
[0135] Business scope characteristics: Represented by standardized business scope text vectors.
[0136] Industry characteristics: This involves industry classification, definitions of sub-sectors, and analysis of the correlation between upstream and downstream industries.
[0137] Product / service characteristics: This includes information such as the distribution of product categories and product name vectors.
[0138] Transaction network characteristics: Describes the composition of upstream and downstream enterprises and the relationship between core trading partners.
[0139] Certification characteristics: Records the type of certificate held and its coverage.
[0140] Financial indicators include the company's size, growth, and stability.
[0141] For text information such as business scope and product names, the following steps are taken for processing:
[0142] Word segmentation and standardization: First, the text is segmented into words, and then the vocabulary is standardized.
[0143] Constructing TF-IDF or Word2Vec vectors: Next, use the TF-IDF method or Word2Vec model to convert the text into vector form.
[0144] Calculate text similarity: Finally, calculate the similarity between these vectors to quantify the degree of similarity between the texts.
[0145] Based on invoice data, the following features were constructed:
[0146] TOP N Product Sales Distribution Characteristics: Identify the top N best-selling products and analyze their sales distribution.
[0147] Product category diversity index: assesses the diversity of product categories sold by a business.
[0148] Sales concentration index: measures the degree to which sales are concentrated on a specific product.
[0149] To build a network of inter-enterprise transaction relationships, the following aspects should be considered:
[0150] Percentage of core trading partners: Analyze the proportion of core trading partners in total transactions.
[0151] Transaction network diversity: Examine the number and types of different transaction objects in the transaction network.
[0152] Geographical distribution of transactions: Study the geographical distribution of transaction activities to understand the market coverage of enterprises.
[0153] The above process ensures a systematic handling of the data from raw data to feature extraction, providing strong support for subsequent model training.
[0154] For example, as shown in Tables 1 to 5.
[0155] Table 1. Examples of Basic Enterprise Characteristics
[0156]
[0157]
[0158] Table 2. Examples of Product Sales Characteristics
[0159]
[0160] Table 3. Examples of Business Scope Characteristics
[0161]
[0162]
[0163] Table 4. Examples of Transaction Relationship Characteristics
[0164]
[0165] Table 5. Examples of Qualification Certificate Features
[0166]
[0167]
[0168] After performing feature importance analysis, the following five features and their corresponding similarity or matching scores were obtained. These scores reflect the relative importance of each feature in the model or decision-making process:
[0169] Product category similarity is the most important feature, with a similarity value of 0.32. This means that the degree of matching in product categories carries significant weight when assessing the similarity between businesses. If two businesses have a high similarity in product categories, they are likely to have a high degree of relevance in other aspects as well.
[0170] The similarity of business scope follows closely behind product category similarity, with a similarity value of 0.28. This indicates that the business scope of a company is also an important indicator for measuring the similarity between companies. The similarity of business scope can reflect commonalities among companies in terms of market positioning, business models, and other aspects.
[0171] A similarity score of 0.22 for overlap in trading partners indicates that the degree of overlap between companies in terms of trading partners is also an important factor to consider. If two companies have many common trading partners, they may have certain connections in terms of supply chain, market demand, etc.
[0172] The similarity score for industry classification consistency is 0.10, which, although relatively low, still influences the assessment of similarity between firms to some extent. Consistency in industry classification can help identify competitive relationships and cooperation opportunities among firms within the same industry.
[0173] The similarity score for qualification certificates was the lowest, at 0.08. Nevertheless, it remains a useful indicator for assessing the similarity between companies. The matching score of qualification certificates can reflect the similarity between companies in terms of compliance, professional capabilities, and other aspects.
[0174] Based on the above feature importance analysis, we can conclude that when assessing the similarity between enterprises, product category similarity and business scope similarity are the two most critical factors, accounting for 32% and 28% of the total similarity, respectively. The overlap of trading partners also has a certain influence, accounting for 22% of the total similarity. While industry classification consistency and qualification certificate matching are relatively less important, they still contribute to the overall assessment results. Therefore, in practical applications, these features should be considered comprehensively to accurately assess the similarity between enterprises.
[0175] To meet the enterprise's requirement for 15-day anonymization of information, a short lifecycle management mechanism for features was designed:
[0176] Identifiable features: These include features that can be directly linked to a specific company, such as the company's tax identification number and company name.
[0177] Anonymized features: Features that have undergone anonymization, including industry categories, operating indicators, and other features that cannot be directly linked to a specific company.
[0178] Processing within 15 days:
[0179] Within 15 days: Enterprise identifiers are stored using reversible encryption, allowing the system to trace back to the specific enterprise.
[0180] 15 days later: Replace with completely random UUIDs to ensure there is no correlation, and retain the feature vectors.
[0181] The computational model in this embodiment is mainly divided into four layers: input layer, transfer learning layer, feature representation layer, and output layer. Each layer contains different modules and processing steps to calculate the similarity between company A and company B.
[0182] In the input layer, the feature vectors of the two companies are first obtained:
[0183] Company A's feature vector contains various attributes and feature information of Company A.
[0184] Company B Feature Vector: Contains various attributes and feature information of Company B.
[0185] These feature vectors are the foundational data for all subsequent calculations.
[0186] In the transfer learning layer, there are two key pre-trained models:
[0187] Industry knowledge pre-training model: By using industry-related knowledge and experience for pre-training, it can capture the characteristics and patterns of a specific industry.
[0188] Enterprise Representation Pre-trained Model: Based on a large amount of enterprise data, it is pre-trained and can extract general enterprise feature representations.
[0189] In addition, there are two feature encoders:
[0190] Feature Encoder A: Encodes the input feature vector of company A to generate a higher-level feature representation (company A representation vector).
[0191] Feature Encoder B: Encodes the input feature vector of Company B to generate a higher-level feature representation (Company B representation vector).
[0192] In this process, transfer knowledge is used to guide the learning process of feature encoders, enabling them to better capture the essence of enterprise characteristics.
[0193] At the feature representation layer, two encoded representation vectors are obtained:
[0194] Company A Representation Vector: Generated by feature encoder A, it contains a high-level feature representation of company A.
[0195] Enterprise B Representation Vector: Generated by feature encoder B, it contains a high-level feature representation of enterprise B.
[0196] These two representation vectors will be passed to the next level for further processing.
[0197] In the output layer, there is a similarity calculation module responsible for calculating various similarity metrics between company A and company B:
[0198] Product category similarity: measures the degree of similarity between two businesses in terms of product categories.
[0199] Business scope similarity: measures the degree of similarity between two companies in terms of their business scope.
[0200] Overall similarity: A comprehensive similarity value is calculated by combining the above two similarity metrics and other possible similarity indicators.
[0201] In the similarity calculation layer, the similarity calculation module receives the representation vectors of enterprise A and enterprise B from the feature representation layer, and calculates the product category similarity, business scope similarity, and comprehensive similarity through a series of calculation methods (such as cosine similarity, Euclidean distance, etc.).
[0202] The entire model achieves a comprehensive assessment of the similarity between company A and company B through multi-layered processing and computation. From the original feature vectors in the input layer, to the feature encoding and application of the pre-trained model in the transfer learning layer, to the high-level feature representation in the feature representation layer, and finally to the various similarity calculations in the output layer, each link is closely connected, together forming an efficient and accurate company similarity calculation system.
[0203] The method in this embodiment employs a deep neural network architecture based on transfer learning, and achieves high-precision calculation of enterprise similarity through pre-trained industry and enterprise representation models:
[0204] To improve model accuracy and address the data sparsity problem, a transfer learning mechanism is introduced:
[0205] Industry knowledge pre-training: Pre-training industry knowledge representation models on large-scale industry data; capturing inter-industry relationships and internal industry characteristics; providing industry-level knowledge transfer;
[0206] Enterprise Representation Pre-training: Pre-trains enterprise representation models using a large amount of publicly available enterprise data; learns representation methods for basic enterprise characteristics; and provides enterprise-level knowledge transfer.
[0207] Knowledge transfer methods: Freeze the parameters of the pre-trained layer and fine-tune only the subsequent layers; the feature adaptation layer transforms the pre-trained features to the target task; the layered fine-tuning strategy gradually unfreezes the pre-trained layer.
[0208] Similarity calculation models include:
[0209] Input layer: The model receives numerical features, category features, and textual features of enterprises as input.
[0210] Feature encoder: Numerical features go directly into the normalization layer; categorical features are converted into dense vectors through the embedding layer; text features are converted into vector representations through the text encoding layer (such as word embedding or pre-trained language model).
[0211] The feature fusion layer concatenates the three feature vectors mentioned above to form a comprehensive enterprise feature vector.
[0212] Fully connected layer-512: The fused feature vectors are first reduced in dimensionality by using a fully connected layer to map them to a 512-dimensional space.
[0213] Batch normalization layer: Performs batch normalization on the 512-dimensional feature vector to accelerate the training process and improve model stability.
[0214] LeakyReLU activation: The LeakyReLU activation function is used to introduce a nonlinear transformation, which enhances the expressive power of the model.
[0215] Dropout-0.3: Applies Dropout technology to randomly discard 30% of neurons to prevent overfitting.
[0216] Fully Connected Layer - 256: Further dimensionality reduction to 256-dimensional space.
[0217] Batch normalization layer: Batch normalization is performed again.
[0218] LeakyReLU activation: Apply the LeakyReLU activation function again.
[0219] Dropout-0.2: Applies Dropout technology to randomly discard 20% of neurons.
[0220] Enterprise Representation Vector - 128 Dimensions: The final result is a 128-dimensional enterprise representation vector, which serves as the input for subsequent similarity calculations.
[0221] The similarity calculation module includes:
[0222] The representation vectors of Company A and Company B are input into the similarity calculation module respectively.
[0223] Absolute difference calculation: Calculate the absolute difference between the representation vectors of two enterprises.
[0224] Cosine similarity calculation: Calculate the cosine similarity between the representation vectors of two enterprises.
[0225] Hadamard product calculation: Calculate the Hadamard product (element-wise multiplication) of two enterprise representation vectors.
[0226] Similarity feature concatenation: The results of the above three similarity calculations are concatenated to form a comprehensive similarity feature vector.
[0227] Fully connected layer-128: Performs the first dimensionality reduction on the similarity feature vectors, using a fully connected layer to map them to a 128-dimensional space.
[0228] ReLU activation: Introducing a nonlinear transformation using the ReLU activation function.
[0229] Dropout-0.2: Applies Dropout technology to randomly discard 20% of neurons.
[0230] Fully Connected Layer-64: Further dimensionality reduction to 64-dimensional space.
[0231] ReLU activation: Reapply the ReLU activation function.
[0232] The output layer includes: Product similarity output: Output the similarity between company A and company B in terms of products.
[0233] Business scope similarity output: Output the similarity between company A and company B in the business scope dimension.
[0234] Overall similarity output: A comprehensive similarity value is calculated by combining the two similarity metrics mentioned above with other possible similarity metrics.
[0235] The entire model achieves a comprehensive assessment of the similarity between company A and company B through multi-layered processing and computation. From the raw features of the input layer, to the advanced feature representation of the feature encoding and fusion layer, to the various similarity calculation methods in the similarity calculation module, and finally to the multiple similarity outputs in the output layer, each link is closely connected, together forming an efficient and accurate company similarity calculation system.
[0236] To achieve effective knowledge transfer, the following strategy is adopted: In the pre-training stage, the model is initially trained with large-scale data to acquire the ability to represent industry knowledge and enterprise characteristics.
[0237] Industry Representation Model: Using large-scale industry data as input, an industry representation model is trained. This model can capture the overall characteristics and patterns of the industry, forming an industry knowledge representation.
[0238] Enterprise Representation Model: Using publicly available enterprise data as input, an enterprise representation model is trained. This model can extract specific characteristics and information of the enterprise, forming an enterprise characteristic representation.
[0239] These two models, from the perspectives of industry and enterprise respectively, provide the basic knowledge and feature representations for the subsequent fine-tuning stage.
[0240] During the fine-tuning phase, the model is further trained and optimized for specific target tasks to adapt to specific application scenarios.
[0241] Target task data: Use data related to the target task as input to fine-tune the model.
[0242] Feature adaptation layer: A feature adaptation layer is added to the model to transfer the industry knowledge representation and enterprise feature representation obtained in the pre-training stage to the target task.
[0243] By using transfer learning, the model can fully utilize the knowledge and features learned in the pre-training stage to improve the performance of the target task.
[0244] Target task model: The model after fine-tuning becomes the target task model, which can be directly applied to specific target tasks and provide accurate prediction and decision support.
[0245] In the fine-tuning process, a phased unfreezing strategy was adopted to gradually unleash the model's potential and avoid overfitting.
[0246] Initial Stage - Fine-tuning the Output Layer Only: In the initial stage of fine-tuning, only the model's output layer is fine-tuned, while the parameters of other parts remain unchanged. This allows for rapid adjustment of the model's output to adapt it to the requirements of the target task.
[0247] In the second stage of fine-tuning, the higher-level parts of the model are gradually unfrozen, allowing the parameters of these layers to be updated. In this way, the model can further optimize the extraction and representation of high-level features while retaining pre-trained knowledge.
[0248] In the final stage of fine-tuning, the entire model is comprehensively fine-tuned, including allowing updates to the parameters of all layers.
[0249] This allows the model to achieve optimal performance on the target task and fully realize its potential.
[0250] The entire process involves training with a large amount of data in the pre-training phase to acquire the ability to represent industry knowledge and enterprise characteristics; through targeted training and feature adaptation in the fine-tuning phase, the model can be adapted to specific target tasks; and through a phased unfreezing strategy, the potential of the model is gradually released to avoid overfitting, ultimately achieving the optimal performance of the model on the target task.
[0251] Network adaptation with anonymization:
[0252] Determine if the company's original data is within the last 15 days: if so, retain the reversible identifier. If not, anonymize the identifier.
[0253] Full Feature Vector / De-identified Feature Vector: For data retaining reversible identifiers, a full feature vector is generated. For data with anonymized identifiers, a de-identified feature vector is generated.
[0254] Full Feature Encoder / Desensitized Feature Encoder: A full feature vector is used to generate a full representation vector. A desensitized feature vector is used to generate a desensitized representation vector.
[0255] Feature representation adaptation: The complete representation vector and the desensitized representation vector are used for different application scenarios.
[0256] Model training and application: Training data update: Update and train the model using new data.
[0257] Real-time similarity calculation: Similarity calculation is performed in real time based on the latest data.
[0258] Historical pattern analysis: Analyzing patterns and regularities in historical data.
[0259] Anonymous model optimization: Optimize models using anonymized data to improve model accuracy and robustness.
[0260] Starting with raw enterprise data, the system determines whether to retain reversible identifiers or anonymize them by assessing the data's temporal attributes, thereby generating corresponding feature vectors. These feature vectors are processed by a feature encoder to generate representation vectors for model training and application. During the model training and application phases, comprehensive processing and utilization of enterprise data are achieved through continuous updates to training data, real-time similarity calculations, analysis of historical patterns, and optimization of the anonymization model.
[0261] To accommodate the short 15-day lifecycle of enterprise data, a special anonymization mechanism was designed in the network.
[0262] To achieve high-accuracy enterprise similarity prediction, a loss function with multi-objective joint optimization was designed:
[0263] Total loss function expression:
[0264] L_total=α*L_product+β*L_scope+γ*L_composite+δ*L_contrastive+ε*L_consistency;
[0265] Where: L_product: product similarity loss (mean squared error); L_scope: business scope similarity loss (mean squared error); L_composite: comprehensive similarity loss (mean squared error); L_contrastive: contrastive learning loss (enhancing the similarity of similar enterprise representations and reducing the similarity of dissimilar enterprise representations); L_consistency: feature consistency loss (ensuring the consistency of feature representations before and after anonymization).
[0266] Contrastive learning loss is used to enhance the model's discriminative ability, and the formula is as follows: Where: v i and v j It is the representation vector of similar firm pairs; v k is the representation vector of other companies in the batch; sim() is the cosine similarity function; τ is the temperature parameter, which controls the difficulty of comparative learning;
[0267] To ensure that anonymization does not affect model performance, a feature consistency loss is introduced: L_consistency = ||f_complete(x) - f_anonymized(x')|| 2 Where: f_complete is the feature encoder for complete data; f_anonymized is the feature encoder for anonymized data; x is the original enterprise data; x' is the anonymized enterprise data.
[0268] The model in this embodiment is trained using a multi-objective joint optimization strategy:
[0269] Gradient normalization: In multi-objective optimization, the loss functions for different tasks may have different scales and ranges, leading to gradients that are too large or too small for some tasks, affecting the overall optimization effect. Therefore, it is necessary to first normalize the gradients of each task to make the gradients of each task comparable.
[0270] Normalize gradients for each task: Apply the normalized gradients to each task to ensure that each task has the same weight and influence during the optimization process.
[0271] Proportional parameter updates: Based on the normalized gradient, the model parameters are updated proportionally to ensure that each task is treated fairly during the parameter update process and to prevent any one task from dominating the entire optimization process.
[0272] Dynamic weight adjustment: In multi-objective optimization, the importance of different tasks may change as training progresses. Therefore, it is necessary to dynamically adjust the weights of each task to adapt to changes in the importance and correlation between tasks.
[0273] Monitor the loss of each task: Monitor the loss value of each task in real time, understand the optimization progress and effect of each task, and provide a basis for dynamic weight adjustment.
[0274] Adaptive weight adjustment: Based on the monitored task loss, the weights of each task are adaptively adjusted so that the model can better balance the relationship between multiple tasks and improve the overall optimization effect.
[0275] Phased training: In multi-objective optimization, the entire training process can be divided into multiple phases, with each phase focusing on different tasks or combinations of tasks, gradually achieving joint optimization of all tasks.
[0276] Pre-training each subtask: Before formal joint optimization, each subtask can be pre-trained to give the model a certain initial performance for each task, laying the foundation for subsequent joint optimization.
[0277] Joint optimization of all tasks: Based on pre-training, a multi-objective optimization strategy is used to jointly optimize all tasks, thereby achieving synergistic improvement of each task and optimization of overall performance.
[0278] The entire multi-objective optimization strategy achieves joint optimization of multiple tasks through various methods such as gradient normalization, dynamic weight adjustment, and phased training. Starting with gradient normalization, it ensures that each task has the same weight and influence during the optimization process; through dynamic weight adjustment and monitoring of the loss of each task, it adaptively adjusts the weight relationship between tasks; finally, through phased training and pre-training of each subtask, it gradually achieves joint optimization of all tasks, reaching the optimal overall performance.
[0279] Specifically, gradient normalization addresses the issue of inconsistent gradient magnitudes across different loss functions.
[0280] Calculate the gradient for each subtask separately;
[0281] The gradients for each task are L2 normalized;
[0282] Gradient normalized by weight combination;
[0283] Dynamic weight adjustment: The weights of each task are dynamically adjusted based on the training progress;
[0284] Monitor the trend of loss changes in each subtask;
[0285] Increase the weight of a task when the rate of decrease in loss slows down.
[0286] Use exponentially decaying moving averages to assess task progress;
[0287] Phased training: First optimize each subtask individually, then optimize them together;
[0288] Phase 1: Pre-train the models for each sub-task separately;
[0289] Phase 2: Freeze some layers and jointly train all tasks;
[0290] Phase 3: Unfreeze all layers and fine-tune the entire model with a low learning rate;
[0291] Data from the most recent 15 days is selected from the identifiable dataset. This data typically contains the latest information and changes, reflecting the current situation and trends. Detailed feature training is then performed using this latest 15-day data to capture the most recent features and patterns.
[0292] Historical anonymized data was selected from the UUID anonymized dataset. This data underwent anonymization to protect user privacy while preserving historical information and patterns. General pattern training was performed using this historical anonymized data to learn long-standing patterns and regularities.
[0293] Knowledge transfer from detailed models to anonymous models: Through knowledge distillation techniques, knowledge and features learned in detailed models are transferred to anonymous models, enabling the anonymous models to also have strong predictive capabilities.
[0294] During model training, new data is continuously introduced, and through continuous learning, the model gradually adapts to new data distributions and changes, thereby improving the model's robustness and generalization ability.
[0295] Integrating predictions from detailed and anonymous models: This method combines the predictions from detailed and anonymous models, taking into account the advantages of both models to obtain more accurate and reliable prediction results.
[0296] The framework for the entire training data management and composite training strategy effectively trains and optimizes the model by making reasonable use of identifiable datasets and UUID anonymized datasets. Starting with detailed feature training from the latest 15 days of data and general pattern training from historical anonymized data, the system captures the latest features and long-standing patterns respectively. Furthermore, through knowledge distillation, continuous learning, and model fusion, the model's performance and reliability are further improved. Ultimately, this achieves effective prediction and decision support for complex scenarios.
[0297] Given the short 15-day lifespan of enterprise data, a special processing strategy is employed during the training process:
[0298] Dual-track training: a detailed feature model is trained using the latest 15 days of identifiable data; a general pattern model is trained using historical random UUID anonymized data.
[0299] Knowledge distillation: Distills knowledge from identifiable data models into UUID anonymous data models; ensuring that model performance does not significantly decrease after anonymization;
[0300] Continuous learning: Using a sliding window strategy to update training data; enabling the model to continuously adapt to changes in data distribution;
[0301] Model fusion: Integrates predictions from detailed feature models and UUID anonymization models; dynamically adjusts the weights of the two models based on data characteristics;
[0302] Based on the above technical solutions, a complete enterprise similarity calculation system architecture was constructed:
[0303] Data Acquisition Layer:
[0304] Basic Enterprise Information: Collect basic information about the enterprise, including name, address, contact information, etc.
[0305] Transaction record data: Collects transaction record data of enterprises, including transaction amount, transaction time, transaction counterparty, etc.
[0306] Business scope data: Collect data on the business scope of enterprises, including main business, product categories, etc.
[0307] Product sales data: Collects the company's product sales data, including sales volume, sales amount, sales regions, etc.
[0308] Qualification Certificate Data: Collect enterprise qualification certificate data, including business licenses, permits, etc.
[0309] 15-day timeliness assessment: If the data is within 15 days, it will be processed as identifiable data. If the data is after 15 days, it will be processed as anonymized data.
[0310] Feature engineering layer: Extracts detailed feature information from identifiable data for subsequent feature vector construction.
[0311] Anonymous feature extraction: Extracting anonymous feature information from anonymized data for subsequent feature vector construction.
[0312] The extracted detailed features and anonymous features are used to construct feature vectors, which are then used as input to the model.
[0313] By using transfer learning, existing knowledge and experience can be transferred to new tasks, thereby improving the training efficiency and performance of the model.
[0314] Calculate the similarity between different products based on their feature vectors.
[0315] Business scope similarity calculation: Based on the feature vector of business scope, calculate the similarity between the business scopes of different enterprises.
[0316] Overall similarity calculation: The overall similarity of enterprises is calculated by combining product similarity and business scope similarity.
[0317] Application Layer: Supply Chain Partner Recommendation: Recommends suitable supply chain partners based on enterprise similarity scores. Competitor Analysis: Analyzes a company's competitors to understand their strengths and weaknesses and formulates corresponding competitive strategies. Industry Ecosystem Building: Constructs an industry ecosystem to promote cooperation and win-win outcomes among enterprises.
[0318] The entire enterprise similarity scoring service framework begins with the data acquisition layer. Through data integration and cleaning, short lifecycle management, and feature engineering, detailed feature vectors and anonymized feature vectors are constructed. Then, through transfer learning models and similarity calculations at the model layer, a comprehensive similarity score for the enterprise is obtained. Finally, at the application layer, various services are provided, including supply chain cooperation recommendations, competitor analysis, and industry ecosystem building. This framework helps enterprises better understand and analyze the market environment and formulate scientific and reasonable business strategies.
[0319] Based on the evaluation on the test set, the method of this embodiment achieved the following performance in enterprise similarity calculation, as shown in Table 6.
[0320] Table 6. Performance
[0321]
[0322] The data protection effectiveness of the short lifecycle processing mechanism is shown in Table 7.
[0323] Table 7. Data Protection Effectiveness of Short Lifetime Processing Mechanisms
[0324] Evaluation index Before processing After processing Enterprise identification rate 98.7% <0.1% Model performance degradation - <3.5% Feature information retention rate 100% 92.3%
[0325] Table 8 shows a comparison between this method and existing technical solutions.
[0326] Table 8. Comparison Results
[0327]
[0328]
[0329] The method in this embodiment is applicable to the following various application scenarios:
[0330] Supply chain partner recommendation: Helps companies find potential supply chain partners by calculating the similarity between companies to recommend the most suitable partners.
[0331] Competitor analysis: Identify competitors within the same industry to provide companies with competitive intelligence and market positioning advice.
[0332] Enterprise profiling: Enrich enterprise profiles by leveraging the characteristic information of similar enterprises to improve understanding of enterprise business models and market strategies.
[0333] Loan risk assessment: By comparing the financial condition and operating performance of similar companies, the risk level of the loan applicant company is assessed.
[0334] Precise marketing targeting: Identify businesses with similar characteristics to existing customers to improve the targeting and efficiency of marketing campaigns.
[0335] The method in this embodiment achieves high-precision enterprise similarity calculation while protecting enterprise data privacy. By setting a 15-day lifecycle for enterprise data and employing a deep learning model with transfer learning and multi-objective optimization, this technology not only meets immediate business needs but also ensures long-term data privacy and security. This method provides strong support for supply chain cooperation, competitor analysis, and other areas.
[0336] Compared with the prior art, the method of this embodiment has the following advantages:
[0337] Balancing privacy protection with business needs: A 15-day anonymization mechanism ensures accurate short-term identification while maintaining long-term data privacy. Tests show that the enterprise identity verification rate decreased from 98.7% to less than 0.1% after processing.
[0338] Multi-dimensional feature fusion improves accuracy: By fusing multi-dimensional features (such as basic information, business scope, and product sales distribution), the accuracy of similarity calculation is improved. Tests show that the overall similarity accuracy reaches 91.7%, an improvement of 9.4% compared to traditional methods.
[0339] Differentiated processing enhances adaptability: Differentiated feature engineering and similarity calculation strategies are designed for different types of enterprises, which improves adaptability to different enterprise types.
[0340] Clear and compliant authorization boundaries: Clearly distinguish the labeling rules between authorized and unauthorized enterprises to ensure the compliant operation of the system.
[0341] Adaptive learning reduces human intervention: By employing transfer learning and multi-objective optimization models, human intervention is reduced while maintaining model performance.
[0342] Significantly improved computational efficiency: Through parallel computing and model compression techniques, the number of samples processed per second has been greatly increased, while reducing the consumption of computing resources.
[0343] Efficient preservation of feature information: Even after anonymization, most feature information can still be retained, supporting continuous model optimization.
[0344] Reduce system maintenance costs: Automatically adapts to new data distributions and industry characteristics, reducing the cost of manual maintenance.
[0345] This embodiment proposes a 15-day enterprise data lifecycle management mechanism, achieving a balance between data availability and privacy protection through dynamic identifier transformation. Different similarity definitions and labeling rules are designed based on different enterprise types to ensure effective calculation of similarity across various enterprise types. A comprehensive enterprise feature extraction and transformation method is constructed to capture the characteristics of various aspects of an enterprise. Industry knowledge pre-training models and enterprise representation pre-training models are introduced to improve the model's adaptability. A multi-objective loss function is designed to jointly optimize multiple similarity indicators, enhancing the model's discriminative ability and robustness. The impact of anonymized data on model training is addressed, ensuring the system can continuously learn and optimize. A dedicated module is designed to output multiple similarity indicators to meet the needs of different business scenarios.
[0346] In one embodiment, an alternative to short-lifetime anonymization mechanisms is:
[0347] Use differential privacy technology instead of simple time-based anonymization; adjust the lifecycle duration to adapt to different business needs; and adopt a federated learning framework for distributed computing.
[0348] Alternatives to enterprise classification and labeling mechanisms:
[0349] We refine enterprise types and adopt more detailed classification standards; we combine a small number of manually labeled samples with a large number of unlabeled samples for semi-supervised learning; and we adopt an active learning strategy to improve labeling efficiency.
[0350] Alternatives to feature engineering systems:
[0351] Automatic feature extraction using unsupervised learning methods; enhancement of feature representation by introducing external knowledge graphs; and direct modeling of the transaction network using graph neural networks.
[0352] Alternatives to network architecture:
[0353] Replace existing multilayer perceptron networks with Transformer architecture; replace traditional neural networks with graph neural networks; and replace deep learning models with traditional machine learning methods such as gradient boosting trees.
[0354] Alternatives to the multi-objective optimization framework:
[0355] Evolutionary algorithms are used instead of gradient descent to optimize multi-objective problems; a hierarchical training strategy is employed to optimize primary and secondary objectives; and a multi-task learning framework is used for joint optimization.
[0356] Alternatives to short-lifecycle data learning strategies:
[0357] Update the model using incremental learning methods; improve the model's ability to adapt to new tasks using meta-learning techniques; and design a data synthesis mechanism to expand the training set.
[0358] Alternatives to enterprise similarity calculation:
[0359] This method uses metric learning to learn distance metrics in the enterprise representation space; it employs an attention mechanism to dynamically weight the importance of different features; and it combines expert rules with machine learning models to balance interpretability and accuracy.
[0360] Alternatives to model deployment and services:
[0361] Model quantization techniques are used to further compress model size; edge computing deployment is used to reduce the load on central servers; and an automatic model update mechanism is designed to reduce manual intervention.
[0362] Alternatives to anonymous data storage:
[0363] Blockchain technology is used to ensure data immutability; homomorphic encryption technology is applied to allow similarity calculations in an encrypted state; and data fragmentation storage is implemented to reduce the risk of leakage.
[0364] Alternatives to enterprise feature representation:
[0365] Hierarchical feature representation is used; temporal features are introduced to capture dynamic changes in enterprise operations; and latent features are extracted based on natural language processing technology.
[0366] These alternatives can be flexibly selected based on specific application scenarios and technological developments, and combined with the method of this embodiment to further improve the accuracy, efficiency, and security of enterprise similarity calculation.
[0367] The aforementioned short-cycle anonymization method for calculating enterprise similarity comprehensively captures enterprise characteristics by acquiring data from multiple enterprises, extracting features, and constructing feature vectors. It assesses the duration of identifiable features within the feature vectors and processes them according to a pre-defined short-cycle anonymization mechanism to protect data privacy while retaining sufficient feature information for subsequent analysis. This approach ensures privacy while meeting business needs. Differentiated feature engineering and similarity calculation strategies are implemented for different types of enterprises to improve the accuracy and adaptability of the analysis. Clear distinctions are made between authorized and unauthorized enterprises' data usage rules to ensure compliant operation. An adaptive learning mechanism reduces manual intervention, enabling the system to automatically adapt to new data distributions and industry characteristics. By inputting the anonymized features and processing results into the similarity calculation model, the final output achieves efficient and secure assessment of inter-enterprise similarity, effectively manages data timeliness, and enhances the system's practicality and maintenance efficiency.
[0368] Figure 3 This is a schematic block diagram of a short-cycle anonymization enterprise similarity calculation device 300 provided in an embodiment of the present invention. Figure 3 As shown, corresponding to the above-described short-cycle anonymization enterprise similarity calculation method, the present invention also provides a short-cycle anonymization enterprise similarity calculation apparatus 300. This short-cycle anonymization enterprise similarity calculation apparatus 300 includes a unit for performing the above-described short-cycle anonymization enterprise similarity calculation method, and the apparatus can be configured in a server. Specifically, please refer to... Figure 3 The short-cycle anonymized enterprise similarity calculation device 300 includes a data acquisition unit 301, a feature vector construction unit 302, a processing unit 303, a calculation unit 304, and an output unit 305.
[0369] The data acquisition unit 301 is used to acquire data from multiple enterprises; the feature vector construction unit 302 is used to extract features and construct feature vectors for each enterprise data to obtain feature vectors; the processing unit 303 is used to determine the duration of identifiable features in the feature vectors and process the corresponding feature vectors according to the determination results to obtain processing results; the calculation unit 304 is used to input the anonymized features in the feature vectors and the processing results into a similarity calculation model for similarity calculation to obtain calculation results; and the output unit 305 is used to output the calculation results.
[0370] In one embodiment, the feature vector construction unit 302 includes:
[0371] Extraction subunits are used to extract different categories of features from each of the enterprise data; construction subunits are used to process, fuse, select features, and perform principal component analysis on the different categories of features to obtain feature vectors.
[0372] In one embodiment, the processing unit 303 includes:
[0373] The system includes a judgment subunit for determining whether the lifecycle of an identifiable feature in the feature vector exceeds a set number of days; a first processing subunit for anonymizing the identifiable feature to generate a random identifier if the lifecycle of the identifiable feature in the feature vector exceeds the set number of days, thus forming a processing result; and a second processing subunit for reversibly encrypting and storing the enterprise identifier of the identifiable feature if the lifecycle of the identifiable feature in the feature vector does not exceed the set number of days, thus obtaining a processing result.
[0374] In one embodiment, the computing unit 304 includes:
[0375] The input subunit is used to input the anonymized features from the feature vector and the processing result into the similarity calculation model; the preprocessing subunit is used to preprocess the anonymized features from the feature vector and the processing result through the corresponding encoder to obtain the encoded features; the integration subunit is used to integrate all the encoded features in the feature fusion layer to form a unified feature representation to obtain the fused features; the representation processing subunit is used to process the fused features sequentially through a fully connected layer containing 512 neurons, a batch normalization layer, and to introduce nonlinear processing using the LeakyReLU activation function, and to apply Dropout technology to randomly discard neurons, and then to extract features through a fully connected layer containing 256 neurons, and then to perform batch normalization again. First, the LeakyReLU activation function is applied again, and Dropout is used to randomly discard neurons to obtain a representation vector. A computation subunit is used to perform multi-type similarity calculations on the representation vector to obtain various similarity features. A concatenation subunit is used to concatenate the various similarity features to form a comprehensive similarity feature vector. A similarity calculation subunit is used to input the similarity feature vector into a fully connected layer containing 128 neurons, apply the ReLU activation function to the output of the fully connected layer for non-linear transformation, apply Dropout to randomly discard neurons, and then pass it through a fully connected layer containing 64 neurons for feature extraction and compression, and finally process it again using the ReLU activation function to obtain the calculation result.
[0376] In one embodiment, the computing subunit includes:
[0377] The first calculation module is used to calculate the absolute difference element-by-element of the representation vectors of multiple enterprises to obtain the absolute difference; the second calculation module is used to calculate the cosine similarity between the representation vectors of multiple enterprises to obtain the cosine similarity; the third calculation module is used to perform the Hadamard product of the representation vectors of multiple enterprises to obtain the Hadamard product; wherein, the multiple similarity features include absolute difference, cosine similarity and Hadamard product.
[0378] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the aforementioned short-cycle anonymization enterprise similarity calculation device 300 and its various units can be found in the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity, these details will not be repeated here.
[0379] The aforementioned short-cycle anonymization enterprise similarity calculation device 300 can be implemented as a computer program, which can, for example... Figure 4 It runs on the computer device shown.
[0380] Please seeFigure 4 , Figure 4 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0381] See Figure 4 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0382] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a short-cycle anonymized enterprise similarity calculation method.
[0383] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0384] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a short-cycle anonymized enterprise similarity calculation method.
[0385] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0386] The processor 502 is used to run the computer program 5032 stored in the memory to implement all the steps of the above-described short-cycle anonymization enterprise similarity calculation method.
[0387] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU) 303. The processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0388] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0389] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform all the steps of the above-described short-cycle anonymization enterprise similarity calculation method.
[0390] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0391] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0392] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0393] The steps in the method of this invention can be adjusted, merged, or deleted according to actual needs. The units in the device of this invention can be merged, divided, or deleted according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit 303, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0394] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0395] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for calculating enterprise similarity using short-cycle anonymization, characterized in that, include: Acquire data from multiple enterprises; For each of the enterprise data, feature extraction and feature vector construction are performed to obtain a feature vector; The duration of identifiable features in the feature vector is determined, and the corresponding feature vector is processed according to the determination result to obtain the processing result; The anonymized features in the feature vector and the processing result are input into the similarity calculation model to calculate the similarity and obtain the calculation result. Output the calculation results.
2. The method for calculating enterprise similarity with short-cycle anonymization according to claim 1, characterized in that, The step of extracting features and constructing feature vectors for each of the enterprise data to obtain feature vectors includes: Extract different categories of features from the data of each of the aforementioned enterprises; The features of different categories are processed, fused, selected, and subjected to principal component analysis to obtain feature vectors.
3. The method for calculating enterprise similarity with short-cycle anonymization according to claim 1, characterized in that, The step of determining the duration of identifiable features in the feature vector and processing the corresponding feature vector based on the determination result to obtain a processing result includes: Determine whether the lifecycle of the identifiable features in the feature vector exceeds a set number of days; If the lifetime of an identifiable feature in the feature vector exceeds a set number of days, the identifiable feature is anonymized to generate a random identifier, thus forming the processing result. If the lifecycle of the identifiable features in the feature vector does not exceed a set number of days, then the enterprise identifier is stored in reversible encryption on the identifiable features to obtain the processing result.
4. The method for calculating enterprise similarity with short-cycle anonymization according to claim 1, characterized in that, The similarity calculation model includes a feature encoder and a similarity calculation module; The feature encoder includes a neural network structure, which includes a fully connected layer, a batch normalization layer, a LeakyReLU activation function, and a Dropout layer. The similarity calculation module includes a similarity feature calculation layer, a fully connected layer, a ReLU activation function, and a Dropout layer.
5. The method for calculating enterprise similarity with short-cycle anonymization according to claim 4, characterized in that, The step of inputting the anonymized features in the feature vector and the processing result into a similarity calculation model for similarity calculation to obtain the calculation result includes: The anonymized features in the feature vector and the processing result are input into the similarity calculation model; The anonymized features in the feature vector and the processing result are preprocessed using the corresponding encoder to obtain the encoded features; All encoded features are integrated in the feature fusion layer to form a unified feature representation, thus obtaining fused features; The fused features are sequentially processed through a fully connected layer containing 512 neurons, a batch normalization layer, and nonlinear processing is introduced using the LeakyReLU activation function. Dropout technology is applied to randomly discard neurons. The features are then extracted through a fully connected layer containing 256 neurons, batch normalization is performed again, and the LeakyReLU activation function is applied again. Dropout technology is applied to randomly discard neurons to obtain the representation vector. The representation vector is subjected to multi-type similarity calculations to obtain various similarity features; The various similarity features are concatenated to form a comprehensive similarity feature vector; The similarity feature vector is input into a fully connected layer containing 128 neurons. The ReLU activation function is used to perform a non-linear transformation on the output of the fully connected layer. The Dropout technique is applied to randomly discard neurons. The output is then processed through a fully connected layer containing 64 neurons for feature extraction and compression. The ReLU activation function is used again to obtain the calculation result.
6. The method for calculating enterprise similarity with short-cycle anonymization according to claim 5, characterized in that, The process of performing multi-type similarity calculations on the representation vector to obtain various similarity features includes: The absolute difference is calculated element-by-element for the representation vectors of multiple enterprises to obtain the absolute difference. Calculate the cosine similarity between the representation vectors of multiple enterprises to obtain the cosine similarity. Perform the Hadamard product of the representation vectors of the multiple enterprises to obtain the Hadamard product; The various similarity features include absolute difference, cosine similarity, and Hadamard product.
7. The method for calculating enterprise similarity with short-cycle anonymization according to claim 1, characterized in that, The total loss function used in training the similarity calculation model is a weighted combination of product similarity loss, business scope similarity loss, comprehensive similarity loss, contrastive learning loss, and feature consistency loss, in order to evaluate and optimize the accuracy and consistency of the model in calculating similarity between enterprises.
8. The method for calculating enterprise similarity with short-cycle anonymization according to claim 1, characterized in that, During the training process of the similarity calculation model, dual-track training is adopted to ensure that the model adapts to the latest and historical data features simultaneously. Knowledge distillation technology is combined to transfer the knowledge of the identified data to the anonymized data. The dataset is updated through continuous learning and sliding window strategies, and model fusion technology is used to dynamically adjust the prediction weights of different models to adapt to changes in data distribution.
9. A short-cycle anonymization enterprise similarity calculation device, characterized in that, include: The data acquisition unit is used to acquire data from multiple enterprises; The feature vector construction unit is used to extract features and construct feature vectors for each of the enterprise data to obtain feature vectors; The processing unit is used to determine the duration of identifiable features in the feature vector and process the corresponding feature vector according to the determination result to obtain the processing result; The calculation unit is used to input the anonymized features in the feature vector and the processing result into the similarity calculation model to calculate the similarity and obtain the calculation result. The output unit is used to output the calculation results.
10. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 7.