Market subject multi-label data fusion method and system

By preprocessing and similarity measurement of enterprise label data and shared data, and using hybrid distance calculation algorithms to achieve cross-industry and cross-departmental integration of data among market entities, the problem of limited data circulation in the existing technology is solved and the accuracy and efficiency of data fusion is improved.

CN119939508APending Publication Date: 2025-05-06天元大数据信用管理有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510019527.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology lacks effective multi-label data fusion methods in the process of open sharing of public data, resulting in limited circulation and application of data across industries and departments.

Method used

By extracting enterprise tag data and sharing data, an enterprise tag set and data set are formed, pre-processed and similarity measurements are performed, and the correlation and similarity between market entities are quantified by using a hybrid distance calculation algorithm to achieve cross-industry and cross-departmental integration of data.

Benefits of technology

It realizes the circulation of data across industries and departments, accelerates data sharing and exchange, improves the accuracy and efficiency of data fusion, and provides a solid data foundation for intelligent decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939508A_ABST
    Figure CN119939508A_ABST
Patent Text Reader

Abstract

The invention discloses a market subject multi-label data fusion method and system, and relates to the technical field of big data analysis and mining. Comprising the steps of 1, extracting enterprise label data, forming an enterprise label set according to enterprise numbers, extracting enterprise shared data, forming an enterprise data set, and combining the enterprise label set and the enterprise data set to form a new data set; 2, preprocessing the data in the new data set; 3, similarity measurement is carried out on the data in the preprocessed new data set, the data types of the data in the new data set are judged, the data types comprise a nominal attribute type, a binary attribute type and a numerical attribute type, and the similarity of the data of different data types in the new data set is calculated respectively; and 4, setting a threshold value of each data type according to a data application scene, and selecting two pieces of data with the distance meeting the threshold value and the highest similarity from the new data set for data fusion until fusion of the data capable of being fused is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a method and system for fusing multi-label data of market entities, and relates to the technical field of big data analysis and mining. Background Art

[0002] The research on existing data fusion technology has greatly promoted the field of multimodal data fusion, but the research on the field of data fusion technology in the current process of public data openness and sharing is relatively insufficient. The main reason is that the field of public data governance and authorized operation is a relatively segmented market. The current public data sharing and opening is developing in depth, and new data and technologies are constantly emerging. In the case of segmented markets, there is no perfect fusion method for multi-label data of market players. Summary of the invention

[0003] In view of the problems of the prior art, the present invention provides a method and system for fusing multi-label data of market entities, quantifies the complex and multi-dimensional correlations and differences between market entities, accurately calculates the similarities between market entities, thereby accurately identifying potential consistent relationships, and fusing market entities that meet specific thresholds, thereby realizing the mutual application of data across industries and departments.

[0004] The specific scheme proposed by the present invention is:

[0005] The present invention also provides a method for fusing multi-label data of market entities, comprising:

[0006] Step 1: Extract enterprise label data, form enterprise label sets according to enterprise numbers, extract enterprise shared data to form enterprise data sets, merge enterprise label sets and enterprise data sets to form new data sets;

[0007] Step 2: Preprocess the data in the new dataset;

[0008] Step 3: Measure the similarity of the preprocessed data in the new dataset:

[0009] Determine the data type of the data in the new data set. The data types include: nominal attribute type, binary attribute type, and numerical attribute type.

[0010] Calculate the similarity of data of different data types in the new data set respectively;

[0011] Step 4: Set the threshold of each data type according to the data application scenario, and select two data with the highest similarity and a distance that meets the threshold in the new data set for data fusion until all the data that can be fused are fused.

[0012] Further, the enterprise label data extracted in step 1 of the method for fusion of multi-label data of market entities includes enterprise name, unified social credit code, enterprise name MD5 code, unified social credit code MD5 code, organization code, tax registration certificate number, registered capital, legal representative, type, establishment date, business address, primary industry, secondary industry, tertiary industry, quaternary industry and registration authority,

[0013] Form enterprise label sets according to enterprise numbers:

[0014] Enterprise 1 = [a 11 ,a 12 ,...,a 1n ], Enterprise 2 = [a 21 ,a 22 ,...,a 2n ],

[0015] Until the enterprise m = [a m1 ,a m2 ,...,a mn ], m and n are both positive integers.

[0016] Furthermore, in step 1 of the method for fusion of multi-label data of market entities, the enterprise data set is formed by comparing the enterprise numbers, including: all missing enterprise labels are set to null values, and the enterprise data set is formed as follows:

[0017] Enterprise 1 = [b 11 ,b 12 ,...,b 1t ], Enterprise 2 = [b 21 ,b 22 ,...,b 2t ],

[0018] Until enterprise s = [b s1 ,b s2 ,...,b st ], s and t are both positive integers.

[0019] Furthermore, step 2 of the method for fusion of multi-label data of market entities includes:

[0020] For the data in the new dataset, remove the features with single values.

[0021] For the data in the new dataset, remove the attributes with low variance.

[0022] Furthermore, in step 3 of the method for fusion of multi-label data of market entities, the similarity of data of different data types in the new data set is calculated respectively, including:

[0023] The feature vectors of two objects i and j are expressed as:

[0024] X i =[X i1 ,X i2 ,……X ip ]

[0025] X j =[X j1 ,X j2 ,……X jp ]

[0026] The similarity calculation formula 1 of the data type is expressed as:

[0027]

[0028] When the data type is a numeric attribute type, Formula 1 can be expressed as:

[0029]

[0030] When the data type is binary attribute or missing, x in Formula 2 if or x jf Missing, or x if =x jf =0, then in formula 1 otherwise

[0031] When the data type is a nominal attribute type, the nominal attribute is vectorized, and the similarity distance is calculated using cosine similarity to obtain the similarity, which can be expressed as:

[0032]

[0033] After obtaining the similarity, it is used for data fusion.

[0034] The present invention also provides a system for fusion of multi-label data of market entities, including a data set construction module, a preprocessing module, a calculation and analysis module, and a fusion module.

[0035] The data set construction module extracts enterprise label data, forms enterprise label sets according to enterprise numbers, extracts enterprise shared data to form enterprise data sets, and merges enterprise label sets and enterprise data sets to form new data sets;

[0036] The preprocessing module preprocesses the data in the new data set;

[0037] The calculation and analysis module measures the similarity of the data in the preprocessed new data set:

[0038] Determine the data type of the data in the new data set. The data types include: nominal attribute type, binary attribute type, and numerical attribute type.

[0039] Calculate the similarity of data of different data types in the new data set respectively;

[0040] The fusion module sets the threshold of each data type according to the data application scenario, and selects two data with the highest similarity and a distance that meets the threshold in the new data set for data fusion until all the data that can be fused are fused.

[0041] Furthermore, the enterprise label data extracted by the data set construction module of the system for multi-label data fusion of market entities includes enterprise name, unified social credit code, enterprise name MD5 code, unified social credit code MD5 code, organization code, tax registration certificate number, registered capital, legal representative, type, establishment date, business address, first-level industry, second-level industry, third-level industry, fourth-level industry and registration authority,

[0042] Form enterprise label sets according to enterprise numbers:

[0043] Enterprise 1 = [a 11 ,a 12 ,...,a 1n ], Enterprise 2 = [a 21 ,a 22 ,...,a 2n ],

[0044] Until the enterprise m = [a m1 ,a m2 ,...,a mn ], m and n are both positive integers.

[0045] Furthermore, the data set construction module of the system for multi-label data fusion of market entities forms an enterprise data set by comparing the enterprise number, including: setting all missing enterprise labels to null values ​​to form the enterprise data set as follows:

[0046] Enterprise 1 = [b 11 ,b 12 ,...,b 1t ], Enterprise 2 = [b 21 ,b 22 ,...,b 2t ],

[0047] Until enterprise s = [b s1 ,b s2 ,...,b st ], s and t are both positive integers.

[0048] Furthermore, the system for multi-label data fusion of market entities has a preprocessing module that removes features with single values ​​from data in a new data set, and removes attributes with low variance from data in the new data set.

[0049] Furthermore, the computing and analysis modules of the system for fusion of multi-label data of market entities respectively calculate the similarity of data of different data types in the new data set, including:

[0050] The feature vectors of two objects i and j are expressed as:

[0051] X i =[X i1 ,X i2 ,……X ip ]

[0052] X j =[X j1 ,X j2 ,……X jp ]

[0053] The similarity calculation formula 1 of the data type is expressed as:

[0054]

[0055] When the data type is a numeric attribute type, Formula 1 can be expressed as:

[0056]

[0057] When the data type is binary attribute or missing, x in Formula 2 if or x jf Missing, or x if =x jf =0, then in formula 1 otherwise

[0058] When the data type is a nominal attribute type, the nominal attribute is vectorized, and the similarity distance is calculated using cosine similarity to obtain the similarity, which can be expressed as:

[0059]

[0060] After obtaining the similarity, it is used for data fusion.

[0061] The benefits of the present invention are specifically embodied in the following aspects:

[0062] Promoting data circulation: By introducing a multi-label system for market entities, the present invention breaks down the barriers caused by different standards and mismatched fields in the traditional data fusion process, realizes the circulation of data across industries and departments, accelerates data sharing and exchange, and promotes the release of data value as a new production factor;

[0063] Improve the accuracy and efficiency of data fusion: By creatively using the hybrid distance calculation algorithm between samples, the present invention can quantify the complex and multi-dimensional correlation and consistency between market entities, and accurately calculate the similarity between market entities. In the practice of different scenarios, it can more accurately identify potential consistent relationships, significantly improving the accuracy and efficiency of data fusion.

[0064] Promote intelligent decision-making in the field of data application: In the process of data application, the label and quantity of samples are a key link in intelligent application. The present invention can break down data barriers, integrate more diverse data labels, data features and data volumes, and provide a solid data foundation for intelligent decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a schematic flow chart of the method of the present invention. DETAILED DESCRIPTION

[0066] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.

[0067] Example 1

[0068] The present invention also provides a method for fusing multi-label data of market entities, comprising:

[0069] Step 1: Extract enterprise label data, form an enterprise label set according to the enterprise number, extract enterprise shared data to form an enterprise data set, merge the enterprise label set and the enterprise data set to form a new data set.

[0070] The enterprise tag data extracted in step 1 includes enterprise name, unified social credit code, enterprise name MD5 code, unified social credit code MD5 code, organization code, tax registration certificate number, registered capital, legal representative, type, establishment date, business address, first-level industry, second-level industry, third-level industry, fourth-level industry and registration authority.

[0071] Form enterprise label sets according to enterprise numbers:

[0072] Enterprise 1 = [a 11 ,a 12 ,...,a 1n ], Enterprise 2 = [a 21 ,a22 ,...,a 2n ],

[0073] Until the enterprise m = [a m1 ,a m2 ,...,a mn ], m and n are both positive integers.

[0074] The enterprise data set is formed by comparing the enterprise numbers, including: all missing enterprise labels are set to null values, and the enterprise data set is formed as follows:

[0075] Enterprise 1 = [b 11 ,b 12 ,...,b 1t ], Enterprise 2 = [b 21 ,b 22 ,...,b 2t ],

[0076] Until enterprise s = [b s1 ,b s2 ,...,b st ], s and t are both positive integers.

[0077] When merging data, assume that the samples in the enterprise label set have n features, the samples in the enterprise data set have p features, assume that all samples have t features in total (t<=n+p), and mark the attributes with empty values ​​as missing. Merge the enterprise label set and the enterprise data set into a new data set containing m+s objects and t features.

[0078] Step 2: Preprocess the data in the new dataset.

[0079] In step 2, the preprocessing may specifically include:

[0080] For the data in the new dataset, remove the features with single values. Single-valued variables have constant values ​​in the entire dataset, such as all null values, or the same value, and do not contribute to the data matching model.

[0081] For the data in the new data set, remove the attributes with low variance. Similar to variables with a single value, the values ​​of low variance variables are not unique, but the overall data changes very little, and the discrimination for enterprise identification is not great. This type of variable can be removed by setting a screening threshold.

[0082] Step 3: Measure the similarity of the preprocessed data in the new dataset:

[0083] Determine the data type of the data in the new data set. The data types include: nominal attribute type, binary attribute type, and numerical attribute type.

[0084] The similarity of data of different data types in the new data set is calculated respectively.

[0085] These may specifically include:

[0086] The feature vectors of two objects i and j are expressed as:

[0087] X i =[X i1 ,X i2 ,……X ip ]

[0088] X j =[X j1 ,X j2 ,……X jp ]

[0089] The similarity calculation formula 1 of the data type is expressed as:

[0090]

[0091] When the data type is a numeric attribute type, Formula 1 can be expressed as:

[0092]

[0093] When the data type is binary attribute or missing, x in Formula 2 if or x jf Missing, or x if =x jf =0, then in formula 1 otherwise

[0094] When the data type is a nominal attribute type, the nominal attribute is vectorized, and the similarity distance is calculated using cosine similarity to obtain the similarity, which can be expressed as:

[0095]

[0096] After obtaining the similarity, it is used for data fusion.

[0097] Step 4: Set the threshold of each data type according to the data application scenario, and select two data with the highest similarity and a distance that meets the threshold in the new data set for data fusion until all the data that can be fused are fused.

[0098] Example 2

[0099] The present invention also provides a system for fusion of multi-label data of market entities, including a data set construction module, a preprocessing module, a calculation and analysis module, and a fusion module.

[0100] The data set construction module extracts enterprise label data, forms enterprise label sets according to enterprise numbers, extracts enterprise shared data to form enterprise data sets, and merges enterprise label sets and enterprise data sets to form new data sets;

[0101] The preprocessing module preprocesses the data in the new data set;

[0102] The calculation and analysis module measures the similarity of the data in the preprocessed new data set:

[0103] Determine the data type of the data in the new data set. The data types include: nominal attribute type, binary attribute type, and numerical attribute type.

[0104] Calculate the similarity of data of different data types in the new data set respectively;

[0105] The fusion module sets the threshold of each data type according to the data application scenario, and selects two data with the highest similarity and a distance that meets the threshold in the new data set for data fusion until all the data that can be fused are fused.

[0106] As the contents of information interaction between modules in the above-mentioned device, execution of readable program process, etc. are based on the same concept as the embodiment of the method of the present invention, the specific contents can be found in the description of the embodiment of the method of the present invention, and will not be repeated here.

[0107] Similarly, the benefits of the system of the present invention are specifically embodied in the following aspects:

[0108] Promoting data circulation: By introducing a multi-label system for market entities, the present invention breaks down the barriers caused by different standards and mismatched fields in the traditional data fusion process, realizes the circulation of data across industries and departments, accelerates data sharing and exchange, and promotes the release of data value as a new production factor;

[0109] Improve the accuracy and efficiency of data fusion: By creatively using the hybrid distance calculation algorithm between samples, the present invention can quantify the complex and multi-dimensional correlation and consistency between market entities, and accurately calculate the similarity between market entities. In the practice of different scenarios, it can more accurately identify potential consistent relationships, significantly improving the accuracy and efficiency of data fusion.

[0110] Promote intelligent decision-making in the field of data application: In the process of data application, the label and quantity of samples are a key link in intelligent application. The present invention can break down data barriers, integrate more diverse data labels, data features and data volumes, and provide a solid data foundation for intelligent decision-making.

[0111] It should be noted that not all steps and modules in the above-mentioned processes and system structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices may be implemented together.

[0112] The above-described embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or changes made by those skilled in the art based on the present invention are within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A method for fusing multi-label data of market entities, characterized by include: Step 1: Extract enterprise label data, form enterprise label sets according to enterprise numbers, extract enterprise shared data to form enterprise data sets, merge enterprise label sets and enterprise data sets to form new data sets; Step 2: Preprocess the data in the new dataset; Step 3: Measure the similarity of the preprocessed data in the new dataset: Determine the data type of the data in the new data set. The data types include: nominal attribute type, binary attribute type, and numerical attribute type. Calculate the similarity of data of different data types in the new data set respectively; Step 4: Set the threshold of each data type according to the data application scenario, and select two data with the highest similarity and a distance that meets the threshold in the new data set for data fusion until all the data that can be fused are fused.

2. According to claim 1, a method for fusion of multi-label data of market entities is characterized by: The enterprise tag data extracted in step 1 includes enterprise name, unified social credit code, enterprise name MD5 code, unified social credit code MD5 code, organization code, tax registration certificate number, registered capital, legal representative, type, establishment date, business address, first-level industry, second-level industry, third-level industry, fourth-level industry and registration authority, Form enterprise label sets according to enterprise numbers: Enterprise 1 = [a 11 ,a 12 ,...,a 1n ], Enterprise 2 = [a 21 ,a 22 ,...,a 2n ], Until the enterprise m = [a m1 ,a m2 ,...,a mn ], m and n are both positive integers.

3. The method for fusion of multi-label data of market entities according to claim 1 is characterized by: In step 1, the enterprise data set is formed by comparing the enterprise numbers, including: all missing enterprise labels are set to null values, and the enterprise data set is formed as follows: Enterprise 1 = [b 11 ,b 12 ,...,b 1t ], Enterprise 2 = [b 21 ,b 22 ,...,b 2t ], Until enterprise s = [b s1 ,b s2 ,...,b st ], s and t are both positive integers.

4. The method for fusion of multi-label data of market entities according to claim 1 is characterized in that Step 2 includes: For the data in the new dataset, remove the features with single values. For the data in the new dataset, remove the attributes with low variance.

5. The method for fusion of multi-label data of market entities according to claim 1 is characterized by: In step 3, the similarity of data of different data types in the new data set is calculated separately, including: The feature vectors of two objects i and j are expressed as: X i =[X i1 ,X i2 ,……X ip ] X j =[X j1 ,X j2 ,……X jp ] The similarity calculation formula 1 of the data type is expressed as: When the data type is a numeric attribute type, Formula 1 can be expressed as: When the data type is binary attribute or missing, x in Formula 2 if or x jf Missing, or x if =x jf =0, then in formula 1 otherwise When the data type is a nominal attribute type, the nominal attribute is vectorized, and the similarity distance is calculated using cosine similarity to obtain the similarity, which can be expressed as: After obtaining the similarity, it is used for data fusion.

6. A system for fusion of multi-label data of market entities, characterized by Including data set construction module, preprocessing module, calculation analysis module, fusion module, The data set construction module extracts enterprise label data, forms an enterprise label set according to the enterprise number, extracts enterprise shared data to form an enterprise data set, and merges the enterprise label set and the enterprise data set to form a new data set; The preprocessing module preprocesses the data in the new data set; The calculation and analysis module measures the similarity of the data in the preprocessed new data set: Determine the data type of the data in the new data set. The data types include: nominal attribute type, binary attribute type, and numerical attribute type. Calculate the similarity of data of different data types in the new data set respectively; The fusion module sets the threshold of each data type according to the data application scenario, and selects two data with the highest similarity and a distance that meets the threshold in the new data set for data fusion until all the data that can be fused are fused.

7. The market subject multi-label data fusion system according to claim 1 is characterized by: The enterprise label data extracted by the dataset construction module includes enterprise name, unified social credit code, enterprise name MD5 code, unified social credit code MD5 code, organization code, tax registration certificate number, registered capital, legal representative, type, establishment date, business address, first-level industry, second-level industry, third-level industry, fourth-level industry and registration authority. Form enterprise label sets according to enterprise numbers: Enterprise 1 = [a 11 ,a 12 ,...,a 1n ], Enterprise 2 = [a 21 ,a 22 ,...,a 2n ], Until the enterprise m = [a m1 ,a m2 ,...,a mn ], m and n are both positive integers.

8. The method for fusion of multi-label data of market entities according to claim 1 is characterized by: The data set construction module forms an enterprise data set by comparing the enterprise numbers, including: setting all missing enterprise labels to null values, forming the enterprise data set as follows: Enterprise 1 = [b 11 ,b 12 ,...,b 1t ], Enterprise 2 = [b 21 ,b 22 ,...,b 2t ], Until enterprise s = [b s1 ,b s2 ,...,b st ], s and t are both positive integers.

9. The system for fusion of multi-label data of market entities according to claim 1 is characterized by: The preprocessing module removes features with single values ​​and attributes with low variance from the data in the new dataset.

10. The system for fusion of multi-label data of market entities according to claim 1 is characterized by: The calculation and analysis modules each calculate the similarity of data of different data types in the new data set, including: The feature vectors of two objects i and j are expressed as: X i =[X i1 ,X i2 ,……X ip ] X j =[X j1 ,X j2 ,……X jp ] The similarity calculation formula 1 of the data type is expressed as: When the data type is a numeric attribute type, Formula 1 can be expressed as: When the data type is binary attribute or missing, x in Formula 2 if or x jf Missing, or x if =x jf =0, then in formula 1 otherwise When the data type is a nominal attribute type, the nominal attribute is vectorized, and the similarity distance is calculated using cosine similarity to obtain the similarity, which can be expressed as: After obtaining the similarity, it is used for data fusion.

Citation Information

Cited By

  • Method and system for improving quality of model training sample

    CN121434790A