Multi-feature fusion-based beacon data deduplication method and system
Through the deduplication method based on multi-feature fusion, duplicate records in the flag data are identified and eliminated, and the problems of low data quality and heavy storage burden in the prior art are solved, and efficient data deduplication and storage optimization are achieved.
Patent Information
- Application Number
- CN202510193145.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively identify and remove duplicate records in the coded data, resulting in low data quality and heavy storage burden.
The deduplication method based on multi-feature fusion is adopted, and duplicate records in the flag data are identified and eliminated through four steps: data acquisition, feature extraction, multi-feature fusion deduplication and duplicate data retention. The specific steps include obtaining the standard text data, extracting key features, customizing the multi-feature fusion deduplication rules and deduplication, and finally storing the identified duplicate data in a dedicated database.
It significantly improves the accuracy and uniqueness of the standard data, reduces storage space requirements, reduces maintenance costs, and improves user experience.
Smart Images

Figure CN120144572A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to a method and system for removing duplicate bid information data based on multi-feature fusion. Background Art
[0002] With the development of Internet information technology, a large amount of bid information data is published on the Internet. These data are extremely important sources of business intelligence for enterprises and individuals. However, in practical applications, due to diverse data sources, inconsistent formats, etc., there are often a large number of duplicate information, which not only wastes storage resources but also brings inconvenience to users. Duplicate bid information data not only increases the complexity of data processing but may also lead enterprises or individuals to make decisions based on incorrect information. Although existing methods can remove some duplicate data to a certain extent, due to their reliance on a single feature, it is difficult to handle complex and changeable bid information data situations, especially when dealing with highly similar but not completely identical bid information, the effect is not good.
[0003] How to effectively identify and eliminate duplicate records in bid information data to improve data quality and reduce storage burden is a technical problem that needs to be solved. Summary of the Invention
[0004] The technical task of the present invention is to provide a method and system for removing duplicate bid information data based on multi-feature fusion to solve the technical problem of how to effectively identify and eliminate duplicate records in bid information data to improve data quality and reduce storage burden in view of the above deficiencies.
[0005] In the first aspect, a method for removing duplicate bid information data based on multi-feature fusion of the present invention includes the following steps:
[0006] Data collection: Obtain the bid text data published daily. The bid text data includes the HTML full text information of the bid and the corresponding URL, and perform preliminary deduplication on the collected URLs to remove redundant bid text data with duplicate URLs.
[0007] Feature extraction: Parse the bid text data based on the UIE text extraction small model, extract the key features of the bid text data, and perform fuzzy recognition based on the HTML full text information, classify the industry field, and label the industry label for the bid text data.
[0008] Multi-feature fusion deduplication: Define a deduplication rule for multi-feature fusion based on the URL and the extracted key features, and identify duplicate data in the bid text data based on the deduplication rule to obtain duplicate bid text data.
[0009] Retention of duplicate data: Store the identified duplicate bid text data in a dedicated database.
[0010] As a preference, during data collection, the text data of daily published bidding information is obtained from the official platform by crawling.
[0011] As a preferred option, key features include the purchaser, project number, project name, tender information release time, tender content, announcement type, budget amount, winning bid amount and winning bidder;
[0012] Correspondingly, the deduplication rules include the following operations:
[0013] Hard deduplication: Deduplication is performed based on the collected URLs to exclude data with duplicate URLs;
[0014] Deduplication of iconic key groups: select project number, announcement type and release time as iconic key groups. If all three are the same, it is considered as duplicate data.
[0015] Titles that meet the conditions for deduplication: For two bidding text data with the same title, text similarity analysis is performed on the two bidding text data. When the similarity between the two bidding text data is greater than the preset threshold, they are determined to be duplicate data. When the similarity between the two bidding text data is less than the preset value threshold, but the announcement type and release time are the same, the bidding text data are regarded as problem data that need further judgment, and the problem data are stored in the duplicate data problem library.
[0016] Preferably, when performing text similarity analysis on two standard message text data, a similarity calculation model constructed based on deep learning is used to calculate the text similarity.
[0017] Preferably, the preset threshold is set based on historical data analysis and user needs, and the deduplication rules and preset threshold support user-defined adjustment.
[0018] In a second aspect, the present invention provides a standard information data deduplication system based on multi-feature fusion, including a data acquisition module, a feature extraction module, a multi-feature fusion deduplication module and a deduplication data retention module;
[0019] The data collection module is used to perform the following: obtain the text data of the daily published newsletters, which include the HTML full text information of the newsletters and the corresponding URLs, and perform preliminary deduplication on the collected URLs to remove redundant newsletter text data with repeated URLs;
[0020] The feature extraction module is used to perform the following: parse the bid information text data based on the UIE text extraction model, extract the key features of the bid information text data, perform fuzzy recognition based on the HTML full text information, classify the industry fields, and annotate the bid information text data with industry labels;
[0021] The multi - feature fusion and duplicate removal module is used to perform the following: customize the duplicate removal rules of multi - feature fusion based on the URL and the extracted key features, identify duplicate bid text data based on the duplicate removal rules for the bid text data, and obtain the duplicate bid text data;
[0022] The duplicate data retention module is used to perform the following: store the identified duplicate bid text data in a dedicated database.
[0023] Preferably, the data acquisition module is used to obtain the daily released bid text data from the official platform by means of crawling.
[0024] Preferably, the key features include the purchaser, project number, project name, bid release time, subject matter, announcement type, budget amount, winning bid amount, and winning bid unit;
[0025] Correspondingly, the duplicate removal rules include the following operations:
[0026] Hard duplicate removal: Remove duplicates according to the collected URL, excluding data with duplicate URLs;
[0027] Signature key group duplicate removal: Select the project number, announcement type, and release time as the signature key group. If all three are the same, it is determined as duplicate data;
[0028] Title - compliant duplicate removal: For two bid text data with the same title, perform text similarity analysis on the two bid text data. When the similarity between the two bid text data is greater than the preset threshold, it is determined as duplicate data. When the similarity between the two bid text data is less than the preset value threshold but the announcement type and release time are the same, the bid text data is regarded as problem data that needs further judgment, and the problem data is stored in the duplicate data problem library.
[0029] Preferably, when performing text similarity analysis on two bid text data, the multi - feature fusion and duplicate removal module is used to calculate the text similarity by using a similarity calculation model constructed based on deep learning.
[0030] Preferably, the preset threshold is set according to historical data analysis and user requirements. There is a duplicate removal rule interface configured in the multi - feature fusion and duplicate removal module, which is used to support users to customize and adjust the duplicate removal rules and the preset threshold through the duplicate removal rule interface.
[0031] The bid data duplicate removal method and system based on multi - feature fusion of the present invention have the following advantages:
[0032] 1. Improve data quality: Significantly improve the accuracy and uniqueness of bid data through multi - feature fusion and duplicate removal;
[0033] 2. Reduce storage costs: Reduce the storage of duplicate data, save storage space, and lower maintenance costs;
[0034] 3. Improve user experience: Help users quickly and accurately obtain the required information and enhance the information screening efficiency;
[0035] 4. Enhance flexibility: Support users to customize duplicate removal rules to meet the duplicate removal requirements of data in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] The present invention will be further described below in conjunction with the drawings.
[0038] Figure 1 FIG. 18 is a flowchart of a method for removing duplicate bid information data based on multi-feature fusion in Embodiment 1;
[0039] Figure 2 FIG. 22 is a flowchart of the duplicate removal rule in the method for removing duplicate bid information data based on multi-feature fusion in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be further described below in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the illustrated embodiments are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0041] The embodiments of the present invention provide a method and system for removing duplicate bid information data based on multi-feature fusion, which are used to solve the technical problem of how to effectively identify and remove duplicate records in bid information data to improve data quality and reduce storage burden.
[0042] Embodiment 1:
[0043] A method for removing duplicate bid information data based on multi-feature fusion according to the present invention includes four steps: data collection, feature extraction, multi-feature fusion duplicate removal, and retention of duplicate data.
[0044] Step S100 Data collection: Obtain the bid information text data released daily. The bid information text data includes the HTML full text information of the bid and the corresponding URL, and perform preliminary duplicate removal on the collected URLs to remove the redundant bid information text data with duplicate URLs.
[0045] As a specific implementation of data collection, daily released bid information text data is obtained from official platforms through crawling means.
[0046] In this embodiment, through means such as crawlers, daily released bid information text data is obtained from official platforms such as the National Government Procurement Network, Public Resource Trading Centers, and Sunshine Procurement Platforms of each province, and the HTML full text of the bid information and the corresponding URLs are collected. Preliminary duplicate removal is performed according to the collected URLs, and redundant data with duplicate URLs is excluded.
[0047] Step S200 Feature Extraction: Based on the UIE text extraction small model, the bid information text data is parsed to extract the key features of the bid information text data, and fuzzy recognition is performed according to the HTML full text information to classify the industry fields and label the industry tags for the bid information text data.
[0048] Among them, the key features include the purchaser, project number, project name, bid information release time, subject matter, announcement type, budget amount, winning bid amount, and winning bid unit
[0049] In this embodiment, the UIE text extraction small model is used to deeply parse the bid information text to obtain the key information of the bid, such as the purchaser, project number, project name, bid information release time, subject matter, announcement type, budget amount, winning bid amount, winning bid unit, etc. Further fuzzy recognition is performed according to the full text information, vector data encoding is performed on the industry classification for the industry fields required by the business, and fuzzy matching is performed with the full text of the bid information to obtain the industry field to which the bid belongs, and the industry tag is automatically assigned to this bid.
[0050] Step S300 Multi-Feature Fusion Duplicate Removal: Based on the URL and the extracted key features, a duplicate removal rule for multi-feature fusion is customized, and duplicate bid information text data is identified based on the duplicate removal rule.
[0051] As a specific implementation of the duplicate removal rule, it includes the following operations:
[0052] (1) Hard Duplicate Removal: Duplicate removal is performed according to the collected URLs to exclude data with duplicate URLs;
[0053] (2) Signature Key Group Duplicate Removal: The project number, announcement type, and release time are selected as the signature key group. If all three are the same, it is determined as duplicate data;
[0054] (3) Title deduplication meets the conditions: For two bidding information text data with the same title, the two bidding information text data are subjected to text similarity analysis. When the similarity between the two bidding information text data is greater than the preset threshold, they are judged as duplicate data. When the similarity between the two bidding information text data is less than the preset value threshold, but the announcement type and release time are the same, the bidding information text data are regarded as problem data that need further judgment, and the problem data are stored in the duplicate data problem library.
[0055] Among them, when performing text similarity analysis on two standard information text data, a similarity calculation model built based on deep learning is used to calculate the text similarity.
[0056] In this embodiment, the preset threshold is set based on historical data analysis and user needs, and the deduplication rules and preset threshold support user-defined adjustment.
[0057] In this embodiment, a hard deduplication rule is determined. First, deduplication is performed based on the collection URL to exclude the case of duplicate collection. At the same time, considering the diversity of data sources, the same data may be published on different websites, resulting in the possibility that bid announcements with different URLs may have the same bid announcement content. However, directly comparing the bid announcement text requires a lot of processing time and memory, so it is considered to select a symbolic key group for deduplication. For example, if the project number, announcement type and release time are the same, it is considered to be duplicate data.
[0058] At the same time, when users see the same bidding information title, they will preconceivedly think that it is duplicate data. However, there are many cases where the bidding information title is the same but not duplicate data, such as:
[0059] (1) After a project is cancelled, the bidding information data with the same title is re-published;
[0060] (2) The same bidding title is used at different project announcement stages;
[0061] (3) The same project number but different procurement contents, etc.
[0062] In this regard, this embodiment formulates a deduplication rule determined according to the composite conditions of the bidding message title, and performs similarity analysis on two bidding message texts with the same bidding message title. When the similarity of the two texts is greater than a certain value, it is considered to be duplicate data. For bidding messages with similarity less than a threshold, it is further determined whether their announcement type and release time are the same. If they are the same, it is considered to be problem data that needs further judgment, and this type of data is stored in the duplicate data problem library.
[0063] After using the small model to first parse the key field information, the key field information is used for preliminary deduplication, which greatly reduces the amount of data for subsequent full-text fuzzy comparison of the large model, improves the efficiency of duplicate data comparison, and reduces GPU usage.
[0064] Step S400: Repeated data retention: Store the identified repeated tender text data in a dedicated database.
[0065] In this embodiment, the identified repeated data is stored in a dedicated database, which is convenient for subsequent data analysis and proofreading.
[0066] The method of this embodiment provides a more comprehensive and accurate method for deduplicating tender data. By calculating the similarity by integrating multiple key attributes and combining an intelligent threshold judgment mechanism, duplicate records are effectively identified and eliminated, thereby improving data quality, reducing storage burden, and enhancing user experience.
[0067] Embodiment 2:
[0068] A tender data deduplication system based on multi-feature fusion of the present invention includes a data acquisition module, a feature extraction module, a multi-feature fusion deduplication module, and a deduplicated data retention module.
[0069] The data acquisition module is used to perform the following: Obtain the tender text data released daily, where the tender text data includes the full HTML information of the tender and the corresponding URL, and perform preliminary deduplication on the collected URLs to eliminate redundant tender text data with duplicate URLs.
[0070] As a specific implementation of the data acquisition module, this module obtains the tender text data released daily from the official platform by means of crawling.
[0071] In this embodiment, by means of crawlers and other means, obtain the tender text data released daily from official platforms such as the National Government Procurement Network, Public Resource Trading Centers, and Sunshine Procurement Platforms of each province, collect the full HTML of the tender and the corresponding URL. Perform preliminary deduplication according to the collected URLs to eliminate redundant data with duplicate URLs.
[0072] The feature extraction module is used to perform the following: Parse the tender text data based on the UIE text extraction small model, extract the key features of the tender text data, and perform fuzzy recognition according to the full HTML information, classify the industry field, and label the industry label for the tender text data.
[0073] Among them, the key features include the purchaser, project number, project name, tender release time, subject matter, announcement type, budget amount, winning bid amount, and winning bid unit
[0074] In this embodiment, the feature extraction module is used to deeply analyze the tender information text by using the UIE text extraction small model to obtain the key information of the tender, such as the purchaser, project number, project name, tender release time, subject matter, announcement type, budget amount, winning bid amount, winning bid unit, etc. Further fuzzy recognition is performed based on the full text information, and vector data encoding is performed on the industry classification for the industry fields required by the business, and fuzzy matching is performed with the full text of the tender to obtain the industry field to which the tender belongs, and the industry label is automatically assigned to this tender.
[0075] The multi-feature fusion and duplicate removal module is used to perform the following: Customize the duplicate removal rules of multi-feature fusion based on the URL and the extracted key features, and identify duplicate data from the tender text data based on the duplicate removal rules to obtain duplicate tender text data.
[0076] As a specific implementation of the duplicate removal rules, it includes the following operations:
[0077] (1) Hard duplicate removal: Remove duplicates according to the collected URL, excluding data with duplicate URLs;
[0078] (2) Signature key group duplicate removal: Select the project number, announcement type, and release time as the signature key group. If all three are the same, it is determined as duplicate data;
[0079] (3) Title compliance duplicate removal: For two tender text data with the same title, perform text similarity analysis on the two tender text data. When the similarity between the two tender text data is greater than the preset threshold, it is determined as duplicate data. When the similarity between the two tender text data is less than the preset value threshold but the announcement type and release time are the same, the tender text data is regarded as problem data that needs further judgment, and the problem data is stored in the duplicate data problem library.
[0080] Among them, when performing text similarity analysis on two tender text data, a similarity calculation model constructed based on deep learning is used to calculate the text similarity.
[0081] In this embodiment, the preset threshold is set according to historical data analysis and user requirements, and the duplicate removal rules and the preset threshold support user-defined adjustment.
[0082] In this embodiment, to determine the hard duplicate removal rules, first remove duplicates according to the collected URL to exclude duplicate collection situations. At the same time, considering the diversity of data sources, the same data may be published on different websites, resulting in tender announcements with different URLs may have the same tender content. However, directly comparing tender texts requires a large amount of processing time and memory. Therefore, it is considered to select a signature key group for duplicate removal. For example, if the project number, announcement type, and release time are all the same, it is considered duplicate data.
[0083] Meanwhile, when users are using it, they will preconceive that the same bid notice title means duplicate data. However, there are many cases where the bid notice titles are the same but the data is not duplicate, such as:
[0084] (1) The bid notice data with the same bid notice title after a certain project is rejected and reissued;
[0085] (2) Different project announcement stages have the same bid notice title;
[0086] (3) The same project number but different procurement contents, etc.
[0087] In response to this, this embodiment formulates a deduplication rule determined according to the composite conditions of the bid notice title, performs a similarity analysis on two bid notice texts with the same bid notice title. When the similarity of the two texts is greater than a certain specific value, it is considered duplicate data. For bid notices with a similarity less than the threshold, further judge whether their announcement types and release times are the same. If both are the same, it is considered problem data that needs further judgment, and this type of data is stored in the duplicate data problem library.
[0088] After using the small model to first parse the keyword field information, the keyword field information is used for preliminary deduplication, which greatly reduces the data volume of the subsequent full-text fuzzy comparison by the large model, improves the duplicate data comparison efficiency, and reduces the GPU usage rate.
[0089] The duplicate data retention module is used to perform the following: store the identified duplicate bid notice text data in a dedicated database.
[0090] In this embodiment, the identified duplicate data is stored in a dedicated database for subsequent data analysis and proofreading.
[0091] The system of this embodiment can implement the deduplication of bid notice text data by executing the method disclosed in Embodiment 1.
[0092] The above has introduced in detail the method and system for deduplicating bid notice data based on multi-feature fusion provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for deduplication of standard information data based on multi-feature fusion, characterized in that: The steps include: Data collection: Obtain the text data of the bid news published daily, which includes the HTML full text information of the bid news and the corresponding URL, and perform preliminary deduplication on the collected URLs to remove redundant bid news text data with repeated URLs; Feature extraction: Analyze the bidding information text data based on the UIE text extraction model, extract the key features of the bidding information text data, perform fuzzy recognition based on the HTML full-text information, classify the industry fields, and annotate the bidding information text data with industry labels; Multi-feature fusion deduplication: Customize the multi-feature fusion deduplication rules based on the URL and the extracted key features, identify duplicate data of the bid information text data based on the deduplication rules, and obtain duplicate bid information text data; Duplicate data retention: The identified duplicate bidding information text data will be stored in a dedicated database.
2. The method for deduplication of standard information data based on multi-feature fusion according to claim 1 is characterized in that: When collecting data, the text data of the daily published bidding information is obtained from the official platform through crawling.
3. The method for deduplication of standard information data based on multi-feature fusion according to claim 1 is characterized in that: Key features include the purchaser, project number, project name, tender information release time, tender content, announcement type, budget amount, winning bid amount and winning bidder; Correspondingly, the deduplication rules include the following operations: Hard deduplication: Deduplication is performed based on the collected URLs to exclude data with duplicate URLs; Deduplication of iconic key groups: select project number, announcement type and release time as iconic key groups. If all three are the same, it is considered as duplicate data. Titles that meet the conditions for deduplication: For two bidding text data with the same title, text similarity analysis is performed on the two bidding text data. When the similarity between the two bidding text data is greater than the preset threshold, they are determined to be duplicate data. When the similarity between the two bidding text data is less than the preset value threshold, but the announcement type and release time are the same, the bidding text data are regarded as problem data that need further judgment, and the problem data are stored in the duplicate data problem library.
4. The method for deduplication of standard information data based on multi-feature fusion according to claim 3 is characterized in that: When performing text similarity analysis on two standard message text data, a similarity calculation model based on deep learning is used to calculate the text similarity.
5. The method for deduplication of standard information data based on multi-feature fusion according to claim 3 is characterized in that: The preset thresholds are set based on historical data analysis and user needs. Deduplication rules and preset thresholds support user-defined adjustments.
6. A standard information data deduplication system based on multi-feature fusion, characterized in that: It includes data collection module, feature extraction module, multi-feature fusion deduplication module and deduplication data retention module; The data collection module is used to perform the following: obtain the text data of the daily published newsletters, which include the HTML full text information of the newsletters and the corresponding URLs, and perform preliminary deduplication on the collected URLs to remove redundant newsletter text data with repeated URLs; The feature extraction module is used to perform the following: parse the bid information text data based on the UIE text extraction model, extract the key features of the bid information text data, perform fuzzy recognition based on the HTML full text information, classify the industry fields, and annotate the bid information text data with industry labels; The multi-feature fusion deduplication module is used to perform the following: customize the multi-feature fusion deduplication rules based on the URL and the extracted key features, identify duplicate data of the bid message text data based on the deduplication rules, and obtain duplicate bid message text data; The duplicate data retention module is used to perform the following: storing the identified duplicate bidding information text data in a dedicated database.
7. The system for deduplication of standard information data based on multi-feature fusion according to claim 6 is characterized in that: The data collection module is used to obtain the text data of daily published bidding information from the official platform through crawling.
8. The system for deduplication of standard information data based on multi-feature fusion according to claim 6 is characterized in that: Key features include the purchaser, project number, project name, tender information release time, tender content, announcement type, budget amount, winning bid amount and winning bidder; Correspondingly, the deduplication rules include the following operations: Hard deduplication: Deduplication is performed based on the collected URLs to exclude data with duplicate URLs; Deduplication of iconic key groups: select project number, announcement type and release time as iconic key groups. If all three are the same, it is considered as duplicate data. Titles that meet the conditions for deduplication: For two bidding text data with the same title, text similarity analysis is performed on the two bidding text data. When the similarity between the two bidding text data is greater than the preset threshold, they are determined to be duplicate data. When the similarity between the two bidding text data is less than the preset value threshold, but the announcement type and release time are the same, the bidding text data are regarded as problem data that need further judgment, and the problem data are stored in the duplicate data problem library.
9. The system for deduplication of standard information data based on multi-feature fusion according to claim 8, characterized in that: When performing text similarity analysis on two standard message text data, the multi-feature fusion deduplication module is used to calculate text similarity using a similarity calculation model built based on deep learning.
10. The system for deduplication of standard information data based on multi-feature fusion according to claim 8, characterized in that: The preset threshold is set based on historical data analysis and user needs. The multi-feature fusion deduplication module is configured with a deduplication rule interface to support users to customize the deduplication rules and preset thresholds through the deduplication rule interface.
Citation Information
Patent Citations
Public resource transaction data-oriented cleaning and duplicate removal method and system
CN110196848A
Massive internet news cleaning system
CN111859070A
Duplicate removal method and system based on text information extraction result and medium
CN112989791A
Big data-based bidding and tendering analysis method, system and equipment and storage medium
CN115080698A
Cited By
Two-stage data deduplication method and system without damaging data timeliness
CN121807819A