A method and system for processing merchant duplicate data
Through the merchant duplicate data processing method and system, using data lifecycle management and identification algorithms, the duplicate data problem in multi-source merchant data is solved, the data quality is improved and the customer experience is enhanced.
Patent Information
- Application Number
- CN202110509245.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-11
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2041-05-11
AI Technical Summary
The existing system failed to effectively handle duplicate data in merchant data from multiple sources, resulting in a decline in data quality, affecting customer experience and team reputation.
By introducing merchant duplicate data processing methods and systems, and utilizing data lifecycle management, trust levels, and duplicate data identification algorithms, data deduplication and quality protection are performed, including duplicate identification, conflict resolution, and manual intervention, and the trust level is adjusted to solve the duplicate data problem.
It improves data quality, enhances customer experience, makes data traceable, maintainable, customizable and intervenible, and solves the problem of duplicate data from multiple data sources.
Smart Images

Figure CN113157682B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for processing merchant duplicate data. Background Art
[0002] Currently, the merchant information on the Car Owner Platform is entered from multiple sources, such as Baifumei, Merchant Cloud, and Chediandian. Different sources may contain the same merchant data, resulting in duplicate merchant information.
[0003] For the entry of data from multiple sources, the existing system does not perform graded processing and lifecycle management on the quality (trustworthiness) of incremental data, nor does it perform graded protection on the quality (trustworthiness) of existing data, and currently does not identify and exclude duplicate data.
[0004] With the accumulation of existing data and the increase in incremental data sources, duplicate data may increase rapidly, resulting in a decline in data quality, confusing page displays, and affecting customer experience and team reputation.
[0005] Therefore, there is an urgent need for a technical solution that can overcome the shortcomings of existing technologies and effectively process merchant duplicate data. Summary of the Invention
[0006] In order to solve the problems existing in the prior art, the present invention proposes a method and system for processing merchant duplicate data. The present invention can eliminate duplicate entries for newly added data and protect the quality of existing data, thereby eliminating the situation where duplicate data of low quality in newly added data is overwritten. By introducing lifecycle management of imported data and merchant data trust levels, as well as identification algorithms and processing strategies that can be custom-assembled for duplicate data, a data duplication solution is provided in scenarios where data has multiple sources.
[0007] In a first aspect of an embodiment of the present invention, a method for processing duplicate merchant data is provided, the method comprising:
[0008] Obtain the merchant's newly added data and update the data batch status according to the intermediate data information of the newly added data;
[0009] Obtain a duplicate identification algorithm based on the data batch status and perform duplicate identification on the newly added data; if the identification result is that the data is duplicated, merge the newly added data and record the data protection level and data status; if the identification result is that the data is not duplicated, record a data conflict identifier;
[0010] Obtaining a conflict resolution strategy based on the data batch status, and intelligently handling the conflict based on the conflict resolution strategy; if the conflict is resolved, then updating, merging, or discarding the data, and recording the data conflict resolution strategy and data status; if the conflict is not resolved, then recording the data status and invoking manual intervention;
[0011] The conflicting data is displayed to the operator, who then manually selects an intervention strategy to intervene in the conflicting data and adjust the confidence level of the conflicting data.
[0012] According to the adjusted trust level, it is determined whether the trust level meets the trust level upgrade requirement; if so, the trust level of the data source of the newly added data is updated; if not, the newly added data is directly saved.
[0013] Furthermore, the data intermediate information includes: original data, source data, batch data, data trustworthiness and status records.
[0014] Furthermore, the newly added data of the merchant is obtained, and the data batch status is updated according to the intermediate information of the newly added data, including:
[0015] Acquire the merchant's newly added data, extract the data intermediate information based on the data source registry of the newly added data, and save it to the data intermediate table; wherein the data source registry contains at least the batch identifier, batch code, batch name, data source identifier, data batch confidence, duplicate identification algorithm, and conflict resolution algorithm; the data intermediate table contains at least the original data information, data source identifier, data batch identifier, data confidence, merge status, conflict identifier, and conflict strategy; the merge status is collection completed, merging in progress, merge conflict resolution in progress, or merge completed;
[0016] Furthermore, obtaining the merchant's newly added data and updating the data batch status according to the intermediate data information of the newly added data also includes:
[0017] The data batch registration table is updated according to the data intermediate information, and the data batch status is updated; which at least includes: batch identification, batch code, batch name, data source identification, data batch trust, progress status, duplicate identification algorithm and conflict resolution algorithm; the progress status is adjusted with the progress of processing, and the progress status is ready to collect, collecting, collection completed, merging, merging conflict resolution or merging completed.
[0018] Furthermore, a duplicate identification algorithm is obtained according to the data batch status to perform duplicate identification on the newly added data, including:
[0019] According to the duplicate identification algorithm, the corresponding algorithm interface implementation class is obtained to perform duplicate identification on the newly added data.
[0020] Furthermore, obtaining a conflict resolution strategy according to the data batch status and performing conflict processing according to the conflict resolution strategy include:
[0021] An intelligent conflict handling strategy is acquired according to the conflict resolution strategy, and the conflict is intelligently handled and resolved using the intelligent conflict handling strategy.
[0022] Furthermore, the conflicting data is presented to the operator, who then manually selects an intervention strategy to intervene in the conflicting data and adjust the confidence level of the conflicting data, including:
[0023] The conflicting data are compared according to the conflict identification, and the comparison results are displayed to the operator, who manually selects an intervention strategy to intervene in the conflicting data, adjusts the confidence level of the conflicting data, and records the data status and intervention behavior.
[0024] Furthermore, the conflicting data is displayed to the operator, who then manually selects an intervention strategy to intervene in the conflicting data and adjust the confidence level of the conflicting data, which also includes:
[0025] After the intervention processing of the conflicting data is completed, the intervention behavior is recorded in the intervention data record table; wherein the intervention data record table at least includes: intervention identification, data source identification, data batch identification, data identification, conflict identification, intervention type, intervention strategy and intervention result; the intervention result is merged or discarded.
[0026] Furthermore, judging whether the adjusted trust level meets the trust level upgrade requirement includes:
[0027] Collecting manual intervention strategy data, and judging whether the trust level meets the trust level upgrade requirement based on the trust level after the manual intervention.
[0028] Furthermore, the duplicate identification algorithm at least includes duplicate identification using the industrial and commercial number, duplicate identification using the industrial and commercial number and address, duplicate identification using the legal person and address, and duplicate identification using the industrial and commercial number, address and main business.
[0029] Furthermore, the conflict resolution strategy includes at least: forced overwriting and updating, forced deletion and addition, aggressive processing based on trust, conservative processing based on trust, and processing based on trust and manual intervention.
[0030] Furthermore, the manual intervention method includes at least: forced updating and forced abandonment; and manual comparison of duplicate data to perform corresponding manual intervention processing.
[0031] In a second aspect of an embodiment of the present invention, a system for processing merchant duplicate data is provided, the system comprising:
[0032] A data acquisition module is used to acquire new data from merchants and update the data batch status according to the intermediate data information of the new data;
[0033] A duplicate identification module is used to obtain a duplicate identification algorithm based on the data batch status and perform duplicate identification on the newly added data; if the identification result is that the data is duplicated, the newly added data is merged and the data protection level and data status are recorded; if the identification result is that the data is not duplicated, a data conflict identifier is recorded;
[0034] A conflict intelligent processing module is used to obtain a conflict resolution strategy based on the data batch status and intelligently handle conflicts according to the conflict resolution strategy; if the conflict is resolved, the data is updated, merged, or discarded, and the data conflict resolution strategy and data status are recorded; if the conflict is not resolved, the data status is recorded and manual intervention is invoked;
[0035] The conflict manual intervention module is used to display conflict data to operators, who can manually select intervention strategies to intervene in the conflict data and adjust the confidence level of the conflict data.
[0036] The trust setting module is used to determine whether the trust meets the trust upgrade requirements based on the adjusted trust; if so, the trust of the data source of the newly added data is updated; if not, the newly added data is directly saved.
[0037] Furthermore, the data intermediate information includes: original data, source data, batch data, data trustworthiness and status records.
[0038] Furthermore, the data acquisition module is specifically used to:
[0039] Acquire the merchant's newly added data, extract the data intermediate information based on the data source registry of the newly added data, and save it to the data intermediate table; wherein the data source registry contains at least the batch identifier, batch code, batch name, data source identifier, data batch confidence, duplicate identification algorithm, and conflict resolution algorithm; the data intermediate table contains at least the original data information, data source identifier, data batch identifier, data confidence, merge status, conflict identifier, and conflict strategy; the merge status is collection completed, merging in progress, merge conflict resolution in progress, or merge completed;
[0040] Furthermore, the data acquisition module is also used to:
[0041] The data batch registration table is updated according to the data intermediate information, and the data batch status is updated; which at least includes: batch identification, batch code, batch name, data source identification, data batch trust, progress status, duplicate identification algorithm and conflict resolution algorithm; the progress status is adjusted with the progress of processing, and the progress status is ready to collect, collecting, collection completed, merging, merging conflict resolution or merging completed.
[0042] Furthermore, the duplicate identification module is specifically used to:
[0043] According to the duplicate identification algorithm, the corresponding algorithm interface implementation class is obtained to perform duplicate identification on the newly added data.
[0044] Furthermore, the conflict intelligent processing module is specifically used to:
[0045] An intelligent conflict handling strategy is acquired according to the conflict resolution strategy, and the conflict is intelligently handled and resolved using the intelligent conflict handling strategy.
[0046] Furthermore, the conflict manual intervention module is specifically used to:
[0047] The conflicting data are compared according to the conflict identification, and the comparison results are displayed to the operator, who manually selects an intervention strategy to intervene in the conflicting data, adjusts the confidence level of the conflicting data, and records the data status and intervention behavior.
[0048] Furthermore, the conflict manual intervention module is also used to:
[0049] After the intervention processing of the conflicting data is completed, the intervention behavior is recorded in the intervention data record table; wherein the intervention data record table at least includes: intervention identification, data source identification, data batch identification, data identification, conflict identification, intervention type, intervention strategy and intervention result; the intervention result is merged or discarded.
[0050] Furthermore, the trust setting module is specifically used to:
[0051] Collecting manual intervention strategy data, and judging whether the trust level meets the trust level upgrade requirement based on the trust level after the manual intervention.
[0052] Furthermore, the duplicate identification algorithm at least includes duplicate identification using the industrial and commercial number, duplicate identification using the industrial and commercial number and address, duplicate identification using the legal person and address, and duplicate identification using the industrial and commercial number, address and main business.
[0053] Furthermore, the conflict resolution strategy includes at least: forced overwriting and updating, forced deletion and addition, aggressive processing based on trust, conservative processing based on trust, and processing based on trust and manual intervention; the manual intervention method includes at least: forced updating and forced abandonment; and corresponding manual intervention processing methods are performed by manually comparing duplicate data.
[0054] In a third aspect of an embodiment of the present invention, a computer device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a method for processing duplicate merchant data when executing the computer program.
[0055] In a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a method for processing merchant duplicate data is implemented.
[0056] The method and system for processing merchant duplicate data proposed in the present invention utilizes an assembleable data duplication identification algorithm and processing strategy to process merchant newly added data, thereby eliminating duplicate entry of newly added data and protecting the quality of existing data, thereby eliminating duplicate data of low quality in newly added data and overwriting it, thereby improving data quality and customer experience. During the processing process, by introducing lifecycle management of imported data and merchant data trust levels, as well as customizable identification algorithms and processing strategies for duplicate data, a data duplication solution is provided for scenarios with multiple data sources, and the solution can be reused for scenarios with other multiple data sources, making data traceable, maintainable, customizable, and intervenible. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0058] Figure 1 The figure is a flowchart of a method for processing duplicate merchant data according to an embodiment of the present invention.
[0059] Figure 2 It is a detailed flow chart of a method for processing merchant duplicate data according to a specific embodiment of the present invention.
[0060] Figure 3 Schematic diagram of a duplicate identification algorithm according to a specific embodiment of the present invention.
[0061] Figure 4FIG. 4 is a schematic diagram of a conflict resolution algorithm according to a specific embodiment of the present invention.
[0062] Figure 5 It is a schematic diagram of a manual intervention method according to a specific embodiment of the present invention.
[0063] Figure 6 2 is a schematic diagram of a system architecture for processing merchant duplicate data according to an embodiment of the present invention.
[0064] Figure 7 It is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0066] Those skilled in the art will appreciate that the embodiments of the present invention may be implemented as a system, apparatus, device, method, or computer program product. Therefore, the present disclosure may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software.
[0067] According to an embodiment of the present invention, a method and system for processing merchant duplicate data are proposed, which can be used in the field of data processing technology; the present invention can deduplicate the newly added data and protect the quality of the existing data, and exclude the duplicate data of low quality in the newly added data and overwrite it; by introducing the life cycle management of imported data and the trust level of merchant data, as well as the identification algorithm and processing strategy that can be custom-assembled for duplicate data, a data duplication solution is provided in the scenario where data has multiple sources.
[0068] The principles and spirit of the present invention are explained in detail below with reference to several representative embodiments of the present invention.
[0069] Figure 1 FIG. 1 is a flow chart of a method for processing merchant duplicate data according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0070] Step S101: Acquire the merchant's newly added data and update the data batch status according to the intermediate data information of the newly added data;
[0071] Step S102: Obtain a duplicate identification algorithm based on the data batch status and perform duplicate identification on the newly added data; if the identification result is that the data is duplicated, merge the newly added data and record the data protection level and data status; if the identification result is that the data is not duplicated, record a data conflict flag;
[0072] Step S103: Obtain a conflict resolution strategy based on the data batch status, and intelligently handle the conflict based on the conflict resolution strategy; if the conflict is resolved, update and merge or discard the data, and record the data conflict resolution strategy and data status; if the conflict is not resolved, record the data status and invoke manual intervention;
[0073] Step S104: displaying the conflicting data to an operator, who then manually selects an intervention strategy to intervene in the conflicting data and adjust the confidence level of the conflicting data;
[0074] Step S105 , judging whether the adjusted trust level meets the trust level upgrade requirement based on the adjusted trust level; if so, updating the trust level of the data source of the newly added data; if not, directly saving the newly added data.
[0075] The method for processing merchant duplicate data proposed in the present invention can process the merchant's new data through data lifecycle management, using an assembleable data duplicate identification algorithm and processing strategy, to achieve duplicate entry of the new data, and perform quality protection on the existing data, eliminating duplicate data of low quality in the new data and causing data overwriting, thereby improving data quality and improving customer experience.
[0076] During the processing, by introducing the lifecycle management of imported data and the trust level of merchant data, as well as the identification algorithm and processing strategy for customizable assembly of duplicate data, a data duplication solution is provided in scenarios with multiple data sources; and the solution can be reused in scenarios with other multiple data sources, making the data traceable, maintainable, customizable, and intervention-friendly.
[0077] In step S101 of this embodiment, the specific process of obtaining the merchant's newly added data and updating the data batch status according to the intermediate data information of the newly added data is as follows:
[0078] Step S1011, obtaining the merchant's new data, extracting the data intermediate information according to the data source registration table of the new data and saving it to the data intermediate table; wherein,
[0079] The data source registration table contains at least batch identification, batch code, batch name, data source identification, data batch trustworthiness, duplicate identification algorithm and conflict resolution algorithm;
[0080] The data intermediate table contains at least the original data information, data source identifier, data batch identifier, data confidence, merge status, conflict identifier, and conflict strategy; the merge status can be collection completed, merging in progress, merge conflict resolution in progress, or merge completed;
[0081] Step S1012: updating the data batch registration table according to the data intermediate information and updating the data batch status;
[0082] Among them, at least include: batch identification, batch code, batch name, data source identification, data batch trustworthiness, progress status, duplicate identification algorithm and conflict resolution algorithm;
[0083] The progress status is adjusted as the processing progresses, and the progress status is preparing to collect, collecting, collecting completed, merging, resolving merge conflicts, or merging completed.
[0084] In this embodiment, the data intermediate information includes: original data, source data, batch data, data credibility and status records.
[0085] In step S102 of this embodiment, a duplicate identification algorithm is obtained according to the data batch status to perform duplicate identification on the newly added data, including:
[0086] According to the duplicate identification algorithm, the corresponding algorithm interface implementation class is obtained to perform duplicate identification on the newly added data.
[0087] Among them, the duplicate identification algorithm at least includes duplicate identification using the industrial and commercial number, duplicate identification using the industrial and commercial number and address, duplicate identification using the legal person and address, and duplicate identification using the industrial and commercial number, address and main business.
[0088] In step S103 of this embodiment, a conflict resolution strategy is obtained according to the data batch status, and conflict processing is performed according to the conflict resolution strategy, including:
[0089] An intelligent conflict handling strategy is acquired according to the conflict resolution strategy, and the conflict is intelligently handled and resolved using the intelligent conflict handling strategy.
[0090] Among them, the conflict resolution strategy includes at least: forced overwriting and updating, forced deletion and addition, aggressive processing based on trust level, conservative processing based on trust level, and processing based on trust level and manual intervention; among them, the manual intervention method includes at least: forced updating and forced abandonment; and corresponding manual intervention processing methods are performed by manually comparing duplicate data.
[0091] In step S104 of this embodiment, the conflicting data is presented to the operator, who manually selects an intervention strategy to intervene in the conflicting data and adjust the confidence level of the conflicting data, including:
[0092] The conflicting data are compared according to the conflict identification, and the comparison results are displayed to the operator, who manually selects an intervention strategy to intervene in the conflicting data, adjusts the confidence level of the conflicting data, and records the data status and intervention behavior.
[0093] After the intervention processing of the conflicting data is completed, the intervention behavior is recorded in the intervention data record table; wherein the intervention data record table at least includes: intervention identification, data source identification, data batch identification, data identification, conflict identification, intervention type, intervention strategy and intervention result; the intervention result is merged or discarded.
[0094] In step S105 of this embodiment, judging whether the adjusted trust level meets the trust level upgrade requirement includes:
[0095] Collecting manual intervention strategy data, and judging whether the trust level after manual intervention meets the trust level upgrade requirements; wherein,
[0096] If the trust level of the data source of the newly added data is reached, the trust level of the data source of the newly added data will be updated;
[0097] If not reached, the newly added data is directly saved.
[0098] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and drawings, this does not require or imply that these operations must be performed in this specific order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0099] In order to explain the above-mentioned method for processing duplicate merchant data more clearly, a specific embodiment is used for illustration below.
[0100] First, design the data table:
[0101] 1. Data source registry (data_source_register), including:
[0102] ID: source identifier;
[0103] code: source code;
[0104] name: source name;
[0105] “***”: other fields (source, data type, method, frequency, registration time, etc.);
[0106] credibility: the trustworthiness of the data source;
[0107] auto_credibility: whether to allow the trust machine to set;
[0108] auto_threshold: automatically set the trust threshold;
[0109] duplicated_recognize_bean: duplicate recognition algorithm;
[0110] duplicated_solution_bean: conflict resolution algorithm.
[0111] 2. Data batch registration table (data_patch_register):
[0112] ID: batch identification;
[0113] code: batch code (source code + source time);
[0114] name: batch name (source name + source time);
[0115] source_id: data source identifier;
[0116] “***”: other fields (batch quantity, completed quantity, etc.);
[0117] credibility: the trustworthiness of the data batch;
[0118] duplicated_recognize_bean: duplicate recognition algorithm;
[0119] duplicated_solution_bean: conflict resolution algorithm;
[0120] Status: Progress status: preparing to collect, collecting, collecting completed, merging, resolving merge conflicts, merging completed.
[0121] 3. Intervention data record form:
[0122] ID: intervention identification;
[0123] source_id: data source identifier;
[0124] patch_id: data batch identifier;
[0125] data_id: data identifier;
[0126] duplicated_id: conflict identifier;
[0127] type: intervention type: intelligent intervention, human intervention;
[0128] duplicated_solution: intervention strategy;
[0129] result: The intervention results are merged and discarded.
[0130] 4. Data intermediate table ({source_name}_temp_info):
[0131] “***”: original data information;
[0132] source_id: data source identifier;
[0133] patch_id: data batch identifier;
[0134] credibility: data trust;
[0135] status: merge status (collection completed, merging in progress, merge conflict resolution in progress, merge completed);
[0136] duplicated_id: conflict identifier;
[0137] duplicated_solution_bean: conflict strategy;
[0138] 5. Merchant data information table (enterprise_info):
[0139] “***”: merchant data field;
[0140] Credibility: Data trustworthiness (degree of protection).
[0141] The following references Figure 2 The method for processing duplicate merchant data of the present invention is described in detail. Figure 2 FIG. 1 is a detailed flow chart of a method for processing merchant duplicate data according to a specific embodiment of the present invention.
[0142] In this embodiment, the terms that need to be explained are:
[0143] Data entry: Data enters the data intermediate table from the source system.
[0144] Data identification and conflict-free merging: Mark duplicate data and directly merge non-duplicate data.
[0145] Intelligent conflict handling: By comparing trust values and protection values, data is discarded, merged, or reserved.
[0146] Conflict manual intervention: Human resources compare duplicate data and make final decisions.
[0147] Intelligent trust setting: Increase the trust value of the data source by intervening in behavioral data analysis.
[0148] Batch progress status: preparing to collect, collecting, collection completed, conflict identification in progress, identification completed, intelligent processing in progress, intelligent processing completed, manual intervention in progress, merging completed.
[0149] Data merging status:,collection completed, conflict identification in progress, identification completed, intelligent processing in progress, intelligent processing completed, human intervention in progress, merging completed.
[0150] like Figure 2 As shown, the specific processing process is:
[0151] Step S1, data entry:
[0152] Step S11, record data batch information (prepare for collection);
[0153] Step S12: save the intermediate data information, including original data, source data, batch data, data reliability and status records;
[0154] Step S13, update the batch status (entry completed).
[0155] Step S2, data identification and conflict-free merging:
[0156] Step S21, obtaining a duplicate recognition algorithm (duplicated_recognize_bean) through batch information;
[0157] Step S22, obtain the corresponding algorithm interface implementation class through duplicated_recognize_bean, and identify whether it is duplicated;
[0158] If repeated, step S23, perform data addition and merging, record data protection level, and record data status (merging completed);
[0159] If there is no duplication, step S24 is to record the data conflict identifier (duplicated_id).
[0160] Step S3, intelligent conflict handling:
[0161] Step S31, obtaining the conflict resolution strategy (duplicated_solution_bean) through batch information;
[0162] Step S32: Obtain an intelligent conflict handling strategy through duplicated_solution_bean and handle the conflict intelligently.
[0163] If resolved, step S33: update and merge the data or discard it, and record the data conflict resolution strategy and data status (merge completed);
[0164] If the problem is not solved, step S34 is to record the data status (prepare for manual intervention).
[0165] Step S4: Conflict manual intervention:
[0166] Step S41, comparing conflicting data by duplicated_id;
[0167] Step S42, manually selecting an intervention strategy;
[0168] Step S43, record data status (merging completed);
[0169] Step S44: record the intervention behavior.
[0170] Step S5: Intelligent setting of trustworthiness:
[0171] Step S51: Collecting statistics on human intervention strategy data to determine whether trust upgrade is satisfied;
[0172] If satisfied, step S52, update the data source trust;
[0173] If not satisfied, step S53, directly save the newly added data.
[0174] The following is a detailed introduction to the duplicate identification algorithm, conflict resolution algorithm, and manual intervention methods involved in the merchant duplicate data processing process.
[0175] refer to Figure 3 , which is a schematic diagram of a duplicate identification algorithm according to a specific embodiment of the present invention.
[0176] like Figure 3 As shown, the duplicate identification algorithm includes:
[0177] Algorithm 1: Business ID;
[0178] Algorithm 2: (Business ID + Address) or (Legal Person + Address);
[0179] Algorithm three: business ID + address + main business.
[0180] refer to Figure 4 , which is a schematic diagram of a conflict resolution algorithm according to a specific embodiment of the present invention.
[0181] like Figure 4 As shown, the conflict resolution algorithm includes:
[0182] Algorithm 1: Forced overwrite update;
[0183] Algorithm 2: Forced deletion of new additions;
[0184] Algorithm 3: More aggressive processing based on trust level;
[0185] Algorithm 4: conservative processing based on trust level;
[0186] Algorithm 5: Processing by trust level + manual intervention.
[0187] refer to Figure 5 , which is a schematic diagram of a manual intervention method according to a specific embodiment of the present invention.
[0188] like Figure 5 As shown, manual intervention methods include:
[0189] Algorithm 1: forced update;
[0190] Algorithm 2: Forced discard.
[0191] Compared with existing technologies, the existing system does not perform hierarchical processing and lifecycle management on the quality (trustworthiness) of incremental data, nor does it perform hierarchical protection on the quality (trustworthiness) of existing data, and currently does not identify and exclude duplicate data.
[0192] Therefore, with the accumulation of existing data and the increase in incremental data sources, duplicate data may increase rapidly, resulting in a decline in data quality, confusing page display, and affecting customer experience and team reputation.
[0193] The method for processing merchant duplicate data proposed in the present invention can eliminate duplicate entries for newly added data and protect the quality of existing data, thereby eliminating data coverage caused by duplicate data of low quality in newly added data. It can reuse solutions for scenarios with other multiple data sources and can achieve traceability, maintainability, customization, and intervention of data.
[0194] During the implementation of the overall solution, data quality was improved and customer experience was enhanced through data lifecycle management and configurable data duplication identification algorithms and processing strategies.
[0195] After introducing the method of the exemplary embodiment of the present invention, next, reference is made to Figure 6 A system for processing merchant duplicate data according to an exemplary embodiment of the present invention is introduced.
[0196] The implementation of the merchant duplicate data processing system can be found in the implementation of the above-mentioned method, and the repeated parts will not be repeated here. The terms "module" or "unit" used below can refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated.
[0197] Based on the same inventive concept, the present invention also proposes a system for processing merchant duplicate data, such as Figure 6 As shown, the system includes:
[0198] The data acquisition module 610 is used to acquire the merchant's newly added data and update the data batch status according to the intermediate information of the newly added data;
[0199] The duplicate identification module 620 is used to obtain a duplicate identification algorithm based on the data batch status and perform duplicate identification on the newly added data; if the identification result is that the data is duplicated, the newly added data is merged and the data protection level and data status are recorded; if the identification result is that the data is not duplicated, a data conflict flag is recorded;
[0200] The intelligent conflict processing module 630 is configured to obtain a conflict resolution strategy based on the data batch status and intelligently process the conflict according to the conflict resolution strategy. If the conflict is resolved, the data is updated, merged, or discarded, and the data conflict resolution strategy and data status are recorded. If the conflict is not resolved, the data status is recorded and manual intervention is invoked.
[0201] The conflict manual intervention module 640 is used to display the conflict data to the operator, who then manually selects an intervention strategy to intervene in the conflict data and adjust the confidence level of the conflict data;
[0202] The trust setting module 650 is used to determine whether the trust meets the trust upgrade requirements based on the adjusted trust; if so, the trust of the data source of the newly added data is updated; if not, the newly added data is directly saved.
[0203] In one embodiment, the data intermediate information includes: original data, source data, batch data, data credibility and status records.
[0204] In one embodiment, the data acquisition module 610 is specifically configured to:
[0205] Acquire the merchant's newly added data, extract the data intermediate information based on the data source registry of the newly added data, and save it to the data intermediate table; wherein the data source registry contains at least the batch identifier, batch code, batch name, data source identifier, data batch confidence, duplicate identification algorithm, and conflict resolution algorithm; the data intermediate table contains at least the original data information, data source identifier, data batch identifier, data confidence, merge status, conflict identifier, and conflict strategy; the merge status is collection completed, merging in progress, merge conflict resolution in progress, or merge completed;
[0206] In one embodiment, the data acquisition module 610 is further configured to:
[0207] The data batch registration table is updated according to the data intermediate information, and the data batch status is updated; which at least includes: batch identification, batch code, batch name, data source identification, data batch trust, progress status, duplicate identification algorithm and conflict resolution algorithm; the progress status is adjusted with the progress of processing, and the progress status is ready to collect, collecting, collection completed, merging, merging conflict resolution or merging completed.
[0208] In one embodiment, the duplicate identification module 620 is specifically configured to:
[0209] According to the duplicate identification algorithm, the corresponding algorithm interface implementation class is obtained to perform duplicate identification on the newly added data.
[0210] In one embodiment, the intelligent conflict processing module 630 is specifically configured to:
[0211] An intelligent conflict handling strategy is acquired according to the conflict resolution strategy, and the conflict is intelligently handled and resolved using the intelligent conflict handling strategy.
[0212] In one embodiment, the conflict manual intervention module 640 is specifically configured to:
[0213] The conflicting data are compared according to the conflict identification, and the comparison results are displayed to the operator, who manually selects an intervention strategy to intervene in the conflicting data, adjusts the confidence level of the conflicting data, and records the data status and intervention behavior.
[0214] In one embodiment, the conflict manual intervention module is further configured to:
[0215] After the intervention processing of the conflicting data is completed, the intervention behavior is recorded in the intervention data record table; wherein the intervention data record table at least includes: intervention identification, data source identification, data batch identification, data identification, conflict identification, intervention type, intervention strategy and intervention result; the intervention result is merged or discarded.
[0216] In one embodiment, the trust setting module 650 is specifically configured to:
[0217] Collecting manual intervention strategy data, and judging whether the trust level meets the trust level upgrade requirement based on the trust level after the manual intervention.
[0218] In one embodiment, the duplicate identification algorithm at least includes duplicate identification using the business number, duplicate identification using the business number and address, duplicate identification using the legal person and address, and duplicate identification using the business number, address and main business.
[0219] In one embodiment, the conflict resolution strategy includes at least: forced overwriting and updating, forced deletion and addition, aggressive processing based on trust, conservative processing based on trust, and processing based on trust and manual intervention; wherein, the manual intervention method includes at least: forced updating and forced abandonment; and performing corresponding manual intervention processing methods by manually comparing duplicate data.
[0220] The merchant duplicate data processing system proposed in the present invention can process the merchant's new data through data lifecycle management, using an assembleable data duplicate identification algorithm and processing strategy, to achieve duplicate entry of the new data, and perform quality protection on the existing data, eliminating duplicate data of low quality in the new data and causing data overwriting, thereby improving data quality and improving customer experience.
[0221] During the processing, by introducing the lifecycle management of imported data and the trust level of merchant data, as well as the identification algorithm and processing strategy for customizable assembly of duplicate data, a data duplication solution is provided in scenarios with multiple data sources; and the solution can be reused in scenarios with other multiple data sources, making the data traceable, maintainable, customizable, and intervention-friendly.
[0222] It should be noted that while the detailed description above mentions several modules of the merchant duplicate data processing system, this division is merely exemplary and not mandatory. In practice, according to embodiments of the present invention, the features and functions of two or more modules described above may be embodied in a single module. Conversely, the features and functions of a single module described above may be further divided and embodied by multiple modules.
[0223] Based on the above invention concept, Figure 7 As shown, the present invention further provides a computer device 700, including a memory 710, a processor 720, and a computer program 730 stored in the memory 710 and executable on the processor 720. When the processor 720 executes the computer program 730, a method for processing duplicate merchant data is implemented. The method includes:
[0224] Obtain the merchant's newly added data and update the data batch status according to the intermediate data information of the newly added data;
[0225] Obtain a duplicate identification algorithm based on the data batch status and perform duplicate identification on the newly added data; if the identification result is that the data is duplicated, merge the newly added data and record the data protection level and data status; if the identification result is that the data is not duplicated, record a data conflict identifier;
[0226] Obtaining a conflict resolution strategy based on the data batch status, and intelligently handling the conflict based on the conflict resolution strategy; if the conflict is resolved, then updating, merging, or discarding the data, and recording the data conflict resolution strategy and data status; if the conflict is not resolved, then recording the data status and invoking manual intervention;
[0227] The conflicting data is displayed to the operator, who then manually selects an intervention strategy to intervene in the conflicting data and adjust the confidence level of the conflicting data.
[0228] According to the adjusted trust level, it is determined whether the trust level meets the trust level upgrade requirement; if so, the trust level of the data source of the newly added data is updated; if not, the newly added data is directly saved.
[0229] Based on the aforementioned inventive concept, the present invention proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the aforementioned method for processing merchant duplicate data is implemented.
[0230] The method and system for processing merchant duplicate data proposed in the present invention utilizes an assembleable data duplication identification algorithm and processing strategy to process merchant newly added data, thereby eliminating duplicate entry of newly added data and protecting the quality of existing data, thereby eliminating duplicate data of low quality in newly added data and overwriting it, thereby improving data quality and customer experience. During the processing process, by introducing lifecycle management of imported data and merchant data trust levels, as well as customizable identification algorithms and processing strategies for duplicate data, a data duplication solution is provided for scenarios with multiple data sources, and the solution can be reused for scenarios with other multiple data sources, making data traceable, maintainable, customizable, and intervenible.
[0231] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0232] The present invention is described with reference to flowcharts and / or block diagrams of methods and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0233] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0234] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0235] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A method for processing merchant duplicate data, characterized in that: This method uses an assembler-based duplicate identification algorithm and processing strategy to process new merchant data and remove duplicate entries for the new data, including: Obtain the merchant's newly added data and update the data batch status according to the intermediate data information of the newly added data; Obtain a duplicate identification algorithm based on the data batch status and perform duplicate identification on the newly added data; if the identification result is that the data is duplicated, merge the newly added data and record the data protection level and data status; if the identification result is that the data is not duplicated, record a data conflict identifier; Obtaining a conflict resolution strategy based on the data batch status, and intelligently handling the conflict based on the conflict resolution strategy; if the conflict is resolved, then updating, merging, or discarding the data, and recording the data conflict resolution strategy and data status; if the conflict is not resolved, then recording the data status and invoking manual intervention; The conflicting data is displayed to the operator, who then manually selects an intervention strategy to intervene in the conflicting data and adjust the confidence level of the conflicting data. According to the adjusted trust level, it is determined whether the trust level meets the trust level upgrade requirement; if so, the trust level of the data source of the newly added data is updated; if not, the newly added data is directly saved; The process of obtaining the merchant's newly added data and updating the data batch status according to the intermediate data information of the newly added data includes: Acquire the merchant's newly added data, extract the data intermediate information based on the data source registry of the newly added data, and save it to the data intermediate table; wherein the data source registry contains at least the batch identifier, batch code, batch name, data source identifier, data batch confidence, duplicate identification algorithm, and conflict resolution algorithm; the data intermediate table contains at least the original data information, data source identifier, data batch identifier, data confidence, merge status, conflict identifier, and conflict strategy; the merge status is collection completed, merging in progress, merge conflict resolution in progress, or merge completed; Update the data batch registry according to the data intermediate information, and update the data batch status; which at least includes: batch identification, batch code, batch name, data source identification, data batch confidence, progress status, duplicate identification algorithm and conflict resolution algorithm; the progress status is adjusted as the processing progresses, and the progress status is ready to collect, collecting, collection completed, merging, merging conflict resolution or merging completed; The conflicting data is presented to the operator, who then manually selects an intervention strategy to intervene in the conflicting data and adjust the trust level of the conflicting data, including: The conflicting data are compared according to the conflict identification, and the comparison results are displayed to the operator, who manually selects an intervention strategy to intervene in the conflicting data, adjusts the confidence level of the conflicting data, and records the data status and intervention behavior.
2. The method for processing merchant duplicate data according to claim 1, characterized in that: The data intermediate information includes: original data, source data, batch data, data trustworthiness and status records.
3. The method for processing merchant duplicate data according to claim 2, characterized in that: Obtaining a duplicate identification algorithm based on the data batch status and performing duplicate identification on the newly added data includes: According to the duplicate identification algorithm, the corresponding algorithm interface implementation class is obtained to perform duplicate identification on the newly added data.
4. The method for processing merchant duplicate data according to claim 3, characterized in that: Obtaining a conflict resolution strategy according to the data batch state, and performing conflict processing according to the conflict resolution strategy, including: An intelligent conflict handling strategy is acquired according to the conflict resolution strategy, and the conflict is intelligently handled and resolved using the intelligent conflict handling strategy.
5. The method for processing merchant duplicate data according to claim 4, characterized in that: The conflicting data is displayed to the operator, who then manually selects an intervention strategy to intervene in the conflicting data and adjust the confidence level of the conflicting data. This also includes: After the intervention processing of the conflicting data is completed, the intervention behavior is recorded in the intervention data record table; wherein the intervention data record table at least includes: intervention identification, data source identification, data batch identification, data identification, conflict identification, intervention type, intervention strategy and intervention result; the intervention result is merged or discarded.
6. The method for processing merchant duplicate data according to claim 1, characterized in that: Determining, based on the adjusted trustworthiness, whether the trustworthiness meets trustworthiness upgrade requirements includes: Collecting manual intervention strategy data, and judging whether the trust level meets the trust level upgrade requirement based on the trust level after the manual intervention.
7. The method for processing merchant duplicate data according to claim 1 or 3, characterized in that: The duplicate identification algorithm at least includes duplicate identification using the business number, duplicate identification using the business number and address, duplicate identification using the legal person and address, and duplicate identification using the business number, address and main business.
8. The method for processing merchant duplicate data according to claim 1 or 4, characterized in that: The conflict resolution strategy includes at least: forced overwriting and updating, forced deletion and addition, aggressive processing based on trust, conservative processing based on trust, and processing based on trust and manual intervention.
9. The method for processing merchant duplicate data according to claim 8, characterized in that: The manual intervention method at least includes: forced update and forced abandonment; and manual comparison of duplicate data and corresponding manual intervention processing.
10. A merchant duplicate data processing system, characterized in that: The system uses configurable duplicate identification algorithms and processing strategies to process new merchant data and remove duplicate entries for the new data, including: A data acquisition module is used to acquire new data from merchants and update the data batch status according to the intermediate data information of the new data; A duplicate identification module is used to obtain a duplicate identification algorithm based on the data batch status and perform duplicate identification on the newly added data; if the identification result is that the data is duplicated, the newly added data is merged and the data protection level and data status are recorded; if the identification result is that the data is not duplicated, a data conflict identifier is recorded; A conflict intelligent processing module is used to obtain a conflict resolution strategy based on the data batch status and intelligently handle conflicts according to the conflict resolution strategy; if the conflict is resolved, the data is updated, merged, or discarded, and the data conflict resolution strategy and data status are recorded; if the conflict is not resolved, the data status is recorded and manual intervention is invoked; The conflict manual intervention module is used to display conflict data to operators, who can manually select intervention strategies to intervene in the conflict data and adjust the confidence level of the conflict data. A trust setting module is used to determine whether the trust meets the trust upgrade requirements based on the adjusted trust; if so, the trust of the data source of the newly added data is updated; if not, the newly added data is directly saved; The data acquisition module is specifically used for: Acquire the merchant's newly added data, extract the data intermediate information based on the data source registry of the newly added data, and save it to the data intermediate table; wherein the data source registry contains at least the batch identifier, batch code, batch name, data source identifier, data batch confidence, duplicate identification algorithm, and conflict resolution algorithm; the data intermediate table contains at least the original data information, data source identifier, data batch identifier, data confidence, merge status, conflict identifier, and conflict strategy; the merge status is collection completed, merging in progress, merge conflict resolution in progress, or merge completed; Update the data batch registry according to the data intermediate information, and update the data batch status; which at least includes: batch identification, batch code, batch name, data source identification, data batch confidence, progress status, duplicate identification algorithm and conflict resolution algorithm; the progress status is adjusted as the processing progresses, and the progress status is ready to collect, collecting, collection completed, merging, merging conflict resolution or merging completed; The conflict manual intervention module is specifically used to: The conflicting data are compared according to the conflict identification, and the comparison results are displayed to the operator, who manually selects an intervention strategy to intervene in the conflicting data, adjusts the confidence level of the conflicting data, and records the data status and intervention behavior.
11. The merchant duplicate data processing system according to claim 10, characterized in that: The data intermediate information includes: original data, source data, batch data, data trustworthiness and status records.
12. The merchant duplicate data processing system according to claim 11, characterized in that: The repeat identification module is specifically used for: According to the duplicate identification algorithm, the corresponding algorithm interface implementation class is obtained to perform duplicate identification on the newly added data.
13. The merchant duplicate data processing system according to claim 12, characterized in that: The conflict intelligent processing module is specifically used for: An intelligent conflict handling strategy is acquired according to the conflict resolution strategy, and the conflict is intelligently handled and resolved using the intelligent conflict handling strategy.
14. The merchant duplicate data processing system according to claim 13, characterized in that: The conflict manual intervention module is also used to: After the intervention processing of the conflicting data is completed, the intervention behavior is recorded in the intervention data record table; wherein the intervention data record table at least includes: intervention identification, data source identification, data batch identification, data identification, conflict identification, intervention type, intervention strategy and intervention result; the intervention result is merged or discarded.
15. The merchant duplicate data processing system according to claim 10, characterized in that: The trust setting module is specifically used to: Collecting manual intervention strategy data, and judging whether the trust level meets the trust level upgrade requirement based on the trust level after the manual intervention.
16. The merchant duplicate data processing system according to claim 10 or 12, characterized in that: The duplicate identification algorithm at least includes duplicate identification using the business number, duplicate identification using the business number and address, duplicate identification using the legal person and address, and duplicate identification using the business number, address and main business.
17. The merchant duplicate data processing system according to claim 10 or 13, characterized in that: The conflict resolution strategy includes at least: forced overwriting and updating, forced deletion and addition, aggressive processing based on trust, conservative processing based on trust, and processing based on trust and manual intervention; wherein, the manual intervention method includes at least: forced updating and forced abandonment; and corresponding manual intervention processing methods are performed by manually comparing duplicate data.
18. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.
19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
A data quality solution based on knowledge
CN102930023A
Data collision processing method and device
CN104113571A