Automatic data conflict resolution method and apparatus, electronic device, and medium
Patent Information
- Application Number
- PCT/CN2025/093099
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-05-07
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025093099_01102026_PF_FP_ABST
Abstract
Description
Automatic data conflict resolution methods, devices, electronic equipment and media
[0001] This application claims priority to Chinese Patent Application No. 202510363177.6, filed on March 26, 2025, entitled “Automatic Data Conflict Resolution Method, Apparatus, Electronic Device and Medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This document relates to the field of computer technology, and in particular to an automatic data conflict resolution method, device, electronic device, and medium. Background Technology
[0003] In multi-departmental collaborations within large enterprises, each department often provides its own answers or information based on different business needs and data sources. This is particularly true during the training data preparation phase of large AI language models, where data from different departments (such as technology, customer service, and sales) frequently presents conflicts and inconsistencies. For example, different answers to the same question may contain different interpretations, different expressions, or different levels of detail. If these conflicting data are not resolved, they can increase data noise during AI model training, thereby affecting model performance.
[0004] Existing conflict resolution methods typically rely on manual intervention or static rules, failing to automatically select the optimal resolution strategy based on different departments, data types, and conflict scenarios. Therefore, an intelligent, automated solution is needed that can dynamically select appropriate rules to resolve conflicts based on their specific type and data context. Summary of the Invention
[0005] The purpose of this invention is to provide an automatic data conflict resolution method, apparatus, electronic device, and medium, which aims to solve the above-mentioned problems in the prior art.
[0006] This invention provides an automatic data conflict resolution method for preparing training data for large AI language models, including:
[0007] Automatically identify data conflicts through text comparison and semantic analysis;
[0008] Based on a rule base and reinforcement learning algorithm, the optimal conflict resolution rule is dynamically selected and executed based on the identified data conflicts.
[0009] The optimal version to be executed is selected according to the conflict resolution rules, and the data is merged to save the key information of the conflict data;
[0010] While ensuring the integrity of the training data for the large AI language model, redundant data is removed from the merged conflicting data, and incremental updates are performed.
[0011] This invention provides an automatic data conflict resolution device for preparing training data for large AI language models, comprising:
[0012] The conflict identification module is used to automatically identify data conflicts through text comparison and semantic analysis;
[0013] The rule engine module is used to dynamically select the optimal conflict resolution rule based on the identified data conflicts, and execute the conflict resolution rule, based on the rule base and reinforcement learning algorithm.
[0014] The version management and data merging module is used to select the optimal version to be executed according to the conflict resolution rules, merge the data, and save the key information of the conflict data.
[0015] The data deduplication and incremental update module is used to remove redundant data from the merged conflicting data and perform incremental updates while ensuring the integrity of the training data for the large AI language model.
[0016] This invention also provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the above-described steps for automatically resolving data conflicts in the preparation of training data for large AI language models.
[0017] This invention also provides a computer-readable storage medium storing an information transmission implementation program. When the program is executed by a processor, it implements the above-described steps for automatically resolving data conflicts in the preparation of training data for large AI language models.
[0018] By employing embodiments of the present invention and applying an intelligent rule engine, data conflicts are automatically detected and resolved, the training dataset of the AI model is optimized, and the integrity and consistency of the data are ensured. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 is a flowchart of an automatic data conflict resolution method for preparing training data for large AI language models according to an embodiment of the present invention.
[0021] Figure 2 is a schematic diagram of an automatic data conflict resolution device for preparing training data for large AI language models according to an embodiment of the present invention.
[0022] Figure 3 is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0023] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0024] Method Implementation Examples
[0025] According to an embodiment of the present invention, an automatic data conflict resolution method for preparing training data for large AI language models is provided. Figure 1 is a flowchart of the automatic data conflict resolution method for preparing training data for large AI language models according to an embodiment of the present invention. As shown in Figure 1, the automatic data conflict resolution method for preparing training data for large AI language models according to an embodiment of the present invention specifically includes:
[0026] Step S101 involves automatically identifying data conflicts through text comparison and semantic analysis; specifically including:
[0027] The text data provided by each department is preprocessed and converted into vector representations. The similarity between vectors from different departments is calculated. Based on the similarity, it is determined whether there is a conflict between the texts. If a conflict is determined, the conflict type is identified and the conflicting data is automatically marked according to the conflict type. The conflict types specifically include: semantic conflict, format conflict, information redundancy, and version conflict.
[0028] Step S102, based on the rule base and reinforcement learning algorithm, dynamically selects the optimal conflict resolution rule based on the identified data conflicts, and executes the conflict resolution rule; specifically including:
[0029] A rule base is established, and the most suitable conflict solution is selected based on the conflict type and the rules in the rule base. For semantic conflicts, the most suitable content is automatically selected and retained according to the target task. For format conflicts, the format is unified or converted according to the target format. The rule base includes static rules and dynamic rules. Static rules are based on department priority, timestamp, and data quality, while dynamic rules are obtained by reinforcement learning optimization through historical conflict resolution data.
[0030] Step S103 involves selecting the optimal version to execute based on the conflict resolution rules, merging the data, and saving key information about the conflict data; specifically including:
[0031] The latest version of the data or the most authoritative data source is selected based on the rule base. In case of version conflicts, the newer data is selected first, or information from multiple versions is combined and merged. When merging data, conflicting data is merged according to the selected rules, and key information is retained. When there are conflicts between technical issues and user operation guidance, the two are merged into a comprehensive answer. The key information specifically includes important technical details or user guidance information.
[0032] Step S104 involves removing redundant data from the merged conflicting data while ensuring the integrity of the AI language large model training data, and then performing incremental updates. Specifically, this includes:
[0033] When processing multi-version data, duplicate content is removed, and only newly added valid data is retained. Duplicate data is detected using hash algorithms and text similarity calculations, and duplicate training data is deleted. Each time the data is updated, it is automatically compared with the existing data and incrementally updated. The newly added data is automatically integrated into the existing training data to ensure that the training data is always up-to-date.
[0034] In summary, the intelligent rule engine algorithm for automatically resolving cross-departmental data conflicts, as described in this invention, can automatically select and execute the optimal resolution strategy based on data conflicts between departments. Through the collaborative work of multiple modules, it ensures efficient and accurate handling of cross-departmental data conflicts when preparing training data for large AI language models, thereby improving data consistency and quality.
[0035] The technical solutions described above in the embodiments of the present invention will be explained in detail below. The specific steps are as follows:
[0036] 1. Text preprocessing: Preprocess the text data provided by various departments, including stop word removal, word segmentation, and standardization.
[0037] Text representation and similarity calculation:
[0038] Use techniques such as TF-IDF, Word2Vec, and BERT to convert text into vector representations.
[0039] Calculate the similarity between responses from different departments, and determine whether there are conflicts between texts based on cosine similarity or other similarity metrics.
[0040] Conflict type identification:
[0041] Semantic conflicts: such as conflicts between technical details and user operation instructions.
[0042] Formatting conflicts: For example, different departments use different text formats or expression styles.
[0043] Information redundancy: such as duplicate data or descriptions.
[0044] Version conflicts: For example, answers from different points in time may have version differences.
[0045] The embodiments of the present invention automatically mark conflicting data through the above steps.
[0046] 2. Use a dynamic rule base and reinforcement learning algorithms to select the optimal conflict resolution strategy. The main steps are as follows:
[0047] The rule base includes static rules and dynamic rules. Static rules are based on departmental priority, timestamps, data quality, etc., while dynamic rules are optimized through reinforcement learning by resolving historical conflicts.
[0048] Static rules: Prioritize based on departmental priority (e.g., technical department, customer service department), timestamp (e.g., latest version priority), and text quality (e.g., level of detail, accuracy, etc.).
[0049] Dynamic rules: Rule selection is continuously optimized through reinforcement learning algorithms. For example, the model automatically adjusts the weights and strategies in the rule base based on the handling effect of historical conflicts.
[0050] Rule Selection and Execution: Based on the conflict type and rules in the rule base, the rule engine selects the most suitable conflict resolution method. For semantic conflicts, the rule engine automatically selects the most suitable content to retain based on the target task (such as technical issues or user operation guidance). For format conflicts, the rule engine either unifies the format or converts it according to the target format.
[0051] 3. Select the optimal version and merge conflicting data. The main steps are as follows:
[0052] Version control: The system selects the latest version of data or the most authoritative data source based on the rule base. In case of version conflicts, the system will prioritize the newer data or combine information from multiple versions.
[0053] Data merging: Conflicting data is merged according to selected rules, ensuring that critical information is preserved. During the merging process, important technical details or user guidance information are avoided from being lost.
[0054] Merging strategy: When there is a conflict between technical issues and user operation instructions, the system will merge the two into a comprehensive answer that is both technically sound and easy for users to understand.
[0055] 4. Remove redundant data and ensure the integrity and timeliness of the training data.
[0056] Data deduplication: When processing multi-version data, duplicate content is removed, retaining only newly added and valid data. Hash algorithms and text similarity calculations are used to detect duplicate data, avoiding redundant training data.
[0057] Incremental update: Each time the data is updated, it is automatically compared with the existing data and an incremental update is performed. Newly added data is automatically integrated into the existing training data to ensure that the training data is always up-to-date.
[0058] Below is a simulation of an intelligent rule engine algorithm written in Python, demonstrating the process of automatically resolving cross-departmental data conflicts. It will use simple modules such as text similarity calculation, rule engine, data merging, and version management to simulate the conflict detection and resolution process.
[0059] When the above code was executed, it simulated the answers provided by the technical and customer service departments, and processed the text data through conflict detection, rule engine, version management, and data merging modules. The following are the output results of the simulation experiment: Text similarity: 0.66, conflict detected, type: format_conflict, conflict resolution strategy: Please merge the two contents, retaining technical details while adding user-friendly guidance. Merged answer: Please download the installation package from the official website and run the installer. During installation, you can choose "default settings" for quick installation, or choose "custom installation" for advanced settings such as proxy settings and network configuration. In this experiment, the answers provided by the two departments had some differences, therefore they were judged as a conflict (text similarity of 0.66, below 0.7). Based on the conflict type, the rule engine selected the "format conflict" resolution strategy and merged the two contents, generating the final merged answer.
[0060] The technical solution of this invention automatically detects and resolves data conflicts in cross-departmental collaboration through collaborative work, ensuring optimal performance in content consistency and quality of the generated AI training dataset. It uses text similarity calculation or semantic analysis techniques to automatically identify conflicts and inconsistencies in data provided by different departments. Based on a static rule base and reinforcement learning algorithms, it dynamically selects the most suitable conflict resolution strategy to resolve issues such as semantic conflicts, format conflicts, information redundancy, and version conflicts. It selects the optimal version of the data according to the conflict resolution strategy and merges the data, ensuring data consistency and preventing the loss of key information. This invention is applicable to the preparation of training data for large-scale AI language models, automatically handling text data conflicts from multiple departments to generate high-quality, highly consistent training datasets. By using text representation methods such as TF-IDF and BERT and similarity calculation algorithms, it automatically identifies conflicts in the text and provides conflict types and solutions. It employs version control technology and intelligent merging strategies, combined with departmental priorities, timestamps, and other information, to ensure the timeliness and integrity of the training data. Furthermore, this invention can dynamically adjust and optimize based on historical data to adapt to the conflict handling needs of different departments and scenarios.
[0061] In summary, the technical solution adopted in this invention can automatically select and execute the optimal resolution strategy based on data conflicts between departments. The core of the algorithm is to ensure efficient and accurate handling of cross-departmental data conflicts during the preparation of large-scale AI language model training data through the collaborative work of multiple modules, thereby improving data consistency and quality.
[0062] Device Example 1
[0063] According to an embodiment of the present invention, an automatic data conflict resolution device for preparing training data for large AI language models is provided. Figure 2 is a schematic diagram of the automatic data conflict resolution device for preparing training data for large AI language models according to an embodiment of the present invention. As shown in Figure 2, the automatic data conflict resolution device for preparing training data for large AI language models according to an embodiment of the present invention specifically includes:
[0064] The conflict identification module 20 is used to automatically identify data conflicts through text comparison and semantic analysis; specifically, it is used for:
[0065] The text data provided by each department is preprocessed and converted into vector representation. The similarity between vectors from different departments is calculated. Based on the similarity, it is determined whether there is a conflict between the texts. If a conflict is determined, the conflict type is identified and the conflict data is automatically marked according to the conflict type. The conflict types specifically include: semantic conflict, format conflict, information redundancy, and version conflict.
[0066] Rule engine module 22 is used to dynamically select the optimal conflict resolution rule based on the identified data conflicts, using a rule base and reinforcement learning algorithm, and to execute the conflict resolution rule; specifically, it is used for:
[0067] A rule base is established, and the most suitable conflict solution is selected based on the conflict type and the rules in the rule base. For semantic conflicts, the most suitable content is automatically selected and retained according to the target task. For format conflicts, the format is unified or converted according to the target format. The rule base includes static rules and dynamic rules. Static rules are based on department priority, timestamp, and data quality, while dynamic rules are obtained by reinforcement learning optimization through historical conflict resolution data.
[0068] Version management and data merging module 24 is used to select the optimal version to execute according to the conflict resolution rules, merge the data, and save key information of the conflict data; specifically, it is used for:
[0069] The latest version of the data or the most authoritative data source is selected based on the rule base. In case of version conflict, the newer data is selected first, or information from multiple versions is combined and merged. When merging data, conflicting data is merged according to the selected rules, and key information is retained. When there is a conflict between technical issues and user operation guidance, the two are merged into a comprehensive answer. The key information specifically includes: important technical details or user guidance information.
[0070] The data deduplication and incremental update module 26 is used to remove redundant data from the merged conflicting data and perform incremental updates while ensuring the integrity of the AI language large model training data. Specifically, it is used for:
[0071] When processing multi-version data, duplicate content is removed, and only newly added valid data is retained. Duplicate data is detected using hash algorithms and text similarity calculations, and duplicate training data is deleted. Each time the data is updated, it is automatically compared with the existing data and incrementally updated. The newly added data is automatically integrated into the existing training data to ensure that the training data is always up-to-date.
[0072] The technical solution provided in this invention can automatically resolve data conflicts from multiple departments during the text data preparation process for training large-scale AI language models through intelligent conflict identification, rule selection, data merging, and deduplication, ensuring the accuracy, consistency, and high quality of training data. This system is suitable for large enterprises, cross-departmental collaboration, and scenarios involving multi-source data, significantly improving the quality of AI training data and thus enhancing the performance of large-scale AI language models.
[0073] The embodiments of the present invention are device embodiments corresponding to the above method embodiments. The specific operation of each module can be understood with reference to the description of the method embodiments, and will not be repeated here.
[0074] Device Example 2
[0075] An embodiment of the present invention provides an electronic device, as shown in FIG3, including: a memory 30, a processor 32, and a computer program stored in the memory 30 and executable on the processor 32. When the computer program is executed by the processor 32, it implements the steps described in the method embodiment.
[0076] Device Example 3
[0077] This invention provides a computer-readable storage medium storing an information transmission implementation program, which, when executed by a processor 32, performs the steps described in the method embodiment.
[0078] The computer-readable storage media described in this embodiment include, but are not limited to, ROM, RAM, disk, or optical disk.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An automatic data conflict resolution method for preparing training data for large AI language models, characterized in that, include: Automatically identify data conflicts through text comparison and semantic analysis; Based on a rule base and reinforcement learning algorithm, the optimal conflict resolution rule is dynamically selected and executed based on the identified data conflicts. The optimal version to be executed is selected according to the conflict resolution rules, and the data is merged to save the key information of the conflict data; While ensuring the integrity of the training data for the large AI language model, redundant data is removed from the merged conflicting data, and incremental updates are performed.
2. The method according to claim 1, characterized in that, Automatic identification of data conflicts through text comparison and semantic analysis includes: The text data provided by each department is preprocessed and converted into vector representations. The similarity between vectors from different departments is calculated. Based on the similarity, it is determined whether there is a conflict between the texts. If a conflict is determined, the conflict type is identified and the conflicting data is automatically marked according to the conflict type. The conflict types specifically include: semantic conflict, format conflict, information redundancy, and version conflict.
3. The method according to claim 1, characterized in that, Based on a rule base and reinforcement learning algorithm, and based on the identified data conflicts, the optimal conflict resolution rule is dynamically selected and executed. Specifically, this includes: A rule base is established, and the most suitable conflict solution is selected based on the conflict type and the rules in the rule base. For semantic conflicts, the most suitable content is automatically selected and retained according to the target task. For format conflicts, the format is unified or converted according to the target format. The rule base includes static rules and dynamic rules. Static rules are based on department priority, timestamp, and data quality, while dynamic rules are obtained by reinforcement learning optimization through historical conflict resolution data.
4. The method according to claim 1, characterized in that, The optimal version is selected for execution based on the conflict resolution rules, and data is merged. Key information regarding the saved conflict data includes: The latest version of the data or the most authoritative data source is selected based on the rule base. In case of version conflicts, the newer data is selected first, or information from multiple versions is combined and merged. When merging data, conflicting data is merged according to the selected rules, and key information is retained. When there are conflicts between technical issues and user operation guidance, the two are merged into a comprehensive answer. The key information specifically includes important technical details or user guidance information.
5. The method according to claim 1, characterized in that, While ensuring the integrity of the training data for the large AI language model, redundant data is removed from the merged conflicting data, and incremental updates are performed. Specifically, this includes: When processing multi-version data, duplicate content is removed, and only newly added valid data is retained. Duplicate data is detected using hash algorithms and text similarity calculations, and duplicate training data is deleted. Each time the data is updated, it is automatically compared with the existing data and incrementally updated. The newly added data is automatically integrated into the existing training data to ensure that the training data is always up-to-date.
6. A device for automatically resolving data conflicts in the preparation of training data for large AI language models, characterized in that, include: The conflict identification module is used to automatically identify data conflicts through text comparison and semantic analysis; The rule engine module is used to dynamically select the optimal conflict resolution rule based on the identified data conflicts, and execute the conflict resolution rule, based on the rule base and reinforcement learning algorithm. The version management and data merging module is used to select the optimal version to be executed according to the conflict resolution rules, merge the data, and save the key information of the conflict data. The data deduplication and incremental update module is used to remove redundant data from the merged conflicting data and perform incremental updates while ensuring the integrity of the training data for the large AI language model.
7. The apparatus according to claim 6, characterized in that, The conflict identification module is specifically used for: The text data provided by each department is preprocessed and converted into vector representation. The similarity between vectors from different departments is calculated. Based on the similarity, it is determined whether there is a conflict between the texts. If a conflict is determined, the conflict type is identified and the conflict data is automatically marked according to the conflict type. The conflict types specifically include: semantic conflict, format conflict, information redundancy, and version conflict. The rules engine module is specifically used for: A rule base is established, and the most suitable conflict solution is selected based on the conflict type and the rules in the rule base. For semantic conflicts, the most suitable content is automatically selected and retained according to the target task. For format conflicts, the format is unified or converted according to the target format. The rule base includes static rules and dynamic rules. Static rules are based on department priority, timestamp, and data quality, while dynamic rules are obtained by reinforcement learning optimization through historical conflict resolution data.
8. The apparatus according to claim 6, characterized in that, The version management and data merging module is specifically used for: The latest version of the data or the most authoritative data source is selected based on the rule base. In case of version conflict, the newer data is selected first, or information from multiple versions is combined and merged. When merging data, conflicting data is merged according to the selected rules, and key information is retained. When there is a conflict between technical issues and user operation guidance, the two are merged into a comprehensive answer. The key information specifically includes: important technical details or user guidance information. The data deduplication and incremental update module is specifically used for: When processing multi-version data, duplicate content is removed, and only newly added valid data is retained. Duplicate data is detected using hash algorithms and text similarity calculations, and duplicate training data is deleted. Each time the data is updated, it is automatically compared with the existing data and incrementally updated. The newly added data is automatically integrated into the existing training data to ensure that the training data is always up-to-date.
9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of automatically resolving data conflict in preparation of training data for large AI language models as described in any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an information transmission implementation program, which, when executed by a processor, implements the steps of automatically resolving data conflict in preparation of training data for large AI language models as described in any one of claims 1 to 5.