Method and device for automatically solving data conflict, electronic equipment and medium
Through text comparison and semantic analysis combined with rule base and reinforcement learning algorithm, conflict resolution rules are dynamically selected and cross-departmental data conflicts are automatically handled, which solves the problem that the optimal solution strategy cannot be selected in the existing technology, and improves the quality and consistency of AI model training data.
Patent Information
- Application Number
- CN202510363177.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-18
AI Technical Summary
Existing conflict resolution methods cannot automatically select the optimal solution strategy based on different departments, different data types and different conflict situations, resulting in increased data noise during AI model training, affecting model performance.
Automatically identify data conflicts through text comparison and semantic analysis, dynamically select the optimal conflict resolution rules based on the rule base and reinforcement learning algorithm, execute conflict resolution rules, and retain key information when merging data, remove redundant data, and perform incremental updates.
It realizes that in the preparation process of AI language big model training data, cross-departmental data conflicts are handled automatically and intelligently, ensuring data integrity and consistency, and improving data quality and model training efficiency.
Smart Images

Figure CN120336339A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and particularly to a method, apparatus, electronic device, and medium for automatically resolving data conflicts. Background Art
[0002] In the multi-department collaboration of large enterprises, each department often provides its own answers or information based on different business requirements and data sources. Especially in the training data preparation stage of large AI language models, data from different departments (such as the technology department, customer service department, sales department, etc.) often have conflict and inconsistency problems. For example, different answers to the same question may contain different explanations, different expression methods, or different levels of detail. If these conflicting data are not resolved, it may lead to an increase in data noise during the training of the AI model, thereby affecting the model performance.
[0003] Existing conflict resolution methods usually rely on manual intervention or static rules, and cannot automatically select the optimal resolution strategy according to different departments, different data types, and different conflict situations. Therefore, an intelligent and automated solution is needed that can dynamically select appropriate rules to resolve conflicts according to the specific type of conflict and data context. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, apparatus, electronic device, and medium for automatically resolving data conflicts, aiming to solve the above problems in the prior art.
[0005] The present invention provides a method for automatically resolving data conflicts for the training data preparation of large AI language models, including:
[0006] Automatically identifying data conflicts through text comparison and semantic analysis;
[0007] Based on a rule base and a reinforcement learning algorithm, dynamically selecting the optimal conflict resolution rule based on the identified data conflicts, and executing the conflict resolution rule;
[0008] Selecting the optimal version to be executed according to the conflict resolution rule, and performing data merging, and saving the key information of the conflict data;
[0009] Removing redundant data from the merged conflict data while ensuring the integrity of the training data of the large AI language model, and performing incremental updates.
[0010] The present invention provides an apparatus for automatically resolving data conflicts for the training data preparation of large AI language models, including:
[0011] A conflict identification module for automatically identifying data conflicts through text comparison and semantic analysis;
[0012] A rule engine module, which is used to dynamically select the optimal conflict resolution rule based on the rule library and reinforcement learning algorithm, based on the identified data conflicts, and execute the conflict resolution rule;
[0013] A version management and data merging module, which is used to select the optimal version to be executed according to the conflict resolution rule, perform data merging, and save the key information of the conflict data;
[0014] A data deduplication and incremental update module, which is used to remove redundant data from the merged conflict data and perform incremental update while ensuring the integrity of the training data of the AI language large model.
[0015] An embodiment of the present invention also provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the above-mentioned method for automatically resolving data conflicts for AI language large model training data preparation are implemented.
[0016] An embodiment of the present invention also provides a computer-readable storage medium, on which an implementation program for information transmission is stored. When the program is executed by a processor, the steps of the above-mentioned method for automatically resolving data conflicts for AI language large model training data preparation are implemented.
[0017] By adopting the embodiment of the present invention, through the application of the intelligent rule engine, data conflicts are automatically detected and resolved, the training data set of the AI model is optimized, and the integrity and consistency of the data are ensured. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of the method for automatically resolving data conflicts for AI language large model training data preparation according to the embodiment of the present invention;
[0020] Figure 2 It is a schematic diagram of the device for automatically resolving data conflicts for AI language large model training data preparation according to the embodiment of the present invention;
[0021] Figure 3 It is a schematic diagram of the electronic device according to the embodiment of the present invention. Detailed Embodiments
[0022] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification in conjunction with the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.
[0023] Method Embodiment
[0024] According to an embodiment of the present invention, there is provided an automatic data conflict resolution method for AI language large model training data preparation. Figure 1 It is a flowchart of the automatic data conflict resolution method for AI language large model training data preparation according to an embodiment of the present invention, as Figure 1 shown. The automatic data conflict resolution method for AI language large model training data preparation according to an embodiment of the present invention specifically includes:
[0025] Step S101, automatically identify data conflicts through text comparison and semantic analysis; specifically including:
[0026] Preprocess the text data provided by each department, convert the text data into vector representations, calculate the similarity between vectors of different departments, determine whether there are conflicts between texts according to the similarity, if it is determined that there are conflicts, identify the conflict types, and automatically mark the conflict data according to the conflict types. Among them, the conflict types specifically include: semantic conflicts, format conflicts, information redundancy, and version conflicts.
[0027] Step S102, based on a rule library and a reinforcement learning algorithm, dynamically select the optimal conflict resolution rule based on the identified data conflicts, and execute the conflict resolution rule; specifically including:
[0028] Establish a rule library, select the most suitable conflict resolution solution according to the conflict type and the rules in the rule library. Among them, for semantic conflicts, automatically select the most suitable content to be retained according to the target task. For format conflicts, unify the format or convert according to the target format. The rule library includes static rules and dynamic rules. Among them, the static rules are based on rules of department priority, timestamp, and data quality, and the dynamic rules are rules optimized through reinforcement learning using historical conflict resolution data.
[0029] Step S103, select the optimal version to be executed according to the conflict resolution rule, perform data merging, and save the key information of the conflict data; specifically including:
[0030] Select the latest version of the data or the most authoritative data source according to the rule library. For version conflicts, preferentially select the updated data, or combine the information of multiple versions. When merging data, merge the conflicting data according to the selected rules and retain the key information. When there are conflicts between technical issues and user operation guides, merge the two into a comprehensive answer. Among them, the key information specifically includes: important technical details or user guide information.
[0031] Step S104, remove redundant data from the merged conflicting data and perform incremental updates while ensuring the integrity of the training data of the AI language model. Specifically, it includes:
[0032] When processing multi-version data, remove duplicate content and only retain the newly added valid data. Use hash algorithms and text similarity calculations to detect duplicate data and delete duplicate training data. When updating the data each time, automatically compare it with the existing data for incremental updates, and automatically integrate the newly added data into the existing training data to ensure that the training data is always in the latest state.
[0033] In summary, by adopting the intelligent rule engine algorithm for automatically resolving cross-department data conflicts in the embodiments of the present invention, the best resolution strategy can be automatically selected and executed according to the data conflicts between departments. Through the collaborative work of multiple modules, it is ensured that when preparing the training data of the AI language model, cross-department data conflicts can be processed efficiently and accurately, improving the consistency and quality of the data.
[0034] The above technical solutions of the embodiments of the present invention are described in detail below. The specific steps are as follows:
[0035] 1. Text preprocessing: Preprocess the text data provided by each department, including removing stop words, word segmentation, standardization, etc.
[0036] Text representation and similarity calculation:
[0037] Use technologies such as TF-IDF, Word2Vec, and BERT to convert the text into vector representations.
[0038] Calculate the similarity between the answers of different departments, and judge whether there are conflicts between texts based on cosine similarity or other similarity measurement criteria.
[0039] Conflict type identification:
[0040] Semantic conflicts: For example, conflicts between technical details and user operation guides.
[0041] Format conflicts: For example, different departments use different text formats or expression styles.
[0042] Information redundancy: For example, duplicate data or descriptions.
[0043] Version conflicts: For example, there may be version differences in answers at different time points.
[0044] Embodiments of the present invention automatically mark conflicting data through the above steps.
[0045] 2. Use a dynamic rule library and a reinforcement learning algorithm to select the optimal conflict resolution strategy. The main steps are as follows:
[0046] The rule library includes static rules and dynamic rules. Static rules are based on department priorities, timestamps, data quality, etc., and dynamic rules are optimized through reinforcement learning using historical conflict resolution data.
[0047] Static rules: Set priorities based on department priorities (such as the technical department, customer service department), timestamps (such as the latest version first), and text quality (such as level of detail, accuracy, etc.).
[0048] Dynamic rules: Continuously optimize rule selection through a reinforcement learning algorithm. For example, the model automatically adjusts the weights and strategies in the rule library based on the handling effects of historical conflicts.
[0049] Rule selection and execution: Select the most suitable conflict resolution solution according to the conflict type and the rules in the rule library. For semantic conflicts, the rule engine will automatically select the most suitable content to retain according to the target task (such as technical issues or user operation guides). For format conflicts, the rule engine will unify the format or perform conversions according to the target format.
[0050] 3. Select the optimal version and merge the conflicting data. The main steps are as follows:
[0051] Version control: Select the latest version of the data or the most authoritative data source according to the rule library. For version conflicts, the system will preferentially select the updated data or combine the information of multiple versions for merging.
[0052] Data merging: Merge the conflicting data according to the selected rules to ensure that key information is retained. During the merging process, avoid losing important technical details or user guidance information.
[0053] Merging strategy: When there are conflicts between technical issues and user operation guides, the system will merge the two into a comprehensive answer that is both technical and easy for users to understand.
[0054] 4. Remove redundant data and ensure the integrity and timeliness of the training data.
[0055] Data deduplication: When processing multi-version data, remove duplicate content and only retain newly added valid data. Use hash algorithms and text similarity calculations to detect duplicate data and avoid duplicate training data.
[0056] Incremental update: When updating data each time, automatically compare it with the existing data for incremental updates. Newly added data will be automatically integrated into the existing training data to ensure that the training data is always up-to-date.
[0057] The following is the simulation code of the intelligent rule engine algorithm written in Python to achieve the automatic resolution of cross-departmental data conflicts. Modules such as simple text similarity calculation, rule engine, data merging, and version management will be used to simulate the conflict detection and resolution process.
[0058] import numpy as np
[0059] from sklearn.feature_extraction.text import TfidfVectorizer
[0060] from sklearn.metrics.pairwise import cosine_similarity
[0061] # Simulated data
[0062] answers = {
[0063] "technical": "Please download the installation package of Software A, run the installation program, and select custom installation. During the installation process, you can configure advanced options such as proxy settings and network settings.",
[0064] "customer_service": "Please download the installation package from the official website and run the installation program. Select 'default settings' during installation and complete the installation according to the on-screen instructions."
[0065] }
[0066] # Rule library (static rules and dynamic rules)
[0067] rule_library = {
[0068] "format_conflict": "Please merge the two contents, retain the technical details and add user-friendly guidance.",
[0069] "content_conflict": "Select the most appropriate version and merge the contents.",
[0070] }}
[0071] # Conflict Identification Module: Identify conflicts based on text similarity calculation
[0072] def detect_conflict(answers):
[0073] vectorizer = TfidfVectorizer()
[0074] tfidf_matrix = vectorizer.fit_transform(answers.values())
[0075] similarity_matrix = cosine_similarity(tfidf_matrix)
[0076] similarity_score = similarity_matrix[0,1]
[0077] if similarity_score < 0.7: # Assume that if the similarity is less than 0.7, it is considered a conflict
[0078] return True, similarity_score
[0079] return False, similarity_score
[0080] # Rule Engine Module: Select conflict resolution strategies based on the rule library
[0081] def resolve_conflict(conflict, conflict_type):
[0082] if conflict:
[0083] if conflict_type == "format_conflict":
[0084] return rule_library["format_conflict"]
[0085] elif conflict_type == "content_conflict":
[0086] return rule_library["content_conflict"]
[0087] return "No conflict, content is consistent"
[0088] # Version management and data merging module: Select the optimal version and merge
[0089] def merge_data(answers, conflict_resolution):
[0090] merged_answer = (
[0091] "Please download the installation package from the official website and run the installation program. During installation, you can select 'Default Settings' to complete the installation quickly,"
[0092] "or select 'Custom Installation' to perform advanced settings such as proxy settings and network configuration." )
[0094] return merged_answer
[0095] # Simulate the conflict detection and resolution process
[0096] def simulate_process():
[0097] conflict, similarity_score = detect_conflict(answers)
[0098] print(f"Text similarity: {similarity_score:.2f}")
[0099] if conflict:
[0100] conflict_type = "format_conflict" if similarity_score < 0.5 else "content_conflict"
[0101] print(f"Conflict detected, type: {conflict_type}")
[0102] conflict_resolution = resolve_conflict(conflict, conflict_type)
[0103] print(f"Conflict resolution strategy: {conflict_resolution}")
[0104] merged_answer = merge_data(answers, conflict_resolution)
[0105] print(f"The merged answer: {merged_answer}")
[0106] else:
[0107] print("No conflict, data is consistent.")
[0108] # Run the simulation
[0109] simulate_process()
[0110] When executing the above code, the answers provided by the technical department and the customer service department were simulated, and the text data was processed through the conflict detection, rule engine, version management, and data merging modules. The following is the output result of the simulation experiment: Text similarity: 0.66, conflict detected, type: format_conflict, conflict resolution strategy: Please merge the contents of both, retain the technical details while adding user-friendly guidance. Merged answer: Please download the installation package from the official website and run the installer. During installation, you can choose "Default settings" to complete the installation quickly, or choose "Custom installation" to perform advanced settings such as proxy settings and network configuration. In this experiment, the answers provided by the two departments had certain differences, so they were judged as conflicting (text similarity was 0.66, lower than 0.7). According to the conflict type, the rule engine selected the resolution strategy of "format conflict" and merged the contents of both to generate the final merged answer.
[0111] The technical solution of the embodiment of the present invention automatically detects and resolves data conflicts in cross-departmental collaboration through collaborative work, ensuring the optimal performance of the generated AI training dataset in terms of content consistency and quality. Using text similarity calculation or semantic analysis technology, it automatically identifies conflicts and inconsistencies in the data provided by different departments. Based on a static rule library and a reinforcement learning algorithm, it dynamically selects the most appropriate conflict resolution strategy to solve problems such as semantic conflicts, format conflicts, information redundancy, and version conflicts. It selects the optimal version of the data according to the conflict resolution strategy and merges the data to ensure data consistency and no loss of key information. The embodiment of the present invention is applicable to the process of preparing training data for AI language large models, automatically processing text data conflicts from multiple departments, and generating high-quality and highly consistent training datasets. By using text representation methods such as TF-IDF and BERT and similarity calculation algorithms, it automatically identifies conflicts in the text and provides conflict types and solutions. It adopts version control technology and intelligent merging strategies, combines information such as department priorities and timestamps, and ensures the timeliness and integrity of the training data. In addition, the embodiment of the present invention can also be dynamically adjusted and optimized according to historical data to adapt to the conflict handling requirements in different departments and scenarios.
[0112] In summary, by adopting the technical solution of the embodiment of the present invention, the best solution strategy can be automatically selected and executed according to the data conflicts between departments. The core of the algorithm is to ensure that cross-departmental data conflicts can be efficiently and accurately processed during the preparation of AI language large model training data through the collaborative work of multiple modules, improving data consistency and quality.
[0113] Device Embodiment 1
[0114] According to an embodiment of the present invention, there is provided an automatic data conflict resolution device for preparing training data for an AI language large model. Figure 2 It is a schematic diagram of the automatic data conflict resolution device for preparing training data for an AI language large model according to an embodiment of the present invention, as Figure 2 shown. The automatic data conflict resolution device for preparing training data for an AI language large model according to an embodiment of the present invention specifically includes:
[0115] A conflict identification module 20, configured to automatically identify data conflicts through text comparison and semantic analysis; specifically for:
[0116] Preprocess the text data provided by each department, convert the text data into vector representations, calculate the similarity between vectors of different departments, determine whether there are conflicts between texts according to the similarity, if it is determined that there are conflicts, identify the conflict type, and automatically mark the conflict data according to the conflict class, where the conflict type specifically includes: semantic conflict, format conflict, information redundancy, and version conflict;
[0117] The rule engine module 22 is used to dynamically select the optimal conflict resolution rule based on the rule library and the reinforcement learning algorithm, and execute the conflict resolution rule based on the identified data conflict; specifically used for:
[0118] Establish a rule library, select the most suitable conflict solution according to the conflict type and the rules in the rule library. Among them, for semantic conflicts, automatically select the most suitable content to retain according to the target task. For format conflicts, unify the format or convert it according to the target format. The rule library includes static rules and dynamic rules. The static rules are based on rules of department priority, timestamp, and data quality, and the dynamic rules are obtained by optimizing reinforcement learning through historical conflict resolution data.
[0119] The version management and data merging module 24 is used to select the optimal version to be executed according to the conflict resolution rule, and perform data merging and save the key information of the conflict data; specifically used for:
[0120] Select the latest version of the data or the most authoritative data source according to the rule library. For version conflicts, preferentially select the updated data, or combine the information of multiple versions for merging. When performing data merging, merge the conflict data according to the selected rule and retain the key information. When there is a conflict between technical issues and user operation guidance, merge the two into a comprehensive answer. The key information specifically includes important technical details or user guidance information.
[0121] The data deduplication and incremental update module 26 is used to remove redundant data from the merged conflict data while ensuring the integrity of the training data of the AI language large model, and perform incremental update. Specifically used for:
[0122] When processing multi-version data, remove duplicate content and only retain the newly added valid data. Use the hash algorithm and text similarity calculation to detect duplicate data and delete duplicate training data. When updating the data each time, automatically compare it with the existing data for incremental update, and automatically integrate the newly added data into the existing training data to ensure that the training data is always in the latest state.
[0123] The technical solution provided by the embodiments of the present invention can automatically solve data conflicts from multiple departments through intelligent conflict identification, rule selection, data merging, and deduplication processing during the preparation of text data for training the AI language large model, ensuring the accuracy, consistency, and high quality of the training data. This system is applicable to scenarios of large enterprises, cross-departmental collaboration, and multi-source data, and can greatly improve the quality of AI training data, thereby enhancing the effect of the AI language large model.
[0124] The embodiments of the present invention are corresponding device embodiments to the above method embodiments. The specific operations of each module can be understood with reference to the description of the method embodiments, and will not be elaborated here.
[0125] Device Embodiment Two
[0126] An embodiment of the present invention provides an electronic device, such as Figure 3 as shown, including: a memory 30, a processor 32, and a computer program stored on the memory 30 and executable on the processor 32. When the computer program is executed by the processor 32, it implements the steps described in the method embodiments.
[0127] Device Embodiment Three
[0128] An embodiment of the present invention provides a computer-readable storage medium. An implementation program for information transmission is stored on the computer-readable storage medium. When the program is executed by the processor 32, it implements the steps described in the method embodiments.
[0129] The computer-readable storage medium described in this embodiment includes, but is not limited to: ROM, RAM, magnetic disk, optical disk, etc.
[0130] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An automatic data conflict resolution method for AI language large model training data preparation, characterized in that including: Automatically identify data conflicts through text comparison and semantic analysis; Based on a rule library and a reinforcement learning algorithm, dynamically select the optimal conflict resolution rule based on the identified data conflicts, and execute the conflict resolution rule; Select the optimal version to be executed according to the conflict resolution rule, perform data merging, and save the key information of the conflict data; Remove redundant data from the merged conflict data while ensuring the integrity of the training data of the AI language large model, and perform incremental updates.
2. The method according to claim 1, wherein Automatically identifying data conflicts through text comparison and semantic analysis specifically includes: Preprocess the text data provided by each department, convert the text data into vector representations, calculate the similarity between vectors of different departments, judge whether there are conflicts between texts according to the similarity, if it is judged that there are conflicts, identify the conflict types, and automatically label the conflict data according to the conflict types. Among them, the conflict types specifically include: semantic conflicts, format conflicts, information redundancy, and version conflicts.
3. The method according to claim 1, characterized in that, Based on a rule library and a reinforcement learning algorithm, dynamically selecting the optimal conflict resolution rule based on the identified data conflicts, and executing the conflict resolution rule specifically includes: Establish a rule library, select the most appropriate conflict resolution solution according to the conflict type and the rules in the rule library. Among them, for semantic conflicts, automatically select the most suitable content to retain according to the target task. For format conflicts, unify the format or convert it according to the target format. The rule library includes static rules and dynamic rules. The static rules are based on rules of department priority, timestamp, and data quality, and the dynamic rules are rules obtained by optimizing reinforcement learning through historical conflict resolution data.
4. The method according to claim 1, characterized in that, Select the optimal version to be executed according to the conflict resolution rule, perform data merging, and save the key information of the conflict data specifically includes: Select the latest version of the data or the most authoritative data source according to the rule library. For version conflicts, preferentially select the updated data, or combine the information of multiple versions for merging. When performing data merging, merge the conflict data according to the selected rule and retain the key information. When there are conflicts between technical issues and user operation guides, combine the two into a comprehensive answer. Among them, the key information specifically includes: important technical details or user guidance information.
5. The method according to claim 1, characterized in that, Removing redundant data from the merged conflict data while ensuring the integrity of the training data of the AI language large model, and performing incremental updates specifically includes: When processing multi-version data, remove duplicate content and only retain newly added valid data. Use the hash algorithm and text similarity calculation to detect duplicate data, delete duplicate training data. When updating the data each time, automatically compare it with the existing data for incremental updates, and automatically integrate the newly added data into the existing training data to ensure that the training data is always in the latest state.
6. An automatic data conflict resolution device for AI language large model training data preparation, characterized in that including: A conflict identification module for automatically identifying data conflicts through text comparison and semantic analysis; A rule engine module for dynamically selecting the optimal conflict resolution rule based on a rule library and a reinforcement learning algorithm based on the identified data conflicts, and executing the conflict resolution rule; A version management and data merging module, which is used to select the optimal version to execute according to the conflict resolution rules, perform data merging, and save the key information of the conflict data; A data deduplication and incremental update module, which is used to remove redundant data from the merged conflict data and perform incremental update while ensuring the integrity of the training data of the AI language large model.
7. The device according to claim 6, wherein: The conflict identification module is specifically used for: Preprocess the text data provided by each department, convert the text data into vector representations, calculate the similarity between vectors of different departments, judge whether there are conflicts between texts according to the similarity, if it is judged that there are conflicts, identify the conflict types, and automatically mark the conflict data according to the conflict types, wherein the conflict types specifically include: semantic conflicts, format conflicts, information redundancy, and version conflicts; The rule engine module is specifically used for: Establish a rule library, and select the most suitable conflict resolution solution according to the conflict types and the rules in the rule library. For semantic conflicts, automatically select the most suitable content to retain according to the target task. For format conflicts, unify the format or convert it according to the target format. The rule library includes static rules and dynamic rules, wherein the static rules are based on rules of department priorities, timestamps, and data quality, and the dynamic rules are rules obtained by optimizing reinforcement learning through historical conflict resolution data.
8. The device according to claim 6, wherein: The version management and data merging module is specifically used for: Select the latest version of the data or the most authoritative data source according to the rule library. For version conflicts, preferentially select the updated data, or combine the information of multiple versions for merging. When performing data merging, merge the conflict data according to the selected rules and retain the key information. When there are conflicts between technical issues and user operation guides, merge the two into a comprehensive answer, wherein the key information specifically includes: important technical details or user guidance information; The data deduplication and incremental update module is specifically used for: When processing multi-version data, remove duplicate content and only retain the newly added valid data. Use the hash algorithm and text similarity calculation to detect duplicate data and delete duplicate training data. When updating the data each time, automatically compare it with the existing data for incremental update, and automatically integrate the newly added data into the existing training data to ensure that the training data is always in the latest state.
9. An electronic device, characterized in that, including: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the data conflict automatic resolution method for AI language large model training data preparation according to any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that, An information transmission implementation program is stored on the computer-readable storage medium. When the program is executed by the processor, it implements the steps of the data conflict automatic resolution method for AI language large model training data preparation according to any one of claims 1 to 5.
Citation Information
Cited By
Traffic instruction conflict detection method and system and computer storage medium
CN121034116A
Model training method and device, electronic equipment, storage medium and program product
CN121936632A
Model training method and device, electronic equipment, storage medium and program product
CN121936632B