Redundant data processing method and system based on pre-training model
Through a redundant data processing method based on a pre-trained model, using deep learning algorithms and temperature value optimization, intelligent identification and automatic cleaning of redundant data are achieved, solving the problems of low data processing accuracy and efficiency in existing technologies, and improving the accuracy and intelligence level of data analysis.
Patent Information
- Application Number
- CN202510185030.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-10-03
AI Technical Summary
Existing redundant data processing methods have problems such as low data processing accuracy and poor data cleaning efficiency, and lack of intelligent and automated processing solutions, which affects the accuracy and depth of data analysis.
A pre-training model-based method is adopted to build a prompt information pre-training model and a redundant data pre-training model, use deep learning algorithms to identify and process redundant data, and adjust and optimize prompt information in combination with temperature values to achieve automated and intelligent data cleaning.
It improves the efficiency of identifying and processing redundant data, reduces the risk of underreporting and misreporting, improves the accuracy and quality of data cleaning, and promotes the intelligence and adaptability of data analysis.
Smart Images

Figure CN120744307A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a redundant data processing method and system based on a pre-training model. Background Art
[0002] In the big data environment, data redundancy is an increasingly serious problem, resulting in reduced data storage and analysis efficiency. Traditional data processing methods often rely on manual screening and algorithmic filtering, which is inefficient and prone to errors. Furthermore, existing technologies lack intelligent, automated processing solutions, which can easily lead to redundant data not being effectively cleaned, thus affecting the accuracy of the results and the depth of data analysis. Therefore, by effectively cleaning redundant information, the accuracy and reliability of analysis results can be improved. This not only helps to reveal potential patterns in the data, but also provides high-quality data input for subsequent machine learning model training, thereby improving model performance and predictive capabilities. It can be seen that existing redundant data processing methods suffer from low data processing accuracy and poor data cleaning efficiency. Summary of the Invention
[0003] The present invention provides a redundant data processing method and system based on a pre-training model to solve the problems of low data processing accuracy and poor data cleaning efficiency in existing redundant data processing methods.
[0004] In order to achieve the above object, the present invention is implemented through the following technical solutions: In a first aspect, the present invention provides a redundant data processing method based on a pre-trained model, comprising: Obtain the data set to be processed, and classify the data in the data set to be processed according to the data type to obtain the corresponding data type to be processed; Constructing a prompt information pre-training model, inputting the data set to be processed and the data type to be processed into the redundant data pre-training model to obtain feature prompt information; Evaluating the characteristic prompt information to obtain an evaluation result, and further optimizing the characteristic prompt information according to the evaluation result to obtain optimized prompt information; Constructing a redundant data pre-training model, inputting the data set to be processed and the optimization prompt information into the redundant data pre-training model to identify redundant data that meets the conditions; The redundant data that meets the conditions is deleted from the data set to be processed to obtain the processed data set.
[0005] Optionally, the data types to be processed include: database type, text file type, and log file type.
[0006] Optionally, constructing a prompt information pre-training model includes: The dataset to be processed and the type of data to be processed corresponding to the dataset to be processed are used as input, and the feature prompt information is used as output to build a prompt information pre-training model; The prompt information pre-training model is used to identify redundant data in the data set to be processed, and the corresponding feature prompt information is output according to the type of data to be processed corresponding to the redundant data.
[0007] Optionally, the evaluating the characteristic prompt information to obtain an evaluation result includes: Setting accuracy scoring standards, completeness scoring standards, and operability scoring standards, and scoring the feature prompt information based on the accuracy scoring standards, the completeness scoring standards, and the operability scoring standards, and using the final score as the evaluation result; The accuracy scoring criteria are: whether the feature prompt information accurately identifies redundant records; the completeness scoring criteria are: whether the feature prompt information covers all potential redundant records; and the operability scoring criteria are: whether the feature prompt information is easy to implement. When the feature prompt information meets any of the scoring criteria, 1 point is scored, with the highest score being 3 points and the lowest being 0 points.
[0008] Optionally, further optimizing the feature prompt information according to the evaluation result to obtain optimized prompt information includes: Sorting the characteristic prompt information according to the evaluation result, and dividing the sorted prompt information into two equal parts according to a certain ratio according to actual needs; One part of the bisection is optimized by introducing the concept of low temperature value, and the other part is optimized by introducing the concept of high temperature value, and the optimized feature prompt information is used as optimization prompt information.
[0009] Optionally, the concept of the low temperature value is: based on the current feature prompt information, the accuracy and certainty of the feature prompt information is improved; the concept of the high temperature value is: based on the current feature prompt information, the accuracy and certainty of the feature prompt information is reduced.
[0010] Optionally, the step of constructing a redundant data pre-training model and inputting the dataset to be processed and the optimization prompt information into the redundant data pre-training model to identify redundant data that meets the conditions includes: Based on the optimization prompt information, the data set to be processed is preliminarily classified, and the filtering algorithm is used to analyze the preliminarily classified data to locate redundant data. The redundant data pre-training model identifies redundant data in the data to be processed based on the redundant data positioning results.
[0011] In a second aspect, an embodiment of the present application provides a redundant data processing system based on a pre-trained model, characterized in that the system includes: A data input module, used for receiving and storing data sets to be processed; Prompt design module, used to generate and optimize prompt information based on the data set to be processed; Large model processing module, used to build and call pre-processing models to perform data analysis and redundancy identification on the data to be processed; Data processing module, used to delete and organize redundant data; The result output module is used to display the processed data set.
[0012] Beneficial effects: The redundant data processing method based on the pre-trained model provided by the present invention can quickly identify and process redundant data by utilizing the automation capabilities of the large model, reduce manual intervention, and improve the efficiency of data cleaning. The large model can identify complex patterns and relationships through deep learning algorithms, thereby accurately determining which data is redundant or repeated, reducing the risk of omissions and misreporting; and the importance of adaptively adjusting the temperature lies in that it can improve the pertinence and quality of the generated text through dynamic feedback mechanisms and contextual analysis. This enables the model to understand the contextual relationships between data, making the identification of redundant data more intelligent; at the same time, utilizing the deep learning capabilities of the large model, continuously learning and optimizing the technology for redundant data processing, and continuously improving the identification and processing capabilities. Through model evaluation and dynamic optimization of the processing flow, the redundant data processing method is increasingly intelligent. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 Flowchart of a redundant data processing method based on a pre-training model according to a preferred embodiment of the present invention; Figure 2 Schematic diagram of the structure of a redundant data processing system based on a pre-training model according to a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0014] The following is a clear and complete description of the technical solutions of the present invention. It should be understood that the embodiments described are only a portion of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0015] Unless otherwise defined, the technical or scientific terms used in the present invention shall have the usual meanings understood by persons of ordinary skill in the field to which the present invention belongs. The words "first", "second" and similar terms used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one" or "a" do not indicate a quantity limitation, but rather indicate the existence of at least one. Words such as "connected" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship also changes accordingly.
[0016] See Figure 1 , an embodiment of the present application provides a redundant data processing method based on a pre-trained model, comprising: Obtain the data set to be processed, and classify the data in the data set to be processed according to the data type to obtain the corresponding data type to be processed; Constructing a prompt information pre-training model, inputting the data set to be processed and the data type to be processed into the redundant data pre-training model to obtain feature prompt information; Evaluating the characteristic prompt information to obtain an evaluation result, and further optimizing the characteristic prompt information according to the evaluation result to obtain optimized prompt information; Constructing a redundant data pre-training model, inputting the data set to be processed and the optimization prompt information into the redundant data pre-training model to identify redundant data that meets the conditions; The redundant data that meets the conditions is deleted from the data set to be processed to obtain the processed data set.
[0017] Optionally, the data types to be processed include: database type, text file type, and log file type.
[0018] Optionally, constructing a prompt information pre-training model includes: The dataset to be processed and the type of data to be processed corresponding to the dataset to be processed are used as input, and the feature prompt information is used as output to build a prompt information pre-training model; The prompt information pre-training model is used to identify redundant data in the data set to be processed, and the corresponding feature prompt information is output according to the type of data to be processed corresponding to the redundant data.
[0019] In the above embodiments, a large pre-trained model is used to generate prompts in order to assist in identifying and processing redundant data.
[0020] Optionally, the evaluating the characteristic prompt information to obtain an evaluation result includes: Setting accuracy scoring standards, completeness scoring standards, and operability scoring standards, and scoring the feature prompt information based on the accuracy scoring standards, the completeness scoring standards, and the operability scoring standards, and using the final score as the evaluation result; The accuracy scoring criteria are: whether the feature prompt information accurately identifies redundant records; the completeness scoring criteria are: whether the feature prompt information covers all potential redundant records; and the operability scoring criteria are: whether the feature prompt information is easy to implement. When the feature prompt information meets any of the scoring criteria, 1 point is scored, with the highest score being 3 points and the lowest being 0 points.
[0021] Optionally, further optimizing the feature prompt information according to the evaluation result to obtain optimized prompt information includes: Sorting the characteristic prompt information according to the evaluation result, and dividing the sorted prompt information into two equal parts according to a certain ratio according to actual needs; One part of the bisection is optimized by introducing the concept of low temperature value, and the other part is optimized by introducing the concept of high temperature value, and the optimized feature prompt information is used as optimization prompt information.
[0022] In the above embodiment, the prompts are evaluated from different perspectives and further optimized based on the evaluation results. The prompt information is scored using the following evaluation criteria, such as using an arithmetic mean to score: Accuracy (A): indicates whether the prompt can accurately identify redundant records (0-1); Completeness (C): whether all potential redundant records are covered (0-1); Actionability (O): How easy is it to implement the provided tips (0-1).
[0023] Optionally, the concept of the low temperature value is: based on the current feature prompt information, the accuracy and certainty of the feature prompt information is improved; the concept of the high temperature value is: based on the current feature prompt information, the accuracy and certainty of the feature prompt information is reduced.
[0024] In the above embodiment, the temperature value when generating the prompt can be adjusted according to the needs of data processing, thereby affecting the diversity and certainty of the generated text. For example, different temperature values can be tried:
[0025] Optionally, the step of constructing a redundant data pre-training model and inputting the dataset to be processed and the optimization prompt information into the redundant data pre-training model to identify redundant data that meets the conditions includes: Based on the optimization prompt information, the data set to be processed is preliminarily classified, and the filtering algorithm is used to analyze the preliminarily classified data to locate redundant data. The redundant data pre-training model identifies redundant data in the data to be processed based on the redundant data positioning results.
[0026] In the above embodiment, depending on the task type, if it is necessary to filter out completely redundant data, one of the generated and corrected optimization prompts is selected and input into the large model. With the help of the prompt, completely redundant data is identified for the initial data, and it is cleaned and refined to make the overall data more concise and clearly expressed.
[0027] like Figure 2 As shown, the embodiment of the present application further provides a redundant data processing system based on a pre-trained model, characterized in that the system includes: A data input module, used for receiving and storing data sets to be processed; Prompt design module, used to generate and optimize prompt information based on the data set to be processed; Large model processing module, used to build and call pre-processing models to perform data analysis and redundancy identification on the data to be processed; Data processing module, used to delete and organize redundant data; The result output module is used to display the processed data set.
[0028] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A redundant data processing method based on a pre-training model, characterized in that: include: Obtain the data set to be processed, and classify the data in the data set to be processed according to the data type to obtain the corresponding data type to be processed; Constructing a prompt information pre-training model, inputting the data set to be processed and the data type to be processed into the redundant data pre-training model to obtain feature prompt information; Evaluating the characteristic prompt information to obtain an evaluation result, and further optimizing the characteristic prompt information according to the evaluation result to obtain optimized prompt information; Constructing a redundant data pre-training model, inputting the data set to be processed and the optimization prompt information into the redundant data pre-training model to identify redundant data that meets the conditions; The redundant data that meets the conditions is deleted from the data set to be processed to obtain the processed data set.
2. The redundant data processing method based on the pre-training model according to claim 1 is characterized in that: The data types to be processed include: database type, text file type, and log file type.
3. The redundant data processing method based on the pre-training model according to claim 1 is characterized in that: The constructing of the prompt information pre-training model includes: The dataset to be processed and the type of data to be processed corresponding to the dataset to be processed are used as input, and the feature prompt information is used as output to build a prompt information pre-training model; The prompt information pre-training model is used to identify redundant data in the data set to be processed, and the corresponding feature prompt information is output according to the type of data to be processed corresponding to the redundant data.
4. The redundant data processing method based on the pre-training model according to claim 1, characterized in that: The evaluating of the characteristic prompt information to obtain an evaluation result includes: Setting accuracy scoring standards, completeness scoring standards, and operability scoring standards, and scoring the feature prompt information based on the accuracy scoring standards, the completeness scoring standards, and the operability scoring standards, and using the final score as the evaluation result; The accuracy scoring criteria are: whether the feature prompt information accurately identifies redundant records; the completeness scoring criteria are: whether the feature prompt information covers all potential redundant records; and the operability scoring criteria are: whether the feature prompt information is easy to implement. When the feature prompt information meets any of the scoring criteria, 1 point is scored, with the highest score being 3 points and the lowest being 0 points.
5. The redundant data processing method based on the pre-training model according to claim 1 is characterized in that: The further optimizing the feature prompt information according to the evaluation result to obtain optimized prompt information includes: Sorting the characteristic prompt information according to the evaluation result, and dividing the sorted prompt information into two equal parts according to a certain ratio according to actual needs; One part of the bisection is optimized by introducing the concept of low temperature value, and the other part is optimized by introducing the concept of high temperature value, and the optimized feature prompt information is used as optimization prompt information.
6. The redundant data processing method based on the pre-training model according to claim 1, characterized in that: The concept of the low temperature value is: based on the current feature prompt information, the accuracy and certainty of the feature prompt information is improved; the concept of the high temperature value is: based on the current feature prompt information, the accuracy and certainty of the feature prompt information is reduced.
7. The redundant data processing method based on the pre-training model according to claim 1 is characterized in that: The constructing of a redundant data pre-training model, inputting the dataset to be processed and the optimization prompt information into the redundant data pre-training model to identify redundant data that meets the conditions, includes: Based on the optimization prompt information, the data set to be processed is preliminarily classified, and the filtering algorithm is used to analyze the preliminarily classified data to locate redundant data. The redundant data pre-training model identifies redundant data in the data to be processed based on the redundant data positioning results.
8. A redundant data processing system based on a pre-trained model, characterized in that: The system comprises: A data input module, used for receiving and storing data sets to be processed; Prompt design module, used to generate and optimize prompt information based on the data set to be processed; Large model processing module, used to build and call pre-processing models to perform data analysis and redundancy identification on the data to be processed; Data processing module, used to delete and organize redundant data; The result output module is used to display the processed data set.