Big data processing system based on large model fine tuning
By designing a big data processing system, combining simulation training and dynamic resource allocation, the performance degradation caused by the adaptation problems of the big model in specific scenarios and data distribution differences is solved, the stability and accuracy of the model in different task scenarios are achieved, and the intelligence level of the system is improved.
Patent Information
- Application Number
- CN202510547026.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
AI Technical Summary
In actual application, large models need to be adaptively adjusted for specific scenarios. Traditional fine-tuning methods rely on a large amount of labeled data and have high demand for computing resources. They also suffer performance degradation due to insufficient model adaptation or data distribution differences when dealing with complex tasks.
A big data processing system based on fine-tuning of large models is designed, including data acquisition module, model optimization module, task allocation module, parameter adjustment module, performance evaluation module, storage unit, feedback module and termination unit. The simulation data set is generated through the simulation training unit, and a variety of optimization algorithms and dynamic resource allocation strategies are combined to achieve the adaptation and stability of the model in specific scenarios.
It improves the adaptability of the large model in specific scenarios, optimizes resource utilization efficiency, ensures the stability and accuracy of the model in different task scenarios, provides comprehensive performance evaluation and interactivity, and ensures the reliability and security of the system.
Smart Images

Figure CN120471104A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing and artificial intelligence technology, and in particular to a big data processing system based on large model fine-tuning. Background Art
[0002] To improve the efficiency and intelligence of big data processing, large models are widely used for data analysis and prediction in related fields. However, in practical applications, large models require adaptive adjustments to specific scenarios. Traditional fine-tuning methods often rely on large amounts of labeled data and require high computing resources. Furthermore, when processing complex tasks, existing methods may suffer performance degradation due to insufficient model adaptation or uneven data distribution. These issues limit the widespread application of large models in real-world big data environments. To address this issue, we propose a big data processing system based on large model fine-tuning. Summary of the Invention
[0003] The purpose of the invention is to provide a big data processing system based on large model fine-tuning, which solves the problems mentioned in the background technology.
[0004] The present invention is implemented as follows: a big data processing system based on large model fine-tuning includes a data acquisition module, a model optimization module, a task allocation module, a parameter adjustment module, a performance evaluation module, a storage unit, a feedback module, and a termination unit; the data acquisition module, model optimization module, task allocation module, parameter adjustment module, and performance evaluation module send signals to the storage unit for recording, the storage unit transmits the recorded data to the feedback module for analysis, and the feedback module transmits the analysis results to the termination unit. The system also includes a simulation training unit that sends simulation signals to the model optimization module, which in turn sends optimization signals to the storage unit for recording. The task allocation module includes a basic task unit, an extended task unit, a priority task unit, and an urgent task unit, which respectively send basic task signals, extended task signals, priority task signals, and urgent task signals to the storage unit for recording. The parameter adjustment module includes a learning rate adjustment unit and a weight update unit, which respectively send learning rate signals and weight signals to the storage unit for recording. The storage unit records the signals in the order of data acquisition module, model optimization module, task allocation module, parameter adjustment module, and performance evaluation module. The data acquisition module sends the raw data signal to the storage unit for recording. The storage unit transmits the recorded raw data signal to the termination unit, which is set with a data quality threshold. The feedback module includes a visualization unit and a prompt unit. The visualization unit uses a graphical interface and receives the recorded data results transmitted by the storage unit.
[0005] The data acquisition module connects to external data sources via a network interface, acquiring raw data from various sources and transmitting the raw data signals to a storage unit for recording. The network interface utilizes the TCP / IP protocol stack for data transmission, ensuring data integrity and real-time performance. After preprocessing, the raw data signals are stored in the storage unit in timestamp order for subsequent module access. The storage unit is internally configured with a cache and a main storage area. The cache is used for temporary storage of frequently accessed data, while the main storage area is used for long-term preservation of historical data records.
[0006] The model optimization module is connected to the simulation training unit via an algorithm library, and is used to receive simulation signals and generate optimization signals. The algorithm library contains multiple optimization algorithms, such as stochastic gradient descent and Adam optimization, for the model optimization module to select based on task requirements. The simulation training unit generates a simulation data set to simulate the data distribution characteristics in actual application scenarios, thereby improving the adaptability of the model optimization module. After being processed by the model optimization module, the optimization signal is transmitted to the storage unit for recording, and the workflow of the task allocation module is triggered at the same time.
[0007] The task allocation module is connected to the basic task unit, extended task unit, priority task unit and emergency task unit through the task scheduler, and is used to allocate computing resources according to task type and priority. The task scheduler adopts a dynamic allocation strategy to adjust the resource allocation ratio according to the current system load and task requirements. The basic task unit is responsible for processing routine data analysis tasks, the extended task unit is used to process multi-dimensional data in complex scenarios, the priority task unit provides a rapid response for high-importance tasks, and the emergency task unit handles sudden task requirements. The task signals generated by each task unit are summarized by the task scheduler and transmitted to the storage unit for recording.
[0008] The parameter adjustment module is connected to the learning rate adjustment unit and the weight update unit through a control interface, and is used to adjust the model parameters according to the feedback results of the performance evaluation module. The learning rate adjustment unit dynamically adjusts the learning rate value by monitoring the model convergence speed to avoid overfitting or underfitting. The weight update unit calculates the weight update value by gradient descent method and transmits the updated weight signal to the storage unit for recording. The parameter adjustment module works in conjunction with the model optimization module to ensure the adaptability of the model in different task scenarios.
[0009] The performance evaluation module is connected to the storage unit via an evaluation interface and is used to quantitatively evaluate the model optimization results. The evaluation interface uses a multi-metric evaluation system, including indicators such as accuracy, recall, and F1 score, to comprehensively measure model performance. The performance evaluation module transmits the evaluation results to the storage unit for recording and triggers the feedback module to generate an analysis report. The feedback module displays the evaluation results through the visualization unit, and the prompt unit uses voice notification to remind users to pay attention to changes in key indicators.
[0010] The termination unit is connected to the storage unit via a monitoring interface and is used to monitor whether data quality thresholds are met. These thresholds include indicators such as data integrity, consistency, and timeliness. When any of these indicators falls below a set threshold, the termination unit triggers the system to stop and generate an exception report. This exception report is displayed through the feedback module, prompting the user to take appropriate measures.
[0011] Compared with the prior art, the present invention realizes the intelligent upgrade of the big data processing system through the above modules and units, and solves the problem that the traditional method relies on a large amount of labeled data and computing resources. The data acquisition module obtains the original data through the network interface, ensuring the diversity and real-time nature of the data sources. The model optimization module generates optimization signals in combination with the simulation training unit, which improves the adaptability of the model in specific scenarios. The task allocation module optimizes resource utilization efficiency through dynamic allocation strategies and meets the needs of multi-task concurrent processing. The parameter adjustment module works in coordination with the learning rate adjustment unit and the weight update unit to ensure the stability and accuracy of the model in different task scenarios. The performance evaluation module quantifies the model performance through a multi-indicator evaluation system to provide users with a comprehensive performance reference. The feedback module displays the analysis results through a visualization unit and a prompt unit, enhancing the interactivity and user experience of the system. The termination unit ensures the reliability and security of the system through data quality threshold monitoring.
[0012] The present invention effectively addresses the deficiencies of the existing technologies mentioned in the background technology through the above-mentioned technical means, and in particular provides a specific and feasible solution to the adaptation problems of large models in practical applications and the performance degradation caused by data distribution differences. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a structural diagram of a big data processing system based on large model fine-tuning of the present invention;
[0014] Figure 2 This is a data flow diagram of a big data processing system based on large model fine-tuning in the financial field application scenario of the present invention.
[0015] The accompanying drawings are numbered as follows:
[0016] 1. Data acquisition module; 2. Model optimization module; 3. Task allocation module; 4. Parameter adjustment module; 5. Performance evaluation module; 6. Storage unit; 7. Feedback module; 8. Termination unit; 9. Simulation training unit; 10. Basic task unit; 11. Extended task unit; 12. Priority task unit; 13. Emergency task unit; 14. Learning rate adjustment unit; 15. Weight update unit; 16. Visualization unit; 17. Prompt unit. DETAILED DESCRIPTION
[0017] The present invention provides a large data processing system based on large model fine-tuning, the overall structure of which is as follows: Figure 1 The system includes a data acquisition module 1, a model optimization module 2, a task allocation module 3, a parameter adjustment module 4, a performance evaluation module 5, a storage unit 6, a feedback module 7 and a termination unit 8. The modules are connected and interacted through signal transmission to achieve intelligent upgrades in big data processing. Figure 1 Detailed description of the specific implementation of the system is given in detail with reference to the accompanying drawings.
[0018] The data acquisition module 1 connects to external data sources via a network interface that utilizes the TCP / IP protocol stack for data transmission. The module acquires raw data from various sources and transmits the raw data signals to storage unit 6 for recording. Storage unit 6 includes a cache area and a main storage area. The cache area temporarily stores frequently accessed data, while the main storage area preserves historical data records over the long term. After preprocessing, the raw data signals are stored in storage unit 6 in timestamp order for easy access by subsequent modules. The data acquisition module 1 is located at the front end of the system, providing basic data support for subsequent modules.
[0019] The model optimization module 2 is connected to the simulation training unit 9 through the algorithm library. The simulation training unit 9 generates a simulation data set and sends a simulation signal to the model optimization module 2. After receiving the simulation signal, the model optimization module 2 generates an optimization signal and transmits the optimization signal to the storage unit 6 for recording. The algorithm library contains a variety of optimization algorithms such as the stochastic gradient descent algorithm and the Adam optimization algorithm for the model optimization module 2 to select according to task requirements. The location of the model optimization module 2 is close to the data acquisition module 1, which is convenient for real-time data reception and optimization processing. The connection between the model optimization module 2 and the storage unit 6 is realized through a dedicated signal channel to ensure that the optimization signal can be transmitted efficiently.
[0020] The task allocation module 3 is connected to the basic task unit 10, the extended task unit 11, the priority task unit 12 and the emergency task unit 13 through the task scheduler. The task scheduler adopts a dynamic allocation strategy to adjust the resource allocation ratio according to the current system load and task requirements. The basic task unit 10 is responsible for processing routine data analysis tasks, the extended task unit 11 is used to process multi-dimensional data in complex scenarios, the priority task unit 12 provides a rapid response for high-importance tasks, and the emergency task unit 13 handles sudden task requirements. The task signals generated by each task unit are summarized by the task scheduler and transmitted to the storage unit 6 for recording. The task allocation module 3 is located after the model optimization module 2 to ensure that the optimized model can reasonably allocate computing resources according to the task type.
[0021] The parameter adjustment module 4 is connected to the learning rate adjustment unit 14 and the weight update unit 15 through a control interface. The learning rate adjustment unit 14 dynamically adjusts the learning rate value by monitoring the model convergence speed to avoid overfitting or underfitting. The weight update unit 15 calculates the weight update value by the gradient descent method and transmits the updated weight signal to the storage unit 6 for recording. The parameter adjustment module 4 works in conjunction with the model optimization module 2 to ensure the adaptability of the model in different task scenarios. The parameter adjustment module 4 is located close to the task allocation module 3 to facilitate real-time adjustment of model parameters according to task requirements.
[0022] The performance evaluation module 5 is connected to the storage unit 6 via an evaluation interface and is used to quantitatively evaluate the model optimization results. This evaluation interface uses a multi-metric evaluation system, including metrics such as precision, recall, and F1 score, to comprehensively measure model performance. The performance evaluation module 5 transmits the evaluation results to the storage unit 6 for recording and triggers the feedback module 7 to generate an analysis report. The performance evaluation module 5 is located after the parameter adjustment module 4 to ensure that the effects of model parameter adjustments can be promptly evaluated.
[0023] Feedback module 7 includes a visualization unit 16 and a prompt unit 17. Visualization unit 16 displays evaluation results using a graphical interface, while prompt unit 17 alerts users to changes in key indicators through voice notification. Feedback module 7 receives and analyzes the recorded data from storage unit 6, transmitting the analysis results to termination unit 8. Feedback module 7 is located near performance evaluation module 5, facilitating the receipt of evaluation results and the generation of feedback information.
[0024] Termination unit 8 is connected to storage unit 6 via a monitoring interface and is used to monitor data quality thresholds. Data quality thresholds include indicators such as data integrity, consistency, and timeliness. When any of these indicators falls below the set threshold, termination unit 8 triggers the system to stop operation and generate an exception report. The exception report is displayed through feedback module 7, prompting the user to take appropriate measures. Termination unit 8 is located at the end of the system and serves as the termination node for the entire system.
[0025] Storage unit 6, the core hub of the entire system, records signals from data acquisition module 1, model optimization module 2, task allocation module 3, parameter adjustment module 4, and performance evaluation module 5. Storage unit 6 transmits the recorded data to feedback module 7 for analysis, which in turn transmits the analysis results to termination unit 8. Storage unit 6's central location facilitates efficient connection and data exchange with other modules.
[0026] The specific operation process of this system is as follows: First, the data acquisition module 1 obtains raw data from an external data source via a network interface and transmits the raw data signal to the storage unit 6 for recording. The storage unit 6 transmits the recorded raw data signal to the model optimization module 2. The model optimization module 2 selects an appropriate optimization algorithm from the algorithm library and generates an optimization signal based on the simulation data set generated by the simulation training unit 9. After the optimization signal is transmitted to the storage unit 6 for recording, the workflow of the task allocation module 3 is triggered. The task allocation module 3 allocates computing resources based on task type and priority through the task scheduler. The task signals generated by each task unit are aggregated by the task scheduler and transmitted to the storage unit 6 for recording. The parameter adjustment module 4 adjusts the model parameters based on the feedback from the performance evaluation module 5 through the learning rate adjustment unit 14 and the weight update unit 15. The adjusted parameter signals are transmitted to the storage unit 6 for recording. The performance evaluation module 5 quantitatively evaluates the model optimization results and transmits the evaluation results to the storage unit 6 for recording, triggering the feedback module 7 to generate an analysis report. The feedback module 7 displays the analysis results through the visualization unit 16 and the prompt unit 17 and transmits them to the termination unit 8. The termination unit 8 monitors the data quality threshold through the monitoring interface. When any indicator is lower than the set threshold, the system is triggered to stop running and generate an abnormality report.
[0027] The various modules and units of this system work closely together through signal transmission to ensure the intelligence and efficiency of the big data processing process. The data acquisition module 1 ensures the diversity and real-time nature of data sources. The model optimization module 2 combines with the simulation training unit 9 to improve the adaptability of the model in specific scenarios. The task allocation module 3 optimizes resource utilization efficiency through dynamic allocation strategies. The parameter adjustment module 4 works together through the learning rate adjustment unit 14 and the weight update unit 15 to ensure the stability and accuracy of the model in different task scenarios. The performance evaluation module 5 quantifies the model performance through a multi-indicator evaluation system to provide users with a comprehensive performance reference. The feedback module 7 displays the analysis results through the visualization unit 16 and the prompt unit 17 to enhance the interactivity and user experience of the system. The termination unit 8 monitors the data quality threshold to ensure the reliability and security of the system.
[0028] In order to better enable relevant personnel in this technical field to fully understand and implement the present invention, the specific implementation principle of the present invention is supplemented below with reference to a specific application scenario.
[0029] In the financial sector, big data processing scenarios involve the system using Data Acquisition Module 1 to acquire raw data from multiple sources, such as bank transaction records, market data, and user behavior logs. This data is transmitted to Storage Unit 6 via the TCP / IP protocol stack and stored in a cache and main storage area in timestamp order. The cache area temporarily stores frequently accessed real-time transaction data, while the main storage area stores historical transaction records for subsequent analysis. Storage Unit 6 transmits the preprocessed data signals to Model Optimization Module 2, providing foundational support for subsequent model optimization.
[0030] The model optimization module 2 selects the stochastic gradient descent algorithm through the algorithm library, and fine-tunes the large model in combination with the simulation data set generated by the simulation training unit 9. The simulation training unit 9 generates a simulated trading scenario based on historical trading data, simulating the data distribution characteristics in actual applications, thereby improving the adaptability of the model in specific scenarios. After the optimization signal is transmitted to the storage unit 6 through a dedicated signal channel for recording, the workflow of the task allocation module 3 is triggered. The task allocation module 3 allocates computing resources according to the task type and priority through the task scheduler. For example, the basic task unit 10 is responsible for processing routine transaction data analysis tasks, the extended task unit 11 is used to analyze complex multi-dimensional market trend data, the priority task unit 12 quickly responds to high-risk transaction monitoring needs, and the emergency task unit 13 handles sudden abnormal trading events. The task signals generated by each task unit are summarized by the task scheduler and transmitted to the storage unit 6 for recording.
[0031] The parameter adjustment module 4 dynamically adjusts model parameters based on feedback from the performance evaluation module 5. The learning rate adjustment unit 14 monitors the model's convergence rate and automatically increases the learning rate when it detects slow convergence; it decreases the learning rate when overfitting occurs. The weight update unit 15 calculates updated weights using gradient descent and transmits the updated weight signals to the storage unit 6 for recording. For example, when handling high-risk transaction monitoring tasks, the parameter adjustment module 4 reduces the learning rate and adjusts the weights to ensure that the model can accurately identify potentially risky transactions.
[0032] The performance evaluation module 5 performs a quantitative evaluation of the optimized model through the evaluation interface. The evaluation interface adopts a multi-index evaluation system, including indicators such as accuracy, recall rate and F1 score, to comprehensively measure the performance of the model. For example, in the transaction data analysis task, accuracy is used to measure the model's ability to identify normal transactions, recall rate is used to evaluate the model's ability to detect abnormal transactions, and F1 score comprehensively considers the performance of both. The performance evaluation module 5 transmits the evaluation results to the storage unit 6 for record, and triggers the feedback module 7 to generate an analysis report. The feedback module 7 displays the evaluation results through the visualization unit 16, and the prompt unit 17 reminds the user to pay attention to changes in key indicators through voice broadcast. For example, when the recall rate of the model is lower than the set threshold, the prompt unit 17 will issue a voice reminder to prompt the user to take corresponding measures.
[0033] Termination Unit 8 monitors data quality thresholds through a monitoring interface to ensure they meet required requirements. For example, if transaction data integrity falls below a set threshold, Termination Unit 8 triggers system shutdown and generates an exception report. This exception report is displayed via Feedback Module 7, prompting the user to take appropriate action. For example, if some transaction data is missing, the system prompts the user to check the data source connection status or re-collect the data.
[0034] Through the above steps, this system achieves an intelligent upgrade of the big data processing process. The data acquisition module 1 ensures the diversity and real-time nature of data sources. The model optimization module 2, combined with the simulation training unit 9, improves the model's adaptability in specific scenarios. The task allocation module 3 optimizes resource utilization efficiency through a dynamic allocation strategy. The parameter adjustment module 4 ensures the stability and accuracy of the model in different task scenarios through the collaborative work of the learning rate adjustment unit 14 and the weight update unit 15. The performance evaluation module 5 quantifies the model performance through a multi-index evaluation system, providing users with a comprehensive performance reference. The feedback module 7 enhances the system's interactivity and user experience through the visualization unit 16 and the prompt unit 17. The termination unit 8 ensures the system's reliability and security through data quality threshold monitoring.
[0035] The present invention effectively addresses the deficiencies of the existing technologies mentioned in the background technology through the above-mentioned technical means, and in particular provides a specific and feasible solution to the adaptation problems of large models in practical applications and the performance degradation caused by data distribution differences.
[0036] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A big data processing system based on large model fine-tuning, characterized in that: The system comprises a data acquisition module (1), a model optimization module (2), a task allocation module (3), a parameter adjustment module (4), a performance evaluation module (5), a storage unit (6), a feedback module (7) and a termination unit (8); the data acquisition module (1), the model optimization module (2), the task allocation module (3), the parameter adjustment module (4) and the performance evaluation module (5) send signals to the storage unit (6) for recording, the storage unit (6) transmits the recorded data to the feedback module (7) for analysis, and the feedback module (7) transmits the analysis results to the termination unit (8).
2. The big data processing system based on large model fine-tuning according to claim 1, characterized in that: It also includes a simulation training unit (9), which sends a simulation signal to the model optimization module (2), and the model optimization module (2) sends an optimization signal to the storage unit (6) for recording.
3. The big data processing system based on large model fine-tuning according to claim 1, characterized in that: The task allocation module (3) includes a basic task unit (10), an extended task unit (11), a priority task unit (12) and an emergency task unit (13), wherein the basic task unit (10), the extended task unit (11), the priority task unit (12) and the emergency task unit (13) respectively send a basic task signal, an extended task signal, a priority task signal and an emergency task signal to a storage unit (6) for recording.
4. The big data processing system based on large model fine-tuning according to claim 1, characterized in that: The parameter adjustment module (4) comprises a learning rate adjustment unit (14) and a weight updating unit (15), wherein the learning rate adjustment unit (14) and the weight updating unit (15) respectively send a learning rate signal and a weight signal to a storage unit (6) for recording.
5. The big data processing system based on large model fine-tuning according to claim 1, characterized in that: The order in which the storage unit (6) records signals is the data acquisition module (1), the model optimization module (2), the task allocation module (3), the parameter adjustment module (4) and the performance evaluation module (5).
6. The big data processing system based on large model fine-tuning according to claim 1, characterized in that: The data acquisition module (1) sends the original data signal to the storage unit (6) for recording, and the storage unit (6) transmits the recorded original data signal to the termination unit (8), and the termination unit (8) is provided with a data quality threshold.
7. The big data processing system based on large model fine-tuning according to claim 1, characterized in that: The feedback module (7) includes a visualization unit (16) and a prompt unit (17). The visualization unit (16) adopts a graphical interface. The visualization unit (16) and the prompt unit (17) receive the recorded data results transmitted by the storage unit (6).