Intelligent Data Processing Device, Medium and Electronic Device Based on Large Model

Through the intelligent data processing device based on large models, modular design and parallel processing technology are adopted to solve the problem of inefficient single-threaded data processing, efficient data processing and storage are achieved, and the overall performance and adaptability of the system are improved.

CN119807823BActive Publication Date: 2025-07-18SHENZHEN HUMENG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510292954.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-18
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In the prior art, single-threaded data processing process is cumbersome and inefficient, making it difficult to efficiently process explosive growth data.

Method used

Using an intelligent data processing device based on large models, five modular designs are adopted: the target module determines the big model goals and designs the overall architecture, the data acquisition module obtains the original data, the setting module performs calculation unit settings and data preprocessing, the data processing module performs parallel processing and stores in the data lake, and the integration module performs testing and verification and adjustment of process strategies.

Benefits of technology

The data processing process is optimized, the efficiency and flexibility of intelligent data processing are improved, and the performance and adaptability of the system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807823B_ABST
    Figure CN119807823B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent data processing device, medium and electronic device based on a large model, belonging to the technical field of data processing; the modules include: a target module: determining a preset large model target and designing the overall architecture of the large model; a data acquisition module: acquiring original data through a preset data source and storing the original data in an original database; a setting module: after setting a calculation unit to preprocess to obtain standard data, matching different processing paths according to the type and processing requirements of the standard data; a data processing module: based on the calculation unit, transmitting the standard data to the corresponding processing path through a high-speed transmission channel, performing parallel compression processing to obtain standard intelligent data, and storing the standard intelligent data in a data lake; an integration module: integrating the data acquisition module, the setting module and the data processing module into the overall architecture of the large model, and configuring them into a test environment for test verification to obtain feedback information, and adjusting the data processing flow and calculation strategy according to the feedback information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to an intelligent data processing device, medium, and electronic device based on a large model. Background Art

[0002] Currently, with the development of technologies such as the Internet, Internet of Things, 5G, and artificial intelligence, the amount of data has grown explosively. This has made the storage, processing, and analysis of data become increasingly complex and important. In order to extract valuable information from the huge amount of data, it is necessary to rely on powerful data processing capabilities and intelligent analysis means; usually, a simple single-threaded data processing process is too cumbersome and inefficient.

[0003] Therefore, the present invention proposes an intelligent data processing device, medium, and electronic device based on a large model. Summary of the Invention

[0004] The present invention provides an intelligent data processing device, medium, and electronic device based on a large model, which are used to determine the large model target and design the overall architecture through five modules: a target module; obtain and store raw data through a data source by a data acquisition module; perform calculation unit setting and data preprocessing by a setting module; perform data transmission, parallel processing, and storage in a data lake based on the calculation unit by a data processing module; integrate each module and perform test verification by an integration module, and adjust the data processing process and calculation strategy according to the feedback. The overall solution aims to optimize the data processing process and improve the efficiency of intelligent data processing.

[0005] On the one hand, the present invention provides an intelligent data processing device, medium, and electronic device based on a large model, including:

[0006] A target module: determine a preset large model target according to the processing requirements of intelligent data, and design the overall architecture of the large model;

[0007] A data acquisition module: obtain raw data through a preset data source and store the raw data in a raw database;

[0008] A setting module: set a calculation unit, and at the same time, after preprocessing the raw data to obtain standard data, match different processing paths according to the type and processing requirements of the standard data;

[0009] A data processing module: transmit the standard data to the corresponding processing path based on the calculation unit through a high-speed transmission channel, perform parallel compression processing to obtain standard intelligent data, and store it in a data lake;

[0010] An integration module: integrate the data acquisition module, setting module, and data processing module into the overall architecture of the large model, and configure them to a test environment for test verification to obtain feedback information, and adjust the data processing process and calculation strategy according to the feedback information.

[0011] On the other hand, the target module includes:

[0012] Requirement determination unit: clarify the application scenario of the intelligent data processing service and determine the processing requirements of the processing device;

[0013] Task type unit: based on the processing requirements of the processing device, clarify the intelligent data processing task type, and preset the large model target according to the task type.

[0014] On the other hand, the target module further includes:

[0015] Architecture unit: match the model type and its architecture according to the task type, and generate the overall architecture of the large model;

[0016] Hardware acceleration unit: configure the preset hardware as the external hardware part of the large model, and perform parallel computing on the large model.

[0017] On the other hand, the data acquisition module includes:

[0018] Data source unit: determine the data source according to the preset source of the intelligent data, and obtain the connection information of all data sources;

[0019] Interface unit: configure the data interface according to the data type of any data source, and connect through an HTTP request based on the connection information to obtain the original data of the data source;

[0020] Data storage unit: design the original database table structure according to the preset structure, and store the original data into the original database with the combined primary key of timestamp - data source identifier.

[0021] On the other hand, the setting module includes:

[0022] Computing unit configuration: set multiple computing units for parallel data processing according to the data volume and processing requirements of the original data;

[0023] Normalization unit: obtain the original data of any data source in the original database, clean the original data, and remove and replace the outliers to obtain the first data;

[0024] After standard normalization of the first data, standard data is obtained, specifically:

[0025] ; where represents the standard data at the i-th time node in the first data, represents the data value at the i-th time node in the first data, represents the mean value of the data values at all time nodes in the first data, represents the standard deviation of the data values at all time nodes in the first data, and n represents that there are n time nodes in the first data. represents the minimum value among the data values at all time nodes in the first data. represents the maximum value among the data values at all time nodes in the first data.

[0026] On the other hand, the setting module further includes:

[0027] Classification unit: Determine the number of classifications according to the number of calculation units, perform feature engineering processing on the original data of any data source to obtain the first feature sequence of the original data;

[0028] If the number of sequences in the first feature sequence is greater than the number of determined classifications, perform principal component analysis dimensionality reduction operation on the first feature sequence to obtain a second feature sequence consistent with the number of determined classifications. Otherwise, directly use the first feature sequence as the second feature sequence;

[0029] Use any feature of the second feature sequence as the classification criterion to match with the calculation unit to obtain the required unit, perform feature selection on the standard data, select the feature with the highest expected value as the classification, and assign it to the corresponding required unit;

[0030] The required unit selects the processing path of the required unit according to data type - processing requirement - processing path.

[0031] On the other hand, the data processing module includes:

[0032] Serialization unit: Serialize the standard data according to the standard data classified under any classification calculation unit to obtain byte stream data, and transmit the byte stream data into the high-speed transmission channel through the network transmission protocol according to the processing path of the required unit. The transmission rate is:

[0033] ; where R represents the transmission rate of the high-speed transmission of the required unit, represents the bandwidth of the high-speed transmission channel, represents the delay of the high-speed transmission channel, represents the byte stream data, represents the high-speed processing efficiency function;

[0034] If the transmission rate is less than the preset threshold, reallocate the high-speed transmission channel. Otherwise, transmit normally;

[0035] Task distribution unit: Through the high-speed transmission channel, the scheduler distributes according to the type and demand of the data, and constructs and allocates data processing tasks according to the load balancing algorithm to obtain a multi-task queue;

[0036] For the multi-task queue, the system performs parallel processing and obtains first intelligent data after processing;

[0037] Compression unit: Encodes and compresses the first intelligent data, constructs a frequency table according to the symbols of the first intelligent data, merges the two symbols with the lowest frequencies to form a new node until the frequency tree is constructed, and generates binary codes for each symbol according to the frequency tree to obtain the standard intelligent data of the first intelligent data. Specifically:

[0038] ; where represents the compression code of the j-th symbol, represents the j-th symbol of the Huffman value, {} represents the compression function, represents that the first intelligent data has a total of m symbols, represents the interval size of the frequency tree from the starting point, represents the interval size of the frequency tree from the end point;

[0039] Data lake unit: Constructs a data lake for intelligent data and transmits the standard intelligent data of all data sources into the data lake.

[0040] On the other hand, the integration module includes:

[0041] Integration unit: Seamlessly integrates the data acquisition module, the setting module, and the data processing module, formulates the interfaces between the modules according to the data stream and the communication protocol, and completes the preliminary integration;

[0042] Performs integration testing on the preliminary integration. After each module meets the expected standards, the integration is completed to obtain a large model. Otherwise, reconfigures the interfaces between the modules;

[0043] Testing unit: Builds a test environment, configures the large model into the test environment, monitors the performance of each module and the structure of intelligent data processing in real time, and generates test logs;

[0044] Feedback unit: Obtains feedback information on the operation of the large model based on the test logs, and adjusts the data processing flow and calculation strategy according to the feedback information.

[0045] Compared with the prior art, the beneficial effects of the present invention are:

[0046] The present invention provides an intelligent data processing device, medium, and electronic device based on a large model, which are used to determine the large model target and design the overall architecture through five modules: a target module; a data acquisition module to obtain and store raw data through a data source; a setting module to perform calculation unit setting and data preprocessing; a data processing module to perform data transmission, parallel processing based on the calculation unit, and store the data in a data lake; and an integration module to integrate each module, perform test verification, and adjust the data processing flow and calculation strategy according to the feedback. The overall solution aims to optimize the data processing flow and improve the efficiency of intelligent data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 FIG. is a schematic structural diagram of an intelligent data processing device based on a large model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0050] Embodiment 1:

[0051] As Figure 1 shown, an intelligent data processing device, medium, and electronic device based on a large model provided by an embodiment of the present invention include:

[0052] A target module: According to the processing requirements of intelligent data, determine a preset large model target and design the overall architecture of the large model;

[0053] A data acquisition module: Obtain raw data through a preset data source and store the raw data in a raw database;

[0054] A setting module: Set a calculation unit. At the same time, after preprocessing the raw data to obtain standard data, match different processing paths according to the type and processing requirements of the standard data;

[0055] Data Processing Module: Based on the computing unit, the standard data is transmitted to the corresponding processing path through the high-speed transmission channel, and after parallel compression processing, the standard intelligent data is obtained and stored in the data lake.

[0056] Integration Module: Integrate the data acquisition module, the setting module, and the data processing module into the overall architecture of the large model, and configure it to the test environment for test verification to obtain feedback information, and adjust the data processing flow and calculation strategy according to the feedback information.

[0057] In this embodiment, the intelligent data is data containing structure and information.

[0058] In this embodiment, the processing requirement refers to the processing methods, steps, and strategies required according to specific tasks or goals during the intelligent data processing process.

[0059] In this embodiment, the preset large model target refers to the actual application requirements, technical requirements, and system goals during the intelligent data processing and analysis process.

[0060] In this embodiment, the overall architecture is the high-level structure of the entire system, ensuring that each module can work in coordination to achieve the preset goal.

[0061] In this embodiment, the preset data source refers to various data sources that have been determined and configured during the large model architecture design stage to provide raw data.

[0062] In this embodiment, the raw data refers to the data that has not been processed or cleaned in the data processing system.

[0063] In this embodiment, the raw database is a database used to store the data collected from various data sources that has not been processed or only preliminarily processed.

[0064] In this embodiment, the computing unit refers to the software component responsible for executing specific computing tasks in the entire data processing architecture. According to different processing requirements and standard data types, it executes the required computing operations and performs data processing, conversion, and optimization.

[0065] In this embodiment, the preprocessing is the operation of cleaning, formatting, converting, and standardizing the raw data.

[0066] In this embodiment, the standard data refers to the data after preprocessing, which has been cleaned, converted, normalized, and sorted according to the predetermined rules, formats, and requirements.

[0067] In this embodiment, the processing path is the specific process route for performing different types of processing, transformation, analysis, etc. on the data.

[0068] In this embodiment, the high-speed transmission channel refers to the communication path technology for quickly and effectively transmitting data within the system.

[0069] In this embodiment, compression processing reduces the volume of data through specific technologies, thereby improving storage efficiency and transmission efficiency.

[0070] In this embodiment, standard intelligent data refers to data that has undergone certain preprocessing and formatting and conforms to specific standards and specifications.

[0071] In this embodiment, the data lake is a storage repository for a large amount of centralized data storage.

[0072] In this embodiment, integration integrates each independently developed module according to the design architecture, enabling them to smoothly perform data flow and information exchange.

[0073] In this embodiment, the test environment refers to an environment used to verify and test whether a system, module, or function works as expected during the development process.

[0074] In this embodiment, feedback information refers to the information generated during the test process when the large model runs in the test environment according to the execution results of the data processing module.

[0075] In this embodiment, the calculation strategy refers to how to select and execute different algorithms, calculation methods, and optimization means according to different data types, processing requirements, and constraints of computing resources during the data processing process to efficiently and accurately complete the data processing task.

[0076] The working principle and beneficial effects of the above technical solution are as follows: Through modular design, the processes of obtaining, preprocessing, processing, and storing intelligent data are optimized, improving data processing efficiency. The integrated module realizes the coordination and optimization of the overall architecture, ensuring the efficiency and flexibility of data processing, and helping to improve the performance of intelligent systems.

[0077] Embodiment 2:

[0078] Based on the above Embodiment 1, the target module includes:

[0079] Requirement determination unit: clarify the application scenario of the intelligent data processing service and determine the processing requirements of the processing device;

[0080] Task type unit: based on the processing requirements of the processing device, clarify the intelligent data processing task type, and preset the large model target according to the task type.

[0081] In this embodiment, the application scenario refers to the actual usage background and requirements of the intelligent data processing service in a specific environment.

[0082] The working principle and beneficial effects of the above technical solution are as follows: The application scenario and processing requirements are determined by the requirement determination unit, and the task type unit defines specific task types according to the requirements and preset the goals of the large model, so as to achieve accurate goal setting and task classification for intelligent data processing. This method improves the pertinence and processing efficiency of the system, and helps to optimize and refine the intelligent data processing service.

[0083] Embodiment 3:

[0084] Based on the above Embodiment 1, the target module further includes:

[0085] Architecture unit: According to the task type, match the model type and its architecture, and generate the overall architecture of the large model;

[0086] Hardware acceleration unit: By configuring the preset hardware as the external hardware part of the large model, perform parallel computing on the large model.

[0087] In this embodiment, the preset hardware refers to the hardware device specially configured to accelerate the large model calculation, such as CPU or GPU.

[0088] In this embodiment, parallel refers to the process of decomposing a computing task into multiple subtasks and executing these subtasks simultaneously.

[0089] The working principle and beneficial effects of the above technical solution are as follows: The architecture unit matches the appropriate model architecture according to the task type to ensure the adaptability and efficiency of the large model design. The hardware acceleration unit uses external hardware for parallel computing, improving the computing efficiency and processing speed of the model, and helping to optimize the intelligent data processing performance.

[0090] Embodiment 4:

[0091] Based on the above Embodiment 1, the data acquisition module includes:

[0092] Data source unit: Determine the data source according to the preset source of the intelligent data and obtain the connection information of all data sources;

[0093] Interface unit: Configure the data interface according to the data type of any data source, and connect through an HTTP request based on the connection information to obtain the original data of the data source;

[0094] Data storage unit: Design the original database table structure according to the preset structure, and store the original data into the original database according to the combined primary key of timestamp - data source identifier.

[0095] In this embodiment, the preset source refers to the data source defined in the system design stage.

[0096] In this embodiment, connection information refers to all the detailed information required when interacting with a data source, including UPL, IP address, protocol, etc.

[0097] In this embodiment, a data interface refers to a communication method for exchanging data between different systems, applications, or components.

[0098] In this embodiment, an HTTP request is a way for a client to send a request to a server to obtain resources.

[0099] In this embodiment, a preset structure refers to the database table structure and data storage method predefined during the system design stage.

[0100] In this embodiment, a composite primary key refers to a primary key composed of multiple columns (fields) used to uniquely identify each row of data in a table.

[0101] The working principle and beneficial effects of the above technical solution are as follows: The data source unit determines the data source, and the interface unit configures a data interface to obtain raw data. The data storage unit designs a suitable database structure to ensure the efficient storage and management of data, improving the accuracy of data processing and the stability of the system.

[0102] Embodiment 5:

[0103] Based on the above Embodiment 1, the setting module includes:

[0104] Calculation unit configuration: According to the data volume and processing requirements of the raw data, multiple calculation units are set for parallel data processing;

[0105] Normalization unit: Obtain the raw data of any data source of the original database, clean the raw data, and obtain the first data after removing and replacing outliers;

[0106] The first data is standardized to obtain the standard data, specifically:

[0107] ; where represents the standard data at the i-th time node in the first data, represents the data value at the i-th time node in the first data, represents the mean value of the data values at all time nodes in the first data, represents the standard deviation of the data values at all time nodes in the first data, n represents that there are n time nodes in the first data, represents the minimum value of the data values at all time nodes in the first data, represents the maximum value of the data values at all time nodes in the first data.

[0108] In this embodiment, the data volume refers to the quantity of the original data involved in the processing.

[0109] In this embodiment, data cleaning improves the quality and applicability of data by identifying and correcting errors, inconsistencies, or missing values in the data.

[0110] In this embodiment, an outlier refers to data in a dataset that significantly deviates from other data points.

[0111] In this embodiment, the first data refers to the original data after data cleaning and removal of outliers.

[0112] In this embodiment, standard normalization converts data with different dimensions or scales to a unified standard scale.

[0113] The working principle and beneficial effects of the above technical solution are as follows: By configuring multiple computing units for parallel data processing, the data processing efficiency is improved. The standardization unit removes outliers and standardizes the data through data cleaning and normalization, ensuring the accuracy and consistency of the data and providing a high-quality data basis for subsequent analysis.

[0114] Embodiment 6:

[0115] Based on the above Embodiment 5, the setting module further includes:

[0116] Classification unit: Determine the number of classifications according to the number of computing units, perform feature engineering processing on the original data of any data source to obtain the first feature sequence of the original data;

[0117] If the number of sequences in the first feature sequence is greater than the number of determined classifications, perform a principal component analysis dimensionality reduction operation on the first feature sequence to obtain a second feature sequence consistent with the number of determined classifications; otherwise, directly use the first feature sequence as the second feature sequence;

[0118] Match any feature of the second feature sequence with the computing unit as the classification criterion to obtain the required unit, perform feature selection on the standard data, select the feature with the highest expected value as the classification, and assign it to the corresponding required unit;

[0119] The required unit selects the processing path of the required unit according to data type - processing requirement - processing path.

[0120] In this embodiment, feature engineering refers to processing and transforming the original data to generate useful features to improve model performance.

[0121] In this embodiment, the first feature sequence refers to a set of features obtained by performing feature engineering processing on the original data.

[0122] In this embodiment, the principal component analysis dimensionality reduction operation maps data from a high-dimensional space to a low-dimensional space while preserving the original feature information of the data as much as possible. Through PCA (principal component analysis), the complexity of the data can be reduced.

[0123] In this embodiment, feature selection refers to selecting the most important and most predictive features from the features of the original data during the machine learning process.

[0124] In this embodiment, the expected value is an index representing the matching degree in feature selection.

[0125] The working principle and beneficial effects of the above technical solution are as follows: By processing feature engineering and principal component analysis dimensionality reduction, the data features are optimized to ensure that the number of classifications is consistent with the computing units. Through feature selection and matching, the data is ensured to be assigned to appropriate units, and the best processing path is selected according to the processing requirements, improving the efficiency and accuracy of data classification and processing.

[0126] Embodiment 7:

[0127] Based on the above Embodiment 6, the data processing module includes:

[0128] Serialization unit: According to the standard data classified under any classification computing unit, the standard data is serialized to obtain byte stream data, and the byte stream data is transmitted into the high-speed transmission channel through the network transmission protocol according to the processing path of the required unit. The transmission rate is:

[0129] ; where R represents the transmission rate of the high-speed transmission of the required unit, represents the bandwidth of the high-speed transmission channel, represents the delay of the high-speed transmission channel, represents the byte stream data, represents the high-speed processing efficiency function;

[0130] If the transmission rate is less than the preset threshold, the high-speed transmission channel is reallocated; otherwise, it is transmitted normally;

[0131] Task distribution unit: Through the high-speed transmission channel, the scheduler distributes according to the type and requirements of the data, and constructs and allocates data processing tasks according to the load balancing algorithm to obtain a multi-task queue;

[0132] For the multi-task queue, the system performs parallel processing and obtains the first intelligent data after processing;

[0133] Compression Unit: Encodes and compresses the first intelligent data, constructs a frequency table based on the symbols of the first intelligent data, merges the two symbols with the lowest frequencies to form a new node until the frequency tree is constructed, generates binary codes for each symbol according to the frequency tree, and obtains the standard intelligent data of the first intelligent data. Specifically:

[0134] ; where represents the compression code of the j-th symbol, represents the j-th symbol of the Huffman value, {} represents the compression function, represents that the first intelligent data has a total of m symbols, represents the interval size from the starting point of the frequency tree, represents the interval size from the ending point of the frequency tree;

[0135] Data Lake Unit: Constructs a data lake for intelligent data and inputs the standard intelligent data of all data sources into the data lake.

[0136] In this embodiment, serialization processing is to convert data from the object form in memory into a format that can be stored or transmitted, so as to facilitate reconstruction or reading in different systems or environments.

[0137] In this embodiment, byte stream data refers to a data stream continuously stored in units of bytes (byte).

[0138] In this embodiment, the network transmission protocol is a rule and standard for regulating the data exchange process in a computer network.

[0139] In this embodiment, the transmission rate refers to the speed at which byte stream data is transmitted through a network transmission channel.

[0140] In this embodiment, the bandwidth refers to the maximum data transmission capacity of a high-speed transmission channel.

[0141] In this embodiment, the high-speed processing efficiency function refers to a function of the influencing factors of data processing speed and efficiency.

[0142] In this embodiment, the preset threshold is a defined standard used to determine whether the transmission rate of the high-speed transmission channel meets the standard.

[0143] In this embodiment, the scheduler reasonably allocates data processing tasks according to the type and requirements of the data.

[0144] In this embodiment, the load balancing algorithm is a technology that evenly distributes the computational task data traffic to multiple processing units.

[0145] In this embodiment, the multi-task queue is a data structure for storing and managing multiple tasks to be processed.

[0146] In this embodiment, the first intelligent data refers to the preliminary processing results obtained after the task distribution unit (scheduler) completes task allocation and through parallel processing.

[0147] In this embodiment, encoding compression converts the original data into a more compact form that occupies less space through a certain encoding algorithm.

[0148] In this embodiment, the frequency table is used to record the occurrence frequency of each symbol (such as characters, bytes, etc.) in the dataset.

[0149] In this embodiment, the frequency tree refers to a tree structure constructed based on the frequency information of different symbols (such as bytes or characters).

[0150] In this embodiment, the standard intelligent data refers to the data after operations such as preprocessing, compression, and encoding.

[0151] In this embodiment, the Huffman value refers to the binary encoding value generated for each symbol by the Huffman coding algorithm.

[0152] The working principle and beneficial effects of the above technical solutions are as follows: Optimize data processing through serialization, compression, task distribution, and data lake management. Use load balancing and parallel computing to improve processing efficiency. At the same time, compress intelligent data through Huffman coding, construct a data lake to centrally manage standard data, and improve data transmission and storage efficiency.

[0153] Embodiment 8:

[0154] Based on the above Embodiment 1, the integration module includes:

[0155] Integration unit: Seamlessly integrate the data acquisition module, setting module, and data processing module, formulate the interfaces between modules according to the data stream and communication protocol, and complete the preliminary integration;

[0156] Conduct integration testing on the preliminary integration. After each module meets the expected standards, complete the integration to obtain the large model. Otherwise, reconfigure the interfaces between each module;

[0157] Testing unit: Set up a test environment, configure the large model into the test environment, monitor the performance of each module and the structure of intelligent data processing in real time, and generate test logs;

[0158] Feedback unit: Obtain the feedback information of the large model operation based on the test logs, and adjust the data processing flow and calculation strategy according to the feedback information.

[0159] In this embodiment, preliminary integration refers to the process of combining and configuring various independent modules or subsystems according to a predetermined design during system development to form a preliminary and executable overall system.

[0160] In this embodiment, integration testing is to verify whether multiple modules can work together properly as expected after integration, ensuring that different parts of the system can effectively interact.

[0161] In this embodiment, a test log refers to a record file generated during software testing, which contains detailed information about the test execution.

[0162] The working principle and beneficial effects of the above technical solution are as follows: By seamlessly integrating data acquisition, setting, and processing modules, integration testing is carried out to ensure the normal function of the modules; by real-time monitoring and feedback adjustment of the data processing flow, the performance and calculation strategy of the large model are improved, and the overall system efficiency and reliability are optimized.

[0163] Embodiment 9:

[0164] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the logical process of the intelligent data processing device based on a large model according to any one of the above.

[0165] The working principle and beneficial effects of the above technical solution are as follows: By storing the computer program in a readable storage medium, the processor executes the intelligent data processing logic based on a large model, automatically executes the data processing flow, improves the calculation efficiency and intelligence level, and optimizes the system performance.

[0166] Embodiment 10:

[0167] An electronic device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the logical process of the intelligent data processing device based on a large model according to any one of the above.

[0168] The working principle and beneficial effects of the above technical solution are as follows: By storing the computer program in the memory, the processor executes the intelligent data processing logic based on a large model, realizes automatic intelligent data processing, improves the device performance and processing efficiency, and optimizes the data processing process.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An intelligent data processing device based on a large model, characterized in that Including: Target module: Determine the preset large model target according to the processing requirements of intelligent data, and design the overall architecture of the large model; Data acquisition module: Obtain the original data through the preset data source, and store the original data in the original database; Setting module: Set the computing unit. At the same time, after preprocessing the original data to obtain the standard data, match different processing paths according to the type and processing requirements of the standard data; Data processing module: Based on the computing unit, transmit the standard data to the corresponding processing path through the high-speed transmission channel, and obtain the standard intelligent data after parallel compression processing, and store it in the data lake; Integration module: Integrate the data acquisition module, the setting module and the data processing module into the overall architecture of the large model, and configure it to the test environment for test verification to obtain feedback information, and adjust the data processing flow and calculation strategy according to the feedback information; Among them, the data processing module includes: Serialization unit: Serialize the standard data according to the standard data classified under any classification computing unit to obtain byte stream data, and transmit the byte stream data into the high-speed transmission channel through the network transmission protocol according to the processing path of the required unit obtained based on the computing unit. The transmission rate is: ; wherein, R represents the transmission rate of the required unit for high-speed transmission, represents the bandwidth of the high-speed transmission channel, represents the delay of the high-speed transmission channel, represents byte stream data, represents a high-speed processing efficiency function; If the transmission rate is less than the preset threshold, reallocate the high-speed transmission channel, otherwise transmit normally; Task distribution unit: Through the high-speed transmission channel, the scheduler distributes according to the type and requirements of the data, and constructs and allocates data processing tasks according to the load balancing algorithm to obtain a multi-task queue; For the multi-task queue, the system performs parallel processing and obtains the first intelligent data after processing; Compression unit: Encode and compress the first intelligent data, construct a frequency table according to the symbols of the first intelligent data, merge the two symbols with the lowest frequency to form a new node until the frequency tree is constructed, and generate binary codes for each symbol according to the frequency tree to obtain the standard intelligent data of the first intelligent data. Specifically: ; among them, represents the compressed encoding of the j-th symbol, represents the j-th symbol 's Huffman value, {} represents the compression function, indicates that the first intelligent data has a total of m symbols, represents the interval size of the frequency tree from the starting point, represents the interval size of the frequency tree from the ending point; Data lake unit: Construct a data lake for intelligent data, and transmit the standard intelligent data of all data sources into the data lake.

2. The intelligent data processing device based on a large model according to claim 1, characterized in that The target module includes: Requirement determination unit: Clarify the application scenario of the intelligent data processing service and determine the processing requirements of the processing device; Task type unit: Based on the processing requirements of the processing device, clarify the intelligent data processing task type, and preset the large model target according to the task type.

3. The intelligent data processing device based on a large model according to claim 1, characterized in that, The target module further includes: Architecture unit: Match the model type and its architecture according to the task type, and generate the overall architecture of the large model; Hardware acceleration unit: Configure the preset hardware as the external hardware part of the large model to perform parallel computing on the large model.

4. The intelligent data processing device based on a large model according to claim 1, wherein The data acquisition module includes: Data source unit: Determine the data source according to the preset source of the intelligent data, and obtain the connection information of all data sources; Interface unit: Configure the data interface according to the data type of any data source, and connect through the HTTP request based on the connection information to obtain the original data of the data source; Data storage unit: Design the original database table structure according to the preset structure, and store the original data into the original database with the combined primary key of timestamp - data source identifier.

5. The intelligent data processing device based on a large model according to claim 1, wherein The setting module includes: Calculation unit configuration: Set multiple calculation units for parallel data processing according to the data volume and processing requirements of the original data; Normalization unit: Obtain the original data of any data source in the original database, perform data cleaning on the original data, and remove and replace outliers to obtain the first data; After performing standard normalization on the first data, standard data is obtained, specifically: ; wherein, represents the standard data at the i-th time node in the first data, represents the data value at the i-th time node in the first data, represents the mean value of the data values at all time nodes in the first data, represents the standard deviation of the data values at all time nodes in the first data, and n represents that there are n time nodes in the first data, represents the minimum value among the data values at all time nodes in the first data, represents the maximum value among the data values at all time nodes in the first data.

6. The intelligent data processing device based on a large model according to claim 5, wherein The setting module further includes: Classification unit: Determine the number of classifications according to the number of calculation units, perform feature engineering processing on the original data of any data source, and obtain the first feature sequence of the original data; If the number of sequences in the first feature sequence is greater than the number of determined classifications, perform principal component analysis dimensionality reduction operation on the first feature sequence to obtain a second feature sequence consistent with the number of determined classifications. Otherwise, directly use the first feature sequence as the second feature sequence; Match any feature of the second feature sequence with the classification standard and the calculation unit to obtain the required unit, perform feature selection on the standard data, select the feature with the highest expected value as the classification, and allocate it to the corresponding required unit; The required unit selects the processing path of the required unit according to the data type - processing requirement - processing path.

7. The intelligent data processing device based on a large model according to claim 1, wherein The integration module includes: Integration unit: Seamlessly integrate the data acquisition module, setting module, and data processing module, formulate the interfaces between the modules according to the data flow and communication protocol, and complete the preliminary integration; Conduct integration testing on the preliminary integration. After each module meets the expected standards, the integration is completed to obtain the large model. Otherwise, reconfigure the interfaces between the modules; Testing unit: Build a test environment, configure the large model into the test environment, monitor the performance of each module and the structure of intelligent data processing in real time, and generate test logs; Feedback unit: Obtain the feedback information of the large model operation based on the test logs, and adjust the data processing flow and calculation strategy according to the feedback information.

Citation Information

Patent Citations

  • Intelligent auditing model construction method and device based on big data and medium

    CN118820812A