Complex data processing method and system based on large model

Through complex data processing methods based on large models, including technical means such as data cleaning, multimodal feature extraction and model training, the problems of low efficiency and large error in complex data processing in the existing technology are solved, and more efficient and accurate data analysis is achieved.

CN119989298AInactive Publication Date: 2025-05-13GUANGZHOU YUEZHENG NETWORK INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081474.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is inefficient and has large errors when processing complex data, and cannot efficiently execute applications with different resource requirements at the same time, resulting in insufficient comprehensiveness and accuracy of data processing.

Method used

A complex data processing method based on large models is adopted to build a complex data analysis model framework through data cleaning, multimodal feature extraction, feature coding, feature fusion and automatic configuration algorithms, and an analysis model that meets preset performance standards is obtained through model training and verification.

Benefits of technology

Improves data accuracy, completeness and consistency, enhances model performance and prediction accuracy, and can efficiently process large and complex data, suitable for real-time or near-real-time data analysis needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989298A_ABST
    Figure CN119989298A_ABST
Patent Text Reader

Abstract

The invention relates to the field of big data analysis, and discloses a complex data processing method and system based on a large model, and the method comprises the steps: carrying out the data cleaning of complex data, defining a multi-modal feature extraction algorithm of the cleaned data, extracting the multi-modal features of the cleaned data, determining a model activation function of the cleaned data, and carrying out the analysis of the model activation function. Performing feature fusion on the encoded data, defining an automatic configuration algorithm of the fused data, and configuring a complex data analysis model framework of the fused data; determining a loss function of the complex data analysis model framework, constructing an initial model of fused data, defining a model training algorithm of the initial model, training the initial model, and verifying the model performance of the training model; and when the model performance meets a preset model performance standard, taking the training model as a complex data analysis model of the fused data, and analyzing the fused data to obtain target data. According to the invention, the comprehensiveness and accuracy of complex data processing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a complex data processing method and system based on a large model, and belongs to the field of big data analysis. Background Art

[0002] Complex data processing refers to a series of operations and transformations on data that are high-dimensional, large-scale, multi-source, multi-modal, unstructured or semi-structured, in order to extract useful information, discover data patterns, support decision making or perform specific tasks. Complex data processing helps to quickly and accurately obtain information from large amounts of data, thereby improving the efficiency and quality of decision-making.

[0003] At present, data processing methods are mainly carried out through traditional data software. Due to the large volume and complexity of the data to be processed, this method cannot efficiently execute applications with different resource requirements at the same time, resulting in low data processing efficiency and large errors.

[0004] Therefore, there is an urgent need for a solution that can improve the comprehensiveness and accuracy of complex data processing. Summary of the invention

[0005] The present invention provides a complex data processing method and system based on a large model, the main purpose of which is to improve the comprehensiveness and accuracy of complex data processing.

[0006] To achieve the above object, the present invention provides a complex data processing method based on a large model, comprising:

[0007] Clarify the processing objectives and processing requirements of complex data, identify missing values, duplicate values ​​and abnormal values ​​of the complex data, and perform data cleaning on the complex data based on the missing values, duplicate values ​​and abnormal values ​​to obtain cleaned data;

[0008] Defining a multimodal feature extraction algorithm for the cleaned data, extracting multimodal features of the cleaned data based on the multimodal feature extraction algorithm, determining a model activation function for the cleaned data, and performing feature encoding on the cleaned data according to the model activation function to obtain encoded data;

[0009] According to the multimodal features, the encoded data is subjected to feature fusion to obtain fused data, an automatic configuration algorithm for the fused data is defined according to the processing target and the processing requirement, and a complex data analysis model framework for the fused data is configured according to the automatic configuration algorithm;

[0010] Determine the loss function of the complex data analysis model framework, construct an initial model of the fused data according to the loss function and the complex data analysis model framework, define a model training algorithm for the initial model, train the initial model using a preset training set according to the model training algorithm to obtain a training model, and verify the model performance of the training model using a preset validation set;

[0011] When the model performance meets the preset model performance standard, the training model is used as the complex data analysis model of the fused data, and the fused data is analyzed using the complex data analysis model to obtain the target data.

[0012] Optionally, performing data cleaning on the complex data based on the missing values, the repeated values, and the abnormal values ​​to obtain cleaned data includes:

[0013] Extracting missing value-related data of the complex data according to the missing value;

[0014] Constructing a regression analysis model for the missing value related data, and analyzing the analysis value of the missing value according to the regression analysis model;

[0015] Filling the complex data using the analysis value to obtain filled data;

[0016] Deleting duplicate items from the filled data according to the repeated values ​​and the processing requirements corresponding to the complex data to obtain deleted data;

[0017] Analyze the outlier type of the outlier, and perform outlier conversion on the deleted data according to the outlier type to obtain converted data;

[0018] The data quality of the converted data is detected, and when the data quality meets a preset standard, the converted data is used as cleaned data.

[0019] Optionally, the multimodal feature extraction algorithm for defining the cleaned data includes:

[0020] Analyze the feature extraction requirements of the cleaned data;

[0021] According to the feature extraction requirements, the data types of the cleaned data are divided;

[0022] According to the data type, configuring a multi-class feature extraction algorithm for the cleaned data;

[0023] Respectively identifying multiple types of algorithm parameters of the multiple types of feature extraction algorithms;

[0024] According to the multiple types of algorithm parameters, the multiple types of feature extraction algorithms are fused to obtain a fused algorithm;

[0025] Using preset test data to test the fused algorithm to obtain the accuracy of the algorithm;

[0026] When the algorithm accuracy meets a preset algorithm accuracy threshold, the fused algorithm is used as a multimodal feature extraction algorithm for the cleaned data.

[0027] Optionally, analyzing data distribution and data characteristics of the cleaned data;

[0028] Based on the data distribution, normalizing the cleaned data to obtain normalized data;

[0029] Determining a model activation function framework of the normalized data according to the data characteristics;

[0030] Determining an output range of the model activation function framework;

[0031] A model activation function of the normalized data is determined according to the output range.

[0032] Optionally, the step of performing feature fusion on the encoded data according to the multimodal features to obtain fused data includes:

[0033] Determining a multimodal feature vector of the multimodal feature;

[0034] Analyzing the multimodal dimension of the multimodal feature according to the multimodal feature vector;

[0035] Determining a fusion dimension of the multimodal features according to the multimodal dimension;

[0036] According to the fusion dimension, the multimodal features are spliced ​​to obtain spliced ​​features;

[0037] According to the spliced ​​features, feature fusion is performed on the encoded data to obtain fused data.

[0038] Optionally, defining an automatic configuration algorithm for the fused data according to the processing target and the processing requirement includes:

[0039] Obtaining a model framework set of the fused data;

[0040] Determining configuration parameters of the fused data according to the processing target and the processing requirement;

[0041] Defining configuration logic of the fused data and the model framework set;

[0042] Calculating the matching degree of the model frames in the model frame set according to the configuration parameters;

[0043] According to the matching degree between the configuration logic and the model framework, an automatic configuration algorithm for the fused data is defined.

[0044] Optionally, determining the loss function of the complex data analysis model framework includes:

[0045] Divide the fused data corresponding to the complex data analysis model framework into multiple document pairs, and determine the number of document pairs in the multiple document pairs;

[0046] Analyzing the multi-document correlation between the multi-document pairs and the complex data analysis model framework;

[0047] A loss function of the complex data analysis model framework is determined based on the multiple document pairs, the number of document pairs, and the relevance of the multiple documents.

[0048] Optionally, the model training algorithm for defining the initial model includes:

[0049] Determining the number of training set groups and model parameters of the initial model;

[0050] Determining an optimization algorithm for the initial model according to a loss function corresponding to the initial model;

[0051] Determining a training iteration mechanism for the initial model according to the number of training set groups and the model parameters;

[0052] Constructing a training feedback mechanism for the initial model, and determining the number of iterations of the training iteration mechanism according to the training feedback mechanism;

[0053] A model training algorithm for the initial model is defined according to the optimization algorithm, the training iteration mechanism, the training feedback mechanism and the number of iterations.

[0054] Optionally, the using a preset validation set to validate the model performance of the training model includes:

[0055] Using the verification set to verify the training model to obtain verification data;

[0056] Calculate the accuracy and recall of the training model according to the verification data;

[0057] Determining an F1 score of the training model according to the precision and the recall;

[0058] Analyzing the confusion matrix and analyzing errors of the training model;

[0059] Based on the confusion matrix, the analysis error, and the F1 score, a model performance of the training model is determined.

[0060] In order to solve the above problems, the present invention further provides a complex data processing system based on a large model, the system comprising:

[0061] A data cleaning module is used to clarify the processing objectives and processing requirements of complex data, identify missing values, duplicate values ​​and abnormal values ​​of the complex data, and perform data cleaning on the complex data based on the missing values, duplicate values ​​and abnormal values ​​to obtain cleaned data;

[0062] A feature extraction module, used to define a multimodal feature extraction algorithm for the cleaned data, extract multimodal features of the cleaned data based on the multimodal feature extraction algorithm, determine a model activation function for the cleaned data, and perform feature encoding on the cleaned data according to the model activation function to obtain encoded data;

[0063] A model framework configuration module, for performing feature fusion on the encoded data according to the multimodal features to obtain fused data, defining an automatic configuration algorithm for the fused data according to the processing target and the processing requirement, and configuring a complex data analysis model framework for the fused data according to the automatic configuration algorithm;

[0064] A model training module is used to determine the loss function of the complex data analysis model framework, construct an initial model of the fused data according to the loss function and the complex data analysis model framework, define a model training algorithm for the initial model, train the initial model according to the model training algorithm using a preset training set to obtain a training model, and verify the model performance of the training model using a preset validation set;

[0065] The model analysis module is used to use the training model as a complex data analysis model for the fused data when the model performance meets the preset model performance standard, and use the complex data analysis model to analyze the fused data to obtain target data.

[0066] The embodiment of the present invention cleans the complex data based on the missing values, the duplicate values ​​and the abnormal values ​​to obtain cleaned data, which can improve the accuracy, completeness and consistency of the data and reduce the impact of data errors on the analysis results; optionally, the embodiment of the present invention can increase the dimension of the data by defining a multimodal feature extraction algorithm for the cleaned data, so that the model can capture more subtle information, thereby improving the performance of the model; the embodiment of the present invention can select the most appropriate model structure and parameters according to the characteristics of the data and task requirements according to the processing objectives and the processing requirements, thereby improving the accuracy of data analysis. The embodiment of the present invention can more accurately learn and predict patterns in complex data by defining a model training algorithm for the initial model, thereby improving overall prediction accuracy; the embodiment of the present invention uses the model training algorithm to train the initial model using a preset training set, and the training model obtained can more accurately learn and predict patterns and relationships in complex data, reducing prediction errors, and finally, when the model performance meets the preset model performance standards, the embodiment of the present invention uses the training model as the complex data analysis model of the fused data to efficiently process large amounts of complex fused data, quickly obtain analysis results, and is suitable for real-time or near real-time data analysis needs. Therefore, the complex data processing method and system based on large models provided by the embodiment of the present invention can improve the comprehensiveness and accuracy of complex data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 A schematic diagram of a flow chart of a complex data processing method based on a large model provided by an embodiment of the present invention;

[0068] Figure 2 A schematic diagram of modules for implementing the large model-based complex data processing method provided in one embodiment of the present invention.

[0069] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION

[0070] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0071] The embodiment of the present application provides a complex data processing method based on a large model. The execution subject of the complex data processing method based on a large model includes but is not limited to at least one of the electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the complex data processing method based on a large model can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.

[0072] Embodiment 1:

[0073] Reference Figure 1 FIG. 1 is a flow chart of a complex data processing method based on a large model provided by an embodiment of the present invention. In this embodiment, the complex data processing method based on a large model includes:

[0074] S1. Clarify the processing objectives and processing requirements of complex data, identify missing values, duplicate values ​​and abnormal values ​​of the complex data, and perform data cleaning on the complex data based on the missing values, duplicate values ​​and abnormal values ​​to obtain cleaned data;

[0075] The embodiments of the present invention can provide a clear processing direction by clarifying the processing objectives and processing requirements of complex data, ensuring that data processing activities are carried out around the ultimate goal and avoiding ineffective work. The processing objectives refer to the specific purpose or result that is desired to be achieved when performing complex data processing. The processing requirements refer to the specific conditions, requirements or constraints that need to be met in the process of achieving the data processing objectives.

[0076] The embodiment of the present invention can provide support for subsequent improvement of data quality by identifying missing values, duplicate values ​​and outliers of the complex data. Wherein, the missing value refers to the fact that there is no record value for the attribute or field of some data in the data set. The duplicate value refers to the existence of completely identical or substantially identical data records in the data set. The outlier value refers to a value that significantly deviates from other observed values ​​in the data set.

[0077] Optionally, as an embodiment of the present invention, the identification of missing values, duplicate values ​​and abnormal values ​​of the complex data may be performed by an integrated learning method.

[0078] The embodiment of the present invention performs data cleaning on the complex data based on the missing values, the duplicate values ​​and the abnormal values ​​to obtain cleaned data, which can improve the accuracy, completeness and consistency of the data and reduce the impact of data errors on the analysis results. The cleaned data refers to a data set that has been processed through a series of data cleaning operations.

[0079] As an embodiment of the present invention, the data cleaning of the complex data based on the missing values, the repeated values ​​and the abnormal values ​​to obtain the cleaned data includes:

[0080] Extracting missing value-related data of the complex data according to the missing value;

[0081] Constructing a regression analysis model for the missing value related data, and analyzing the analysis value of the missing value according to the regression analysis model;

[0082] Filling the complex data using the analysis value to obtain filled data;

[0083] Deleting duplicate items from the filled data according to the repeated values ​​and the processing requirements corresponding to the complex data to obtain deleted data;

[0084] Analyze the outlier type of the outlier, and perform outlier conversion on the deleted data according to the outlier type to obtain converted data;

[0085] The data quality of the converted data is detected, and when the data quality meets the preset standards, the converted data is used as the cleaned data. The missing value-related data refers to data that is directly or indirectly related to the missing values ​​in the data set. The regression analysis model refers to a statistical model used to predict missing data points using existing complete data. The analysis value refers to the predicted value or estimated value obtained after applying the regression analysis model, which is used to fill the missing values ​​in the data set. The filled data refers to the data set obtained after the analysis value is used to fill the missing values ​​in the original data set. The deleted data refers to the data set obtained after processing the filled data and removing the duplicates therein. The outlier type refers to those data points in the data set that do not conform to the normal data distribution or pattern. The converted data refers to the data set obtained after identifying and processing the outliers in the data set.

[0086] Optionally, the regression analysis model for constructing the missing value related data can be constructed by a multiple interpolation method.

[0087] S2. Define a multimodal feature extraction algorithm for the cleaned data, extract multimodal features of the cleaned data based on the multimodal feature extraction algorithm, determine a model activation function for the cleaned data, and perform feature encoding on the cleaned data according to the model activation function to obtain encoded data.

[0088] The embodiment of the present invention can increase the dimension of the data by defining a multimodal feature extraction algorithm for the cleaned data, so that the model can capture more subtle information, thereby improving the performance of the model. The multimodal feature extraction algorithm refers to a series of calculation methods and processes for extracting useful features from a data set containing multiple data types (i.e., multimodal data).

[0089] As an embodiment of the present invention, the multimodal feature extraction algorithm for defining the cleaned data includes:

[0090] Analyze the feature extraction requirements of the cleaned data;

[0091] According to the feature extraction requirements, the data types of the cleaned data are divided;

[0092] According to the data type, configuring a multi-class feature extraction algorithm for the cleaned data;

[0093] Respectively identifying multiple types of algorithm parameters of the multiple types of feature extraction algorithms;

[0094] According to the multiple types of algorithm parameters, the multiple types of feature extraction algorithms are fused to obtain a fused algorithm;

[0095] Using preset test data to test the fused algorithm to obtain the accuracy of the algorithm;

[0096] When the algorithm accuracy meets a preset algorithm accuracy threshold, the fused algorithm is used as a multimodal feature extraction algorithm for the cleaned data.

[0097] Among them, the feature extraction requirements refer to the specific requirements for the nature, type, quantity and quality of the features to be extracted from the data set. The data type refers to the way data is classified in the computer. The multi-class feature extraction algorithm refers to a set of methods for extracting features from data sets, which contain multiple types of features, such as numerical, categorical, textual, etc. The multi-class algorithm parameters refer to the variables and configuration options that need to be set when implementing the multi-class feature extraction algorithm. The algorithm accuracy refers to the accuracy of the algorithm when performing a specific task. The preset algorithm accuracy threshold refers to a pre-set accuracy standard or limit, which is used to determine whether the performance of the model meets specific business requirements or reaches an acceptable performance level.

[0098] Optionally, the multiple feature extraction algorithms are fused according to the multiple algorithm parameters, and the fused algorithms can be fused by a feature level fusion method.

[0099] The embodiment of the present invention uses the multimodal feature extraction algorithm to extract the multimodal features of the cleaned data, which can provide an explanation of the data from different angles, and help to better understand the meaning and relationship behind the data. The multimodal feature refers to the information representation extracted from different data sources or different types of data.

[0100] The embodiment of the present invention can more effectively learn complex features from the data and improve the expressiveness of the features by determining the model activation function of the cleaned data. The model activation function refers to a function used to introduce nonlinear factors and determine whether a neuron should be activated.

[0101] As an embodiment of the present invention, determining the model activation function of the cleaned data includes:

[0102] Analyzing data distribution and data characteristics of the cleaned data;

[0103] Based on the data distribution, normalizing the cleaned data to obtain normalized data;

[0104] Determining a model activation function framework of the normalized data according to the data characteristics;

[0105] Determining an output range of the model activation function framework;

[0106] Determine a model activation function of the normalized data according to the output range, wherein the model activation function includes:

[0107]

[0108] Among them, gyh(s) represents the model activation function, s represents the normalized data, and e represents the exponential function with e as the base.

[0109] Among them, the data distribution refers to the frequency of the arrangement and occurrence of values ​​or categories in the data set according to a certain rule or pattern. The data characteristics refer to the specific attributes and characteristics of the data set. The normalized data refers to the data that adjusts the original data to a uniform scale or range through specific mathematical transformations to facilitate comparison and analysis. The model activation function framework refers to a series of function frameworks used to determine whether a neuron should be activated. The output range refers to the possible interval of the output value of the activation function after receiving the input signal.

[0110] The embodiment of the present invention performs feature encoding on the cleaned data according to the model activation function to obtain the encoded data, which enables the model to capture the complex relationships in the data, and the feature-encoded data can better represent these relationships. The encoded data refers to data that has been processed by feature encoding.

[0111] S3. According to the multimodal features, the encoded data are subjected to feature fusion to obtain fused data. According to the processing objectives and the processing requirements, an automatic configuration algorithm for the fused data is defined. According to the automatic configuration algorithm, a complex data analysis model framework for the fused data is configured.

[0112] The embodiment of the present invention performs feature fusion on the encoded data according to the multimodal features to obtain fused data that can more comprehensively describe the data and enhance the model's ability to understand the data, wherein the fused data refers to data processed by the multimodal feature fusion technology.

[0113] As an embodiment of the present invention, the step of performing feature fusion on the encoded data according to the multimodal features to obtain fused data includes:

[0114] Determining a multimodal feature vector of the multimodal feature;

[0115] Analyzing the multimodal dimension of the multimodal feature according to the multimodal feature vector;

[0116] Determining a fusion dimension of the multimodal features according to the multimodal dimension;

[0117] According to the fusion dimension, the multimodal features are spliced ​​to obtain spliced ​​features;

[0118] According to the spliced ​​features, feature fusion is performed on the encoded data to obtain fused data.

[0119] Among them, the multimodal feature vector refers to a set of features extracted from data of different modalities, and each feature vector represents a specific attribute or information of a modal data. The multimodal dimension refers to the number of features of each individual modal feature vector, or the length of each modal feature vector. The fusion dimension refers to the dimension of the fused feature vector during the multimodal feature fusion process. The spliced ​​feature refers to a new feature formed by combining feature vectors from different modalities in a certain way.

[0120] Optionally, determining the fusion dimension of the multimodal features according to the multimodal dimension may be determined by an adaptive fusion method.

[0121] The embodiment of the present invention defines the automatic configuration algorithm of the fused data according to the processing target and the processing requirement, and can select the most suitable model structure and parameters according to the characteristics of the data and the task requirements, thereby improving the accuracy of data analysis. The automatic configuration algorithm refers to an algorithm that can automatically select and adjust the parameters, structure and processing flow of the data analysis model.

[0122] As an embodiment of the present invention, the automatic configuration algorithm of the fused data is defined according to the processing target and the processing requirement, including:

[0123] Obtaining a model framework set of the fused data;

[0124] Determining configuration parameters of the fused data according to the processing target and the processing requirement;

[0125] Defining configuration logic of the fused data and the model framework set;

[0126] Calculating the matching degree of the model frames in the model frame set according to the configuration parameters;

[0127] According to the matching degree between the configuration logic and the model framework, an automatic configuration algorithm for the fused data is defined.

[0128] The model framework set refers to a set of predefined or dynamically generated data model structures. The configuration parameters refer to a set of parameters used to define and adjust the data analysis model framework. The configuration logic refers to a set of rules or criteria for guiding how to select and adjust the configuration parameters of the model framework according to specific data processing goals and requirements. The model framework matching degree refers to a quantitative indicator used to evaluate the degree of adaptability between a given model framework and specific data processing tasks and requirements.

[0129] The embodiment of the present invention can automatically select the most appropriate model structure and parameters by configuring the complex data analysis model framework of the fused data according to the automatic configuration algorithm, thereby improving the prediction accuracy, generalization ability and robustness of the model. The complex data analysis model framework refers to a predefined structure or blueprint for analyzing and processing complex data.

[0130] S4. Determine the loss function of the complex data analysis model framework, construct an initial model of the fused data based on the loss function and the complex data analysis model framework, define a model training algorithm for the initial model, train the initial model using a preset training set according to the model training algorithm to obtain a training model, and verify the model performance of the training model using a preset validation set.

[0131] The embodiment of the present invention can guide the update direction of the model parameters by determining the loss function of the complex data analysis model framework, and can effectively find the optimal or approximately optimal solution of the parameters through the gradient descent algorithm. Among them, the loss function refers to a mathematical function used to quantify the difference between the model prediction result and the actual result.

[0132] As an embodiment of the present invention, the determining of the loss function of the complex data analysis model framework includes:

[0133] Divide the fused data corresponding to the complex data analysis model framework into multiple document pairs, and determine the number of document pairs in the multiple document pairs;

[0134] Analyzing the multi-document correlation between the multi-document pairs and the complex data analysis model framework;

[0135] Determine a loss function of the complex data analysis model framework according to the multiple document pairs, the number of the document pairs, and the relevance of the multiple documents, wherein the loss function includes:

[0136]

[0137] Among them, G represents the loss function, M represents the number of document pairs, i represents the document i corresponding to multiple documents, j represents the document j corresponding to multiple documents, x(i) represents the document i correlation corresponding to the multi-document correlation, x(j) represents the document j correlation corresponding to the multi-document correlation, f represents the logical function, u ij Represents the indicator function. When the relevance of document i is greater than that of document j, then u ij is 1, otherwise it is 0, lg represents the logarithmic function with base 10.

[0138] The multi-document pair refers to a combination formed by pairing documents in a data set in pairs during data processing and analysis. The number of document pairs refers to the total number of document pairs formed during data processing or model training. The multi-document relevance refers to an indicator or measure that measures the degree of correlation between two or more documents in terms of content, subject, semantics or other specific attributes.

[0139] The embodiment of the present invention can customize the model for specific data analysis and tasks by constructing the initial model of the fused data according to the loss function and the complex data analysis model framework, ensuring that the model output is consistent with the business objectives. The initial model refers to the basic framework and structure of the model constructed according to specific data analysis tasks and requirements before starting model training.

[0140] Optionally, constructing the initial model of the fused data based on the loss function and the complex data analysis model framework can be constructed by a transfer learning method.

[0141] The embodiment of the present invention can more accurately learn and predict patterns in complex data by defining a model training algorithm for the initial model, thereby improving the overall prediction accuracy. The model training algorithm refers to a series of processes and methods for adjusting and optimizing model parameters so that the model can achieve a predetermined performance target on a given data set.

[0142] As an embodiment of the present invention, the model training algorithm for defining the initial model includes:

[0143] Determining the number of training set groups and model parameters of the initial model;

[0144] Determining an optimization algorithm for the initial model according to a loss function corresponding to the initial model;

[0145] Determining a training iteration mechanism for the initial model according to the number of training set groups and the model parameters;

[0146] Constructing a training feedback mechanism for the initial model, and determining the number of iterations of the training iteration mechanism according to the training feedback mechanism;

[0147] A model training algorithm for the initial model is defined according to the optimization algorithm, the training iteration mechanism, the training feedback mechanism and the number of iterations.

[0148] Among them, the number of training sets refers to the number of batches into which the entire training data set is divided during the training process of the machine learning model. The model parameters refer to the learnable variables within the model in a machine learning or deep learning task. The optimization algorithm refers to a class of algorithms used to adjust model parameters to minimize the loss function during the training process of a machine learning or deep learning model. The training iteration mechanism refers to a cyclic execution method in the training process of a machine learning or deep learning model, which gradually optimizes the model parameters through multiple iterations until a certain performance standard or stop condition is reached. The training feedback mechanism refers to a systematic method for monitoring and evaluating model performance during one or more rounds of model training iterations.

[0149] Optionally, the training iteration mechanism for determining the initial model based on the number of training set groups and the model parameters can be determined by a Bayesian optimization method.

[0150] The embodiment of the present invention trains the initial model using a preset training set according to the model training algorithm to obtain a training model that can more accurately learn and predict patterns and relationships in complex data and reduce prediction errors. The training model refers to the final model obtained after a certain training cycle, using a specific algorithm and data set to adjust and optimize the parameters of the initial model.

[0151] The embodiment of the present invention can objectively evaluate the performance of the training model on unknown data by using a preset validation set to verify the model performance of the training model, wherein the model performance refers to the performance metric of the model on a specific task.

[0152] As an embodiment of the present invention, the using a preset validation set to validate the model performance of the training model includes:

[0153] Using the verification set to verify the training model to obtain verification data;

[0154] Calculate the accuracy and recall of the training model according to the verification data;

[0155] Determining an F1 score of the training model according to the precision and the recall;

[0156] Analyzing the confusion matrix and analyzing errors of the training model;

[0157] Based on the confusion matrix, the analysis error, and the F1 score, a model performance of the training model is determined.

[0158] Among them, the validation data refers to a set of data generated in the process of evaluating the training model using the validation set. The accuracy rate refers to the proportion of samples correctly predicted by the model on a given data set. The recall rate refers to an important indicator for measuring the performance of a classification model, reflecting the ability of the model to correctly identify all positive samples. The F1 score refers to an indicator used to measure the average performance of the accuracy and recall rate of a classification model. The confusion matrix refers to a specific table layout used to visualize algorithm performance. The analytical error refers to the deviation, inaccuracy or error that occurs during the data analysis process.

[0159] S5. When the model performance meets the preset model performance standard, the training model is used as the complex data analysis model of the fused data, and the fused data is analyzed using the complex data analysis model to obtain the target data.

[0160] The embodiment of the present invention can efficiently process a large amount of complex fused data and quickly obtain analysis results by using the training model as the complex data analysis model of the fused data when the model performance meets the preset model performance standard, and is suitable for real-time or near real-time data analysis needs. The complex data analysis model refers to an algorithm model that can process and analyze data sets with complex structures, high dimensions, and intricate relationships.

[0161] The embodiment of the present invention uses the complex data analysis model to analyze the fused data to obtain target data, which can extract valuable information and insights from a large amount of complex data, and provide powerful data support and decision-making assistance for various industries and fields.

[0162] The embodiment of the present invention cleans the complex data based on the missing values, the duplicate values ​​and the abnormal values ​​to obtain cleaned data, which can improve the accuracy, completeness and consistency of the data and reduce the impact of data errors on the analysis results; optionally, the embodiment of the present invention can increase the dimension of the data by defining a multimodal feature extraction algorithm for the cleaned data, so that the model can capture more subtle information, thereby improving the performance of the model; the embodiment of the present invention can select the most appropriate model structure and parameters according to the characteristics of the data and task requirements according to the processing objectives and the processing requirements, thereby improving the accuracy of data analysis. The embodiment of the present invention can more accurately learn and predict patterns in complex data by defining a model training algorithm for the initial model, thereby improving overall prediction accuracy; the embodiment of the present invention uses the model training algorithm to train the initial model using a preset training set, and the training model obtained can more accurately learn and predict patterns and relationships in complex data, reducing prediction errors, and finally, when the model performance meets the preset model performance standards, the embodiment of the present invention uses the training model as the complex data analysis model of the fused data to efficiently process large amounts of complex fused data, quickly obtain analysis results, and is suitable for real-time or near real-time data analysis needs. Therefore, the complex data processing method and system based on large models provided by the embodiment of the present invention can improve the comprehensiveness and accuracy of complex data processing.

[0163] Embodiment 2:

[0164] like Figure 2 The figure shows a functional module diagram of a complex data processing system based on a large model according to the present invention.

[0165] The complex data processing system 200 based on a large model described in the present invention can be installed in an electronic device. According to the functions implemented, the complex data processing system based on a large model can include a data cleaning module 201, a feature extraction module 202, a model framework configuration module 203, a model training module 204 and a model analysis module 205. The module described in the present invention can also be called a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.

[0166] In the embodiment of the present invention, the functions of each module / unit are as follows:

[0167] The data cleaning module 201 is used to clarify the processing objectives and processing requirements of complex data, identify missing values, duplicate values ​​and abnormal values ​​of the complex data, and perform data cleaning on the complex data based on the missing values, duplicate values ​​and abnormal values ​​to obtain cleaned data;

[0168] The feature extraction module 202 is used to define a multimodal feature extraction algorithm for the cleaned data, extract multimodal features of the cleaned data based on the multimodal feature extraction algorithm, determine a model activation function for the cleaned data, and perform feature encoding on the cleaned data according to the model activation function to obtain encoded data;

[0169] The model framework configuration module 203 is used to perform feature fusion on the encoded data according to the multimodal features to obtain fused data, define an automatic configuration algorithm for the fused data according to the processing target and the processing requirement, and configure a complex data analysis model framework for the fused data according to the automatic configuration algorithm;

[0170] The model training module 204 is used to determine the loss function of the complex data analysis model framework, construct an initial model of the fused data according to the loss function and the complex data analysis model framework, define a model training algorithm for the initial model, train the initial model according to the model training algorithm using a preset training set to obtain a training model, and verify the model performance of the training model using a preset validation set;

[0171] The model analysis module 205 is used to use the training model as a complex data analysis model for the fused data when the model performance meets the preset model performance standard, and use the complex data analysis model to analyze the fused data to obtain target data.

[0172] In detail, each module in the complex data processing system 200 based on the large model in the embodiment of the present invention is used in the same manner as described above. Figure 1 The same technical means as the complex data processing method based on large models described in the text can produce the same technical effects, so I will not go into details here.

[0173] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A complex data processing method based on a large model, characterized in that: The method comprises: Clarify the processing objectives and processing requirements of complex data, identify missing values, duplicate values ​​and abnormal values ​​of the complex data, and perform data cleaning on the complex data based on the missing values, duplicate values ​​and abnormal values ​​to obtain cleaned data; Defining a multimodal feature extraction algorithm for the cleaned data, extracting multimodal features of the cleaned data based on the multimodal feature extraction algorithm, determining a model activation function for the cleaned data, and performing feature encoding on the cleaned data according to the model activation function to obtain encoded data; According to the multimodal features, the encoded data is subjected to feature fusion to obtain fused data, an automatic configuration algorithm for the fused data is defined according to the processing target and the processing requirement, and a complex data analysis model framework for the fused data is configured according to the automatic configuration algorithm; Determine the loss function of the complex data analysis model framework, construct an initial model of the fused data according to the loss function and the complex data analysis model framework, define a model training algorithm for the initial model, train the initial model using a preset training set according to the model training algorithm to obtain a training model, and verify the model performance of the training model using a preset validation set; When the model performance meets the preset model performance standard, the training model is used as the complex data analysis model of the fused data, and the fused data is analyzed using the complex data analysis model to obtain the target data.

2. The complex data processing method based on a large model as claimed in claim 1, characterized in that: The step of performing data cleaning on the complex data based on the missing values, the repeated values, and the abnormal values ​​to obtain cleaned data includes: Extracting missing value-related data of the complex data according to the missing value; Constructing a regression analysis model for the missing value related data, and analyzing the analysis value of the missing value according to the regression analysis model; Filling the complex data using the analysis value to obtain filled data; Deleting duplicate items from the filled data according to the repeated values ​​and the processing requirements corresponding to the complex data to obtain deleted data; Analyze the outlier type of the outlier, and perform outlier conversion on the deleted data according to the outlier type to obtain converted data; The data quality of the converted data is detected, and when the data quality meets a preset standard, the converted data is used as cleaned data.

3. The complex data processing method based on a large model as claimed in claim 1, characterized in that: The multimodal feature extraction algorithm for defining the cleaned data includes: Analyze the feature extraction requirements of the cleaned data; According to the feature extraction requirements, the data types of the cleaned data are divided; According to the data type, configuring a multi-class feature extraction algorithm for the cleaned data; Respectively identifying multiple types of algorithm parameters of the multiple types of feature extraction algorithms; According to the multiple types of algorithm parameters, the multiple types of feature extraction algorithms are fused to obtain a fused algorithm; Using preset test data to test the fused algorithm to obtain the accuracy of the algorithm; When the algorithm accuracy meets a preset algorithm accuracy threshold, the fused algorithm is used as a multimodal feature extraction algorithm for the cleaned data.

4. The complex data processing method based on a large model as claimed in claim 1, characterized in that: Analyzing data distribution and data characteristics of the cleaned data; Based on the data distribution, normalizing the cleaned data to obtain normalized data; Determining a model activation function framework of the normalized data according to the data characteristics; Determining an output range of the model activation function framework; A model activation function of the normalized data is determined according to the output range.

5. The complex data processing method based on a large model as claimed in claim 1, characterized in that: The step of fusing the encoded data according to the multimodal features to obtain fused data includes: Determining a multimodal feature vector of the multimodal feature; Analyzing the multimodal dimension of the multimodal feature according to the multimodal feature vector; Determining a fusion dimension of the multimodal features according to the multimodal dimension; According to the fusion dimension, the multimodal features are spliced ​​to obtain spliced ​​features; According to the spliced ​​features, feature fusion is performed on the encoded data to obtain fused data.

6. The complex data processing method based on a large model as claimed in claim 1, characterized in that: Defining the automatic configuration algorithm of the fused data according to the processing target and the processing requirement includes: Obtaining a model framework set of the fused data; Determining configuration parameters of the fused data according to the processing target and the processing requirement; Defining configuration logic of the fused data and the model framework set; Calculating the matching degree of the model frames in the model frame set according to the configuration parameters; According to the matching degree between the configuration logic and the model framework, an automatic configuration algorithm for the fused data is defined.

7. The complex data processing method based on a large model as claimed in claim 1, characterized in that: The determining of the loss function of the complex data analysis model framework comprises: Divide the fused data corresponding to the complex data analysis model framework into multiple document pairs, and determine the number of document pairs in the multiple document pairs; Analyzing the multi-document correlation between the multi-document pairs and the complex data analysis model framework; A loss function of the complex data analysis model framework is determined based on the multiple document pairs, the number of document pairs, and the relevance of the multiple documents.

8. The complex data processing method based on a large model as claimed in claim 1, characterized in that: The model training algorithm for defining the initial model includes: Determining the number of training set groups and model parameters of the initial model; Determining an optimization algorithm for the initial model according to a loss function corresponding to the initial model; Determining a training iteration mechanism for the initial model according to the number of training set groups and the model parameters; Constructing a training feedback mechanism for the initial model, and determining the number of iterations of the training iteration mechanism according to the training feedback mechanism; A model training algorithm for the initial model is defined according to the optimization algorithm, the training iteration mechanism, the training feedback mechanism and the number of iterations.

9. The complex data processing method based on a large model as claimed in claim 1, characterized in that: The method of using a preset validation set to validate the model performance of the training model includes: Using the verification set to verify the training model to obtain verification data; Calculate the accuracy and recall of the training model according to the verification data; Determining an F1 score of the training model according to the precision and the recall; Analyzing the confusion matrix and analyzing errors of the training model; Based on the confusion matrix, the analysis error, and the F1 score, a model performance of the training model is determined.

10. A complex data processing system based on a large model, characterized in that: The system comprises: A data cleaning module is used to clarify the processing objectives and processing requirements of complex data, identify missing values, duplicate values ​​and abnormal values ​​of the complex data, and perform data cleaning on the complex data based on the missing values, duplicate values ​​and abnormal values ​​to obtain cleaned data; A feature extraction module, used to define a multimodal feature extraction algorithm for the cleaned data, extract multimodal features of the cleaned data based on the multimodal feature extraction algorithm, determine a model activation function for the cleaned data, and perform feature encoding on the cleaned data according to the model activation function to obtain encoded data; A model framework configuration module, for performing feature fusion on the encoded data according to the multimodal features to obtain fused data, defining an automatic configuration algorithm for the fused data according to the processing target and the processing requirement, and configuring a complex data analysis model framework for the fused data according to the automatic configuration algorithm; A model training module is used to determine the loss function of the complex data analysis model framework, construct an initial model of the fused data according to the loss function and the complex data analysis model framework, define a model training algorithm for the initial model, train the initial model according to the model training algorithm using a preset training set to obtain a training model, and verify the model performance of the training model using a preset validation set; The model analysis module is used to use the training model as a complex data analysis model for the fused data when the model performance meets the preset model performance standard, and use the complex data analysis model to analyze the fused data to obtain target data.