Intelligent fusion data analysis method and system based on LLM (Logistics Language Model)
By building training sets, data processing, noise reduction, classification and iterative training, combined with the LLM large language model, the problem of insufficient processing capabilities of multi-source fusion data is solved, efficient intelligent analysis is achieved, and the accuracy and reliability of data processing are improved.
Patent Information
- Application Number
- CN202511054128.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art lacks processing capabilities when processing multi-source converged data, and has limitations, making it difficult to achieve efficient intelligent analysis.
By constructing a training set, data processing, noise reduction, classification and iterative training are carried out, and combined with the LLM large language model for analysis and matching, we realize intelligent analysis of multi-source fusion data.
It improves the accuracy of intelligent analysis and data processing of multi-source fusion data, reduces noise interference, and enhances the accuracy of data classification.
Smart Images

Figure CN120561702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent analysis technology, and in particular to a method and system for intelligent analysis of fused data based on an LLM large language model. Background Art
[0002] Large language models demonstrate an extremely natural way of understanding and generating text, making them excellent in simulating human language behavior and having wide applicability. However, as the capabilities of large models improve, the ability to analyze and discern model-generated content needs to be improved.
[0003] The prior art, such as the invention patent application with announcement number: CN119441391A, discloses an intelligent analysis system for start-ups based on a large language model, and its method includes: a data collection module for collecting start-up data; a data preprocessing module for preprocessing start-up data; a vector database and knowledge graph construction module for constructing a vector database and a knowledge graph based on the preprocessed start-up data; a context information construction module for performing in-depth semantic analysis on the relevant information of the enterprise to be analyzed input by the user, and fusing the semantic analysis results with the data of the vector database and the knowledge graph to obtain context information; an intelligent analysis module for performing intelligent analysis of the context information based on the start-up research and analysis auxiliary large model, obtaining the development status data of the enterprise to be analyzed, and completing the intelligent analysis of start-ups based on the large language model.
[0004] From the above solutions, it can be seen that current data processing is often only targeted at a certain field, and has insufficient processing capabilities when facing multi-source fusion data, and has certain limitations. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for intelligent analysis of fusion data based on the LLM large language model, which solves the problems existing in the background technology.
[0006] To solve the above technical problems, the present invention adopts the following technical solution: The present invention provides a fusion data intelligent analysis method based on the LLM large language model, which specifically includes the following steps: S1. Collect historical multi-source fusion data in the network to build a training set, and process the historical multi-source fusion data in the training set through a data processing method to obtain processed historical multi-source fusion data; S2. De-noise the processed historical multi-source fusion data by using a data de-noising method to obtain de-noised historical multi-source fusion data; S3. Classify the denoised historical multi-source fusion data using a data classification method and output classification results; the classification results include: historical multi-source fusion data category, historical multi-source fusion data center, and classified historical multi-source fusion data; S4. Training the classified historical multi-source fusion data through iterative training to obtain trained historical multi-source fusion data; S5. Collect user input data in real time, and analyze and match the collected user input data through the LLM large language model to obtain multi-source fusion data that matches the user input data; Preferably, the process of collecting historical multi-source fusion data in the network to construct a training set, and processing the historical multi-source fusion data in the training set by a data processing method to obtain the processed historical multi-source fusion data comprises the following steps: S11. Completing the historical multi-source fusion data in the training set by using a data completion algorithm to obtain completed historical multi-source fusion data; S12. The historical multi-source fusion data after aggregation and completion is obtained to obtain the processed historical multi-source fusion data.
[0007] Preferably, the method of completing the historical multi-source fusion data in the training set by using a data completion algorithm to obtain the completed historical multi-source fusion data includes the following steps: The network is divided into multiple sub-regions, and the historical multi-source fusion data collected by the j-th sub-region in the i-th collection period is expressed as , and the historical multi-source fusion data that has not been collected is represented as 0, and the real data is represented as ; Set an identification variable Mark whether the area has been collected; when When , it means that the i-th acquisition cycle of the j-th sub-region has been completed. When , it means that the i-th acquisition cycle of the j-th sub-region is not completed; Set the historical multi-source fusion data collected by the j-th sub-region in the i-th collection cycle to be equal to the corresponding region's real data × identification variable; Summarize the data collected from each sub-region and the corresponding identification variables after multiple rounds of collection cycles, and construct a matrix; set the matrix Represents all collected historical multi-source fusion data, matrix represents the aggregated real data, and the matrix C represents all the collected identification variables; Setting Matrix =Matrix Matrix C, where represents the element-wise product; The summarized historical multi-source fusion data is completed through the completion algorithm; When the matrix When there is no missing, the matrix Can decompose two matrices; Setting Matrix The approximate matrix after the uncompleted operation can be minimized by the gradient descent method With the matrix gap.
[0008] Preferably, the denoising of the processed historical multi-source fusion data by a data denoising method to obtain the denoised historical multi-source fusion data comprises the following steps: S21, decomposing the data using a wavelet basis function selected based on the data characteristics to obtain wavelet decomposition coefficients corresponding to each historical multi-source fusion data; S22, estimating the noise in each historical multi-source fusion data and setting a threshold; S23, filtering the wavelet coefficients of each historical multi-source fusion data obtained in S21 through thresholds and various threshold functions; S24, reconstructing the filtered historical multi-source fusion data in S23; The reconstructed historical multi-source fusion data is obtained by performing dot multiplication of the wavelet packet coefficients with the low-pass filter coefficients and the high-pass filter coefficients and then adding them together. The historical multi-source fusion data after reconstruction is set as the historical multi-source fusion data after noise reduction.
[0009] Preferably, estimating the noise in each historical multi-source fusion data and setting the threshold value comprises the following steps: The signal-to-noise ratio formula is as follows:
[0010] in, represents the signal-to-noise ratio, represents the historical multi-source fusion data after wavelet decomposition, Represents the first set of historical multi-source fusion data, Indicates the Group historical multi-source fusion data, Represents historical multi-source fusion data without wavelet decomposition; Summarize the signal-to-noise ratio calculated from each set of historical multi-source fusion data, and use the mean algorithm to estimate the noise in each sub-region within each period; set the signal-to-noise ratio threshold, and compare the estimated noise with the set signal-to-noise ratio threshold. When the estimated noise is greater than the set signal-to-noise ratio threshold, it indicates that the noise interference is weak, otherwise it is strong. Set the signal-to-noise ratio threshold to threshold.
[0011] Preferably, the step of classifying the denoised historical multi-source fusion data by a data classification method and outputting the classification results comprises the following steps: S31, summarizing the historical multi-source fusion data after noise reduction to construct a sample set, setting each sample to represent a set of historical multi-source fusion data after noise reduction; randomly selecting R samples, and constructing a decision tree based on the selected R samples, setting the selected R samples as the samples at the root node of the decision tree; S32. When each sample has w attributes, randomly select one attribute from these w attributes as the classification attribute of the node; Set each classification attribute to be selected only once, and each classification will only generate two child nodes; The newly generated child node is judged. If the data in the newly generated child node is empty, it means that the classification fails. If the data in the newly generated child node is not empty and the newly generated child node is an intermediate node, the class represented by the intermediate node is set to the class with the most categories in the training data. S33, iteratively classifying the divided child nodes using step S32 until the samples in the child nodes cannot be classified any more, and constructing a decision tree based on the classified nodes; S34. Summarize each sub-node and node classification attribute in the decision tree to obtain classified historical multi-source fusion data.
[0012] Preferably, the training of the classified historical multi-source fusion data by iterative training to obtain the trained historical multi-source fusion data comprises the following steps: Convolutional neural networks include: convolutional layers, pooling layers, and fully connected layers; S41, inputting the classified historical multi-source fusion data into a convolutional neural network, and performing feature extraction through the convolution layer in the convolutional neural network; S42, after the convolution layer extracts features from the historical multi-source fusion data, the extracted features are processed through the pooling layer, and the features are downsampled through the pooling layer to reduce the data dimension; S43, extract features from historical multi-source fusion data through continuous convolution and pooling until the extracted features converge, the convolution stops, and the extracted features are summarized and input to the fully connected layer; and the fully connected layer performs mapping output; Assume that the output historical multi-source fusion data features are the historical multi-source fusion data after training.
[0013] Preferably, the step of inputting the classified historical multi-source fusion data into a convolutional neural network and performing feature extraction through a convolutional layer in the convolutional neural network comprises the following steps: First, the input classified historical multi-source fusion data is linearly mapped through the fully connected layer. After the mapping is completed, the data encoding is recorded and the attention weight is set according to the data encoding. Set the convolution kernel size, convolution kernel step size, pooling layer size and initial attention weight, and perform convolution operation on the input historical multi-source fusion data based on the set parameters, and complete the feature extraction of the input historical multi-source fusion data through the convolution operation.
[0014] Preferably, the real-time collection of user input data and the analysis and matching of the real-time collected user input data by the LLM large language model to obtain multi-source fusion data matching the user input data include the following steps: Collecting user input data in real time, and training the user input data collected in real time based on step S4 to obtain trained user input data; The trained user input data and the trained historical multi-source fusion data are analyzed and matched through the LLM large language model; The LLM large language model analysis and matching formula is: ;
[0015] Among them, KT is the user input data after training, KN is the historical multi-source fusion data after training, Represents the analysis and matching results of the trained user input data and the trained historical multi-source fusion data; Set a threshold for the analysis matching result. When the analysis matching result between the trained user input data and the trained historical multi-source fusion data is greater than or equal to the threshold, the user input data is determined to be related to the trained historical multi-source fusion data. When the analysis matching result between the trained user input data and the trained historical multi-source fusion data is less than the threshold, the user input data is determined to be unrelated to the trained historical multi-source fusion data. The type of the user input data is determined based on the analysis matching results.
[0016] The present invention also discloses a fusion data intelligent analysis system based on the LLM large language model, which is used to implement a fusion data intelligent analysis method based on the LLM large language model. The system includes: a data collection module, a data processing module, a data noise reduction module, a data classification module, a data training module and a data analysis and matching module; The data collection module is used to collect historical multi-source fusion data and user input data in the network; The data processing module is used to process the collected historical multi-source fusion data; The data denoising module is used to denoise the processed historical multi-source fusion data; The data classification module is used to classify the historical multi-source fusion data after noise reduction; The data training module is used to train various types of classified historical multi-source fusion data; The data analysis and matching module is used to analyze and match user input data and historical multi-source fusion data.
[0017] The beneficial effects of the present invention are: The present invention constructs a training set by collecting historical multi-source fusion data in the network, processes the historical multi-source fusion data in the training set by a data processing method, and reduces the noise of the processed historical multi-source fusion data by a data denoising method; for the historical multi-source fusion data after noise reduction, the historical multi-source fusion data after noise reduction is classified by a data classification method, and the classified historical multi-source fusion data is trained by an iterative training method to obtain trained historical multi-source fusion data, and finally collects user input data in real time, and analyzes and matches the user input data collected in real time by an LLM large language model to obtain multi-source fusion data matching the user input data, thereby improving the accuracy of intelligent analysis of fusion data.
[0018] The present invention divides the network into multiple sub-areas and collects the fused data in each sub-area through multiple rounds of collection cycles; at the same time, the data missing in the collection process is supplemented by a data completion algorithm, thereby improving the reliability of data collection.
[0019] The present invention decomposes the data by using wavelet basis functions selected according to data features, estimates the noise in each historical multi-source fusion data, sets a threshold, and finally filters the data through a threshold function based on the set threshold, thereby reducing the interference of noise on the collected historical multi-source fusion data.
[0020] The present invention constructs a sample set by summarizing the historical multi-source fusion data after noise reduction, and completes the classification of the historical multi-source fusion data after noise reduction by randomly selecting samples to construct a decision tree, thereby improving the accuracy of data processing.
[0021] The present invention trains the classified historical multi-source fusion data by combining convolutional neural networks with attention mechanisms, thereby improving the analysis accuracy of historical multi-source fusion data. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 This is a flow chart of the fusion data intelligent analysis method of the present invention. DETAILED DESCRIPTION
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0025] In a specific embodiment of the present invention, Reference Figure 1 As shown, the present invention provides a fusion data intelligent analysis method based on the LLM large language model, comprising the following steps: S1. Collect historical multi-source fusion data in the network to build a training set, and process the historical multi-source fusion data in the training set through a data processing method to obtain processed historical multi-source fusion data; S2. De-noise the processed historical multi-source fusion data by using a data de-noising method to obtain de-noised historical multi-source fusion data; S3. Classify the denoised historical multi-source fusion data using a data classification method and output classification results; the classification results include: historical multi-source fusion data category, historical multi-source fusion data center, and classified historical multi-source fusion data; S4. Training the classified historical multi-source fusion data through iterative training to obtain trained historical multi-source fusion data; S5. Collect user input data in real time, and analyze and match the collected user input data through the LLM large language model to obtain multi-source fusion data that matches the user input data; Further, refer to Figure 1 As shown, the historical multi-source fusion data in the network is collected to construct a training set, and the historical multi-source fusion data in the training set is processed by a data processing method to obtain the processed historical multi-source fusion data, which includes the following steps: S11. Completing the historical multi-source fusion data in the training set by using a data completion algorithm to obtain the completed historical multi-source fusion data, and constructing a training set based on the completed historical multi-source fusion data; The network is divided into multiple sub-regions, and the historical multi-source fusion data collected by the j-th sub-region in the i-th collection period is expressed as , and the historical multi-source fusion data that has not been collected is represented as 0, and the real data is represented as ; Furthermore, set an identification variable Mark whether the area has been collected; when When , it means that the i-th acquisition cycle of the j-th sub-region has been completed. When , it means that the i-th acquisition cycle of the j-th sub-region is not completed; Furthermore, the historical multi-source fusion data collected by the j-th sub-region in the i-th collection period is set to be equal to the corresponding region's real data × the identification variable; Furthermore, the data collected from each sub-region after multiple rounds of collection cycles and the corresponding identification variables are summarized and a matrix is constructed; the matrix is set Represents all collected historical multi-source fusion data, matrix represents the aggregated real data, and the matrix C represents all the collected identification variables; Setting Matrix =Matrix Matrix C, where represents the element-wise product; Furthermore, the summarized historical multi-source fusion data is supplemented through the completion algorithm; When the matrix When there is no missing, the matrix Can decompose two matrices; Setting Matrix The approximate matrix after the uncompleted operation can be minimized by the gradient descent method With the matrix the gap; The completion process formula is as follows: ; in, represents the completion loss function, represents the completed approximate matrix, For the two matrices after decomposition, r represents the rth column after decomposition, and k is smaller than the rows and columns of the matrix before decomposition; S12, summarizing and completing the historical multi-source fusion data to obtain processed historical multi-source fusion data; Further, refer to Figure 1As shown, the noise reduction method is used to reduce the noise of the processed historical multi-source fusion data. The noise reduction method is used to obtain the historical multi-source fusion data. The following steps are included: S21, decomposing the data using a wavelet basis function selected based on the data characteristics to obtain wavelet decomposition coefficients corresponding to each historical multi-source fusion data;
[0026] in, For historical multi-source fusion data, is a low-pass filter, is the filter coefficient, is the wavelet coefficient; S22, estimating the noise in each historical multi-source fusion data and setting a threshold; The signal-to-noise ratio formula is as follows:
[0027] in, represents the signal-to-noise ratio, represents the historical multi-source fusion data after wavelet decomposition, Represents the first set of historical multi-source fusion data, Indicates the Group historical multi-source fusion data, Represents historical multi-source fusion data without wavelet decomposition; Summarize the signal-to-noise ratio calculated from each set of historical multi-source fusion data, and use the mean algorithm to estimate the noise in each sub-region within each period; set the signal-to-noise ratio threshold, and compare the estimated noise with the set signal-to-noise ratio threshold. When the estimated noise is greater than the set signal-to-noise ratio threshold, it indicates that the noise interference is weak, otherwise it is strong. Furthermore, the signal-to-noise ratio threshold is set as the threshold value; S23, filtering the wavelet coefficients of each historical multi-source fusion data obtained in S21 through thresholds and various threshold functions; Threshold function:
[0028] Among them, λ is the set signal-to-noise ratio threshold, sign is the sign function, It is the filtered historical multi-source fusion data; S24, reconstructing the filtered historical multi-source fusion data in S23; The reconstructed historical multi-source fusion data is obtained by performing dot multiplication of the wavelet packet coefficients with the low-pass filter coefficients and the high-pass filter coefficients and then adding them together. The historical multi-source fusion data after reconstruction is set as the historical multi-source fusion data after noise reduction; Further, refer to Figure 1 As shown, for the historical multi-source fusion data after denoising, classifying the historical multi-source fusion data after denoising by data classification and outputting the classification results includes the following steps: S31, summarizing the historical multi-source fusion data after noise reduction to construct a sample set, setting each sample to represent a set of historical multi-source fusion data after noise reduction; randomly selecting R samples, and constructing a decision tree based on the selected R samples, setting the selected R samples as the samples at the root node of the decision tree; S32. When each sample has w attributes, randomly select one attribute from these w attributes as the classification attribute of the node; Set each classification attribute to be selected only once, and each classification will only generate two child nodes; Furthermore, the newly generated child node is judged. If the data in the newly generated child node is empty, it indicates that the classification has failed. If the data in the newly generated child node is not empty and the newly generated child node is an intermediate node, the class represented by the intermediate node is set to the class with the most categories in the training data. S33, iteratively classifying the divided child nodes using step S32 until the samples in the child nodes cannot be classified any more, and constructing a decision tree based on the classified nodes; S34, summarizing each sub-node and node classification attribute in the decision tree to obtain classified historical multi-source fusion data; Further, refer to Figure 1 As shown, the classified historical multi-source fusion data is trained by iterative training to obtain the trained historical multi-source fusion data, which includes the following steps: Convolutional neural networks include: convolutional layers, pooling layers, and fully connected layers; S41, inputting the classified historical multi-source fusion data into a convolutional neural network, and performing feature extraction through the convolution layer in the convolutional neural network; First, the input classified historical multi-source fusion data is linearly mapped through the fully connected layer. After the mapping is completed, the data encoding is recorded and the attention weight is set according to the data encoding. The attention weight formula is as follows: ; in, , , , is a trainable parameter matrix; express Dimensions, represents the attention weight, represents the activation function, Z represents the data encoding sequence; Set the convolution kernel size, convolution kernel step size, pooling layer size, and initial attention weight, and perform a convolution operation on the input historical multi-source fusion data based on the set parameters to complete feature extraction of the input historical multi-source fusion data through the convolution operation; The convolution operation formula is as follows:
[0029] in, Represents the input historical multi-source fusion data, represents the weight of the corresponding convolution kernel, b represents the bias value, represents the output features; S42, after the convolution layer extracts features from the historical multi-source fusion data, the extracted features are processed through the pooling layer, and the features are downsampled through the pooling layer to reduce the data dimension; S43, extract features from historical multi-source fusion data through continuous convolution and pooling until the extracted features converge, the convolution stops, and the extracted features are summarized and input to the fully connected layer; and the fully connected layer performs mapping output; Assume that the output historical multi-source fusion data features are the historical multi-source fusion data after training; Further, refer to Figure 1 As shown, real-time collection of user input data and analysis and matching of the real-time collected user input data by the LLM large language model to obtain multi-source fusion data that matches the user input data include the following steps: Collecting user input data in real time, and training the user input data collected in real time based on step S4 to obtain trained user input data; Furthermore, the trained user input data and the trained historical multi-source fusion data are analyzed and matched through the LLM large language model; The LLM large language model analysis and matching formula is: ; Among them, KT is the user input data after training, KN is the historical multi-source fusion data after training, Represents the analysis and matching results of the trained user input data and the trained historical multi-source fusion data; Set a threshold for the analysis matching result. When the analysis matching result between the trained user input data and the trained historical multi-source fusion data is greater than or equal to the threshold, the user input data is determined to be related to the trained historical multi-source fusion data. When the analysis matching result between the trained user input data and the trained historical multi-source fusion data is less than the threshold, the user input data is determined to be unrelated to the trained historical multi-source fusion data. Further, determining the type of the user input data based on the analysis and matching results; In a specific embodiment, the fusion data intelligent analysis system based on the LLM large language model is used to implement a fusion data intelligent analysis method based on the LLM large language model, and the system includes: a data collection module, a data processing module, a data noise reduction module, a data classification module, a data training module, and a data analysis and matching module; The data collection module is used to collect historical multi-source fusion data and user input data in the network; The data processing module is used to process the collected historical multi-source fusion data; The data denoising module is used to denoise the processed historical multi-source fusion data; The data classification module is used to classify the historical multi-source fusion data after noise reduction; The data training module is used to train various types of classified historical multi-source fusion data; The data analysis and matching module is used to analyze and match user input data and historical multi-source fusion data.
[0030] It should be noted that The above contents are merely examples and explanations of the concept of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in a similar manner. As long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, they should all fall within the scope of protection of the present invention.
Claims
1. A fusion data intelligent analysis method based on LLM large language model, characterized by: The following steps are involved: S1. Collect historical multi-source fusion data in the network to build a training set, and process the historical multi-source fusion data in the training set through a data processing method to obtain processed historical multi-source fusion data; S2. De-noise the processed historical multi-source fusion data by using a data de-noising method to obtain de-noised historical multi-source fusion data; S3. Classify the denoised historical multi-source fusion data using a data classification method and output the classification results. The classification results include: historical multi-source fusion data categories, historical multi-source fusion data centers, and classified historical multi-source fusion data; S4. Training the classified historical multi-source fusion data through iterative training to obtain trained historical multi-source fusion data; S5. Collect user input data in real time, and analyze and match the real-time collected user input data through the LLM large language model to obtain multi-source fusion data that matches the user input data.
2. The method for intelligent analysis of fusion data based on the LLM large language model according to claim 1 is characterized in that: The historical multi-source fusion data in the acquisition network is used to construct a training set, and the historical multi-source fusion data in the training set is processed by a data processing method to obtain the processed historical multi-source fusion data, which includes the following steps: S11. Completing the historical multi-source fusion data in the training set by using a data completion algorithm to obtain completed historical multi-source fusion data; S12. The historical multi-source fusion data after aggregation and completion is obtained to obtain the processed historical multi-source fusion data.
3. The method for intelligent analysis of fusion data based on the LLM large language model according to claim 2 is characterized in that: The method of completing the historical multi-source fusion data in the training set by using the data completion algorithm to obtain the completed historical multi-source fusion data includes the following steps: The network is divided into multiple sub-regions, and the historical multi-source fusion data collected by the j-th sub-region in the i-th collection period is expressed as , and the historical multi-source fusion data that has not been collected is represented as 0, and the real data is represented as ; Set an identification variable Mark whether the area has been collected; when When , it means that the i-th acquisition cycle of the j-th sub-region has been completed. When , it means that the i-th acquisition cycle of the j-th sub-region is not completed; Set the historical multi-source fusion data collected by the j-th sub-region in the i-th collection cycle to be equal to the corresponding region's real data × identification variable; Summarize the data collected from each sub-region and the corresponding identification variables after multiple rounds of collection cycles, and construct a matrix; set the matrix Represents all collected historical multi-source fusion data, matrix represents the aggregated real data, and the matrix C represents all the collected identification variables; Setting Matrix =Matrix Matrix C, where represents the element-wise product; The summarized historical multi-source fusion data is completed through the completion algorithm; When the matrix When there is no missing, the matrix Can decompose two matrices; Setting Matrix The approximate matrix after the uncompleted operation can be minimized by the gradient descent method With the matrix gap.
4. The method for intelligent analysis of fusion data based on the LLM large language model according to claim 1 is characterized in that: The denoising of the processed historical multi-source fusion data by using a data denoising method to obtain the denoised historical multi-source fusion data comprises the following steps: S21, decomposing the data using a wavelet basis function selected based on the data characteristics to obtain wavelet decomposition coefficients corresponding to each historical multi-source fusion data; S22, estimating the noise in each historical multi-source fusion data and setting a threshold; S23, filtering the wavelet coefficients of each historical multi-source fusion data obtained in S21 through thresholds and various threshold functions; S24, reconstructing the filtered historical multi-source fusion data in S23; The reconstructed historical multi-source fusion data is obtained by performing dot multiplication of the wavelet packet coefficients with the low-pass filter coefficients and the high-pass filter coefficients and then adding them together. The historical multi-source fusion data after reconstruction is set as the historical multi-source fusion data after noise reduction.
5. The method for intelligent analysis of fusion data based on the LLM large language model according to claim 4 is characterized in that: The process of estimating the noise in each historical multi-source fusion data and setting the threshold value includes the following steps: Summarize the signal-to-noise ratio calculated from each set of historical multi-source fusion data, and use the mean algorithm to estimate the noise in each sub-region within each period; set the signal-to-noise ratio threshold, and compare the estimated noise with the set signal-to-noise ratio threshold. When the estimated noise is greater than the set signal-to-noise ratio threshold, it indicates that the noise interference is weak, otherwise it is strong. Set the signal-to-noise ratio threshold to threshold.
6. The method for intelligent analysis of fusion data based on the LLM large language model according to claim 1, characterized in that: The method of classifying the denoised historical multi-source fusion data by a data classification method and outputting the classification results includes the following steps: S31, summarizing the historical multi-source fusion data after noise reduction to construct a sample set, setting each sample to represent a set of historical multi-source fusion data after noise reduction; randomly selecting R samples, and constructing a decision tree based on the selected R samples, setting the selected R samples as the samples at the root node of the decision tree; S32. When each sample has w attributes, randomly select one attribute from these w attributes as the classification attribute of the node; Set each classification attribute to be selected only once, and each classification will only generate two child nodes; The newly generated child node is judged. If the data in the newly generated child node is empty, it means that the classification fails. If the data in the newly generated child node is not empty and the newly generated child node is an intermediate node, the class represented by the intermediate node is set to the class with the most categories in the training data. S33, iteratively classifying the divided child nodes using step S32 until the samples in the child nodes cannot be classified any more, and constructing a decision tree based on the classified nodes; S34. Summarize each sub-node and node classification attribute in the decision tree to obtain classified historical multi-source fusion data.
7. The method for intelligent analysis of fusion data based on the LLM large language model according to claim 1 is characterized in that: The training of the classified historical multi-source fusion data by iterative training to obtain the trained historical multi-source fusion data includes the following steps: Convolutional neural networks include: convolutional layers, pooling layers, and fully connected layers; S41, inputting the classified historical multi-source fusion data into a convolutional neural network, and performing feature extraction through the convolution layer in the convolutional neural network; S42, after the convolution layer extracts features from the historical multi-source fusion data, the extracted features are processed through the pooling layer, and the features are downsampled through the pooling layer to reduce the data dimension; S43, extract features from historical multi-source fusion data through continuous convolution and pooling until the extracted features converge, the convolution stops, and the extracted features are summarized and input to the fully connected layer; and the fully connected layer performs mapping output; Assume that the output historical multi-source fusion data features are the historical multi-source fusion data after training.
8. The method for intelligent analysis of fusion data based on the LLM large language model according to claim 7 is characterized in that: The process of inputting the classified historical multi-source fusion data into a convolutional neural network and performing feature extraction through the convolutional layer in the convolutional neural network includes the following steps: First, the input classified historical multi-source fusion data is linearly mapped through the fully connected layer. After the mapping is completed, the data encoding is recorded and the attention weight is set according to the data encoding. Set the convolution kernel size, convolution kernel step size, pooling layer size and initial attention weight, and perform convolution operation on the input historical multi-source fusion data based on the set parameters, and complete the feature extraction of the input historical multi-source fusion data through the convolution operation.
9. The method for intelligent analysis of fusion data based on the LLM large language model according to claim 1, characterized in that: The method of collecting user input data in real time and analyzing and matching the user input data collected in real time through the LLM large language model to obtain multi-source fusion data that matches the user input data includes the following steps: Collecting user input data in real time, and training the user input data collected in real time based on step S4 to obtain trained user input data; The trained user input data and the trained historical multi-source fusion data are analyzed and matched through the LLM large language model; Set a threshold for the analysis matching result. When the analysis matching result between the trained user input data and the trained historical multi-source fusion data is greater than or equal to the threshold, the user input data is determined to be related to the trained historical multi-source fusion data. When the analysis matching result between the trained user input data and the trained historical multi-source fusion data is less than the threshold, the user input data is determined to be unrelated to the trained historical multi-source fusion data. The type of the user input data is determined based on the analysis matching results.
10. A system for implementing the fusion data intelligent analysis method based on the LLM large language model according to any one of claims 1 to 9, characterized in that: The system includes: data collection module, data processing module, data noise reduction module, data classification module, data training module and data analysis and matching module; The data collection module is used to collect historical multi-source fusion data and user input data in the network; The data processing module is used to process the collected historical multi-source fusion data; The data denoising module is used to denoise the processed historical multi-source fusion data; The data classification module is used to classify the historical multi-source fusion data after noise reduction; The data training module is used to train various types of classified historical multi-source fusion data; The data analysis and matching module is used to analyze and match user input data and historical multi-source fusion data.
Citation Information
Patent Citations
Image processing resource allocation method and storage medium
CN116739881A
Sleep classification method, system and equipment based on feature fusion and medium
CN117414107A
Hydroelectric equipment on-line monitoring and diagnosis system
CN119179919A
Method for complementing unknown measurement information of power distribution network based on data flow mapping
CN119474693A