Multi-source data fusion processing method, processing device, processing platform, medium and product
By calling task management services and function scheduling services in the data analysis system, selecting matching models from multiple candidate neural network models to process multi-source data fusion tasks, solving the problems of complex operation and poor maintenance of existing systems, and achieving the effects of simple user operations and high system maintenance.
Patent Information
- Application Number
- CN202411823943.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-06
AI Technical Summary
Existing data analysis systems have problems of complex operations or poor maintenance, especially in multi-source data fusion and machine learning model training.
By calling task management services and functional scheduling services, the target neural network model matching the user's request is selected from multiple candidate neural network models to handle the data analysis task of multi-source data fusion.
It realizes a big data analysis system with simple user operations and high maintenance, and can combine the user's required models from multiple candidate neural network models to handle target tasks, improving the scalability and maintainability of the system.
Smart Images

Figure CN119938256A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a multi-source data fusion processing method, processing device, processing platform, medium and product. Background Art
[0002] Multi-source data fusion refers to the technology of using relevant means to integrate all the information obtained from investigation and analysis, and to conduct a unified evaluation of the information to finally obtain unified information.
[0003] With the advancement of computer information technology, the generation and storage of data have become increasingly important. Big data analysis can help users process and analyze large amounts of heterogeneous data, obtain valuable results, and promote the rapid development of science and technology, production, and society. However, the analysis of big data requires a certain threshold, especially the use of machine learning models for data analysis, which involves professional knowledge in areas such as data preprocessing and model training. At the same time, it also takes a certain amount of time and effort to integrate, analyze, and store a large amount of different types of data.
[0004] Currently, two common methods are to use the Application Programming Interface (API) provided by various programming languages, quickly build data preprocessing and machine learning models by calling the algorithms provided in the API, and then store the prediction data after training. However, this method requires a certain computer and advanced mathematics foundation, which is difficult for users who do not understand computer technology.
[0005] Another method is to use an existing data analysis system. Users only need to upload data to obtain the results of model analysis, which is relatively simple to operate. This method processes the data uploaded by users by calling a trained conventional single baseline model in the background and returns the calculation results to the user. The model scalability of this method is not strong, and only conventional baseline models can be used.
[0006] It can be seen that the current data analysis system has problems of complex operation or poor maintainability. Summary of the invention
[0007] The present application provides a multi-source data fusion processing method, processing device, processing platform, medium and product to solve the problem of complex operation or poor maintainability of data analysis systems in related technologies. The present application can combine the model required by the user from multiple candidate neural network models to process the target task. The user operation is simple and has high maintainability.
[0008] The technical solution of this application is implemented as follows:
[0009] A multi-source data fusion processing method, the method comprising:
[0010] Invoke the task management service to generate task configuration information in response to the acquired user request;
[0011] Calling the task management service to create a target task according to the task configuration information;
[0012] The function scheduling service is called to select a target neural network model matching the task configuration information from a plurality of candidate neural network models to process the target task; the target neural network model is used to perform data analysis of multi-source data fusion.
[0013] A multi-source data fusion processing device, comprising:
[0014] A first processing module, configured to call a task management service to generate task configuration information in response to an acquired user request;
[0015] The first processing module is used to call the task management service to create a target task according to the task configuration information;
[0016] The second processing module is used to call the function scheduling service to select a target neural network model that matches the task configuration information from multiple candidate neural network models to process the target task; the target neural network model is used to perform data analysis of multi-source data fusion.
[0017] A multi-source data fusion processing platform comprises: a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the steps of the multi-source data fusion processing method as described above.
[0018] A storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the multi-source data fusion processing method as described above.
[0019] A computer product includes a computer program, and when the computer program is executed by a processor, the steps of the multi-source data fusion processing method as described above are implemented.
[0020] The present application provides a multi-source data fusion processing method, which generates task configuration information by calling a task management service in response to an acquired user request; calling a task management service to create a target task according to the task configuration information; calling a function scheduling service to select a target neural network model that matches the task configuration information from multiple candidate neural network models to process the target task; the target neural network model is used to perform data analysis of multi-source data fusion. In this way, the problem of complex operation or poor maintainability of the data analysis system in the related art is solved. The present application can combine the model required by the user from multiple candidate neural network models to process the target task, and the user operation is simple and has a high degree of maintainability. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A schematic diagram of a multi-source data fusion processing method provided in an embodiment of the present application Figure 1 ;
[0022] Figure 2 A schematic diagram of a multi-source data fusion processing method provided in an embodiment of the present application Figure 2 ;
[0023] Figure 3 A schematic diagram of the structure of a multi-source data fusion platform provided in an embodiment of the present application;
[0024] Figure 4 A schematic diagram of a scenario flow of multi-source data fusion provided in an embodiment of the present application;
[0025] Figure 5 A schematic diagram of the structure of a multi-source data fusion processing device provided in an embodiment of the present application;
[0026] Figure 6 A schematic diagram of the structure of a multi-source data fusion processing platform provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0028] It should be understood that the "embodiments of the present application" or "the aforementioned embodiments" mentioned throughout the specification mean that specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, "in the embodiments of the present application" or "in the aforementioned embodiments" appearing throughout the specification may not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. In the various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0029] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present application are merely to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0030] Reference Figure 1 As shown, the multi-source data fusion processing method provided by the present application includes the following steps:
[0031] Step 101: Invoke a task management service in response to an acquired user request to generate task configuration information.
[0032] In actual application, the execution subject of this embodiment can be a multi-source data fusion processing platform, which has functions such as multi-source data fusion processing, data communication and program operation. Generally, the operation of each component in the multi-source data fusion processing platform can be driven by a core controller / processor, so the execution subject of this embodiment can also be a controller / processor in the multi-source data fusion processing platform.
[0033] In actual application, the multi-source data fusion processing platform provides corresponding services to users in the form of microservices to make users use more conveniently and quickly, thereby strengthening the construction of the data mining service system. The microservices provided by the multi-source data fusion processing platform include at least two services: task management service and function scheduling service; the microservices provided by the multi-source data fusion processing platform can also include the following five services: data import service, data analysis service, data visualization service, machine learning service, and data storage service. Among them, the task management service provides the management function of task information. The function scheduling service connects and manages different tasks in the form of workflows, and schedules and executes functions (such as the functions supported by the above five services).
[0034] In actual application, the user request may be a new task request created by the user using a client / terminal, etc. The user request is a request initiated based on user needs in a multi-source data fusion scenario. Task configuration information is used to create a target task.
[0035] Step 102: Call the task management service to create a target task according to the task configuration information.
[0036] In actual application, after the multi-source data fusion processing platform calls the task management service to obtain the user request, it generates task configuration information in response to the obtained user request, and then creates the target task according to the task configuration information, that is, creates a new task.
[0037] Step 103, calling the function scheduling service to select a target neural network model that matches the task configuration information from multiple candidate neural network models to process the target task; the target neural network model is used to perform data analysis of multi-source data fusion.
[0038] In actual application, the task management service sends the user request to the function scheduling service, and the multi-source data fusion processing platform calls the function scheduling service to select the target neural network model that matches the task configuration information from multiple candidate neural network models to process the target task. In this way, the multi-source data fusion processing platform can combine the user's required model from multiple candidate neural network models based on the task configuration information of different target tasks to process the target task. In this way, the user does not need to have a certain computer and advanced data foundation, and the operation is simple; in addition, since the multi-source data fusion processing platform provides multiple candidate neural network models and supports the combination of user-required models to process target tasks, it achieves the effect of customized data processing according to the task configuration information corresponding to different user requests, and the flexible combination of multiple candidate neural network models according to user needs enhances the scalability and maintainability of the model.
[0039] The present application provides a multi-source data fusion processing method, which generates task configuration information by calling a task management service in response to an acquired user request; calling a task management service to create a target task according to the task configuration information; calling a function scheduling service to select a target neural network model that matches the task configuration information from multiple candidate neural network models to process the target task; the target neural network model is used to perform data analysis of multi-source data fusion. In this way, the problem of complex operation or poor maintainability of the data analysis system in the related art is solved. The present application can combine the model required by the user from multiple candidate neural network models to process the target task, and the user operation is simple and has a high degree of maintainability.
[0040] In some embodiments, step 101 calls the task management service in response to the obtained user request to generate task configuration information, which can be done by Figure 2 The steps shown achieve:
[0041] Step 201 : Calling a task management service in response to a user request, extracting different types of target files from different databases, and selecting target data information from the target files.
[0042] In actual application, different types of files stored in different databases are pre-stored files. When extraction is needed (for example, during model training), the required target data information is obtained from the database.
[0043] Step 202: Call the task management service to generate task configuration information according to the file type of the target file and the data type of the target data information.
[0044] In actual application, the task configuration information includes the file type of the target file and the data type of the target data information.
[0045] In some embodiments, step 201 calls the task management service in response to the user request to extract different types of target files from different databases. Before selecting the target data information from the target file, the following steps may also be performed:
[0046] First, the function scheduling service is called to receive the initial data uploaded by the user, and the initial data information is preprocessed to obtain the processed target data information.
[0047] In actual application, regarding the initial data, this application can obtain more comprehensive and accurate data from different data sources in the multi-source data fusion scenario, and can effectively improve the quality and accuracy of data through cross-validation and supplementation of multiple data sources. This helps enterprises make correct judgments based on more reliable data during decision-making and operation, and reduce the risk of wrong decisions.
[0048] In actual application, it takes a certain amount of time and effort to fuse, analyze and store a large number of different types of data. This application can call the function scheduling service to receive the initial data uploaded by the user, perform preprocessing operations on the initial data information, and obtain the processed target data information; that is, the user only needs to upload data regularly / irregularly, and the multi-source data fusion processing platform can call the function scheduling service to receive the initial data uploaded by the user (including all data uploaded by the user), and perform the following preprocessing operations on the initial data information: data grouping, data sorting, data intersection, data screening, data exclusion, data deduplication, data display, data renaming, data filling, numerical processing, column deletion and data sampling, etc.; thereby obtaining the processed target data information.
[0049] In actual application, the types of target data information include but are not limited to: text data information, image data information, and video data information.
[0050] Secondly, the function scheduling service is called to store the target data information in different databases according to different file types, generating different types of target files.
[0051] In actual application, the function scheduling service is called to schedule the data storage service, and the target data information is stored in different databases in a hierarchical manner according to different file types to generate different types of target files. In this way, different file types are stored in the corresponding databases through the data storage service. When extraction is needed (for example, during model training), the required data can be obtained from the database. The visualization service and machine learning service can also be scheduled to combine the user's required model for customized training.
[0052] In actual applications, different databases include but are not limited to key-value storage databases (Redis), relational databases (such as MySQL), and databases based on distributed file storage (such as MongoDB).
[0053] In some embodiments, step 103 calls a function scheduling service to select a target neural network model that matches the task configuration information from a plurality of candidate neural network models to process the target task, including:
[0054] First, call the function scheduling service to select multiple neural network models that match the task configuration information from multiple candidate neural network models.
[0055] In actual application, the function scheduling service is called to select multiple neural network models that match the above file type and the data type of the above target data information from multiple candidate neural network models to combine the model required by the user.
[0056] Secondly, the function scheduling service is called with the target data information as input data and the expected data information included in the user request as output data. Multiple neural network models are trained according to multi-dimensional evaluation indicators, and the training results are displayed on the web page interface; the training results include the names of multiple neural network models, training duration, and multi-dimensional evaluation indicators.
[0057] In actual application, after the user's required model is combined, the function scheduling service is called to train the combined model with the target data information as input data and the expected data information included in the user's request as output data. During the training process, multi-dimensional evaluation indicators are comprehensively considered for training. When all the models that can be combined have completed model training, the training ends and the training results are returned to the user on the front end for the user to choose.
[0058] Finally, the function scheduling service is called in response to the user's selection operation of the target neural network model from the multiple neural network models displayed, and the target neural network model is selected to process the target task.
[0059] In actual application, when the model accuracy reaches the user's expectations, the user downloads the trained model file stored in the system through the export function, uses the target neural network model to process the target task, obtains the predicted data and stores it.
[0060] In some embodiments, the multi-dimensional evaluation index includes: the accuracy confidence consistency of each model training in multiple neural network models, the accuracy of each model training, and the average precision mean of each model training.
[0061] In actual application, users can use the machine learning service to perform at least one model training operation on the algorithms recommended for all files and data types, and comprehensively consider the following multi-dimensional evaluation indicators to select the best recommended model for combined training: the accuracy confidence consistency of each model training in multiple neural network models (set its weight to 0.4), the accuracy of each model training (set its weight to 0.4), and the average precision mean of each model training (set its weight to 0.2).
[0062] In actual application, the weighted values of the above multi-dimensional factors are calculated, and the first two models with the highest final weights and that can be combined are selected in turn as the next model to be trained. If the highest weights of the models are equal and there are more than two, the two models that consume the least time during the model training process are selected. After each model training is completed, the training model name, training time, and multiple evaluation indicators are stored in the database. When all the models that can be combined have completed the model training, the training ends and the training results are returned to the user on the front end. When the model accuracy meets the user's expectations, the user can download the trained model file stored in the system through the export function.
[0063] In some embodiments, the multiple candidate neural network models include: a BERT-based text classification model, a lightweight neural network model corresponding to image data, and a time offset module attention mechanism model corresponding to video data.
[0064] In actual application, for text data information, the multi-source data fusion processing platform will determine the text category field based on text similarity, match the trained BERT-based text classification model such as CategoryBERT model belonging to the category, and then automatically match the long short-term memory network (LSTM) and bidirectional long short-term memory network (BiLSTM) model. Among them, CategoryBERT is a domain-specific language model based on the BERT model, which is pre-trained according to the data category submitted by the user. According to the category to which the user-submitted data belongs, the MLM pre-training of the stored data set of the category is first performed, so that the model learns to recognize and understand the text content, context and category-related semantic information in the category. CategoryBERT will introduce a pre-segmentation step in the prediction stage. The pre-segmentation is to use the segmentation dictionary extracted by the category training, and process according to the frequently appearing words in the category, so that the model pays more attention to and understands the language features related to the category, so as to improve the pre-training effect and performance of the model on the category data.
[0065] In actual applications, for image data information, the multi-source data fusion processing platform will match lightweight neural network models such as the MobileNet-SE model based on the file suffix type, png, jpg, etc., and then automatically match convolutional neural network models (such as LeNet, AlexNet, VGG, ResNet and Inception). The MobileNet-SE model adds the SE module to MobileNet, which is used to adaptively adjust the weight of each channel. The SE module consists of two main steps: Squeeze and Excitation. In the Squeeze stage, the feature map is compressed using a global average pooling operation, and the features of each channel are compressed into a scalar value to obtain the global statistical information of the features of each channel. The formula is as follows:
[0066]
[0067] Among them, Z i represents the feature value on the i-th channel, obtained by the global average pooling operation. The global average pooling operation compresses the feature map of each channel into a scalar value, that is, averages all positions of each channel. j,k,i Represents the feature value of the i-th channel at position (j, k) of the input feature map, j and k represent the row and column indexes in the feature map respectively. Among them, H represents the height of the feature map, W represents the width of the feature map, It means averaging the summation results to obtain the global eigenvalue of the i-th channel.
[0068] The Excitation phase adjusts the weight of each channel by actively learning the interaction relationship of each channel. The SE module is added after each depth-separable convolution of MobileNet to enhance the network's expressiveness on different channels. In this way, during the training process, the network can automatically learn the importance of each channel and adjust the corresponding weight according to its importance. The formula is as follows:
[0069] s = σ(FC 2(ReLU(FC 1(z))))
[0070] Among them, FC1 and FC2 represent two FC networks respectively, σ represents the Sigmoid activation function, and s is a vector of dimension C, in which each element represents the feature significance of the corresponding channel.
[0071] In actual application, for video data information, the multi-source data fusion processing platform will match the time offset module attention mechanism model such as the TSM-Attention model according to the file type mp4. TSM is a module for feature transformation in the time dimension. It introduces the perturbation of the timing information by shifting the input feature map in time with different amplitudes, which helps to extract richer time features. In order to further enhance the model's ability to pay attention to key frames in video sequences, the TSM-Attention of this application introduces an attention mechanism to dynamically adjust the weight of each time step. Specifically, it uses a self-attention mechanism to calculate the attention weight by learning the relationship between each time step. The formula is as follows:
[0072]
[0073] in, represents the attention weight between time steps t and t', Y t represents the output features after weighted summation of the attention mechanism, The feature map at time step t' is a multidimensional array obtained by extracting features from the input video frame using models such as convolutional neural networks (CNN). The three-dimensional convolutional neural network (3D-CNN), FasterR-CNN, YOLO and other models are automatically matched later. This application reduces the temporal grouping, and the temporal transformation module in the TSM model divides the input video into multiple groups, performs temporal transformations in these groups, and introduces an attention mechanism to focus on the content of different areas in the video, thereby improving classification accuracy and generalization ability.
[0074] The multi-source data fusion method provided in this application can aggregate scattered data into a centralized storage system, simplify data access and management, and reduce the complexity and time cost of data processing. By using unified data processing tools and algorithms, enterprises can more efficiently clean, integrate and analyze data, speed up the decision-making process, and improve the efficiency of business operations.
[0075] In some embodiments, users can also upload models to perform data training and analysis. In the data analysis process, after the duplicate data in the data is removed successfully through data deduplication, the front end returns the data deduplication success to the user and puts the deduplicated data back into the database. If the task fails, it returns data deduplication failure. Data filling fills the missing values in the data with the average value of the data. If successful, it returns data filling success and puts the filled data back into the database. Data digitization converts the text labels of the data set into numerical labels. If the task succeeds, the converted data is stored in the database in an overwritten form.
[0076] In order to better evaluate the relationship between model accuracy and confidence, this application proposes a new evaluation indicator, namely accuracy confidence consistency (abbreviated as: ACU). The formula is as follows:
[0077]
[0078] Among them, y_pred represents the confidence of the model's predicted category for a certain sample, y_true represents the true category of the sample, c represents the confidence coefficient. Generally, c can be 1 or normalized according to the confidence range, and n represents the number of samples.
[0079] In some embodiments, the data analysis service first summarizes and collects information on all data uploaded or selected by the user. This information includes: whether the data type is compliant, whether the uploaded data needs to be deduplicated, and the accuracy of the last training operation performed by the machine learning service (null if there has never been any).
[0080] See also Figure 3 As shown, Figure 3It is a structural diagram of a multi-source data fusion platform, which includes a display layer, a business interface layer, a business logic layer, a data access layer, and a database. Among them, the display layer includes the following functional modules: web platform, data upload, data visualization, and data download. The business interface layer includes the following functional modules: unified back-end business interface (webAPI). The business logic layer includes the following functional modules: data module, processing module, and algorithm module. The data access layer includes the following functional modules: data persistence. The database includes the following types of databases: Redis, MySQL, and MongoDB. The multi-source data fusion platform designed and constructed by this application stores and manages the information data generated by different tasks, including new tasks, historical task information query, data mining historical task scheduling and other functions. The multi-source data fusion method provided by this application provides multiple functions in the form of microservices, and adopts an interactive visual operation interface, so that users can simply and quickly analyze data resources and integrate stored data resources.
[0081] Here, we further explain the microservices provided by the multi-source data fusion processing platform. The data import service can parse the data uploaded by users and the model code. It can upload single data, merge multiple data and upload them. It can also store the processed data in the database and export it to csv, excel, sql and other formats. The data visualization service provides the function of visualizing data, including pie charts, bar charts, line charts, scatter plots and radar. Figure 5 Data visualization function. Machine learning service, according to different data types, provides different pre-trained models designed in this application, such as CategoryBERT, MobileNet-SE, TSM-Attention, etc. It also provides common baseline model training functions for machine learning and deep learning, and matches neural network models according to training data and target value types. Machine learning has multiple functions such as classification decision trees, confidence intervals, neural networks, linear regression and K-Means to choose from. Data storage service, to realize the storage function of user uploaded data, training data, prediction data and other data, and to store the training data uploaded by users in different databases according to different file types by categories and levels. When users are training, different types of files can be extracted from different databases according to task requirements.
[0082] This application provides operation functions through the above seven microservices. The entire data fusion and storage process is mainly responsible for the task management service, which provides management of task information. After creating a new task, users can use the function scheduling service to execute different functions in the form of a workflow in the corresponding order. The specific principle of the entire data fusion and storage process is that the user uses the task management service to select data and target values to create a new task, and the request is sent to the function scheduling service. The function scheduling service schedules the machine learning service to match the appropriate neural network model according to the data type selected by the user, and makes a judgment based on the file suffix and data type.
[0083] See also Figure 4 As shown, Figure 4 The following is a flow chart of the multi-source data fusion scenario. When the user uploads data, the multi-source data fusion processing platform calls the function scheduling service to receive the data uploaded by the user, and then calls the function scheduling service to schedule the data analysis service to perform pre-processing operations such as data exclusion, data deduplication, and data filling. The data storage service is scheduled to store the processed target data information in different databases according to different file types and categories. When extraction is required, data is called from the database to perform multi-source data fusion, and the machine learning service is scheduled to perform audit network model training and prediction. The final prediction data is stored in the database.
[0084] From the above, it can be seen that this application proposes three pre-training models. MobileNet-SE uses deep separable convolution. This convolution structure effectively reduces the amount of calculation and the number of parameters, and realizes efficient reasoning on mobile devices. The SE module is introduced, which can adaptively adjust the weights of the channel feature map, thereby improving the expression ability and classification performance of the model. The TSM module is adopted in TSM-Attention. This module introduces the perturbation of the timing information through the time translation operation, which helps to extract rich temporal features. TSM-Attention introduces a self-attention mechanism, which dynamically adjusts the attention weights by learning the relationship between different time steps in the video sequence, thereby improving the modeling ability of key frames. CategoryBERT classification task optimization, CategoryBERT is optimized based on the BERT model, including special optimization of classification tasks, such as classification loss function, use of classification-specific pre-training data sets, etc.
[0085] It should be noted that accuracy refers to the correctness or precision of a prediction model on a set of data. Generally, the higher the accuracy of the model, the more reliable its results. However, accuracy cannot fully reflect the pros and cons of a model, because the prediction results of the model may have different confidence levels. Confidence refers to the prediction of the model on the degree to which a certain prediction result is correct. Therefore, the model evaluated by the accuracy confidence consistency (ACU) proposed in this application has higher reliability.
[0086] The multi-source data fusion processing method provided by the present application has at least the following beneficial effects: greatly improving user participation and operability. Users can visually select model algorithms and their data through the Web page, and perform personalized model training according to different data sets. Even users who do not have a certain computer and advanced mathematics foundation can create workflow tasks through the task management service to obtain the required model training results. The data storage is intelligent and highly secure. Users can upload various types of data. The data storage service selects different databases for file storage according to different file types, and each user only has the authority to operate the data, which ensures the security of the data. The task progress can be observed. Because the tasks created by the task management service are executed in the form of workflows, the success and failure of each stage will return the corresponding status information to the user, and the user can intuitively observe the progress of the task execution through the Web page. The algorithm is diverse and extensible. In addition to the pre-trained models of text data, image data, and video data designed by the present application, and the commonly used baseline deep learning algorithms and machine learning algorithms, users can also upload model codes and train them by themselves, or combine multiple models for training, which greatly improves the flexibility and prediction accuracy of the algorithm and meets the needs of personalized training for users.
[0087] The present application also provides a multi-source data fusion processing device, Figure 5 As shown, the multi-source data fusion processing device 500 includes:
[0088] The first processing module 501 is used to call the task management service to generate task configuration information in response to the obtained user request;
[0089] The first processing module 501 is used to call the task management service to create a target task according to the task configuration information;
[0090] The second processing module 502 is used to call the function scheduling service to select a target neural network model that matches the task configuration information from multiple candidate neural network models to process the target task; the target neural network model is used to perform data analysis of multi-source data fusion.
[0091] In some embodiments, the first processing module 501 is used to call the task management service in response to a user request, extract different types of target files from different databases, and select target data information from the target files; call the task management service to generate task configuration information based on the file type of the target file and the data type of the target data information.
[0092] In some embodiments, the second processing module 502 is used to call the function scheduling service to receive the initial data uploaded by the user, perform preprocessing operations on the initial data information, and obtain processed target data information; call the function scheduling service to store the target data information in different databases in a hierarchical manner according to different file types, and generate different types of target files.
[0093] In some embodiments, the second processing module 502 is used to call a function scheduling service to select multiple neural network models that match the task configuration information from multiple candidate neural network models; call the function scheduling service to use the target data information as input data and the expected data information included in the user request as output data, train the multiple neural network models according to multi-dimensional evaluation indicators, and display the training results on the web page interface; the training results include the names of the multiple neural network models, the training duration, and the multi-dimensional evaluation indicators; call the function scheduling service in response to the user's selection operation of the target neural network model from the multiple neural network models displayed, and select the target neural network model to process the target task.
[0094] In some embodiments, the multi-dimensional evaluation index includes: the accuracy confidence consistency of each model training in multiple neural network models, the accuracy of each model training, and the average precision mean of each model training.
[0095] In some embodiments, the multiple candidate neural network models include: a BERT-based text classification model, a lightweight neural network model corresponding to image data, and a time offset module attention mechanism model corresponding to video data.
[0096] The embodiment of the present application provides a multi-source data fusion processing platform, which can be applied to Figure 1 In a multi-source data fusion processing method provided in a corresponding embodiment, referring to Figure 6 As shown, the multi-source data fusion processing platform 600 includes: a processor 601, a memory 602 and a communication bus 603, wherein: the communication bus 603 is used to realize the communication connection between the processor 601 and the memory 602;
[0097] The processor 601 is used to execute the multi-source data fusion processing program stored in the memory 602 to implement the following steps:
[0098] Invoke the task management service to generate task configuration information in response to the acquired user request;
[0099] Call the task management service to create the target task according to the task configuration information;
[0100] The function scheduling service is called to select a target neural network model that matches the task configuration information from multiple candidate neural network models to process the target task; the target neural network model is used to perform data analysis of multi-source data fusion.
[0101] In some embodiments of the present application, the processor 601 is used to execute the multi-source data fusion processing program stored in the memory 602 to implement the following steps:
[0102] Calling the task management service to respond to user requests, extracting different types of target files from different databases, and selecting target data information from the target files;
[0103] The task management service is called to generate task configuration information according to the file type of the target file and the data type of the target data information.
[0104] In some embodiments of the present application, the processor 601 is used to execute the multi-source data fusion processing program stored in the memory 602 to implement the following steps:
[0105] Call the function scheduling service to receive the initial data uploaded by the user, perform preprocessing operations on the initial data information, and obtain the processed target data information;
[0106] The function scheduling service is called to store the target data information hierarchically into different databases according to different file types, generating different types of target files.
[0107] In some embodiments of the present application, the processor 601 is used to execute the multi-source data fusion processing program stored in the memory 602 to implement the following steps: calling the function scheduling service to select multiple neural network models that match the task configuration information from multiple candidate neural network models;
[0108] The function scheduling service is called with the target data information as input data and the expected data information included in the user request as output data, and multiple neural network models are trained according to the multi-dimensional evaluation indicators, and the training results are displayed on the web page interface; the training results include the names of the multiple neural network models, the training duration, and the multi-dimensional evaluation indicators;
[0109] The calling function scheduling service responds to the user's selection operation of a target neural network model from the displayed multiple neural network models, and selects the target neural network model to process the target task.
[0110] In some embodiments of the present application, the multidimensional evaluation indicators include: the accuracy confidence consistency of each model training in multiple neural network models, the accuracy of each model training, and the average precision mean of each model training.
[0111] In some embodiments of the present application, multiple candidate neural network models include: a BERT-based text classification model, a lightweight neural network model corresponding to image data, and a time offset module attention mechanism model corresponding to video data.
[0112] The processor can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., among which the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0113] It should be noted that, for the description of the same steps and the same contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.
[0114] An embodiment of the present application provides a computer storage medium storing one or more programs, which can be executed by one or more processors to implement the following. Figure 1 Steps shown.
[0115] It should be noted that, for the description of the same steps and the same contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.
[0116] It should be noted that the above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM) and other memories; it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0117] An embodiment of the present application provides a computer product, including a computer program, which can be executed by a processor 601 of a multi-source data fusion processing platform 600 to complete the steps of the aforementioned method.
[0118] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0119] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0120] In addition, all functional units in the embodiments of the present application can be integrated into one processing module, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above integrated unit can be implemented in the form of hardware or in the form of hardware plus software functional units. A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memory (ROM), random access memory (RAM), disks or optical disks, and other media that can store program codes.
[0121] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0122] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0123] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0124] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A multi-source data fusion processing method, characterized in that: include: Invoke the task management service to generate task configuration information in response to the acquired user request; Calling the task management service to create a target task according to the task configuration information; Calling a function scheduling service to select a target neural network model matching the task configuration information from a plurality of candidate neural network models to process the target task; The target neural network model is used to perform data analysis of multi-source data fusion.
2. The method according to claim 1, characterized in that The calling task management service generates task configuration information in response to the acquired user request, including: Invoking the task management service in response to the user request, extracting different types of target files from different databases, and selecting target data information from the target files; The task management service is called to generate the task configuration information according to the file type of the target file and the data type of the target data information.
3. The method according to claim 2, characterized in that The calling of the task management service in response to the user request to extract different types of target files from different databases, before selecting target data information from the target files, further includes: Calling the function scheduling service to receive the initial data uploaded by the user, performing a preprocessing operation on the initial data information, and obtaining processed target data information; The function scheduling service is called to store the target data information in different databases in a hierarchical manner according to different file types, and to generate target files of different types.
4. The method according to claim 2, characterized in that: The calling function scheduling service selects a target neural network model matching the task configuration information from a plurality of candidate neural network models to process the target task, including: Calling the function scheduling service to select a plurality of neural network models matching the task configuration information from the plurality of candidate neural network models; Calling the function scheduling service with the target data information as input data and the expected data information included in the user request as output data, training the multiple neural network models according to the multi-dimensional evaluation index, and displaying the training results on the web page interface; the training results include the names of the multiple neural network models, the training duration, and the multi-dimensional evaluation index; The function scheduling service is called in response to a user's selection operation of the target neural network model from the multiple neural network models displayed, and the target neural network model is selected to process the target task.
5. The method according to claim 4, characterized in that The multidimensional evaluation index includes: the accuracy confidence consistency of each model training in the multiple neural network models, the accuracy of each model training, and the average precision mean of each model training.
6. The method according to claim 4, characterized in that The multiple candidate neural network models include: a BERT-based text classification model, a lightweight neural network model corresponding to image data, and a time offset module attention mechanism model corresponding to video data.
7. A multi-source data fusion processing device, characterized in that: include: A first processing module, configured to call a task management service to generate task configuration information in response to an acquired user request; The first processing module is used to call the task management service to create a target task according to the task configuration information; The second processing module is used to call the function scheduling service to select a target neural network model that matches the task configuration information from multiple candidate neural network models to process the target task; the target neural network model is used to perform data analysis of multi-source data fusion.
8. A multi-source data fusion processing platform, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory, and executing the steps of the multi-source data fusion processing method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the multi-source data fusion processing method as described in any one of claims 1 to 6.
10. A computer product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the multi-source data fusion processing method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Power hybrid model construction method and system, electronic equipment and storage medium
CN120671563A
Industrial research report generation method based on multi-source heterogeneous data fusion
CN121658951A