Digital platform construction method and system based on data medium platform
By collecting and preprocessing structured and unstructured data on the data middle platform, forming comprehensive feature vectors, and establishing deep learning models based on these vectors, the shortcomings of data fusion and analysis in the existing technology are solved, and the deep fusion and intelligent analysis of data are realized, and more accurate and reliable decision support is provided.
Patent Information
- Application Number
- CN202411968294.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
Existing data management and analysis technologies are difficult to achieve comprehensive integration and in-depth analysis of structured and unstructured data, resulting in serious data silos, lack of flexibility and intelligence, and unable to adapt to complex and changeable business needs.
By collecting structured and unstructured data on the data middle platform, pre-processing them separately, and forming comprehensive feature vectors through feature extraction and fusion, a data analysis model is established based on these vectors, and a gradient descent deep learning algorithm is used for training to achieve deep fusion and intelligent analysis of the data.
It has achieved deep integration of different types of data, improved the flexibility and intelligence level of data analysis, can adapt to complex and changeable business scenarios, provided more accurate and reliable decision-making support, enhanced the intelligence level of the entire process, and provided enterprises with a flexible and efficient data management and business support platform.
Smart Images

Figure CN119989256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management and processing technology, and in particular to a method and system for building a digital platform based on a data middle platform. Background Art
[0002] With the rapid development of information technology, data center, as a new type of data management and application architecture, has gradually become the core support for the digital transformation of enterprises. The data center aims to integrate and optimize various data resources within the enterprise, and support the rapid iteration and innovation of the business platform by providing a unified data service interface.
[0003] Existing data management and analysis technologies mainly focus on processing a single type of structured data, and support for unstructured data is relatively weak. Traditional methods usually use independent processing flows to process structured and unstructured data respectively, resulting in serious data silos and difficulty in achieving comprehensive data fusion and in-depth analysis. In addition, most existing data analysis models rely on static rules or simple statistical methods, lack flexibility and intelligence, and cannot adapt to complex and changing business needs. Summary of the invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for building a digital platform based on a data middle platform to solve the problem of difficulty in achieving comprehensive data fusion and in-depth analysis.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for constructing a digital platform based on a data middle platform, which comprises:
[0008] Collect structured data and unstructured data and pre-process them respectively;
[0009] The preprocessed data are subjected to feature extraction respectively and fused to form a comprehensive feature vector;
[0010] A data analysis model is established based on the comprehensive feature vector, and the analyzed information is obtained after the comprehensive vector is input into the data analysis model;
[0011] Categorize the parsed information and provide real-time feedback on event results;
[0012] Record and monitor the results of processed events.
[0013] As a preferred solution of the digital platform construction method based on the data middle platform described in the present invention, wherein: the structured data includes tabular data, time series data, geographic spatial data, transaction data, log data, and statistical survey data;
[0014] The unstructured data includes text file data, image data and video data.
[0015] As a preferred solution of the digital platform construction method based on the data middle platform described in the present invention, the structured data and unstructured data are collected and pre-processed respectively, which specifically includes the following steps:
[0016] Use the ETL tool Talend Open Studio to extract, transform and load unstructured data;
[0017] Use OCR technology to recognize text content in paper documents;
[0018] Extract image features through computer vision algorithms;
[0019] Apply speech-to-text technology to convert audio into text.
[0020] As a preferred solution of the digital platform construction method based on the data middle platform described in the present invention, the following steps are specifically included: feature extraction is performed on the pre-processed data respectively, and fusion is performed to form a comprehensive feature vector.
[0021] Extract numerical features, categorical features, and time series features of structured data using standardization, one-hot encoding, and timestamp parsing;
[0022] Use OpenCV and convolutional neural networks to extract low-level and high-level features of image data;
[0023] Use CountVectorizer, TfidfVectorizer and word embedding models to extract word frequency statistics, TF-IDF and word vector representation of text data;
[0024] The extracted features are combined to form a comprehensive feature vector.
[0025] As a preferred solution of the digital platform construction method based on the data middle platform described in the present invention, wherein: a data analysis model is established based on a comprehensive feature vector, and the comprehensive vector is input into the data analysis model to obtain the analyzed information, which specifically includes the following steps:
[0026] During the initialization and training process, the weights and biases are iteratively updated using the gradient descent algorithm until convergence, obtaining global weights and biases and specific category weights and biases;
[0027] Combine the comprehensive feature vector with the global weights and biases to get the base score;
[0028] Combine the comprehensive feature vector with the class-specific weights and biases to get a class-specific score;
[0029] Based on the basic score and the score of a specific category, the index result of each special category is obtained through the index function;
[0030] The index result of each category is divided by the sum of the index results of all specific categories, and the obtained result is normalized to obtain the probability distribution of the feature vector;
[0031] The information is obtained based on the probability distribution, i.e., the probability that the comprehensive feature vector belongs to each specific category;
[0032] The classification results based on the probability information, the confidence of the classification, and the relative likelihood between specific categories.
[0033] As a preferred solution of the digital platform construction method based on the data middle platform described in the present invention, the parsed information is classified and processed, and the processed event results are fed back in real time, which specifically includes the following steps:
[0034] Define key events based on the parsed information and business requirements, and set listeners to capture key events;
[0035] The listener captures key events and sends them to Apache Kafka, which receives and distributes the events and builds a message queue.
[0036] Based on the built message queue, define the data parsing method and business rules for each message; set up the processor and use the defined data parsing method to convert the message into structured data;
[0037] Operate on structured data according to defined business rules to achieve business goals and generate final processing results;
[0038] Feedback to the user processor real-time processing results.
[0039] As a preferred solution of the digital platform construction method based on the data middle platform described in the present invention, the processed event results are recorded and monitored and user privacy is protected, which specifically includes the following steps:
[0040] Based on the processing results fed back to the user, trigger rules are defined for each processing result according to business needs, feedback actions are generated based on the defined trigger rules, and corresponding actions are executed through the FaaS platform;
[0041] When the FaaS platform executes corresponding actions, the built-in logging mechanism captures and generates log entries, stores the log entries in ELKStack, uses the performance monitoring tool Prometheus to record and monitor key indicators, and stores the key indicators in Prometheus;
[0042] Use federated learning to manage log entries and key indicators, select departments and teams as data contributors, and configure the central server and its security protocols;
[0043] Each data contributor pre-processes log entries and key indicators locally to form key features. At the same time, the central server provides an initial model and starts local training of the initial model based on the key features.
[0044] Each data contributor uses local data to train the initial model, updates parameters and applies differential privacy technology, and uploads parameters to the central server through an encrypted channel for FedAvg aggregation;
[0045] Through multiple rounds of iterative optimization of local training, parameter uploading and global aggregation, a global model that protects user privacy is generated;
[0046] The privacy global model ensures user privacy through local data processing, differential privacy, encrypted communication, and parameter aggregation.
[0047] In a second aspect, the present invention provides a digital platform construction system based on a data middle platform, comprising:
[0048] The preprocessing module collects structured data and unstructured data and performs preprocessing;
[0049] The feature fusion module extracts features from the preprocessed data and fuses them to form a comprehensive feature vector;
[0050] Data parsing module, which establishes a data parsing model based on the comprehensive feature vector to obtain parsed information;
[0051] The processing feedback module classifies the parsed information and feeds back the real-time processed event results to the user;
[0052] The monitoring module triggers other business actions based on the processed event results, and records and monitors them.
[0053] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the method for building a digital platform based on a data middle platform as described in the first aspect of the present invention.
[0054] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the method for building a digital platform based on a data middle platform as described in the first aspect of the present invention.
[0055] The beneficial effects of the present invention are as follows: in the invention, data collection is carried out through OCR technology, computer vision algorithm, and speech-to-text technology, so that unstructured data can be processed in different ways, breaking the limitation of a single data type in traditional data processing, and multimodal feature extraction is performed on the preprocessed data, so that structured data and unstructured data can be effectively converted into representative feature vectors, thereby realizing the deep fusion of different types of data. The data analysis model established based on the comprehensive vector adopts the gradient descent deep learning algorithm, which has higher flexibility and intelligence level, can adapt to complex and changeable business scenarios, provide more accurate and reliable decision support, enhance the intelligence level of the entire process, and provide enterprises with a flexible and efficient data management and business support platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0057] Figure 1 This is a flow chart of the method for building a digital platform based on the data middle platform in Example 1.
[0058] Figure 2 Schematic diagram of the data analysis model of the feature vector in Example 1. DETAILED DESCRIPTION
[0059] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0060] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0061] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0062] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, and provides a method for building a digital platform based on a data middle station, comprising the following steps:
[0063] S1. Collect structured data and unstructured data.
[0064] The specific steps include:
[0065] Structured data refers to data that can be pre-defined in a data model or has fixed fields, such as tabular data, time series data, geospatial data, transaction data, log data, statistical survey data, stored in SQL databases or NoSQL databases, and collected using ETL tools, SQL queries, and API interfaces;
[0066] Unstructured data refers to data that does not follow a predefined pattern or format, such as text file data, image data, and audio data. The content and format are flexible and changeable. It is stored in HDFS, local files, scanned using file tools, web crawlers, cameras, or sensors.
[0067] Select different collection methods based on the characteristics of different types of data to ensure the quality and integrity of the data, reduce manual intervention, and improve the level of automation.
[0068] S2. Preprocess the collected structured data and unstructured data.
[0069] The specific steps include:
[0070] Structured data is extracted from the data source using the ETL tool Talend Open Studio and converted. During the conversion process, data cleaning, format unification, and data standardization are performed. The processed data is efficiently migrated to the target platform and a loading report is generated. Through the automated ETL process, the quality and consistency of the data are ensured, manual intervention is reduced, and data processing efficiency is improved;
[0071] Unstructured data includes text data, image data and audio data. Text data is scanned using a high-resolution scanner and parameters are adjusted to ensure image clarity and readability. Scanned images are enhanced and segmented into pages to improve OCR recognition accuracy. OCR tools are selected and parameters are configured to optimize character recognition accuracy. Generated text is corrected and restored to the original format to ensure the integrity and accuracy of the content. High-quality scanning and OCR processing improves text recognition accuracy, reduces the workload of manual proofreading, and improves the availability and reliability of text data.
[0072] Image data uses computer vision algorithms for edge detection, corner detection, and local feature description to accurately identify key structures and details in images, extract global features, and provide an overall description of the image. Convolutional neural networks (CNNs) are used to extract high-level abstract features and capture complex patterns and semantic information of images. Advanced image processing technology enhances the depth and breadth of image analysis, improves the accuracy and robustness of image recognition, and provides a foundation for subsequent applications.
[0073] The audio data is recognized using a speech-to-text tool and recognition parameters are configured to optimize the recognition effect. The voice signal that has undergone noise reduction and volume standardization is converted to ensure the consistency of audio quality. After obtaining the preliminary text output, text correction is performed, punctuation and formatting are restored, and the final text is more readable and standardized. Efficient audio processing and speech recognition technology improves the accuracy and readability of text conversion and reduces the need for manual editing.
[0074] S3. Extract features from the preprocessed data respectively and fuse them to form a comprehensive feature vector.
[0075] For the numerical features of structured data, first determine which columns in the structured data contain numerical data, process missing values in these columns to ensure data integrity, and convert numerical features into standard normal distribution or scale them to a specific range through standardization methods to eliminate dimensional differences, make numerical features of different scales comparable, and improve model performance and convergence speed;
[0076] For the categorical features of structured data, identify the columns containing categorical data, handle missing values, ensure data integrity, and then use one-hot encoding to convert categorical features into binary vectors to avoid introducing unnecessary sequential relationships, so that the model can better understand categorical information and improve the accuracy and interpretability of classification tasks;
[0077] For the time series features of structured data, first identify the timestamp column and parse it into a datetime object to extract more useful time information and ensure the consistency and operability of the time series data. Then extract useful components to capture the periodicity and trend in time. Generate additional time series features based on business needs and calculate the difference between timestamps to capture the time interval information of events, which helps analyze the frequency and duration of events and supports more sophisticated time series analysis.
[0078] For image data in unstructured data, OpenCV is used for low-level feature extraction, including edge and corner detection, key point identification and generation of local feature descriptors, and global features such as color histogram, texture and shape are extracted to provide an overall description of the image, enhance the breadth and depth of image analysis, and convolutional neural networks are used to extract high-level features. The pre-trained model is loaded and configured, and the image input is forward propagated to calculate the feature map. Finally, high-level feature vectors with strong expression and generalization capabilities are extracted from specific layers, which significantly improves the accuracy and robustness of image recognition and classification.
[0079] For text data in unstructured data, use the natural language processing tool CountVectorizer to extract word frequency statistics, and use the fit_transform method provided by CountVectorizer to convert text data into a word frequency matrix, where each row represents a document, each column represents a vocabulary item, and the values in the matrix represent the number of times the vocabulary item appears in the document. Use the text feature extraction method TfidfVectorizer to extract TF-IDF features, and then use the fit_transform method to convert text data into a TF-IDF weighted matrix, where each row represents a document, each column represents a vocabulary item, and the values in the matrix represent the importance of the vocabulary item in the document, highlighting important vocabulary items, suppressing the influence of common vocabulary, and improving the quality and discrimination of text representation. Use the word embedding model to extract word vector representation features, load the word embedding model and generate word vectors, find the corresponding vector for each word in the document, and use average or other methods to aggregate into document vector representations, process unknown words to ensure that all words have corresponding vectors, avoid feature loss caused by unknown words, and maintain the integrity and consistency of text data;
[0080] After extracting numerical, image, text and other features of structured and unstructured data, ensure that all features are at the same scale and have the same number of samples to avoid problems caused by scale differences or sample inconsistencies, align feature names or indexes, and then merge these features into a comprehensive feature vector by column through horizontal splicing, providing a comprehensive data view and supporting multimodal data analysis and the construction of complex models.
[0081] S4. Establish a data analysis model based on the comprehensive feature vector, and input the comprehensive vector into the data analysis model to obtain the analyzed information.
[0082] The specific steps include:
[0083] The comprehensive vector that fuses structured and unstructured data features is input into the data parsing model to define the model's initial weight matrix W and bias vector b, as well as the category-specific weight matrix W m and the bias vector b m ;
[0084] Use the gradient descent algorithm for iterative updates, gradually optimizing the global and category-specific weights and bias parameters. In each iteration, calculate the loss function between the predicted value and the actual label under the current parameters, and adjust the parameters according to the gradient of the loss function. Iterate until the loss function converges, that is, the loss no longer decreases significantly or reaches the preset maximum number of iterations, so as to obtain the optimal global and category-specific weights and bias parameters. The gradient descent algorithm can improve the generalization ability of the model, enabling it to perform well on unseen data;
[0085] Combine the comprehensive feature vector G with the global weight matrix W and the bias vector b to calculate the global score,
[0086] BaseScore=W·G+bBaseScore=W·G+b,
[0087] For each category m, the comprehensive feature vector G is combined with the category-specific weight matrix W m and the bias vector b m Combined to calculate a category-specific score,
[0088] Class-SpecificScorem=Wm·G+bm,
[0089] The above approach not only takes into account the impact of global features, but also makes personalized adjustments for different categories, thus enhancing the model's ability to distinguish different categories.
[0090] Apply an exponential function to the global score, the category-specific score, and get an exponential result for each category.
[0091] exp(BaseScore) and exp(Class-SpecificScorem)e;
[0092] Then the index result of each category is divided by the sum of the index results of all categories for normalization to ensure that the sum of the probabilities of all categories is 1 to avoid negative results or out-of-range results. The final probability distribution R formula is obtained, which is expressed as:
[0093]
[0094] Among them, R m represents the probability of the mth category, W m represents the weight matrix of the mth category, b m represents the bias vector of the mth category, and M represents the total number of categories;
[0095] According to the probability distribution R, the possibility of the input information belonging to each category is clarified, and the category with the highest probability is selected as the final classification result. The maximum probability value represents the model's confidence in the classification result. By comparing the probabilities of different categories, the relative probability between categories is evaluated, which helps to understand other possible classification options and their probabilities.
[0096] S5. Classify and process the parsed information, and provide real-time feedback on the event results after classification.
[0097] The specific steps include:
[0098] First, define key events according to business needs to ensure that subsequent steps are targeted and in line with business goals. Define trigger conditions for key events, set thresholds and rules to ensure that only key events that meet the conditions are triggered to reduce false alarm rates.
[0099] Set up the listener through the ETL tool Talend. First, ensure that the cluster of the distributed stream processing platform Apache Kafka has been correctly configured and is ready to receive data from Talend. Create a new job in Talend Studio and name it. Then configure the data input source to obtain the parsed information in real time, apply predefined rules to identify key events, use tFilterRow or tMap components to set conditions such as classification results and confidence thresholds, ensure that only key events that meet the conditions are captured, improve the accuracy and efficiency of event processing, and then serialize the identified key events into a unified message format, and then send them to the Apache Kafka topic through the tKafkaProducer component to ensure the efficiency and reliability of data transmission. Finally, test to verify whether the message received by the Apache Kafka topic is correct, optimize performance as needed, and deploy to the production environment to ensure that the job can run stably in the actual environment and ensure business continuity.
[0100] Use the above Apache Kafka topics and clusters to receive and distribute event messages, and then build a message queue. The method of building a message queue is as follows:
[0101] First, configure Apache Kafka dependencies and set server parameters to ensure high availability and scalability. After starting the service, create Apache Kafka topics according to business needs, set the number of partitions and replication factors to ensure effective management and load balancing of different types of events. Then configure the tKafka Producer component in the Talend job to ensure consistent message formats and efficient transmission. At the same time, configure the consumer component in the ETL tool to ensure timely processing of messages and avoid duplicate consumption. The message queue built in this step has high performance and high reliability, which can meet complex business needs.
[0102] Define specific parsing methods for each type of message to ensure that it can be correctly extracted and converted into structured data, ensure the consistency and accuracy of data parsing, and define specific rules and operations for processing structured data according to business needs to ensure that business operations comply with business logic and improve the efficiency and accuracy of business processes;
[0103] Define the data parsing method and business rules for each message and set the processor. The specific method is as follows:
[0104] Create a new job named EventProcessor in Talend Studio to process key events. First, add and configure the tKafkaConsumer component to subscribe to the Kafka topic and set the connection parameters to ensure efficient message reception. Then use the tConvertType or tMap component to parse and validate the message and convert it into a structured data object;
[0105] Based on business needs, apply business rules through tJavaRow or tMap components, perform operations such as updating the database, triggering an alert, or starting a workflow, and call external services through tRESTClient. After processing is completed, construct a response object and use tLogRow to record key steps and results.
[0106] Finally, we choose to feed back the processing results to users through API response, email notification or platform log, and add error handling logic to ensure stability and reliability. The entire process uses Talend's visual interface and component library to simplify development and improve maintainability and scalability.
[0107] S6. Record and monitor the results of processed events and protect user privacy.
[0108] The specific steps include:
[0109] According to business needs, we clarify different types of processing results and their corresponding business operations, define specific trigger rules for each processing result, select the Fass platform of the cloud computing service model, set the function logic to implement the trigger rules, and ensure that the correct operation can be performed according to the processing results. When the processor generates the final processing result, the corresponding FaaS function is called according to the predefined trigger rules. The FaaS platform ensures that the function can be executed quickly after receiving the trigger signal, providing a low-latency response.
[0110] The built-in logging mechanism of the FaaS platform automatically captures log entries during the execution of each function, including input and output, execution time, and error information. ELKStack in the open source log management platform is used to centrally manage log data. Prometheus in the open source monitoring toolkit is used to collect and analyze key indicators for performance monitoring, such as function execution time, request success rate, resource utilization, etc., to ensure that the platform performance is in the best state.
[0111] Federated learning is used to analyze logs and monitoring indicators. First, the architecture is defined, departments or teams are selected as data contributors (such as IT operations, business departments, etc.), and the TLS central server and its security protocol are specified to ensure the security of all data transmission. Then, each data contributor pre-processes, cleans, and converts the logs and monitoring data in its local environment and extracts key features suitable for machine learning. It also ensures that the necessary computing resources and software environment are available to support the subsequent local training process.
[0112] The central server then generates an initial global model and distributes it to each data contributor through a secure channel. After each data contributor downloads the initial model, they start the first round of training in the local environment using the preprocessed data. During the local training process, each data contributor uses the local dataset to train the downloaded initial model, updates the model parameters, and applies differential privacy technology to protect privacy.
[0113] After completing local training, each data contributor securely uploads the updated model parameters to the central server through the TLS encrypted communication channel. The central server aggregates the received parameters using the FedAvg aggregation algorithm to generate a new global model. This process is repeated for multiple rounds. In each round, the central server redistributes the latest global model to each data contributor, and each data contributor continues to perform local training based on the new model. Through multiple rounds of iterations, the model performance is gradually improved to ensure that the final generated global model is efficient and protects user privacy.
[0114] After multiple rounds of iterative optimization, the central server aggregates all updates and generates a final global model. This model uses local data processing to keep all original logs and monitoring data in the local environment of each data contributor, and does not upload them to the central server or any other external entity to ensure data is not leaked.
[0115] Apply differential privacy technology to add appropriate noise to the updated model parameters to prevent specific original data or user information from being inferred from the model parameters;
[0116] The application of encrypted communication enables all participants to communicate with the central server through TLS encrypted channels, ensuring that model parameters cannot be eavesdropped or tampered with during transmission;
[0117] Apply parameter aggregation so that the central server only receives updated model parameters and generates a new global model using secure aggregation algorithms such as FedAvg, avoiding direct sharing of raw data;
[0118] The model performance is gradually improved by applying multiple rounds of iterative optimization. Only a small amount of updated parameters are uploaded in each round, which reduces the amount of data transmitted each time and further reduces privacy risks.
[0119] Efficient data collaboration and user privacy protection are achieved through the above technologies and mechanisms.
[0120] This embodiment also provides a digital platform construction system based on a data middle platform, including:
[0121] The preprocessing module collects structured data and unstructured data and performs preprocessing respectively;
[0122] The feature fusion module extracts features from the preprocessed data and fuses them to form a comprehensive feature vector;
[0123] A data analysis module establishes a data analysis model based on the comprehensive feature vector, and obtains analyzed information after inputting the comprehensive vector into the data analysis model;
[0124] The processing feedback module classifies the parsed information and provides real-time feedback on the results of the processed events;
[0125] The monitoring module records and monitors the event results after classification and processing, and protects user privacy. This embodiment also provides a computer device, which is suitable for the digital platform construction method based on the data middle station, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions, so as to realize the digital platform construction method based on the data middle station as proposed in the above embodiment.
[0126] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0127] The present embodiment also provides a storage medium on which a computer program is stored. When the program is executed by the processor, the method for building a digital platform based on a data middle platform as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, disk or optical disk.
[0128] In summary, the present invention uses: OCR technology, computer vision algorithm, and speech-to-text technology to enable unstructured data to be collected and processed in different ways, breaking the limitation of a single data type in traditional data processing, and performing multimodal feature extraction on preprocessed data, so that structured data and unstructured data can be effectively converted into representative feature vectors, thereby achieving deep fusion of different types of data. The data parsing model established based on the comprehensive vector adopts a gradient descent deep learning algorithm, which has higher flexibility and intelligence, can adapt to complex and changeable business scenarios, provide more accurate and reliable decision support, enhance the intelligence level of the entire process, and provide enterprises with a flexible and efficient data management and business support platform.
[0129] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for building a digital platform based on a data center, characterized in that ,include: Collect structured data and unstructured data and pre-process them respectively; The preprocessed data are subjected to feature extraction respectively and fused to form a comprehensive feature vector; A data analysis model is established based on the comprehensive feature vector, and the analyzed information is obtained after the comprehensive vector is input into the data analysis model; Classify the parsed information and provide real-time feedback on the event results after classification; The event results after classification are recorded and monitored, and user privacy is protected.
2. The method for constructing a digital platform based on a data middle station according to claim 1, characterized in that: The structured data includes tabular data, time series data, geospatial data, transaction data, log data, and statistical survey data; The unstructured data includes text file data, image data and audio data.
3. The method for constructing a digital platform based on a data middle platform according to claim 2, characterized in that: Collect structured data and unstructured data and pre-process them separately, which includes the following steps: Use the ETL tool Talend Open Studio to extract, transform and load unstructured data. Use OCR technology to recognize text content in paper documents; Extract image features through computer vision algorithms; Apply speech-to-text technology to convert audio into text.
4. The method for constructing a digital platform based on a data middle platform according to claim 3, characterized in that: The preprocessed data are subjected to feature extraction and fused to form a comprehensive feature vector, which specifically includes the following steps: Extract numerical features, categorical features, and time series features of structured data using standardization, one-hot encoding, and timestamp parsing; Use OpenCV and convolutional neural networks to extract low-level and high-level features of image data; Use CountVectorizer, TfidfVectorizer and word embedding models to extract word frequency statistics, TF-IDF and word vector representation of text data; The extracted features are fused to form a comprehensive feature vector.
5. The method for constructing a digital platform based on a data middle platform according to claim 4, characterized in that: Establishing a data analysis model based on the comprehensive feature vector, inputting the comprehensive vector into the data analysis model to obtain the analyzed information, specifically includes the following steps: During the initialization and training process, the weights and biases are iteratively updated using the gradient descent algorithm until convergence, obtaining global weights and biases and specific category weights and biases; Combine the comprehensive feature vector with the global weights and biases to get the base score; Combine the comprehensive feature vector with the class-specific weights and biases to get a class-specific score; Based on the basic score and the score of a specific category, the index result of each category is obtained through an exponential function; The index result of each category is divided by the sum of the index results of all specific categories, and the obtained result is normalized to obtain the probability distribution of the feature vector; The information is obtained based on the probability distribution, i.e., the probability that the comprehensive feature vector belongs to each specific category; The classification results based on the probability information, the confidence of the classification, and the relative likelihood between specific categories.
6. The method for constructing a digital platform based on a data middle platform according to claim 5, characterized in that: The parsed information is classified and processed, and the event results after classification are fed back in real time. The specific steps include the following: Define key events based on the parsed information and business requirements, and set listeners to capture key events; The listener captures key events and sends them to Apache Kafka, which receives and distributes the events and builds a message queue. Based on the built message queue, define the data parsing method and business rules for each message; set up the processor and use the defined data parsing method to convert the message into structured data; Operate on structured data according to defined business rules to achieve business goals and generate final processing results; Feedback to the user processor real-time processing results.
7. The method for constructing a digital platform based on a data middle platform according to claim 6, characterized in that: Record and monitor the processed event results and protect user privacy, which includes the following steps: Based on the processing results fed back to the user, trigger rules are defined for each processing result according to business needs, feedback actions are generated based on the defined trigger rules, and corresponding actions are executed through the FaaS platform; When the FaaS platform executes corresponding actions, the built-in logging mechanism captures and generates log entries, stores the log entries in ELKStack, uses the performance monitoring tool Prometheus to record and monitor key indicators, and stores the key indicators in Prometheus; Use federated learning to manage log entries and key indicators, select departments and teams as data contributors, and configure the central server and its security protocols; Each data contributor pre-processes log entries and key indicators locally to form key features. At the same time, the central server provides an initial model and starts local training of the initial model based on the key features. Each data contributor uses local data to train the initial model, updates parameters and applies differential privacy technology, and uploads parameters to the central server through an encrypted channel for FedAvg aggregation; Through multiple rounds of iterative optimization of local training, parameter uploading and global aggregation, a global model that protects user privacy is generated; The privacy global model ensures user privacy through local data processing, differential privacy, encrypted communication, and parameter aggregation.
8. A digital platform construction system based on a data middle platform, based on the digital platform construction method based on a data middle platform according to any one of claims 1 to 7, characterized in that: include, The preprocessing module collects structured data and unstructured data and performs preprocessing respectively; The feature fusion module extracts features from the preprocessed data and fuses them to form a comprehensive feature vector; A data analysis module establishes a data analysis model based on the comprehensive feature vector, and obtains analyzed information after inputting the comprehensive vector into the data analysis model; The processing feedback module classifies the parsed information and provides real-time feedback on the event results after classification; The monitoring module records and monitors the processed event results and protects user privacy.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for building a digital platform based on a data middle platform as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for building a digital platform based on a data middle platform as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Minority-oriented AI identification individual adaptive diet recommendation method
CN121354821A
Intelligent fish protector fish weighing and returning method and system based on deep learning
CN121458228A