OpenCSG intelligent data processing platform

Through large-model-driven automated annotation, data standardization, encryption processing and compliance detection, the OpenCSG intelligent data processing platform solves the problems of long calculation time, manual adjustment of data formats, and low data security in the existing technology, achieving efficient, flexible and secure data processing.

CN119989031AInactive Publication Date: 2025-05-13SHANGHAI CHUANZHISHEN TECH CO LTD

Patent Information

Application Number
CN202510458544.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, there are problems such as long calculation time, manual adjustment of data format, and low data security.

Method used

It provides an OpenCSG intelligent data processing platform, including data acquisition and integration module, data analysis module, data output and application module, and security and scale block. Through large-scale model-driven automated annotation, data standardization, encryption processing and compliance detection, the platform solves the problems of low data processing efficiency, complex format adjustment and insufficient security.

Benefits of technology

Through automated labeling driven by large models, data labeling efficiency is significantly improved, high-precision labeling results are generated, and a variety of types of labeling tasks are adapted to; data standardization and encryption processing improve the flexibility and security of data processing; compliance detection ensures that data processing complies with preset regulatory requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989031A_ABST
    Figure CN119989031A_ABST
Patent Text Reader

Abstract

The invention discloses an OpenCSG intelligent data processing platform comprising a data acquisition and integration module which can acquire data from various data sources and perform standardization and cleaning processing on the data according to the format and structure of the data; the data analysis module is used for analyzing the data by adopting an OpenCSG large model technology and generating a data labeling result; the data output and application module can send a data labeling result to a user and / or other application programs and generate a dynamic interactive data report according to demand information input by the user; and the security and conformity block can be used for encrypting the data and automatically detecting whether the data processing process meets the requirements of preset laws and regulations or not. According to the method, through automatic labeling driven by the large model, the workload of manual labeling is greatly reduced, and the efficiency of data labeling is remarkably improved. And meanwhile, a high-precision labeling result can be generated, and the method is suitable for various types of labeling tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of data processing, and in particular relates to an OpenCSG intelligent data processing platform. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies, data processing has become a core requirement of all industries. Whether it is finance, medical care, retail or manufacturing, they all rely on efficient data processing to improve decision-making quality and business efficiency. As the scale of data continues to grow, traditional data processing methods can no longer meet the needs of modern business. Enterprises need more intelligent and efficient solutions to process massive and diverse data.

[0003] However, existing data processing technologies still have many shortcomings, which affect the efficiency of data utilization. First, traditional tools often require a long computing time when processing large-scale data, and it is difficult to support real-time or near real-time data processing, resulting in a lag in data value mining. Secondly, faced with data from different sources and formats, traditional methods often require tedious manual adjustments, which not only increases the complexity of operations but also reduces the flexibility of the system. In addition, data security and privacy protection also face huge challenges during data transmission and storage, especially in cross-platform and cloud environments. How to effectively ensure data security has become a key issue.

[0004] In addition, the integration complexity of traditional data processing tools is high, and enterprises often need to invest a lot of time and cost when introducing new technologies or connecting with existing systems. Therefore, there is an urgent need for an intelligent data processing solution that is both efficient and easy to integrate to reduce operation and maintenance costs and improve data processing capabilities. Summary of the invention

[0005] The present invention provides an OpenCSG intelligent data processing platform to solve the problems in the prior art such as long calculation time, manual adjustment of data format, and low data security. In order to solve the above technical problems, the embodiments of the present invention disclose the following technical solutions: One aspect of the present invention provides an OpenCSG intelligent data processing platform, comprising: The data collection and integration module is configured to obtain data from various data sources and standardize and clean the data according to the data format and structure; The data analysis module is configured to analyze the data using the OpenCSG large model technology and generate data annotation results; The data output and application module is configured to send the data annotation results to the user and / or other applications, and to generate dynamic interactive data reports according to the required information input by the user; The security and compliance module is configured to encrypt data and automatically detect whether the data processing process complies with preset regulatory requirements.

[0006] Optionally, the data acquisition and integration module includes a data acquisition submodule and a preprocessing submodule, wherein: The data collection submodule is configured to crawl data from CSGHub, preset datasets, and preset networks; The preprocessing submodule is configured to clean the data, convert the format, detect abnormal data, and standardize the data.

[0007] Optionally, the preprocessing submodule includes a standardization unit, a data conversion unit, a data cleaning unit, a data enhancement unit and a review and evaluation unit, wherein: The standardization unit is configured to encode the data and unify the format; The data conversion unit is configured to convert the data into a preset format suitable for model training; The data cleaning unit is configured to delete abnormal data and fill in missing data; The data enhancement unit is configured to perform enhancement processing on the data using a preset enhancement method; The review and evaluation unit is configured to detect data quality and generate data optimization information.

[0008] Optionally, the data analysis module includes a model training and tuning submodule, an entity extraction submodule, a data clustering submodule, a data annotation submodule and a security scanning submodule, wherein: The model training and tuning submodule is configured to train and fine-tune the LLM large language model based on the data obtained by the data acquisition and integration module, and to optimize the response and decision-making capabilities of the LLM model using the RAG retrieval enhancement generation technology to obtain the OpenCSG large model; The entity extraction submodule is configured to extract key entities in the data; The data clustering submodule is configured to perform cluster analysis on data characteristics; The data annotation submodule is configured to automatically annotate data based on the OpenCSG large model technology; The security scanning submodule is configured to check for potential security risks in the data processing process.

[0009] Optionally, the data analysis module also includes an algorithm recommendation submodule, which is configured to collect and store user operation records, and based on machine learning or knowledge graphs, use user operation records to dynamically generate data processing strategy recommendation information.

[0010] Optionally, the data output and application module includes a model optimization submodule, a data sharing submodule, an application submodule, a data storage submodule, a data preview submodule, an upload and download submodule, and a user management submodule, wherein: The model optimization submodule is configured to optimize the training of the OpenCSG large model using the data annotation results; The data sharing submodule is configured to share data annotation results through CSGHub; An application submodule, configured to generate reports and / or decision information based on the data annotation results; The data storage submodule is configured to store data annotation results and provide version management; A data preview submodule is configured to display data annotation results to users; The upload and download submodule is configured to receive data uploaded by the user and send the downloaded data to the user; The user management submodule is configured to control user permissions and manage user identities.

[0011] Optionally, the security and compliance module includes: A data encryption submodule is configured to encrypt data before transmission; The permission verification submodule is configured to verify the user's permission; The compliance review submodule is configured to detect whether the data processing process complies with preset regulations and generate a compliance review report.

[0012] Optionally, the preprocessing submodule includes an abnormal data detection unit configured to identify abnormal data.

[0013] The present invention discloses an OpenCSG intelligent data processing platform, including a data acquisition and integration module, which is used to obtain data from multiple data sources, and standardize and clean the data according to the data format and structure; a data analysis module, which is used to analyze the data using the OpenCSG large model technology and generate data annotation results; a data output and application module, which is used to send the data annotation results to users and / or other applications, and generate dynamic interactive data reports based on the demand information input by the user; a security and compliance module, which is used to encrypt the data and automatically detect whether the data processing process meets the preset regulatory requirements. The present invention greatly reduces the workload of manual annotation and significantly improves the efficiency of data annotation through large model-driven automatic annotation. At the same time, it can generate high-precision annotation results that are suitable for various types of annotation tasks, whether it is sentiment analysis, entity recognition or complex text classification, it can be completed efficiently to meet different data processing needs.

[0014] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.

[0016] Figure 1 A schematic diagram of the structure of an OpenCSG intelligent data processing platform provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0018] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0019] Figure 1 A schematic diagram of the structure of an OpenCSG intelligent data processing platform provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, it includes a data collection and integration module 1, a data analysis module 2, a data output and application module 3, and a security and compliance module 4.

[0020] 1. Data collection and integration module The data collection and integration module 1 can access data from various data sources and standardize and clean the data according to the data format and structure.

[0021] In one embodiment disclosed in the present invention, the data acquisition and integration module 1 includes a data acquisition submodule and a preprocessing submodule.

[0022] 1. The data collection submodule can crawl data from CSGHub, preset data sets and preset networks, and support obtaining data from different types of data sources to ensure data diversity and integrity, including but not limited to: (1) Structured data sources: such as relational databases (MySQL, PostgreSQL, SQL Server), data warehouses (Amazon Redshift, Google BigQuery), etc.

[0023] (2) Semi-structured data sources: data files in formats such as JSON, XML, YAML, and Parquet.

[0024] (3) Unstructured data sources: such as text (TXT, DOCX, PDF), images (JPEG, PNG), audio (WAV, MP3), etc.

[0025] (4) API data interface: supports RESTful API, GraphQL, and WebSocket to access third-party system data, such as ERP, CRM, IoT platforms, etc.

[0026] (5) CSGHub data source: directly connect to the CSGHub platform to access shared data sets or open data resources.

[0027] (6) Web crawling: supports targeted crawling of web pages, social media, and news data, and parsing and extracting useful information.

[0028] 2. The preprocessing submodule is used to clean the data, convert the format, detect abnormal data, and standardize the data.

[0029] In one embodiment disclosed in the present invention, the preprocessing submodule includes a standardization unit, a data conversion unit, a data cleaning unit, a data enhancement unit and a review and evaluation unit.

[0030] (1) The standardization unit is used to encode the data obtained by the data acquisition submodule and unify the format to solve the inconsistency problem of multi-source heterogeneous data.

[0031] For example, UTF-8 can be used for text encoding to avoid garbled characters, and categorical variables can be digitized, such as One-Hot encoding and label encoding.

[0032] Format standardization includes but is not limited to the following methods: unified time format, unified numerical units, standardized field naming, etc.

[0033] Data type conversion includes but is not limited to the following methods: converting string values ​​into floating-point numbers or integers; parsing JSON and XML data structures and converting them into structured data formats.

[0034] (2) The data conversion unit is used to convert the data processed by the standardization unit into a preset format suitable for model training.

[0035] For example, you can use methods such as TF-IDF, Word2Vec, and BERT Embeddings to convert text data into vectors.

[0036] (3) The data cleaning unit is used to delete abnormal data and fill in missing data.

[0037] In one embodiment disclosed in the present invention, the preprocessing submodule further includes an abnormal data detection unit configured to identify abnormal data, wherein the abnormal data at least includes missing values, duplicate values, etc.

[0038] The abnormal data detection unit can use Z-Score and IQR (interquartile range) methods to detect abnormal data, and use Isolation Forest, LOF (local outlier factor) and other algorithms to identify missing values.

[0039] The data cleaning unit can use KNN, regression models, context interpolation, BERT filling methods, etc. to supplement missing information.

[0040] In addition, the abnormal data detection unit can use hash value comparison and cosine similarity to identify nearly duplicate data, and the data cleaning unit can delete the identified duplicate data.

[0041] (4) The data enhancement unit is used to enhance the cleaned data using a preset enhancement method to increase the diversity of the data and improve the generalization ability of the model.

[0042] For example, the EDA (Easy Data Augmentation) method can be used, including synonym replacement, random insertion, random deletion, or OpenCSG can be used for semantic expansion to generate diversified text data.

[0043] (5) The review and evaluation unit is used to detect data quality and provide data optimization suggestions to ensure data availability and stability.

[0044] For example, check whether there are large-scale missing data, too many outliers, etc., or check whether the data format, unit, and field name are consistent. You can also detect abnormal distribution by calculating statistical information such as the mean, variance, skewness, and kurtosis of the data, and use the Kolmogorov-Smirnov test and Shapiro-Wilk test to evaluate whether the data conforms to a specific distribution (such as a normal distribution).

[0045] The review and evaluation unit can provide solutions for abnormal data processing, such as deleting outliers, resampling data, etc.

[0046] 2. Data Analysis Module The data analysis module 2 is used to analyze the data using the OpenCSG large model technology and generate data annotation results, including a model training and tuning submodule, an entity extraction submodule, a data clustering submodule, a data annotation submodule and a security scanning submodule.

[0047] (1) The model training and tuning submodule can train and fine-tune the LLM large language model based on the data obtained by the data collection and integration module, and use the RAG retrieval enhancement generation technology to optimize the response and decision-making capabilities of the LLM model to obtain the OpenCSG large model.

[0048] In the model training phase, the data is first preprocessed to ensure that the data format, distribution, and quality meet the training requirements. Subsequently, supervised learning, unsupervised learning, reinforcement learning and other methods are used to conduct preliminary training on the LLM to learn the semantic patterns and logical relationships in the data.

[0049] Fine-tuning is an important part of the model training and tuning submodule, which mainly uses technologies such as instruction tuning and RLHF (reinforcement learning with human feedback) to make the OpenCSG large model perform better on specific tasks. At the same time, the model training and tuning submodule integrates RAG (Retrieval-Augmented Generation) technology to enhance the knowledge depth of LLM with external knowledge base and real-time retrieval mechanism, making it more accurate and reliable in data analysis and decision support.

[0050] (2) The entity extraction submodule is responsible for identifying and extracting key entities from unstructured or semi-structured data, providing a basis for subsequent analysis and modeling.

[0051] In text data processing, this module uses Named Entity Recognition (NER) technology to extract core information such as names of people, places, institutions, time, quantity, and professional terms from the data. For example, in the medical field, this module can identify disease names, drug names, medical indicators, etc.; in the financial field, it can extract company names, stock market terms, financial data, etc.

[0052] For non-text data (such as images, videos, audio, etc.), this module combines computer vision (CV) and speech recognition (ASR) technology to extract object categories and scene information from images, or identify key speech content from audio. Through deep learning models (such as BERT, RoBERTa, ViT, Whisper), this submodule can automatically mine key information from massive data, greatly improving the efficiency and accuracy of data processing.

[0053] (3) The data clustering submodule can perform cluster analysis on data characteristics.

[0054] The data clustering submodule uses clustering algorithms such as K-Means, DBSCAN, Hierarchical Clustering, and GMM (Gaussian Mixture Model) to automatically group data according to its feature vectors. For example, in an e-commerce recommendation system, this module can classify users into different consumer groups based on their purchasing behavior; in medical diagnosis, patients can be classified based on the similarity of their symptoms for precise treatment.

[0055] In addition, the data clustering submodule supports high-dimensional clustering, and combines dimensionality reduction techniques (such as PCA, t-SNE, and UMAP) to extract the most representative features in large-scale, multi-dimensional data, improving the interpretability and computational efficiency of clustering. Ultimately, the output results of this submodule will be used in a variety of application scenarios such as automatic labeling, personalized recommendations, and anomaly detection.

[0056] (4) The data annotation submodule can combine the analysis capabilities of the OpenCSG large model to automatically annotate data and improve the intelligence of data management.

[0057] The data annotation submodule can replace the traditional manual annotation method and use active learning and self-supervised learning to make the data annotation process more efficient and accurate. It can be applied to a variety of data types, including text, images, videos, voice, etc. For example, in natural language processing (NLP) tasks, this module can annotate text for sentiment classification, intent recognition, text summarization, etc.; in computer vision tasks, it can automatically mark object categories, bounding boxes, segmentation masks, etc.

[0058] In order to ensure the accuracy of data labeling, this sub-module also has a built-in manual review mechanism, which can automatically screen low-confidence labeling through confidence scoring, anomaly detection and other means, and submit it for manual review, thereby improving the quality of the final data labeling.

[0059] (5) The security scanning submodule is used to detect security risks in the data processing process to ensure data compliance, privacy protection and system stability.

[0060] The security scanning submodule supports automatic sensitive information identification and can detect whether the data contains protected sensitive content such as personal information (PII), financial data, medical data, etc. At the same time, the module uses differential privacy and federated learning technology to protect user privacy during data analysis and achieve modeling and analysis without exposing the original data.

[0061] In addition, the module can monitor data flows in real time, detect security threats such as malicious attacks, data tampering, and abnormal access, and ensure the security and stability of the entire data processing system.

[0062] In one embodiment disclosed in the present invention, the data analysis module 2 also includes an algorithm recommendation submodule for collecting and storing user operation records, and dynamically generating data processing strategy recommendation information based on user operation records based on machine learning or knowledge graphs.

[0063] The algorithm recommendation submodule can collect and store user operation records to form a complete user behavior data set. These operation record data are derived from various interactive operations of users on the data processing platform, such as: the selected data source type (structured, semi-structured, unstructured), data preprocessing methods (missing value filling, data cleaning, feature engineering, etc.), selected analysis algorithms (clustering, classification, regression, topic modeling, etc.), generated data annotation types (text classification, target detection, time series prediction, etc.), and user processing preferences on different data sets (such as the tendency to use specific models or parameters).

[0064] Operation records are stored in a preset time series database (Time Series Database) or NoSQL database (such as MongoDB, Cassandra) for subsequent analysis and mining.

[0065] In addition, the system can also build detailed user behavior characteristics based on information such as the frequency, duration, and success rate of user operations, providing richer contextual information for personalized recommendations.

[0066] After collecting the user's operation data, the algorithm recommendation submodule uses the machine learning model to perform pattern recognition and dynamically generate recommended information for data processing strategies, such as: (1) Supervised learning and unsupervised learning Supervised learning methods (such as decision trees, random forests, and XGBoost) are used to predict the data processing strategy that users are most likely to adopt. For example, based on historical operation data, it is predicted whether users are more inclined to use K-means or DBSCAN for clustering.

[0067] Unsupervised learning methods (such as K-means and AutoEncoder) are used to discover similar user groups and provide personalized strategy recommendations to new users based on similar users.

[0068] (2) Reinforcement Learning The algorithm recommendation submodule can use the reinforcement learning (RL) mechanism to continuously adjust the recommendation strategy based on the user's historical operations and feedback. For example, the system can use Multi-Armed Bandit or Deep Q-Learning to dynamically select the optimal algorithm recommendation strategy.

[0069] Each recommendation will bring a user feedback (such as whether the user accepts the recommended algorithm), which will serve as a reward signal for reinforcement learning to continuously optimize the recommendation model.

[0070] (3)RAG (Retrieval-Augmented Generation) enhanced recommendation The algorithm recommendation sub-module can combine RAG (retrieval augmented generation) technology to call external knowledge bases in real time, such as the latest research papers, GitHub open source projects, and industry best practices, to recommend more cutting-edge data processing strategies to users.

[0071] For example, the system can automatically recommend new deep learning architectures (such as Transformer, Graph Neural Networks) or optimization methods (such as Self-Supervised Learning) to users based on their operation history.

[0072] In addition to machine learning, the algorithm recommendation sub-module can also use knowledge graph technology to perform semantic correlation analysis on data processing strategies and provide more intuitive and intelligent recommendation solutions.

[0073] Knowledge graphs can be built based on multiple sources, such as public data sources (such as academic paper libraries, corporate databases), user historical operation data (algorithms and parameters used by different users in different scenarios), and data processing best practices (such as excellent solutions for Kaggle competitions).

[0074] After building the knowledge graph, the system can query based on a graph database (such as Neo4j) to achieve intelligent recommendations based on semantic associations.

[0075] Knowledge reasoning can be used to discover potential associations between different data types and optimal algorithms. For example, it can automatically identify that time series data is suitable for LSTM or Prophet, while text data is more suitable for BERT or TF-IDF.

[0076] When users input new data, the system can use graph neural networks (GNNs) to search for the optimal path in the knowledge graph and recommend the best data processing strategy for users.

[0077] Based on machine learning and knowledge graphs, the algorithm recommendation submodule can dynamically generate data processing strategy recommendation information and adjust the recommended content according to real-time conditions.

[0078] For new users, the algorithm recommendation submodule can provide default recommendation solutions based on the behavior patterns of similar users. For long-term users, the system can continuously optimize the recommended algorithms, parameters, and data processing procedures based on their operating habits and feedback.

[0079] The algorithm recommendation submodule can not only recommend algorithms, but also automatically adjust hyperparameters. For example, if the user often uses XGBoost for classification tasks, the system can recommend the optimal learning rate, regularization parameter, and decision tree depth based on historical data.

[0080] In clustering analysis tasks, the algorithm recommendation submodule can dynamically adjust the K value of K-means, or recommend more suitable clustering methods such as DBSCAN (density-based) or GMM (Gaussian mixture model).

[0081] At the same time, users can accept, reject or adjust the recommended strategy, and the system will dynamically update the recommendation model based on user feedback. For example, if the system recommends LSTM for time series forecasting, but the user chooses ARIMA, the system will record the choice and optimize the recommendation logic in similar tasks in the future.

[0082] (III) Data output and application module The data output and application module 3 is used to send the data annotation results to the user and / or other applications, and to generate dynamic interactive data reports according to the required information input by the user.

[0083] The data output and application module 3 includes a model optimization submodule, a data sharing submodule, an application submodule, a data storage submodule, a data preview submodule, an upload and download submodule, and a user management submodule.

[0084] The model optimization submodule can use the processed data to optimize the training of the OpenCSG large model.

[0085] 1. Model optimization submodule The model optimization submodule is responsible for optimizing the training of the OpenCSG large model using the data annotation results to improve the generalization ability and adaptability of the model.

[0086] The model optimization submodule can use online learning or incremental updates to enable the OpenCSG large model to dynamically adapt to new data and improve long-term stability. The model optimization submodule can also use transfer learning to use existing model weights to fine-tune new domain data, reduce training time and improve model performance, and optimize data selection strategies based on user feedback to improve the quality of training data and enable the model to learn more efficiently.

[0087] The model optimization submodule combines the retrieval-augmented generation (RAG) technology to improve the model's ability to understand problems in specific fields. For example, in medical data analysis applications, RAG can be combined with the knowledge base to provide doctors with more accurate diagnostic suggestions.

[0088] The model optimization submodule uses MLFlow and TensorBoard to monitor the model training process, including key indicators such as loss function changes, accuracy, and recall. It provides an automated hyperparameter optimization (HPO) mechanism to optimize the model training effect.

[0089] 2. Data sharing submodule The data sharing submodule can share data annotation results through CSGHub, promoting efficient reuse and collaboration of data resources.

[0090] (1) CSGHub data sharing mechanism Users can publish annotated data on CSGHub for researchers and developers to use. At the same time, it also supports companies and individuals to set data permissions to control who can access, download or modify the data. Each shared dataset will come with a version history to ensure data traceability and consistency.

[0091] (2) Data permissions and access control The data sharing submodule adopts RBAC (role-based access control) and ABAC (attribute-based access control) to ensure data security. It also provides data watermarking and encryption mechanisms to prevent unauthorized abuse.

[0092] (3) API interface support CSGHub provides RESTful API and GraphQL API to facilitate developers to integrate data sharing functions. For example, AI researchers can automatically obtain the latest annotated data through the API and train custom models.

[0093] 3. Application submodule The application submodule can generate reports and / or decision-making information based on the labeled data to help users understand the value of the data more intuitively.

[0094] The application submodule combines data visualization tools (such as ECharts, D3.js, Tableau) to generate interactive data reports. It supports drag-and-drop report customization, and users can freely combine data dimensions to generate personalized reports.

[0095] The application submodule can be combined with the OpenCSG large model to perform intelligent analysis on data and provide advanced functions such as predictive analysis, anomaly detection, trend analysis, etc. For example, in the financial risk control scenario, the system can generate risk assessment reports based on transaction data to assist banks in making decisions.

[0096] The application submodule supports multiple ways to view and download analysis reports, including web, mobile, and custom APIs. It is compatible with smart terminals (such as smart watches and car systems) to expand application scenarios.

[0097] 4. Data storage submodule The data storage submodule is used to store data annotation results and provide version management to ensure data security and traceability.

[0098] The data storage submodule uses distributed storage (such as HDFS, MinIO, Amazon S3) to ensure high data availability and scalability. It provides tiered storage for hot and cold data to optimize storage costs.

[0099] Each time the data is modified, a new version snapshot will be generated, and users can trace back to historical versions. In addition, the data storage submodule also provides data differential storage to reduce storage redundancy and improve efficiency.

[0100] The data storage submodule uses AES-256 encryption to protect stored data, complies with regulations such as GDPR and CCPA, and provides data deletion and anonymization options.

[0101] 5. Data preview submodule The data preview submodule can display data annotation results to users and provide multiple visualization methods.

[0102] The data preview submodule supports multiple data display formats such as tables, charts, 3D views, heat maps, etc., and integrates Jupyter Notebook to support real-time data analysis using Python code.

[0103] The data preview submodule provides SQL query and natural language query (NLQ, Natural Language Query), and users can query data in daily language. For example, if the user enters "show orders with sales exceeding 10,000 yuan in July 2023", the system will automatically generate SQL and return the results.

[0104] 6. Upload and download submodule The upload and download submodule is responsible for receiving data uploaded by users and sending the downloaded data to users to ensure the convenience of data flow.

[0105] The upload and download submodule supports multiple formats (CSV, JSON, Parquet, Excel, XML), and provides batch upload and automatic format detection functions to improve data import efficiency.

[0106] Users can choose compression formats (ZIP, GZIP) to reduce file size. At the same time, streaming download is provided, which is suitable for large-scale data download. In addition, developers can directly upload / download data through the API to realize automated data processing.

[0107] 7. User management submodule The user management submodule can control user permissions and manage user identities to ensure data security and compliance.

[0108] The user management submodule adopts RBAC (role-based access control) and supports different permission levels such as administrators, analysts, auditors, and ordinary users. Access permissions can be set for different data sets, for example, internal enterprise data and external shared data can be managed separately.

[0109] The user management submodule supports OAuth 2.0, LDAP, SAML and other authentication methods, and provides MFA (SMS verification code, Google Authenticator) to improve security.

[0110] In addition, the user management submodule can also record user login, download, and data modification operations, support retrospective analysis, and generate security audit reports to meet corporate compliance requirements.

[0111] (IV) Security and Compliance Module The security and compliance module 4 is used to encrypt data and automatically detect whether the data processing process complies with preset regulatory requirements, including a data encryption submodule, an authority verification submodule and a compliance review submodule.

[0112] 1. Data encryption submodule The data encryption submodule is used to encrypt data to ensure the confidentiality and integrity of data during transmission, storage, and use, and to prevent unauthorized access or tampering.

[0113] The data encryption submodule uses AES-256 symmetric encryption to ensure data storage security and prevent external attacks from stealing data. It also uses RSA and ECC public key encryption to ensure the security of data during transmission.

[0114] The data encryption submodule supports the TLS1.3 encryption protocol to ensure that data will not be attacked by man-in-the-middle (MITM) when communicating with the API interface. For cloud storage (such as AWS S3, Google Cloud Storage), the data encryption submodule provides object-level encryption, and only authorized users can decrypt data.

[0115] The data encryption submodule supports symmetric + asymmetric hybrid encryption, combining AES and RSA to improve performance while taking into account security, and adopts blockchain-based distributed key management (KMS, Key Management System) to ensure that keys are securely stored and not easily cracked.

[0116] 2. Permission verification submodule The permission verification submodule is used to verify user permissions to ensure that only authorized users can access, modify or manage data to avoid data leakage or abuse.

[0117] The permission verification submodule assigns permissions according to roles, such as administrator, data analyst, auditor, and ordinary user, to avoid abuse of permissions. For example, ordinary users can only view data, while data analysts can download data and administrators can modify data permissions.

[0118] The permission verification submodule performs permission control based on user attributes (such as IP address, geographic location, and device type) to improve security. For example, users can only access sensitive data within the company's internal IP range.

[0119] The permission verification submodule supports 2FA (two-factor authentication), combining password + SMS verification code / Google Authenticator for identity authentication.

[0120] The permission verification submodule records detailed logs of user login, access, and data modification, supports audit tracking, and uses AI anomaly detection algorithms to identify suspicious behaviors (such as users accessing a large amount of sensitive data in a short period of time) and trigger security alerts.

[0121] 3. Compliance review submodule The compliance review submodule is used to detect whether the data processing process complies with preset regulations and generate a compliance review report to ensure that data processing complies with national, industry, and corporate compliance standards.

[0122] The compliance review submodule provides automated compliance checks based on preset laws and regulations to ensure that the data processing process complies with international standards. It also automatically detects compliance with relevant regulations when collecting, storing, and sharing data, such as whether the data is anonymized and whether the user has authorized the use of the data.

[0123] The compliance review submodule supports automatic generation of compliance reports, including information such as data processing flow, access logs, encryption policies, audit logs, etc. For cross-border data transmission, it automatically checks whether it complies with the data protection regulations of various countries.

[0124] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. An OpenCSG intelligent data processing platform, characterized in that: include: The data collection and integration module is configured to obtain data from various data sources and standardize and clean the data according to the data format and structure; The data analysis module is configured to analyze the data using the OpenCSG large model technology and generate data annotation results; The data analysis module includes a model training and tuning submodule, an entity extraction submodule, a data clustering submodule, a data annotation submodule and a security scanning submodule, wherein: The model training and tuning submodule is configured to train and fine-tune the LLM large language model based on the data obtained by the data acquisition and integration module, and to optimize the response and decision-making capabilities of the LLM model using the RAG retrieval enhancement generation technology to obtain the OpenCSG large model; The entity extraction submodule is configured to extract key entities in the data; The data clustering submodule is configured to perform cluster analysis on data characteristics; The data annotation submodule is configured to automatically annotate data based on the OpenCSG large model technology; The security scanning submodule is configured to check for potential security risks during data processing; The data output and application module is configured to send the data annotation results to the user and / or external application, and to generate dynamic interactive data reports according to the required information input by the user; The security and compliance module is configured to encrypt data and automatically detect whether the data processing process complies with preset regulatory requirements.

2. The platform according to claim 1, characterized in that The data acquisition and integration module includes a data acquisition submodule and a preprocessing submodule, wherein: The data collection submodule is configured to crawl data from CSGHub, preset datasets, and preset networks; The preprocessing submodule is configured to clean the data, convert the format, detect abnormal data, and standardize the data.

3. The platform according to claim 2, characterized in that: The preprocessing submodule includes a standardization unit, a data conversion unit, a data cleaning unit, a data enhancement unit and a review and evaluation unit, wherein: The standardization unit is configured to encode the data and unify the format; The data conversion unit is configured to convert the data into a preset format suitable for model training; The data cleaning unit is configured to delete abnormal data and fill in missing data; The data enhancement unit is configured to perform enhancement processing on the data using a preset enhancement method; The review and evaluation unit is configured to detect data quality and generate data optimization information.

4. The platform according to claim 1, characterized in that The data analysis module also includes an algorithm recommendation submodule, which is configured to collect and store user operation records, and based on machine learning or knowledge graphs, use the user's operation records to dynamically generate data processing strategy recommendation information.

5. The platform according to claim 1, characterized in that: The data output and application module includes a model optimization submodule, a data sharing submodule, an application submodule, a data storage submodule, a data preview submodule, an upload and download submodule, and a user management submodule, wherein: The model optimization submodule is configured to optimize the training of the OpenCSG large model using the data annotation results; The data sharing submodule is configured to share data annotation results through CSGHub; An application submodule, configured to generate reports and / or decision information based on the data annotation results; The data storage submodule is configured to store data annotation results and provide version management; A data preview submodule is configured to display data annotation results to users; The upload and download submodule is configured to receive data uploaded by the user and send the downloaded data to the user; The user management submodule is configured to control user permissions and manage user identities.

6. The platform according to claim 1, characterized in that: The Security and Compliance Module includes: A data encryption submodule is configured to encrypt data before transmission; The permission verification submodule is configured to verify the user's permission; The compliance review submodule is configured to detect whether the data processing process complies with preset regulations and generate a compliance review report.

7. The platform according to claim 2, characterized in that: The preprocessing submodule includes an abnormal data detection unit configured to identify abnormal data.

Citation Information

Patent Citations

  • OpenCSG large model intelligent data annotation system

    CN119807414A

  • System for financial analysis based on artificial intelligence with easy report interpretations

    WO2024261779A2

Cited By

  • Artificial intelligence large model data set management method and system

    CN120578711A

  • System security and data protection mechanism method

    CN120930171A

  • A system security and data protection mechanism method

    CN120930171B