Computer data management system based on big data
By employing a distributed data acquisition, intelligent storage scheduling, and multi-level security protection system, combined with a deep reinforcement learning model, the existing data management system addresses issues such as insufficient data value mining, rigid resource scheduling, and lagging security and compliance in multi-source heterogeneous data processing, thereby achieving efficient data management and autonomous decision-making optimization.
Patent Information
- Application Number
- CN202510837157.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing data management systems suffer from insufficient data value mining, rigid resource scheduling, and lagging security and compliance issues when processing multi-source heterogeneous data, especially in high-frequency real-time analysis, storage cost optimization, and dynamic security protection.
It employs a distributed data acquisition module, an intelligent storage scheduling engine, a fusion computing framework, and a multi-level security protection system, combined with a deep reinforcement learning model and a dynamic policy hub, to achieve cross-modal data association, dynamic resource scheduling, and adaptive security protection.
It improved the completeness of data value mining, optimized resource utilization, reduced storage costs and security risks, and achieved closed-loop optimization of the system's autonomous decision-making.
Smart Images

Figure HDA0005461033510000011
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of big data processing, and more particularly, to a computer data management system based on big data. BACKGROUND
[0002] In the current big data application scenario, enterprises need to process multi-source heterogeneous data (including log stream, Internet of Things sensor data and unstructured documents), and support real-time business decision-making. The existing data management system usually adopts Lambda architecture to realize batch-flow unified processing, and relies on static rule configuration to store the strategy and security mechanism. With the explosive growth of data volume and the increasing compliance requirements, traditional systems face serious challenges in high-frequency real-time analysis, storage cost optimization and dynamic security protection. The main shortcomings of the existing technology are as follows:
[0003] 1. Inadequate data value mining
[0004] The existing system lacks deep semantic analysis capability for unstructured data (such as user comments, image description text), and only relies on keyword matching to generate indexes, resulting in broken cross-modal data association. For example, in the e-commerce scenario, user behavior logs and product description texts cannot be effectively associated, resulting in missing feature dimensions of personalized recommendation models. At the same time, cold and hot data recognition relies on artificial preset rules and cannot adapt to business traffic fluctuations. High-frequency access to historical order data may be stored in a low-speed storage layer, significantly increasing query delay.
[0005] 2. Resource scheduling rigidity
[0006] In the traditional architecture, the flow processing and batch processing resource pools are physically isolated. In the event of a sudden traffic scenario, real-time computing resources are insufficient, and offline computing nodes cannot be dynamically allocated to assist (such as GPU resources), causing flow data to accumulate. On the other hand, storage strategy adjustment relies on the experience of operation and maintenance personnel and cannot automatically optimize hierarchical strategies according to changes in data access patterns, resulting in invalid occupation of high-performance storage layers by low-frequency data, increasing storage costs by more than 30%.
[0007] 3. Security and compliance lag
[0008] Static data desensitization rules cannot adapt to changing access scenarios (such as different desensitization strengths required for internal audits and external API calls), often resulting in excessive desensitization affecting analysis or insufficient desensitization leading to information leakage. In addition, regulatory policy updates (such as GDPR cross-border transmission restrictions) need to be manually configured into the system, and there is a policy implementation window period. Access control lacks behavior pattern analysis capability, and abnormal batch data export behavior relies only on fixed threshold alarms, with a false alarm rate of up to 40%.
[0009] Therefore, a computer data management system based on big data is proposed to solve the above problems. SUMMARY
[0010] To overcome the above-mentioned defects of the prior art, embodiments of the present application provide a big data-based computer data management system to solve the problems raised in the above background.
[0011] To achieve the above object, the present application provides the following technical scheme: a big data-based computer data management system, comprising:
[0012] A distributed data acquisition module is configured to receive structured data streams through an extensible API gateway, while deploying a distributed web crawler to capture unstructured data, and embedding an integrity check label in the data injection stage;
[0013] An intelligent storage scheduling engine adopts a hybrid architecture of columnar storage database and distributed object storage, with an access frequency analyzer built-in to realize automatic migration of hot and cold data;
[0014] A fusion computing framework integrates a stream processing engine and a batch processing computing cluster, and coordinates real-time feature extraction and offline data processing through a unified task scheduler;
[0015] An adaptive strategy hub generates data partition rules, compression levels and encryption schemes based on a deep reinforcement learning model, and its decision is based on continuous updates of storage load monitoring indicators;
[0016] A multi-level security protection system deploys a dynamic decryption unit based on attributes at the data access interface layer, and associates a real-time risk perception module to trigger an access interception mechanism.
[0017] Preferably, the distributed data acquisition module includes an intelligent throttling controller and a semantic enhancement processing unit, the intelligent throttling controller dynamically adjusts the number of thread concurrency according to the data source response delay and error rate, and starts an exponential backoff acquisition strategy when continuous timeout errors occur, and the semantic enhancement processing unit is to parse entity attributes in text data through a pre-trained language model, and to construct a cross-modal index graph associated with structured data.
[0018] Preferably, the intelligent storage scheduling engine includes: a data value evaluator is based on a time series prediction algorithm to analyze the access pattern in the past three months, and preloads the predicted high-frequency access data set to the memory acceleration layer, an elastic redundancy configurator automatically switches the redundancy mode according to the data sensitivity level, and implements three-copy storage for key business data and erasure code compression storage for archived data.
[0019] Preferably, the fusion computing framework implements a resource dynamic routing mechanism and an incremental view updater, the resource dynamic routing mechanism automatically allocates idle batch processing nodes to join the real-time computing cluster when the flow processing throughput exceeds a threshold, and the incremental view updater merges the incremental results generated by the flow processing with the full-amount data set of the batch processing in a time window to generate a consistent business view.
[0020] Preferably, the adaptive policy hub includes a policy sandbox verification environment and a compliance adapter, the policy sandbox verification environment performs storage policy changes by simulating real data load, collects the I / O delay change rate as a policy performance evaluation index, and the compliance adapter regularly acquires external data regulation updates and automatically prohibits data cross-border transmission strategies that conflict with existing regulations.
[0021] Preferably, the multi-level security protection system further includes: a behavior pattern analysis engine constructs a user operation baseline model through session logs, triggers secondary authentication for large batch export operations at irregular times, and a context-aware desensitizer dynamically adjusts the masking strength of sensitive fields according to the access terminal type and network environment security rating.
[0022] Preferably, it further includes: a containerized resource coordinator predicts future 2-hour computing load based on a sliding window, and deploys or recycles Docker container instances in advance; and a cost-aware scheduler is associated with a multi-cloud service price API, and preferentially schedules to the most cost-effective computing area under the premise of guaranteeing task SLA.
[0023] Preferably, it is further integrated with: a full-life-cycle meta database and a data health monitor, the full-life-cycle meta database records the blood relationship and change trajectory at the field level of the data table, and supports version backtracking operation, and the data health monitor periodically scans storage nodes to check data block hash values, and triggers cross-node replica repair for failed data.
[0024] Preferably, the adaptive policy hub adopts a multi-dimensional reward and punishment model and a cross-scenario policy migration module, the multi-dimensional reward and punishment model synchronously optimizes storage space utilization, query response percentile, and number of security vulnerabilities during reinforcement learning training, and the cross-scenario policy migration module extracts policy feature vectors of existing business scenarios as generation constraint conditions for initial policies of new businesses.
[0025] Preferably, it further includes a policy visualization governance platform and a policy impact prediction module, the policy visualization governance platform provides a policy rule topology graph display interface and supports administrators to drag and modify data sharding policy parameters, and the policy impact prediction module simulates and calculates the expected storage cost saving rate and query performance fluctuation range before policy deployment to generate a risk assessment matrix.
[0026] The technical effects and advantages of the present application are as follows:
[0027] 1. Full-dimensional data value release
[0028] The semantic enhancement processing unit of the distributed data collection module deeply analyzes unstructured text, constructs a cross-modal index map, and opens up the correlation channels of e-commerce comments, product descriptions, and other multi-source data, making the feature dimension completeness of the recommendation model improve by 60%. Combined with the heat prediction and automatic layering of the intelligent storage scheduling engine, the response delay of high-frequency access data is compressed to within 200ms, while reducing cold data storage costs by 45%, achieving dynamic balance between data value mining and resource costs.
[0029] 2. Elastic resource intelligent scheduling
[0030] Relying on the resource dynamic routing mechanism of the fusion computing framework, batch processing GPU resources are automatically scheduled to join real-time computing during peak flow data periods, ensuring that more than 95% of flow processing tasks are completed within the SLA time limit. In cooperation with the load prediction capability of the containerized resource coordinator, container instances are scaled up or down 2 hours in advance as needed, avoiding resource idleness and waste, and making the utilization rate of computing resources in sudden traffic scenarios improve to more than 85%, reducing the need for operational intervention by 70%.
[0031] 3. Dynamic security compliance protection
[0032] Based on the context-aware desensitizer of the multi-level security protection system, the masking strength is dynamically adjusted according to the access scenario (such as displaying the complete mobile phone number for internal audit and only displaying the first 3 digits for external API), making the conflict between data availability and security decrease by 90%. In cooperation with the compliance adapter, the regulation library is synchronized in real time, automatically intercepting illegal data transmission requests, and shortening the delay of the effectiveness of compliance policies such as GDPR from the traditional 72 hours to instant effectiveness, significantly reducing the legal risk of enterprises.
[0033] 4. Self-determination decision loop optimization
[0034] The multi-dimensional reward and punishment model of the adaptive strategy hub synchronously optimizes storage cost, query performance, and security indicators, shortening the storage strategy adjustment period from weekly to hourly. Through the strategy impact prediction module, administrators can optimize strategy parameters on the visual interface, reducing decision-making error rate by 80%. The system continuously evolves through strategy sandbox verification and reinforcement learning, forming a complete autonomous closed loop of "perception - decision - verification - optimization". BRIEF DESCRIPTION OF DRAWINGS
[0035] Figure 1 The system framework diagram of the present application. DETAILED DESCRIPTION
[0036] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0037] As shown in the accompanying drawings, Figure 1 (1) a computer data management system based on big data, comprising:
[0038] A distributed data acquisition module is configured to receive structured data streams through an extensible API gateway, while deploying a distributed web crawler to capture unstructured data, and embedding an integrity check tag in the data injection stage;
[0039] An intelligent storage scheduling engine adopts a hybrid architecture of columnar storage database and distributed object storage, and is built-in with an access frequency analyzer to realize automatic migration of hot and cold data;
[0040] A fusion computing framework integrates a stream processing engine and a batch processing computing cluster, and coordinates real-time feature extraction and offline data processing through a unified task scheduler;
[0041] An adaptive strategy hub generates data partition rules, compression levels and encryption schemes based on a deep reinforcement learning model, and its decision is based on the continuous update of storage load monitoring indicators;
[0042] A multi-level security protection system is provided, which deploys an attribute-based dynamic decryption unit at a data access interface layer and associates a real-time risk perception module to trigger an access interception mechanism. A distributed data collection module receives order logs (JSON format) pushed by Kafka through an Nginx-built API gateway cluster, and a Scrapy crawler cluster is deployed to crawl e-commerce comments (HTML text). An MD5 check tag is added when data is injected. An intelligent storage scheduling engine is combined with a ClickHouse column store database and a MinIO object store. A Python script built-in scans HDFS access logs every 5 minutes, and data sets with access times > 100,000 in the past 24 hours are migrated to an SSD storage pool. A fusion computing framework uses Flink to consume Kafka data streams in real time to perform order amount statistics, and schedules Spark batch processing to calculate user portraits every morning. Both share YARN resource queues. A self-adaptive strategy hub uses TensorFlow to implement a PPO reinforcement learning model, which takes storage utilization (0-100%), average query delay (milliseconds), and security alert times as input parameters, and outputs the best data shard number (such as 128 pieces) and an AES-256 encryption identifier. The multi-level security protection system integrates a CP-ABE encryption library at the RESTful interface layer, and when the risk score of the access request source IP is > 7 (full score 10), dynamic desensitization is triggered.
[0043] (2) The distributed data collection module includes an intelligent throttling controller and a semantic enhancement processing unit. The intelligent throttling controller dynamically adjusts the number of thread concurrency according to the response delay and error rate of the data source, and starts an exponential backoff collection strategy when continuous timeout errors occur. The semantic enhancement processing unit parses entity attributes in text data through a pre-trained language model and constructs a cross-modal index graph associated with structured data. An adaptive frequency controller uses Prometheus to monitor data source HTTP status codes. When 500 error codes occur for 10 consecutive collections and the average delay is > 3 seconds, the collection interval is adjusted from 1 minute to: new interval = original interval x (1 + 0.3 x random floating factor), with a maximum of 10 minutes. The semantic analysis unit loads the bert-base model of HuggingFace to perform named entity recognition on comment text (such as "Apple phone" -> product entity), associates the identified entity with the product_id field of the MySQL order table, and generates a Neo4j graph query statement: MATCH(user)-[:PURCHASED]->(product) WHERE product.name="Apple phone" RETURN user.id.
[0044] (3) The intelligent storage scheduling engine includes: a data value evaluator based on a time series prediction algorithm analyzes the access pattern in the past three months, preloads the high-frequency access data set to the memory acceleration layer, and an elastic redundancy configurator automatically switches the redundancy mode according to the data sensitivity level, implements three-copy storage for key business data, and adopts erasure code compression storage for archived data, wherein the heat-aware layering device adopts Keras to build an LSTM model, the input feature dimension includes: past 7-day access times (integer), recent 1-hour access growth rate (percentage), and business label priority (1-5 level), and the output 72-hour access heat value (0.0-1.0); when the predicted value > 0.8, the Kubernetes API is called to mount the data Pod to the Intel Optane memory disk, and when the predicted value < 0.3, the AWS CLI is triggered to convert the S3 storage bucket to GLACIER cold storage; the cross-domain redundancy manager scans the HBase table field content, and if more than 5 fields of credit card numbers (regular matching \d{16}) are detected, the 3-copy strategy is enabled in the Ceph storage cluster, otherwise the erasure code strategy (10 data blocks + 4 check blocks) is adopted.
[0045] (4) The fusion computing framework implements a resource dynamic routing mechanism and an incremental view updater, the resource dynamic routing mechanism automatically allocates idle batch processing nodes to join the real-time computing cluster when the stream processing throughput exceeds the threshold, and the incremental view updater aligns and merges the incremental results generated by the stream processing with the batch full data set in the time window to generate a consistent business view, wherein the resource dynamic allocation module monitors the Flink Checkpoint queue backlog, when the backlog > 500,000, the YARN API is called to reduce the number of Spark job executors by 50%, and the released containers are relabeled as flink-taskmanager to join the real-time cluster; the incremental-full collaborative device designs a time window alignment algorithm: at 02:00 every day, Spark SQL calculates the full amount of user purchases R_full, and at the same time, the incremental orders ΔR of Flink stream processing are aggregated every hour, and the verification formula when merging is: | (current full value + today's incremental sum) - yesterday's full value | / yesterday's full value < 0.001, if the threshold is exceeded, an alarm is triggered.
[0046] (5) The adaptive policy hub includes a policy sandbox verification environment and a compliance adapter. The policy sandbox verification environment performs storage policy changes by simulating real data load, collects I / O delay change rate as a policy performance evaluation indicator, and the compliance adapter is a method of regularly obtaining external data regulation updates and automatically disabling data cross-border transmission policies that conflict with existing regulations. The policy simulator deploys a ClickHouse test instance in a separate Docker container, loads 10% of the production data sample, performs policy changes (such as increasing the number of shards from 64 to 128), and records the query delay change rate = (new delay - old delay) / old delay. The compliance checker calls the EU GDPR website RSS feed at 00:00 every Monday, and if the keyword "data transfer restriction" is parsed, it iterates through existing policies, adds a freeze tag to policies that match "cross-border transmission = allowed", and notifies the enterprise WeChat robot through a webhook.
[0047] (6) The multi-level security protection system further includes: a behavior pattern analysis engine builds a user operation baseline model through session logs, triggers secondary authentication for large-scale export operations at irregular times, and a context-aware desensitizer dynamically adjusts the masking strength of sensitive fields according to access terminal types and network environment security ratings. The behavior tracing component collects 30-day operation logs to build a baseline model, calculates the mean daily query frequency μ = 85 times and the standard deviation σ = 12 times, and when the user's single-hour query frequency is greater than μ + 3σ (i.e. 121 times), it jumps to the Keycloak identity authentication service and requires face recognition. The dynamic desensitization engine configures a rule chain: when the request header contains X-Access-Scenario = external_api, apply the regular replacement rule: mobile.replaceAll("(\\d{3})\\d{4}(\\d{4})","$1****$2").
[0048] (7) It also includes: a containerized resource coordinator that predicts future 2-hour load based on a sliding window and deploys or recycles Docker container instances in advance, and a cost-aware scheduler that associates with multi-cloud service price APIs to prioritize scheduling to the most cost-effective computing area while ensuring task SLA. The elastic scaling controller predicts CPU load based on the Prophet time series algorithm, inputs the load curve for the past 7 days, and outputs the predicted value for the next 2 hours. When the predicted peak value is greater than 80%, call the Kubernetes API to expand the Deployment replica number from 10 to 15. The cost optimizer integrates AWS / Azure price APIs and calculates the formula: unit task cost vCPU hour unit price x estimated execution time. When the analysis task priority is "medium", schedule it to the us-east-1 region with the lowest current unit price.
[0049] (8) Further integration: a full life cycle metadata database that records data table field level blood relationship and change trajectory, supports version backtracking operation, and a data health monitor that periodically scans storage nodes to check data block hash values and triggers cross-node copy repair for failed data, wherein the unified metadata center uses Apache Atlas to store data bloodlines and record Hive table field level bloodline SQL: CREATE TABLE order_summary AS SELECT user_id, sum(amount) FROM orders; The automatic repair agent executes the HDFS fsck command every 6 hours, copies a healthy copy from the same rack copy node for data blocks that return a CORRUPT state, and if repair fails for 3 consecutive times, an alarm is raised.
[0050] (9) The adaptive strategy hub uses a multi-dimensional reward and punishment model and a cross-scenario strategy migration module, the multi-dimensional reward and punishment model synchronously optimizes storage space utilization, query response percentile and number of security vulnerabilities during reinforcement learning training, and the cross-scenario strategy migration module extracts the strategy feature vector of an existing business scenario as a generation constraint condition for the initial strategy of a new business, wherein the multi-objective optimization model defines the reward function: reward value = 0.4 × (1 - current storage cost / budget upper limit) + 0.4 × (1 - P99 delay / 1000 ms) - 0.2 × number of security vulnerabilities; The transfer learning module freezes the weights of the fully connected layer of the strategy model trained for the logistics business, and only fine-tunes the input layer to adapt to the feature dimension of the retail business.
[0051] (10) It also includes a strategy visualization governance platform and a strategy impact prediction module, the strategy visualization governance platform provides a strategy rule topology graph display interface, supports administrators to drag and modify data sharding strategy parameters, and the strategy impact prediction module simulates and calculates the expected storage cost saving rate and query performance fluctuation range before strategy deployment, and generates a risk assessment matrix, wherein the strategy visualization governance platform uses Echarts to draw a strategy topology graph, and when the administrator drags the sharding slider from 128 to 256, the backend immediately simulates execution: calculate the storage cost change = (new number of shards / original number of shards) × original cost × expansion factor 1.2; The strategy impact prediction module outputs the risk matrix: high cost risk: marked red when the cost increase is > 15%, query performance risk: marked yellow when P99 delay > 800 ms.
[0052] Embodiment one:
[0053] Step 1: Multi-source data collection and enhancement processing
[0054] ① Dynamic frequency control:
[0055] Data collection module scans each data source status (response latency, error rate) every 5 minutes
[0056] When the error rate is > 5% for 3 consecutive times, start the exponential backoff strategy:
[0057] Next collection interval = current interval x (1 + random coefficient [0.2, 0.5])
[0058] Reset to the baseline frequency after the network is stable (error rate < 1% for 10 minutes)
[0059] ② Cross-modal index construction:
[0060] Unstructured text extracts entities (product name / user ID) through BERT pre-training model
[0061] Correlate structured data fields (such as order ID) to build a triple index:
[0062] {Entity A, Relationship R, Entity B} → {User U123, Purchase, Product G789}
[0063] Index written to graph database Neo4j, supporting multi-hop queries
[0064] Step 2: Intelligent storage hierarchy and redundancy
[0065] ① Hotness prediction model:
[0066] Use LSTM time series model, input features:
[0067] [7-day access frequency, 24-hour change rate, business priority label]
[0068] Output future 72-hour access hotness score (0-100 points), updated every hour
[0069] Execution strategy:
[0070] IF score > 80 → store in memory acceleration layer (Redis cluster)
[0071] IF 30 ≤ score ≤ 80 → store in SSD columnar storage (Apache Parquet)
[0072] IF score < 30 → migrate to object storage (AWS S3)
[0073] ② Elastic redundancy configuration:
[0074] Data sensitivity analyzer scans field content (including regular matching of ID cards / bank cards, etc.)
[0075] Highly sensitive data (matched to 5 or more sensitive fields) automatically enables triple-replica storage
[0076] Low-sensitive data adopts Reed-Solomon error correction code (10+4 strategy: 10 data blocks + 4 check blocks)
[0077] ① Dynamic resource routing:
[0078] Real-time monitoring of Kafka topic backlog (Pending Messages)
[0079] When the backlog exceeds the threshold (such as > 50,000), perform:
[0080] 1. Suspend low-priority batch processing tasks
[0081] 2. Mark the released GPU container as "streaming available"
[0082] 3. Register new TaskManager to Flink cluster
[0083] ② Incremental view update:
[0084] Stream processing engine outputs incremental result ΔR (such as new order amount) every 10 seconds
[0085] Batch processing engine performs full calculation R_full every morning
[0086] Merge algorithm:
[0087] Current view V_t = R_full + Σ(ΔR_{t-23:59} to ΔR_t)
[0088] Check condition: |V_t - (V_{t-1} + ΔR_t)| < 0.1%
[0089] Step 4: Strategy generation and verification
[0090] ① Reinforcement learning policy iteration:
[0091] State space: Contains 12-dimensional vector including storage load rate, query P99 delay, security event count, etc. Action space: {partition number, compression algorithm selection, encryption key length}
[0092] Reward function:
[0093] Reward = α*(1 - storage cost ratio) + β*(1 - delay normalized value) - γ*risk exposure
[0094] (α = 0.4, β = 0.4, γ = 0.2 are empirical weight coefficients)
[0095] Training process: Simulate 1000 times of policy execution in sandbox environment, select the policy with the highest Q value
[0096] ② Automatic compliance adaptation:
[0097] Crawl GDPR and CCPA regulatory update pages at 0:00 every Monday
[0098] Keyword matching (such as "cross-border transfer prohibition") triggers strategy reconstruction
[0099] Conflict policy freeze process:
[0100] 1. Mark the original policy as "invalid"
[0101] 2. Generate a temporary basic policy (all cross-border transfers are prohibited by default)
[0102] 3. Send an alert to the administrator console
[0103] Step 5: Dynamic Security Protection
[0104] ①Behavioral baseline modeling:
[0105] Collect 30 days of normal operation logs (user A queries ≤ 100 times per day)
[0106] Build a Gaussian distribution model:
[0107] Normal range of operation frequency = mean μ ± 3 times standard deviation σ
[0108] When user B's single-hour query is greater than μ+3σ, the face recognition secondary authentication is triggered.
[0109] ②Context desensitization rules:
[0110] Mobile phone number display rules for access scenarios
[0111] Internal financial audit complete display (138****1234→1380011234)
[0112] External API call
[0113] Keep the first 3 digits (138********)
[0114] use
[0115] High-risk network environment
[0116] Completely shielded (***********)
[0117] Finally should be explained a few points are: first, in the description of the present application, it should be pointed out that, unless otherwise specified and limited, the term "installation", "connected", "connection" should be broad, can be mechanical or electrical connection, but also can be two elements inside the communication, can be directly connected, "up", "down", "left", "right" and so on, only for indicating the relative position relationship, when the absolute position of the described object changes, the relative position relationship may change;
[0118] Second: the present application discloses the embodiment in the drawing, only relates to the structure involved in the present application, other structures can refer to the usual design, in the case of no conflict, the same embodiment and different embodiments of the present application can be combined with each other;
[0119] Finally: the above only for the preferred embodiment of the present application, and not for limiting the present application, any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application, should be included in the protection scope of the present application.
Claims
1. A computer data management system based on big data, characterized in that: include: A distributed data collection module, configured to receive structured data streams through a scalable API gateway, deploy distributed web crawlers to capture unstructured data, and embed integrity check tags during the data injection phase; An intelligent storage scheduling engine uses a hybrid architecture of column-based storage database and distributed object storage, with a built-in access frequency analyzer to achieve automatic migration of hot and cold data; A fusion computing framework that integrates a stream processing engine and a batch computing cluster, coordinating real-time feature extraction and offline data processing through a unified task scheduler; Adaptive strategy center, which generates data partitioning rules, compression levels, and encryption schemes based on deep reinforcement learning models. Its decisions are continuously updated based on storage load monitoring indicators. A multi-level security protection system deploys an attribute-based dynamic decryption unit at the data access interface layer and associates it with a real-time risk perception module to trigger an access interception mechanism.
2. A computer data management system based on big data according to claim 1, characterized in that: The distributed data acquisition module includes an intelligent throttling controller and a semantic enhancement processing unit. The intelligent throttling controller dynamically adjusts the number of concurrent threads according to the data source response delay and error rate, and starts an exponential backoff acquisition strategy when timeout errors occur continuously. The semantic enhancement processing unit parses entity attributes in text data through a pre-trained language model to construct a cross-modal index map associated with structured data.
3. A computer data management system based on big data according to claim 1, characterized in that: The intelligent storage scheduling engine includes: a data value evaluator that analyzes the access patterns of the past three months based on a time series prediction algorithm and preloads the predicted high-frequency access data set into the memory acceleration layer; an elastic redundancy configurator that automatically switches the redundancy mode according to the data sensitivity level, implements three-copy storage for key business data, and uses erasure code compression storage for archived data.
4. A computer data management system based on big data according to claim 1, characterized in that: The fusion computing framework implements a dynamic resource routing mechanism and an incremental view updater. When the stream processing throughput exceeds a threshold, the dynamic resource routing mechanism automatically allocates idle batch processing nodes to join the real-time computing cluster. The incremental view updater aligns and merges the incremental results generated by stream processing with the full batch processing data set in a time window to generate a consistent business view.
5. A computer data management system based on big data according to claim 1, characterized in that: The adaptive policy hub includes a policy sandbox verification environment and a compliance adapter. The policy sandbox verification environment executes storage policy changes by simulating real data loads and collects the I / O delay change rate as a policy effectiveness evaluation indicator. The compliance adapter regularly obtains external data regulation updates and automatically prohibits cross-border data transmission policies that conflict with current regulations.
6. A computer data management system based on big data according to claim 1, characterized in that: The multi-level security protection system further includes: a behavioral pattern analysis engine constructing a user operation baseline model through session logs, triggering secondary authentication for large-scale export operations at unconventional times, and a context-aware desensitizer dynamically adjusting the masking intensity of sensitive fields based on the access terminal type and network environment security rating.
7. A computer data management system based on big data according to claim 1, characterized in that: Also includes: The containerized resource coordinator predicts the computing load for the next two hours based on a sliding window and deploys or recycles Docker container instances in advance. The cost-aware scheduler associates the multi-cloud service price API and prioritizes scheduling to the cost-optimal computing area while ensuring the task SLA.
8. A computer data management system based on big data according to claim 1, characterized in that: Further integration: full life cycle metadata database and data health monitor. The full life cycle metadata database records the lineage relationship and change trajectory at the data table field level and supports version backtracking operations. The data health monitor periodically scans the storage node to verify the data block hash value and triggers cross-node replica repair for data that fails verification.
9. A computer data management system based on big data according to claim 1, characterized in that: The adaptive strategy center adopts a multi-dimensional reward and punishment model and a cross-scenario strategy migration module. The multi-dimensional reward and punishment model simultaneously optimizes storage space utilization, query response percentile and the number of security vulnerabilities during reinforcement learning training. The cross-scenario strategy migration module extracts the strategy feature vector of the existing business scenario as the generation constraint condition of the initial strategy of the new business.
10. A computer data management system based on big data according to claim 1, characterized in that: It also includes a policy visualization governance platform and a policy impact prediction module. The policy visualization governance platform provides a policy rule topology display interface, supporting administrators to drag and drop to modify data sharding policy parameters. The policy impact prediction module simulates and calculates the expected storage cost savings rate and query performance fluctuation range before policy deployment to generate a risk assessment matrix.
Citation Information
Cited By
Automobile data cross-border detection system based on multi-source data fusion
CN121167218A
Large-scale education data migration method based on dynamic service routing
CN121478745A
Data governance strategy dynamic execution system based on active metadata and AI recommendation
CN121478755A