Enterprise informatization management integrated platform based on big data

By building an enterprise information management integration platform based on Hadoop, the problems of data silos, system scalability, and low cross-departmental collaboration efficiency have been solved, achieving efficient data processing, resource scheduling, and intelligent analysis, thereby improving the efficiency of enterprise information management.

CN119806724BActive Publication Date: 2025-11-04SHANDONG ZHONGRUN INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202411785976.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-11-04
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Enterprise information management suffers from problems such as data silos, insufficient system scalability, low efficiency of cross-departmental collaboration, and limited data analysis capabilities, which prevent the full realization of data value and lead to a lack of scientific and accurate management decisions.

Method used

An enterprise information management integration platform is built on the Hadoop ecosystem. It adopts a distributed computing and storage architecture, including a parallel data processing hub, an intelligent data mapping engine, an elastic resource scheduler, a distributed storage center, and a business analysis engine. Through technologies such as swarm intelligence algorithms, adaptive wavelet compression, biological cell division mechanisms, adaptive grid partitioning technology, and multi-scale feature fusion, it achieves efficient and unified data processing, resource scheduling, and analysis.

Benefits of technology

It improves data processing efficiency, enhances system scalability, optimizes data sharing mechanisms, strengthens analysis and mining capabilities, improves storage access performance, and supports the efficient utilization of enterprise data resources and intelligent management decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806724B_ABST
    Figure CN119806724B_ABST
Patent Text Reader

Abstract

The application provides an enterprise informatization management integrated platform based on big data, which is constructed based on a Hadoop distributed computing framework, integrates MapReduce parallel processing, Spark real-time computing, HDFS distributed storage and other technologies, realizes data parallel processing through a swarm intelligence algorithm and a pheromone transmission mechanism, carries out heterogeneous data mapping by using a hierarchical self-encoding network, realizes resource dynamic scheduling in combination with a biological cell division mechanism and an entropy value balance mechanism, optimizes storage management by using adaptive grid partitioning and data temperature layering technology, and constructs multi-scale feature fusion and knowledge reasoning based on Spark MLlib to support decision analysis. The platform overcomes problems in traditional enterprise informatization management, such as low data processing efficiency, limited system expansion and difficult data sharing, and provides reliable technical support for enterprise digital transformation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of enterprise information management, and particularly relates to an enterprise information management integrated platform based on big data. BACKGROUND

[0002] Enterprise information management is an important support for enterprise development. In recent years, with the acceleration of digital transformation, enterprises have generated massive data in production, management, sales and other links. How to efficiently manage and utilize these data has become the core challenge of enterprise informationization construction.

[0003] In terms of data processing, the current enterprise information management is difficult to realize efficient data collection, storage and analysis, and the data value cannot be fully utilized. The lack of unified data standards and management mechanisms leads to data islands between business systems, making data sharing difficult and unable to meet the needs of enterprise fine management.

[0004] In terms of system architecture, the existing enterprise information management platform generally adopts traditional architecture design, which lacks scalability. With the expansion of enterprise scale and business volume, the system performance is difficult to cope with the increasing data processing demand. In addition, the integration between systems is difficult, the maintenance cost is high, which limits the ability of enterprises to respond to rapid development.

[0005] In terms of business collaboration, cross-departmental data sharing and business collaboration is inefficient. Due to the lack of unified management view and efficient collaboration mechanism, the information barrier between departments is serious, and resources are difficult to share, which weakens the overall operational efficiency of the enterprise. At the same time, the enterprise decision lacks reliable data support, which reduces the scientificity and accuracy of management decision.

[0006] In terms of data analysis application, the data analysis capability of the existing system is limited. In the face of complex business scenarios, it is difficult to realize the depth mining and value discovery of data, and it is difficult to provide strong support for enterprise decision-making. In addition, the system has high use threshold, and the individualized needs of users at different levels are difficult to effectively meet.

[0007] Therefore, it is urgent to develop a new type of enterprise information management integrated platform to solve the above problems. The platform should have the following characteristics: efficient data processing capability, flexible system scalability, convenient business collaboration mechanism and intelligent data analysis function. Through these characteristics, the platform will realize the efficient utilization of enterprise data resources, and comprehensively improve the information management efficiency, and provide strong support for the digital transformation of enterprises. SUMMARY

[0008] The application provides an enterprise information management integrated platform based on big data, which is constructed based on Hadoop ecosystem, adopts distributed computing and storage architecture, and includes the following modules:

[0009] 1. Parallel data processing hub

[0010] Heterogeneous data generated by enterprise information systems is processed based on a Hadoop distributed computing framework; a MapReduce parallel computing model is adopted, swarm intelligence algorithm and pheromone transmission mechanism are integrated, multiple data processing agent nodes are constructed, task allocation is performed based on historical pheromone concentration, adaptive load balancing is realized; adaptive wavelet compression technology is adopted to store enterprise historical data, and unified processing of structured business data, semi-structured log data and unstructured document data is realized.

[0011] 2. Intelligent data mapping engine

[0012] Real-time processing of enterprise heterogeneous system data is realized based on a Spark real-time computing framework; a distributed stream computing architecture is adopted, a hierarchical self-encoding network is integrated to build a feature extraction model, and adaptive mapping of enterprise heterogeneous system data is realized through a multi-layer encoder-decoder structure; a progressive feature selection mechanism is adopted, and data mapping rules are optimized through iterative feature screening to build an enterprise data standard system.

[0013] 3. Elastic resource scheduler

[0014] Based on a YARN resource scheduling framework, a microkernel architecture is adopted; a biological cell division mechanism is integrated, computing resources are abstracted as cell units, and cell unit division or combination operations are triggered by setting a load threshold; an entropy value balancing mechanism is integrated, dynamic expansion of business modules is realized, and resource prediction is performed through an LSTM model to support pre-allocation and dynamic adjustment of computing resources.

[0015] 4. Distributed storage center

[0016] The distributed storage center adopts a HDFS distributed file system and a data lake architecture to build a unified data storage platform; an adaptive grid partitioning technology is integrated, grid units are built based on data access heat, and dynamic optimization of grid units is realized; deterministic chaotic mapping is adopted for storage strategy optimization, concurrent access is supported by building a phase space model of storage state, and hierarchical storage of hot and cold data is realized through a data temperature layering mechanism.

[0017] 5. Business analysis engine

[0018] The business analysis engine builds an analysis model based on a Spark MLlib framework and integrates multi-scale feature fusion technology; multi-layer feature extraction networks are constructed to realize extraction and fusion of business features, and an enterprise operation monitoring model is constructed using graph projection technology; a knowledge reasoning engine is integrated, a domain knowledge base and reasoning rules are constructed, and intelligent decision analysis for enterprise business scenarios is supported.

[0019] The beneficial effects of the present application are:

[0020] 1. Improve data processing efficiency

[0021] The data processing agent node network is constructed through the swarm intelligence algorithm and pheromone transmission mechanism, the task adaptive allocation based on the big data environment is realized, the data collection and processing efficiency is significantly improved, and the data value is fully utilized.

[0022] 2. Enhance system expansion capability

[0023] Based on the biological cell division mechanism and the entropy value balance mechanism, the resource scheduling is carried out, the adaptive division and combination operation of the cell unit is realized, the accurate dynamic expansion of the computing resource in the enterprise information environment is realized, and the problem that the system is difficult to cope with the business growth demand is solved.

[0024] 3. Optimize data sharing mechanism

[0025] The hierarchical self-encoding network combined with the progressive feature selection mechanism is adopted, the adaptive mapping of large-scale heterogeneous system data is realized, the data island is broken, the cross-department data sharing and business cooperation efficiency is improved, and a unified enterprise data standard system is constructed.

[0026] 4. Strenghten analysis and mining capability

[0027] Through the innovative combination of multi-scale feature fusion technology and graph projection technology, combined with the knowledge reasoning engine, the deep mining and analysis of enterprise operation data are realized, and intelligent support is provided for enterprise management decision-making.

[0028] 5. Improve storage access performance

[0029] The deterministic chaotic mapping and adaptive grid partitioning technology are combined, the storage strategy and data distribution in the big data environment are optimized, the performance of the enterprise-level storage system is improved, and the high-concurrency data access demand is supported. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 The architecture schematic diagram of the enterprise information management integrated platform is shown.

[0031] Figure 2 The workflow diagram of the enterprise information management integrated platform is shown. DETAILED DESCRIPTION

[0032] The exemplary embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0033] In combination Figure 1 , an enterprise information management integrated platform based on big data is provided, which includes the following functional modules:

[0034] 1. Parallel data processing hub

[0035] The parallel data processing hub is built based on the Hadoop distributed computing framework, by integrating improved swarm intelligence algorithm and pheromone transmission mechanism, combined with adaptive wavelet compression technology, to realize efficient processing and intelligent storage of enterprise heterogeneous data. The hub mainly processes structured business data, semi-structured log data and unstructured document data generated in enterprise information systems. The specific implementation includes:

[0036] A1). Swarm intelligence task allocation

[0037] In the parallel data processing hub, the task allocation algorithm is improved based on the idea of swarm intelligence to meet the complexity requirements of enterprise-level tasks. In the specific implementation, the improvement of the algorithm mainly reflects in task modeling, priority evaluation, dependency management and pheromone transmission optimization.

[0038] a). Task modeling

[0039] To meet the diverse needs of enterprise tasks, the parallel data processing hub adopts a unified task modeling method, and each data processing task is represented as a five-tuple , where:

[0040] : task unique identifier, in UUID format;

[0041] : task priority vector, describing the urgency, resource occupation and business importance of the task;

[0042] : resource demand vector, describing the computing, storage and network resources required by the task, ,

[0043] where, is the computing resource requirement, expressed in CPU core number and memory capacity; is the storage resource requirement, expressed in data size; is the network resource requirement, expressed in bandwidth requirement;

[0044] : data dependency set, describing the dependency relationship between tasks;

[0045] : task status identifier, including "to be processed", "processing" and "completed".

[0046] b). Priority evaluation

[0047] The task priority vector is defined as The task was comprehensively evaluated from three dimensions: time urgency, resource consumption, and task importance.

[0048] Urgency The calculation is performed by normalizing the difference between the task deadline and the current time.

[0049] in:

[0050] Task deadline; Current time; : The maximum task execution time supported by the system.

[0051] Resource utilization Based on the combined demand for different resource types, a weighting coefficient is used to balance the importance of each resource type:

[0052] in: Resource type Weighting coefficients; The demand for this type of resource; Total system resources.

[0053] Task importance Dynamically adjust based on task type priority and business priority:

[0054] in: : Weighting coefficient for task type; Task type priority; Task priority.

[0055] The above three dimensions , , Satisfy the normalization condition:

[0056] .

[0057] c) Dependency Management

[0058] In enterprise-level data processing, the dependencies between tasks are often complex and dynamic; therefore, the system uses directed acyclic graph data. In the formula, the dependencies between tasks are represented. A set of task nodes; Let be a set of dependent edges, where the edges Indicates task Must be in the task The previous tasks are completed. By analyzing the dependencies dynamically, the system can arrange the execution order of tasks reasonably, avoiding data conflicts and execution deadlocks.

[0059] d). Optimization of pheromone transmission mechanism

[0060] The dynamic updating mechanism of pheromones combines the efficiency of task processing :

[0061] Wherein:

[0062] : Pheromone decay coefficient, used to control the influence of historical information;

[0063] : Pheromone intensity constant;

[0064] : Task processing efficiency, the calculation formula is:

[0065] Wherein, , are the estimated and actual processing time of the task respectively; are the estimated and actual resource consumption of the task respectively.

[0066] When assigning tasks, the processing node is selected to execute the task The probability of calculation is:

[0067] Wherein:

[0068] : Pheromone factor and heuristic factor respectively;

[0069] : Node fitness, the calculation formula is:

[0070] Wherein, is the comprehensive processing capacity of the node; is the current load rate of the node; is the transmission overhead of the task to the node .

[0071] B1). Adaptive wavelet compression storage

[0072] In view of the characteristics of large amount of enterprise data and diversified data characteristics, the parallel data processing hub adopts adaptive wavelet compression technology to store data efficiently. Through hierarchical decomposition, threshold compression and dynamic adjustment of data, the technology realizes the balance between storage space and data characteristics, ensures the retention of important information and the compression of redundant data.

[0073] a) Wavelet decomposition

[0074] Input data sequence First, multi-scale decomposition is performed using discrete wavelet transform:

[0075] in:

[0076] Indicates the first Layer, First The wavelet coefficients of each component contain information about the input data at different frequencies and locations.

[0077] wavelet basis functions Defined as:

[0078] In the formula, The scale parameter indicates the level of decomposition, used to capture low-frequency (trend) or high-frequency (detail) components of the data; The translation parameter represents the position on the time axis. Wavelet decomposition separates the global trend and local details in the data in a hierarchical manner. The low-frequency part represents the global trend of the data and is the key feature; the high-frequency part contains noise and detailed information, which is selectively retained according to the threshold.

[0079] b) Adaptive threshold calculation

[0080] To remove high-frequency noise and compress redundant information, a compression threshold is set. The adaptive calculation formula is:

[0081] in:

[0082] For the first The standard deviation of the wavelet coefficients of a layer reflects the degree of fluctuation in the data of that layer;

[0083] For the first The number of wavelet coefficients in a layer indicates the scale of the data in that layer.

[0084] c) Wavelet coefficient processing

[0085] After threshold calculation, the wavelet coefficients The coefficients are processed using the soft thresholding method and are then compressed. :

[0086] in:

[0087] This is the sign function for the wavelet coefficients, preserving positive and negative information; The effective information retained after subtracting the threshold value.

[0088] Through threshold screening, compressing the scale of coefficients, and reducing the redundant part of data storage. The part below the threshold value is directly set to zero, regarded as noise or redundant information; for the coefficients above the threshold value, only moderate adjustment is made to ensure that important detailed information is not lost.

[0089] d). Dynamic compression ratio adjustment

[0090] The system dynamically adjusts the compression ratio according to the importance of the data and the storage requirements:

[0091] Where: H (X) and H (Y) are the information entropy of the original and compressed data, respectively; is the compressed coefficient.

[0092] 2. Intelligent data mapping engine

[0093] The intelligent data mapping engine is built based on the Spark real-time computing framework, and through the integration of hierarchical auto-encoding network and progressive feature selection mechanism, it realizes the real-time processing and mapping of enterprise heterogeneous system data; the engine adopts a distributed stream computing architecture to build an enterprise unified data standard system. The specific implementation includes:

[0094] A2). Hierarchical auto-encoding network

[0095] The hierarchical auto-encoding network extracts the features of heterogeneous data through a multi-layer encoder-decoder structure and completes the feature mapping. The network includes an input layer, multiple encoding layers, a feature layer, multiple decoding layers, and an output layer. Its core process includes the following three parts:

[0096] a). Encoding process

[0097] For input data (real number set vector space), through multiple layers of encoders, deep features are gradually extracted, and the formula is as follows:

[0098] Where:

[0099] : First hidden layer output;

[0100] : Weight matrix, randomly initialized and then learned through optimization;

[0101] : Bias vector;

[0102] : Activation function, using ReLU function:

[0103] Recursive computation of the subsequent encoding layers:

[0104] where, is the layer index, , is the output of the th layer of the encoder; is the output of the th layer of the encoder; is the weight matrix of the th layer of the encoder. is the bias vector of the th layer of the encoder.

[0105] Feature layer output:

[0106] where, is the layer index, is the output of the th layer of the encoder; is the output of the th layer of the encoder; is the weight matrix of the th layer of the encoder. is the bias vector of the

[0107] th layer of the encoder.

[0108] The decoding process gradually converts the feature layers into an approximation of the original input data , as follows:

[0109] where:

[0110] is the layer index, decreasing in order from the feature layers to the output layers;

[0111] , is the output of the th layer of the decoder; is the output of the th layer of the decoder;

[0112] is the weight matrix of the th layer of the decoder;

[0113] is the bias vector of the th layer of the decoder;

[0114] : activation function, using the ReLU function.

[0115] Final reconstruction output:

[0116]

[0117] c). Network optimization

[0118] The objective function is to minimize the reconstruction error and regularize the weights:

[0119] where:

[0120] is the squared Euclidean distance between the input data and the reconstructed output .

[0121] is the regularization coefficient to control the model complexity;

[0122] is the squared Frobenius norm of the layer weight matrix.

[0123] This network extracts deep features layer by layer, suitable for processing structural and semantic differences between different data sources, and provides standardized feature representation for subsequent feature selection and mapping rule construction.

[0124] B2). Progressive feature selection

[0125] Progressive feature selection is based on feature importance score, dynamically filters key features, effectively reduces redundancy and improves model efficiency.

[0126] a). Feature importance score

[0127] For a feature set , the importance score of each feature is calculated as follows:

[0128] where:

[0129] : correlation score, measures the correlation between the feature and the target variable;

[0130] : complexity score, reflects the simplicity of the feature value distribution;

[0131] : information gain, measures the information contribution of the feature to the target variable;

[0132] : weight coefficient, satisfies .

[0133] b). Feature score calculation

[0134] ​Relevance score :

[0135] where: is the Pearson correlation coefficient of the feature with the target variable . is the total number of features.

[0136] Complexity score :

[0137] where: is the information entropy of the feature . is the number of samples, i.e., the total number of feature values.

[0138] Information gain :

[0139] where: : the information entropy of the target variable . : the conditional entropy given the feature .

[0140] c). Iterative selection strategy

[0141] In progressive feature selection, the feature selection process is carried out step by step through a forward search strategy. In each iteration, the feature with the highest selection score is selected and added to the selected feature set , until the predetermined selection criteria or stopping condition is met. The formula is as follows:

[0142] where: is the feature selected in the th iteration; is the feature set, containing all candidate features; is the selected feature set, which is dynamically updated in each iteration; is the importance score of the feature; indicates the feature with the highest selection score.

[0143] This mechanism automatically selects the optimal feature subset in dynamic adjustment, reducing dimension redundancy and improving feature quality and mapping efficiency.

[0144] C2). Mapping rule optimization

[0145] a). Mapping rule construction

[0146] In the intelligent data mapping engine, a mapping rule set is constructed based on the selected features​ Each rule represents a mapping relationship from input data to output data. The optimization process of the rule set aims to maximize the overall performance of the mapping rules and improve the accuracy and efficiency of the mapping by adjusting the rule weights . Each rule is represented as:

[0147] where: is the rule condition vector, defining the mapping trigger condition; is the rule operation vector, defining the mapped data standard; is the rule score, with the calculation formula:

[0148] where, is the number of samples correctly mapped by the rule; is the total number of samples; is the number of rule conditions, representing the complexity of the conditions; is the total number of features, representing the total number of selected features in the current model.

[0149] b). Rule set optimization

[0150] Optimize the rule score through linear programming, with the objective function:

[0151] Constraints:

[0152] where: is the rule weight, representing the importance of the rule in the final decision; is the number of rules in the rule set.

[0153] Through the optimization of mapping rules, the accuracy and stability of data standardization mapping are improved, and the diverse needs of complex scenarios are adapted.

[0154] 3. Elastic resource scheduler

[0155] The elastic resource scheduler is built based on the YARN resource scheduling framework and adopts a microkernel architecture design. By integrating biological cell division mechanisms, entropy balance mechanisms, and resource prediction models, it realizes adaptive scheduling and dynamic expansion of computing resources. In a distributed environment, through multi-dimensional resource management and intelligent scheduling, it optimizes the use of computing resources and ensures the balance and efficiency of resource supply. The specific implementation includes:

[0156] A3). Resource management model

[0157] In the process of computing resource management, the resource units in the distributed environment are abstracted as cell units with division characteristics. Each cell unit is expressed as a four-tuple:

[0158] wherein:

[0159] : resource vector, describing the base resource capacity of the cell unit, , wherein,

[0160] represents the CPU resource capacity; represents the memory resource capacity; represents the storage resource capacity;

[0161] : state vector, reflecting the resource usage status of the cell unit, , wherein,

[0162] represents the real-time usage rate of the CPU resource; represents the real-time usage rate of the memory resource; represents the real-time usage rate of the storage resource;

[0163] : load threshold vector, defining the boundary condition of resource usage, , wherein, represents the lower limit threshold of the load; represents the upper limit threshold of the load;

[0164] : cell energy function, used to evaluate the running state of the resource unit, and the formula is:

[0165] wherein, is a weight coefficient, and satisfies the normalization condition ; is the resource usage rate fluctuation value; is the fluctuation threshold; , are the usage rate and total capacity of the i-th resource, respectively, , corresponding to the CPU, memory, and storage; represents the maximum value function. B3). Cell division mechanism

[0166] The cell division mechanism is responsible for managing the expansion and contraction of the resource unit. Its judgment criteria are as follows:

[0167]

[0168] When the cell energy exceeds the upper limit threshold ​When the resource utilization is below the lower threshold, a split operation is triggered, i.e., the resource unit is divided into smaller units to expand the resource; when the resource utilization is below the lower threshold , a merge operation is triggered to combine multiple low-load resource units into a larger resource unit to improve resource utilization and reduce management overhead; when the resource utilization is within the threshold interval, the current resource unit configuration is maintained.

[0169] C3). Entropy balancing mechanism

[0170] To ensure the overall balance of resources, an entropy evaluation mechanism is introduced, and the resource entropy is defined as:

[0171] wherein, represents the resource proportion of the resource unit, which is obtained by weighted calculation:

[0172] wherein, is the weight coefficient of the resource type; respectively represent the CPU, memory and storage resource capacity of the th resource unit; respectively represent the CPU, memory and storage resource capacity of the th resource unit.

[0173] The system further evaluates the resource distribution state by the resource balancing degree :

[0174] wherein: is the maximum entropy value, representing the entropy value when the resources are completely uniformly distributed, wherein, is the number of resource units, and the optimization goal is to minimize the balancing degree deviation, and the formula is:

[0175] wherein:

[0176] is the target balancing degree;

[0177] The constraint condition is:

[0178] wherein:

[0179] is the total resource capacity of the system; is the minimum resource allocation unit of a single resource unit.

[0180] D3). Resource demand prediction

[0181] To achieve future resource demand prediction, the system employs LSTM (Long Short-Term Memory) neural networks. LSTM is suitable for time series data and can capture long-term dependencies for dynamic resource demand prediction.

[0182] The input of LSTM is historical resource usage, load level, and historical resource allocation information:

[0183] Where: represents the resource usage at time step . represents the load level at time step . represents the historical resource allocation information at time step .

[0184] The output of LSTM is the predicted resource demand , i.e., the required computing resources at future time step .

[0185] The training of LSTM model is optimized by minimizing the Mean Squared Error (MSE) loss function:

[0186] Where:

[0187] is the actual resource demand; is the predicted resource demand.

[0188] During optimization, the Adam optimizer is used to accelerate convergence and improve training efficiency.

[0189] E3). Resource scheduling and dynamic expansion

[0190] During resource scheduling, the cell division mechanism is responsible for the expansion and contraction control of individual resource units, the entropy balance mechanism ensures the balance of global resource distribution, and the resource prediction model LSTM provides the basis for future resource demand prediction. The three work together to dynamically adjust resource allocation, ensuring that the system can automatically expand under high load and release unnecessary resources under low load.

[0191] 4. Distributed storage center

[0192] The distributed storage center is based on HDFS (Hadoop Distributed File System) and data lake architecture, building an efficient and reliable unified data storage platform. To further optimize storage efficiency and access performance, adaptive grid partitioning technology, deterministic chaos mapping, and data temperature layered storage mechanism are adopted. These technologies work together to achieve efficient storage and concurrent access management of data. The specific implementation includes:

[0193] A4). Adaptive grid partitioning

[0194] In terms of storage optimization, adaptive grid partitioning technology achieves fine-grained control over data distribution by dividing the storage space into multiple dynamically adjusted grid cells. The division of grid cells is based on the distribution density function of data blocks , whose formula is:

[0195] Where:

[0196] represents the storage location of the th data block; represents the access heat of the data block, which affects the storage density; is a two-dimensional Dirac function, representing the data block at the position; is the total number of data blocks.

[0197] B4). Data block access heat calculation

[0198] The access heat of the data block is calculated by the weighted combination of three indicators: access frequency, correlation, and timeliness. The specific formula is:

[0199] Where:

[0200] is the access frequency of the data block, indicating the frequency of accessing the data block; is the correlation of the data block, calculated by the cosine similarity between two data blocks, indicating the tightness of the relationship between the data block and other data blocks; is the timeliness of the data block, reflecting the activity level of the data block since the last access;

[0201] is the weight coefficient, and satisfies , ensuring the balance of the three indicators.

[0202] By dynamically adjusting the size of the grid cell, the storage system can effectively respond to the density and activity of data access, optimize resource allocation, and the adjustment of the grid cell is controlled by the following function:

[0203] Where:

[0204] and are the minimum and maximum sizes of the grid cell, respectively; is the density response coefficient, controlling the response degree of the grid cell size to the data distribution density.

[0205] C4). Deterministic chaotic mapping

[0206] To further optimize the distribution of data blocks, a two-dimensional Logistic mapping method is employed, which ensures the uniform distribution of data block positions in the storage grid. The two-dimensional position update formula for data blocks is as follows:

[0207] where:

[0208] and are the position coordinates of the data block at the current time;

[0209] and are control parameters, with a value range of to ensure chaotic characteristics.

[0210] The position mapping of data blocks is performed by the following function:

[0211] where:

[0212] is the unique identifier of the data block; and are the number of rows and columns of the storage grid, respectively;

[0213] is the floor operation, ensuring that the data block is mapped to the appropriate grid cell; , are the original mapping coordinates of the data block .

[0214] D4). Phase space model of storage state

[0215] To comprehensively describe the state of the storage system, a phase space model of the storage state is constructed, which is defined as:

[0216] where: is the coordinate of the storage grid; is the access speed of the data block, representing the access frequency of the data block per unit time; is the vector space of the data block access speed.

[0217] To describe the dynamic evolution of the storage system, the following trajectory equation is employed:

[0218] where: is the system state matrix, representing the internal changes of the system state; is the control matrix, representing the influence of external control; is a storage operation vector, representing the system's control strategy for data storage.

[0219] E4). Concurrent access control and stability

[0220] In order to ensure that the storage system can run stably under high concurrent access, Lyapunov stability criterion is used for control. Lyapunov function is defined as:

[0221] The derivative thereof satisfies the following condition:

[0222] wherein: is a symmetric positive definite matrix, used to define the stability of the system; is a storage operation vector, representing external control of the system behavior.

[0223] Through this stability criterion, resource contention and system collapse can be effectively avoided, ensuring the stable operation of the system under high concurrency.

[0224] F4). Data temperature hierarchical storage

[0225] The data temperature hierarchical storage mechanism classifies data blocks through the activity index to ensure that data storage at different levels has different priorities. The activity of a data block is calculated by the formula:

[0226] wherein:

[0227] is the access frequency of the data block;

[0228] is the maximum access frequency;

[0229] is the current time, is the last access time of the data block;

[0230] is the maximum timeliness of all data blocks;

[0231] is the size of the data block;

[0232] is the maximum size of all data blocks;

[0233] is a weight coefficient, and satisfies , indicating the importance of each index.

[0234] The temperature stratification rules for data are as follows:

[0235] Wherein:

[0236] And are temperature thresholds, representing the boundaries of hot, warm and cold layers respectively.

[0237] G4). Inter-layer data migration strategy

[0238] Inter-layer data migration is controlled by a state transition matrix, which is specifically as follows:

[0239] Wherein:

[0240] represent the probabilities of data migrating from the hot layer to the hot, warm and cold layers respectively;

[0241] represent the probabilities of data migrating from the warm layer to the hot, warm and cold layers respectively;

[0242] represent the probabilities of data migrating from the cold layer to the hot, warm and cold layers respectively.

[0243] Through this migration strategy, the system can dynamically adjust the storage location of data according to its activity level, thereby optimizing the utilization efficiency of storage resources.

[0244] 5. Business analysis engine

[0245] The business analysis engine is built based on the Spark MLlib framework, integrating multi-scale feature fusion technology, graph projection technology and knowledge reasoning engine to realize intelligent analysis of enterprise business data. Through the organic synergy of these three core technologies, the business analysis engine can extract multi-level features from data, construct business state graphs, and support intelligent decision-making and analysis through knowledge reasoning. The specific implementation includes:

[0246] A5). Multi-scale feature fusion

[0247] Multi-scale feature fusion technology performs deep learning on input data through multi-level convolution and pooling operations to extract richer feature information, constructing a multi-layer feature extraction network, where the feature representation output by each layer is represented as a feature map sequence :

[0248] Wherein, represents the feature map of the th layer, is the number of feature layers. ​

[0249] Feature mapping of each layer The specific calculation formula is determined by both convolution and pooling operations, and is as follows:

[0250] in: Input data; Sigmoid is chosen as the activation function. For the first Layer weight matrix; For the first Layer bias terms; convolution kernel matrix Performing convolution operations yields convolutional feature maps, which are used to extract local features from the data. It is the convolution kernel matrix, defined as:

[0251] In the formula, It is the size of The convolution kernel matrix, each The kernel element determines the weights of each region during the convolution operation.

[0252] This is a downsampling method, typically used to reduce the spatial dimensionality of data while retaining important feature information. Here, a dynamic weighted pooling strategy is employed, with the pooling formula as follows:

[0253] In the formula, Here are the pooling weights, representing the weight of each element during the pooling process. The specific calculation is as follows:

[0254] , , Input data The Middle The values ​​of the p-th and q-th elements; This is the sensitivity parameter for pooling, controlling how sensitive pooling is to features. A larger value indicates a higher sensitivity. This makes pooling more sensitive to larger input values.

[0255] To extract more meaningful information from features at multiple levels, an attention mechanism is used for feature fusion. The fused features are represented as follows: :

[0256] In the formula, The feature map extracted at layer j is the feature representation obtained after the input data has undergone different convolution and pooling operations; n is the total number of feature layers, representing the number of layers to be fused. Let be the attention weight, representing the th Layer feature mapping in the first The importance of the layer is calculated by the following formula:

[0257] In the formula, is the score function of feature fusion, and the calculation formula is as follows:

[0258] In the formula, is a learnable parameter vector; is a weight matrix for fusing features; indicates that the feature mapping of the first layer and the feature mapping of the first layer are spliced to serve as the input of the attention mechanism; is a hyperbolic tangent activation function, which is used to perform nonlinear transformation on the feature mapping.

[0259] Finally, through the attention mechanism, important features will get higher weights, ensuring that more prominent features get more attention in feature fusion.

[0260] B5). Graph projection

[0261] Graph projection technology is used for modeling and analysis of enterprise business data, representing and analyzing entities in the enterprise and their relationships, expressing business data in a graph structure, including graph construction, graph embedding, and graph projection matrix construction, as follows:

[0262] a). Graph construction

[0263] Business graph is composed of three parts:

[0264] : The node set in the graph represents various entities in the business, and each node represents a business entity, which is the basic unit of the graph;

[0265] : The edge set between nodes represents the relationship between entities, and each edge represents some interaction or relationship between two entities. The edge is the link between nodes;

[0266] : The attribute set possessed by each node and edge provides a detailed description of the entity and relationship, and is the basis for subsequent analysis.

[0267] b). Graph embedding target

[0268] The goal of graph embedding is to map each entity node Mapping into a vector space, so that the relationship between entities can be measured by the similarity of vectors. The objective function in the graph embedding process is as follows:

[0269] Wherein:

[0270] : embedding vector of node , the embedding vector is the result of low-dimensional representation of node , representing the position and relationship of the node in the graph structure;

[0271] : set of nodes adjacent to node , representing other nodes with direct relationship with ;

[0272] : conditional probability, representing the appearance probability of node after given the embedding vector of node , this probability reflects the similarity or correlation between node and , calculated by the following formula:

[0273] In the formula, and are the embedding vectors of node and node respectively; is the exponential value of the inner product of embedding vectors and , used to enhance the weight of strong relationship; is the normalization term, used to ensure that all probabilities sum to 1, to avoid too large or too small probability values. By optimizing this objective function, we can obtain the graph embedding vector, so that similar nodes are closer in the embedding space and have higher relationship probability.

[0274] c). Graph projection matrix

[0275] Graph projection matrix is constructed according to the structure of the business graph, mainly used to project the high-dimensional representation of the graph into a low-dimensional space for graph embedding and subsequent analysis. The construction of the graph projection matrix uses the following formula:

[0276] Wherein: is the degree matrix, representing the degree of each node, i.e. the number of edges directly connected to the node;

[0277] is the adjacency matrix, representing the relationship between nodes in the graph, if node with node has an edge connection, then , otherwise 0; is the inverse square root of the degree matrix , used to normalize the adjacency matrix, which can eliminate the influence of different node degrees, making the influence of each node more balanced in projection.

[0278] Business status vector after graph embedding It can be calculated by the following formula:

[0279] In the formula, is the feature vector matrix of graph projection, obtained by spectral decomposition, which contains the low-dimensional representation of the graph structure; is the embedding vector of node , representing the position of the node in the low-dimensional space; represents mapping the embedding vector to the low-dimensional space to obtain , that is, the business status vector of the node.

[0280] C5). Knowledge reasoning engine

[0281] The knowledge reasoning engine realizes intelligent analysis and decision support of enterprise business knowledge by constructing ontology model and reasoning rules. Based on knowledge graph technology, the engine formalizes the expression of various business objects and their relationships in the enterprise, and mines the implicit business rules and correlation through the reasoning mechanism.

[0282] a). Ontology model and triple set

[0283] The ontology model is formally defined by triple set :

[0284] Among them:

[0285] : represents a business entity, is the head entity, is the tail entity, representing different entity objects in the enterprise;

[0286] : represents a set of business entities, including all business entities, defining all business objects involved in the knowledge graph of the enterprise;

[0287] : represents a set of all possible relationship types, defining all possible relationship types between entities in the graph.

[0288] b). Confidence evaluation of reasoning rules

[0289] To evaluate the effectiveness and reliability of the inference rules, the engine calculates the confidence of each rule , which is specifically formulated as:

[0290] Wherein:

[0291] The numerator: represents the number of entity pairs that meet the rule , i.e., the number of head and tail entity pairs that meet the rule;

[0292] The denominator: represents the total number of head entities involved in the rule , i.e., all possible head entities that have at least one tail entity associated with them.

[0293] The confidence reflects the applicability of the rule in actual business scenarios. A high-confidence rule indicates that it is more common in the knowledge graph and is applicable to more actual situations, so it will have a higher weight in the inference process.

[0294] c). Forward chain reasoning

[0295] The inference process uses a forward chain reasoning algorithm, and the matching probability of a rule is calculated as:

[0296] Wherein:

[0297] represents all rules in the rule set, used for comparison and selection of the best rule; represents the head entity, which is the current entity in the inference path; represents the tail entity, which is the target entity in the inference path;

[0298] : is the scoring function of the rule , used to measure the matching degree of the rule given the head entity and the tail entity , and the rule scoring function is defined as:

[0299] Wherein:

[0300] represents the transpose of the rule weight vector, used to measure the importance of the rule in the current business scenario. Through model training optimization, the weight vector can be automatically adjusted to adapt to actual business needs; representing the embedding vectors of head entity and tail entity are spliced to form a unified feature vector, which helps the system to capture the relationship features between entities; representing the bias term of rule , used to adjust the baseline value of rule matching, ensuring fairness in the reasoning process of different rules.

[0301] In combination Figure 2 , the workflow of the enterprise information management integrated platform of the present application is as follows:

[0302] S1. Data acquisition and parallel processing

[0303] Based on the parallel data processing hub, the system first realizes the acquisition and processing of enterprise heterogeneous data; a processing agent node network is constructed by using a swarm intelligence algorithm, and a pheromone transmission mechanism is used to realize adaptive allocation of tasks, ensuring load balancing of data processing tasks; for different types of data, adaptive wavelet compression technology is used for unified processing: structured business data is subjected to feature extraction and compression, semi-structured log data is subjected to pattern recognition and normalization processing, and unstructured document data is subjected to semantic analysis and structured conversion; this stage establishes the basic framework of enterprise data processing.

[0304] S2. Data mapping and feature extraction

[0305] After the intelligent data mapping engine obtains the preprocessed data, it realizes deep feature extraction through a hierarchical self-encoding network; the multi-layer encoder-decoder structure can automatically learn the intrinsic feature representation of the data, realizing semantic-level mapping of heterogeneous system data; the progressive feature selection mechanism filters out the most representative feature set through an iterative optimization process, and constructs an enterprise unified data standard system accordingly; this stage realizes standardized expression and feature extraction of enterprise data.

[0306] S3. Resource scheduling and dynamic expansion

[0307] The elastic resource scheduler continuously monitors the system resource state and performs dynamic resource management based on a biological cell division mechanism; by abstracting computing resources as cell units, the system can automatically trigger resource splitting or merging operations according to the load condition; the entropy balance mechanism ensures the balance of resource distribution, avoiding resource hotspots; at the same time, through the LSTM time series prediction model, the resource demand trend is predicted in advance, and resource reservation and adjustment are performed in advance, ensuring the smooth operation and elastic expansion capability of the system.

[0308] S4. Distributed storage and management

[0309] The distributed storage center adopts HDFS file system and data lake architecture to build a unified storage platform; adaptive grid partitioning technology dynamically adjusts the size of the storage unit according to data access heat, and deterministic chaos mapping optimizes the data distribution strategy to improve concurrent access performance; the system stores high-frequency access hot data in the high-performance storage layer through the data temperature layering mechanism, and migrates low-frequency access cold data to the low-cost storage layer, realizing efficient utilization and intelligent management of storage resources.

[0310] S5. Business analysis and decision support

[0311] The business analysis engine builds an analysis model based on the Spark MLlib framework, realizes multi-level expression of data features through multi-scale feature fusion technology, projects business entities and relationships into the feature space in combination with graph projection technology, constructs an enterprise operation monitoring model, and then performs deep mining and correlation analysis on business data based on the ontology model and inference rules through the knowledge reasoning engine, to support multi-dimensional decision support including product correlation analysis, customer behavior prediction, process optimization suggestions and the like.

[0312] Through the above description of the drawings, the enterprise information management integrated platform based on big data provided by the application is constructed based on the Hadoop ecosystem, and unified management and analysis of enterprise heterogeneous data are realized through the collaborative work of five core modules of parallel data processing hub, intelligent data mapping engine, elastic resource scheduler, distributed storage center and business analysis engine. The platform innovatively applies swarm intelligence algorithm and pheromone transmission mechanism to parallel data processing, adopts biological cell division mechanism to realize dynamic scheduling of resources, optimizes data storage through adaptive grid partitioning and deterministic chaos mapping, and integrates multi-scale feature fusion and knowledge reasoning technology to support intelligent decision-making. The platform effectively solves the problems of low data processing efficiency, poor system expansibility and difficult data sharing in traditional enterprise information management, and provides reliable big data technical support for enterprise digital transformation.

[0313] It is worth noting that the protection scope of the present application is not limited to the specific embodiments described in the above embodiments. Equivalent modifications or modifications made by those skilled in the art based on the inspiration of the present application should be covered within the protection scope of the present application, as long as they do not deviate from the spirit of the technical solution of the present application.

Claims

1. An enterprise informatization management integrated platform, characterized in that, It comprises a parallel data processing hub, an intelligent data mapping engine, an elastic resource scheduler, a distributed storage center and a business analysis engine. The parallel data processing hub adopts a Hadoop distributed computing framework, sets a swarm intelligence algorithm and a pheromone transmission mechanism to construct a data processing agent node network, and realizes unified processing of enterprise heterogeneous data in combination with an adaptive wavelet compression technology. The intelligent data mapping engine adopts a Spark real-time computing framework, sets a hierarchical self-encoding network with a multi-layer encoder-decoder structure, and realizes adaptive mapping of enterprise heterogeneous system data in combination with a progressive feature selection mechanism; the progressive feature selection mechanism selects an optimal feature subset step by step through a forward search strategy, and the adaptive mapping refers to mapping heterogeneous data from different business systems into standardized data conforming to enterprise unified data standards. The elastic resource scheduler adopts the YARN scheduling framework, abstracting computing resources into cellular units. It triggers division or merging operations based on the cell energy function and utilizes an entropy balancing mechanism to dynamically adjust resources; the cell energy function... Used to quantitatively assess the load status of resource units, defined as The first item Indicates the overall resource utilization rate. For the first Usage of class resources For the first Total capacity of the resource class These correspond to CPU, memory, and storage resources, respectively; the second item Indicates the degree of uneven resource utilization. For CPU utilization, For memory usage; the third item Indicates the degree of resource fluctuation. The standard deviation of resource utilization over the time window. Maximum allowable fluctuation threshold; weighting coefficient , , satisfy ;when When cell division is triggered, The cell merging operation is triggered at time, in which The upper limit threshold, The lower threshold is used as the threshold value; the entropy balance mechanism evaluates the balance of resource allocation by calculating the resource distribution entropy. The distributed storage center adopts a HDFS distributed file system, optimizes data distribution in combination with an adaptive grid partitioning technology and a deterministic chaotic mapping, and realizes hierarchical storage of cold and hot data through a data temperature layering mechanism. The business analysis engine adopts a Spark MLlib framework, integrates multi-scale feature fusion technology and graph projection technology to construct an enterprise operation monitoring model, and sets a knowledge reasoning engine to support business decision analysis; the graph projection technology refers to a dimensionality reduction technology for projecting a knowledge graph composed of enterprise business entities from a high-dimensional relationship space to a low-dimensional vector space. The modules of the enterprise informatization management integrated platform work cooperatively, including: inputting enterprise original data into the parallel data processing hub for preprocessing; transmitting the preprocessed data to the intelligent data mapping engine for standardized mapping; dynamically adjusting resources by the elastic resource scheduler, and storing the standardized mapped data into the distributed storage center according to the dynamic adjustment result; and analyzing the data read from the distributed storage center by the business analysis engine and feeding back to the intelligent data mapping engine.

2. The enterprise informatization management integrated platform according to claim 1, characterized in that, The swarm intelligence algorithm in the parallel data processing hub represents a data processing task as a five tuple , including: Task unique identifier ; Priority vector describing task urgency, resource occupation, and business importance ; Resource vector describing compute, storage, and network resource requirements ; A set of inter-task dependency relationships is described ; Task identification including pending, in-process, and completed status .

3. The enterprise informatization management integrated platform according to claim 1, characterized in that: The adaptive wavelet compression technology in the parallel data processing hub comprises: performing multi-scale decomposition on an input data sequence by using a discrete wavelet transform; dynamically calculating a compression threshold according to a standard deviation of wavelet coefficients; processing the wavelet coefficients by a soft threshold method; adjusting a compression ratio based on an information entropy ratio of original data and compressed data.

4. The enterprise informatization management integrated platform according to claim 1, characterized in that: The hierarchical self-encoding network in the intelligent data mapping engine comprises: performing feature extraction by a multi-layer encoder structure with a ReLU activation function; realizing data reconstruction by a multi-layer decoder structure symmetrical to the encoder; adopting a target function combining reconstruction error and weight regularization for optimization.

5. The enterprise informatization management integrated platform according to claim 1, characterized in that: The cell unit in the elastic resource scheduler is defined as a four-tuple , including: Resource vector describing CPU, memory and storage resource capacities wherein is a CPU resource capacity, is a memory resource capacity, is a storage resource capacity; State vector describing real-time usage of CPU, memory and storage resources wherein is the CPU usage, is the memory usage, is the storage usage; Load threshold vector defining CPU, memory, and storage elasticity boundaries wherein is a lower threshold value, is an upper threshold value; A cell energy function for comprehensively evaluating CPU, memory and storage usage.

6. The enterprise informatization management integrated platform according to claim 1, characterized in that: The data optimization method in the distributed storage center comprises: constructing a grid cell based on a data block distribution density function and access heat; updating data block positions by using a two-dimensional Logistic mapping; realizing temperature layering storage of data according to a data activity index.

7. The enterprise informatization management integrated platform according to claim 1, characterized in that: The knowledge reasoning engine in the business analysis engine comprises: constructing an ontology model by using a triple set; screening reasoning rules by a confidence evaluation mechanism; Generating an inference path based on a forward chain inference algorithm. Generating an inference path based on a forward chain inference algorithm.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN111414381A

  • Electric power big data rapid visualization analysis method and system based on three-dimensional scene

    CN113806606A

Cited By

  • Enterprise informatization management integration platform based on big data

    CN121979873A