Multi-dimensional data intelligent analysis method based on one-table communication

Through the multi-dimensional data intelligent analysis method based on one table, the problem of insufficient depth of multi-source heterogeneous data integration and analysis is solved, efficient data integration and intelligent analysis are realized, real-time feedback and optimization mechanism are provided, and it is suitable for complex business scenarios.

CN119988991APending Publication Date: 2025-05-13CHONGQING INSPUR GOVERNMENT CLOUD MANAGEMENT & OPERATION CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411979495.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology has shortcomings in the integration of multi-source heterogeneous data, analysis depth, real-time feedback and visualization capabilities, and it is difficult to meet the needs of complex business scenarios.

Method used

Using a multidimensional data intelligent analysis method based on one table, a unified table view is generated through data cleaning, standardization and semantic enhancement, a semantic association weight matrix is ​​calculated, a dynamic multidimensional data analysis model is constructed, and interactive charts and path recommendation functions are provided.

Benefits of technology

It realizes efficient integration of multi-source data, improves analysis depth and real-time response capabilities, provides intelligent visual support, and can dynamically optimize analysis models to adapt to data and business changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988991A_ABST
    Figure CN119988991A_ABST
Patent Text Reader

Abstract

The invention relates to a one-table-communication-based multi-dimensional data intelligent analysis method, which comprises the following steps: S1, collecting data from various data sources, and generating a unified table view; s2, calculating a semantic association weight matrix among the data fields, and capturing potential association among the data fields through semantic tags and feature representation; s3, constructing a multi-dimensional data analysis model according to the field weight and the business rule, and capturing a complex nonlinear relationship among the data; s4, generating an interactive chart, supporting drill-up and drill-down, slice analysis and linkage display, and recommending an optimal analysis path; s5, dynamically optimizing and adjusting the multi-dimensional data analysis model; a multi-dimensional data intelligent analysis system based on one-table communication comprises a data acquisition and preprocessing module, a semantic analysis module, a multi-dimensional feature modeling module, a visual presentation module and a real-time feedback and optimization module. The method has the advantages that the problem of heterogeneous data integration is solved, the analysis depth and the real-time response capability are improved, and intelligent visual support is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data analysis, and in particular to a multi-dimensional data intelligent analysis method based on one table. Background Art

[0002] As enterprises deepen their digital transformation, the explosive growth of multi-source, heterogeneous data places higher demands on data analysis. However, existing technologies have significant shortcomings in data integration, analytical depth, real-time feedback, and visualization capabilities, making it difficult to meet the needs of complex business scenarios. Specifically, the following are some of the key issues:

[0003] 1. Difficulty integrating multi-source heterogeneous data. Inconsistent data source formats and semantics lead to difficulties in field matching, insufficient cross-domain semantic associations, and difficulty integrating dynamically changing data efficiently using traditional manual rules or static configuration methods.

[0004] 2. Analytical models lack depth and adaptability. Existing analytical models primarily rely on fixed dimensions and predefined rules, failing to capture complex nonlinear relationships between multidimensional data and lacking the ability to dynamically adjust to data and business needs.

[0005] 3. Lack of real-time feedback and optimization mechanisms: Analytical systems are often unable to update models based on real-time data, resulting in delayed analysis results. Furthermore, weak feedback mechanisms make it difficult to continuously improve model performance.

[0006] 4. Insufficient visualization and intelligence: Existing tools mostly focus on static chart display, lack intelligent analysis path recommendations and flexible multi-dimensional drill-down support, and are unable to meet the needs of complex scenarios.

[0007] Although ETL tools, BI platforms and some machine learning methods have been used to integrate, analyze and display data, their overall adaptability and real-time dynamic optimization capabilities are limited, making it difficult to simultaneously meet the comprehensive needs of multi-source data integration, dynamic analysis and intelligent interaction. Summary of the Invention

[0008] In response to the above-mentioned deficiencies in the existing technology, the present invention proposes a multi-dimensional data intelligent analysis method based on a one-table-based approach, which aims to solve the problem of heterogeneous data integration, improve analysis depth and real-time response capabilities, and provide intelligent visualization support.

[0009] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0010] A multidimensional data intelligent analysis method based on Yibiaotong includes the following steps: S1: collecting data from multiple data sources, and cleaning, standardizing and semantically enhancing the data to generate a unified table view; S2: based on the unified table view described in step S1, calculating the semantic association weight matrix between data fields, and capturing the potential association between data fields through semantic labels and feature representations; S3: extracting key features in the unified table view, building a multidimensional data analysis model based on field weights and business rules, and capturing complex nonlinear relationships between data; S4: based on the output of the multidimensional data analysis model described in step S3, generating interactive charts that support drill-down, slice analysis and linkage display, and recommending the optimal analysis path based on user historical operations and data characteristics; S5: collecting user feedback data and real-time new data, and dynamically optimizing and adjusting the multidimensional data analysis model.

[0011] As an optimization, in step S1, when generating the unified table view, raw data is collected from multiple data sources through a unified data interface, and the raw data includes structured data, semi-structured data and unstructured data; during the preprocessing process, the system cleans and standardizes the collected data, removes outliers, fills missing values, and ensures the consistency of data format and range; at the same time, field mapping and semantic labeling technology are used to semantically enhance the data, and finally a unified table view is constructed.

[0012] As an optimization, in step S2, the calculation of the semantic association weight matrix includes: based on the semantic label set and feature vector of the field, using a semantic similarity algorithm to generate the semantic association relationship between the fields, and optimizing the weight matrix in real time through a dynamic adjustment mechanism.

[0013] As an optimization, in step S3, key features are extracted based on the unified table view and modeled using a multidimensional data analysis model. The model combines the importance weights of fields and business rules to capture complex nonlinear relationships between data and adapt to business needs in different scenarios. During the modeling process, model parameters are dynamically adjusted based on real-time feedback and newly added data.

[0014] As an optimization, in step S4, the interactive charts generated include the following types: drill-down analysis charts: allowing users to gradually drill down from high-dimensional overview data to specific dimensions, supporting the exploration of multi-level data; data slicing analysis charts: by filtering data of specific dimensions or ranges, corresponding subset analysis charts are generated; linkage analysis charts: when the user selects a data point in a chart, other related charts will be updated synchronously to display the related information of the data point.

[0015] As an optimization, in step S5, the system collects user operation behavior and feedback data, and optimizes the analysis model in combination with real-time new data; by dynamically adjusting model parameters, the system can continuously improve the accuracy and real-time performance of the analysis, while reducing the complexity of the model to ensure its stable performance in changing scenarios.

[0016] A multidimensional data intelligent analysis system based on Yibiaotong includes a data acquisition and preprocessing module, a semantic analysis module, a multidimensional feature modeling module, a visualization presentation module and a real-time feedback and optimization module; the data acquisition and preprocessing module is used to acquire data from multiple data sources, and perform cleaning, standardization and semantic enhancement processing on the data to generate a unified table view; the semantic analysis module is used to calculate the semantic association weight matrix between data fields in the unified table view, and dynamically adjust the weights in combination with real-time new data; the multidimensional feature modeling module is used to extract key features from the unified table view and construct a multidimensional data analysis model based on field weights and business rules; the visualization presentation module is used to generate interactive charts, support up and down drilling, slice analysis and linkage display, and guide users to perform intelligent analysis through a path recommendation function; the real-time feedback and optimization module is used to acquire user feedback data and real-time new data, and dynamically adjust the analysis model in combination with the optimization objective function.

[0017] As an optimization, the data acquisition and preprocessing module includes a data interface unit, a field mapping unit and a data cleaning unit to realize multi-source data acquisition, cleaning and standardization, and finally generate a unified table; the semantic analysis module generates an association weight matrix through a semantic calculation unit, and updates the weight in real time through a dynamic adjustment unit to capture the semantic relationship between fields; the multidimensional feature modeling module includes a feature extraction unit and a model training unit, which builds a multidimensional analysis model and extracts key features based on a unified table view; the visualization module provides users with up and down drill-down analysis and intelligent path recommendation functions through a chart generation unit and a path recommendation unit, and intuitively displays data patterns; the feedback optimization module obtains user feedback data through a feedback acquisition unit, and dynamically adjusts model parameters in combination with a parameter optimization unit to ensure that the system continues to adapt to dynamically changing scenario requirements.

[0018] The present invention has the following advantages:

[0019] (1) Efficient multi-source heterogeneous data integration: This invention uses the one-table-through technology to build a semantically enhanced unified table view, achieving unified integration of structured, semi-structured, and unstructured data. Through dynamic field mapping and real-time semantic enhancement, the system significantly improves the efficiency and accuracy of data integration, solving the problem that traditional methods rely on manual matching and cannot adapt to dynamic changes.

[0020] (2) Intelligent semantic association analysis: This invention uses a semantic weight matrix combined with a deep learning model to accurately capture the potential associations between cross-domain data fields and optimize semantic relationships in real time through a dynamic adjustment mechanism. Compared with traditional static rules, the system can automatically adapt to new data and domain changes, significantly improving the intelligence level and analysis accuracy of semantic associations.

[0021] (3) Dynamic multidimensional data modeling: This paper proposes a dynamic multidimensional data analysis model that, through dimension weight adjustment and business rule-driven, can flexibly adapt to various scenarios and capture nonlinear relationships between complex data. Compared with traditional fixed rule models, this paper realizes adaptive optimization of the model when data changes, ensuring the real-time and accuracy of the analysis results.

[0022] (4) Intelligent visual interaction: The system provides multi-dimensional interactive charts and path recommendation functions, supports drilling down, data slicing, and linkage display; dynamically generates the optimal analysis path based on user behavior and data characteristics, helping users quickly understand data patterns, improving analysis efficiency and user experience;

[0023] (5) Real-time feedback and continuous optimization: The system is designed with a feedback optimization module that dynamically adjusts model parameters by combining user operation data and real-time new data, thereby achieving the ability to continuously optimize the analysis model. This mechanism not only improves the adaptability of the system in dynamic scenarios, but also solves the problem of low feedback utilization efficiency in traditional analysis methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a flow chart of a multi-dimensional data intelligent analysis method based on one table according to the present invention.

[0025] Figure 2 This is a framework diagram of a multi-dimensional data intelligent analysis system based on OneTable according to the present invention. DETAILED DESCRIPTION

[0026] The present invention will be described in further detail below with reference to the accompanying drawings. It should be understood that in the description of the present invention, the directions or positional relationships indicated by directional terms such as "upper" and "lower" and "top" and "bottom" are generally based on the directions or positional relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. Unless otherwise indicated, these directional terms do not indicate or imply that the devices or components referred to must have a specific direction or be constructed and operated in a specific direction, and therefore should not be construed as limiting the scope of protection of the present invention. The directional terms "inside" and "outside" refer to the inside and outside relative to the outline of each component itself.

[0027] Example 1

[0028] like Figure 1As shown, a multidimensional data intelligent analysis method based on Yibantong includes the following steps: S1: collecting data from multiple data sources, and cleaning, standardizing and semantically enhancing the data to generate a unified table view; S2: based on the unified table view described in step S1, calculating the semantic association weight matrix between data fields, and capturing the potential association between data fields through semantic labels and feature representations; S3: extracting key features in the unified table view, building a multidimensional data analysis model based on field weights and business rules, and capturing complex nonlinear relationships between data; S4: based on the output of the multidimensional data analysis model described in step S3, generating interactive charts that support drill-down, slice analysis and linkage display, and recommending the optimal analysis path based on user historical operations and data characteristics; S5: collecting user feedback data and real-time new data, and dynamically optimizing and adjusting the multidimensional data analysis model.

[0029] In this embodiment, in step S1, when generating the unified table view, original data is collected from multiple data sources through a unified data interface, and the original data includes structured data, semi-structured data and unstructured data; during the preprocessing process, the system cleans and standardizes the collected data, removes outliers, fills missing values, and ensures the consistency of data format and range; at the same time, field mapping and semantic labeling technology are used to semantically enhance the data, and finally a unified table view is constructed.

[0030] Specifically, data cleaning includes the following steps: deleting records with too many missing values, filling in some missing values, eliminating outliers, and standardizing the data to meet the requirements of unified table construction;

[0031] Specifically, A1: Cleaning formula and parameter explanation

[0032] For a single field X, the cleaned data is represented as:

[0033] X clean =γ·X+(1-γ)·median(X)

[0034] In this formula: X clean : cleaned field data; X: original field data; median(X): median of field data X, used to replace outliers or fill missing values; γ: smoothing factor, ranging from 0 ≤ γ ≤ 1, used to balance the ratio between original data and median;

[0035] In practical applications: If the field data contains many outliers or missing values, γ can be set to a smaller value (such as 0.2) and rely more on median substitution; if the field data quality is high, γ can be set to a larger value (such as 0.8) to retain more characteristics of the original data;

[0036] A2: Outlier Detection and Processing

[0037] During the cleaning process, outliers are detected based on the following formula:

[0038] Outlier(X) = {x i ∈X||x i -μ|>k·σ}

[0039] Where: x i is a single data value of field X; μ is the mean of field data X; σ is the standard deviation of field data X; k is the multiple threshold for outlier detection, usually k = 3;

[0040] For the detected outlier Outlier(X), replace it with the following formula:

[0041]

[0042] A3: Data Standardization

[0043] To unify the value range of the field, the system standardizes the field data; the standardization formula is:

[0044]

[0045] Where: X norm is the standardized field data; min(X) and max(X) are the minimum and maximum values ​​of field X respectively;

[0046] A4: Build a unified table view

[0047] After cleaning and standardization, the data is mapped to a unified table view through field mapping and semantic enhancement technology. uni , the formula is:

[0048]

[0049] Where: D i : the i-th data source; A set of semantic tags that define semantic features for each field; A set of field mapping rules used to map fields to a unified table;

[0050] Transform: Field mapping and semantic enhancement function, defined as follows:

[0051] Match field x and rule set Returns the mapped value if the match succeeds.

[0052] In this embodiment, in step S2, the calculation of the semantic association weight matrix includes: based on the semantic label set and feature vector of the field, using a semantic similarity algorithm to generate the semantic association relationship between the fields, and optimizing the weight matrix in real time through a dynamic adjustment mechanism.

[0053] Specifically, based on the unified table view, the system calculates the semantic relationship between data fields through the semantic association analysis module and constructs a semantic association network between fields; through the semantic labels and feature representations of the fields, the system can capture the potential association relationships between cross-domain data, and dynamically adjust the association weights in combination with real-time data updates, thereby ensuring the timeliness and accuracy of the analysis.

[0054] Specifically, B1: semantic weight matrix calculation

[0055] The calculation formula of the semantic weight matrix R is:

[0056]

[0057] Where: R ij : represents the semantic association weight between field i and field j; Represents the semantic label sets of field i and field j respectively; Represents a set of semantic tags and Similarity; V i , V j : are the feature vectors of field i and field j, respectively, generated using the word embedding method Word2Vec; || V i ||,||V j ||: Field feature vector V i and V j The modulus of is used for normalization;

[0058] Semantic similarity calculation:

[0059]

[0060] Where: m: the number of semantic labels; The kth semantic label of field i and field j; The cosine similarity of two vectors x and y.

[0061] Feature vector generation:

[0062] Field feature vector V i The word vector method Word2Vec is used to generate the feature vector. The formula for generating the feature vector is:

[0063]

[0064] Where: n: the number of semantic labels of field i; Semantic Tags Embedded vector representation of .

[0065] B2: Dynamic adjustment of weight matrix

[0066] To ensure the accuracy and timeliness of the semantic weight matrix R in a real-time data environment, the present invention optimizes R through a dynamic adjustment mechanism. When the system receives new data, the dynamic adjustment formula is as follows:

[0067]

[0068] in: The semantic association weight between field i and field j at time t+1; The semantic association weight of field i and field j at time t; ΔR: the incremental adjustment value of the weight due to the new data, calculated as follows:

[0069]

[0070] The result of calculating the semantic similarity of the newly added data to fields i and j; α: dynamic adjustment coefficient, which controls the impact of the newly added data on the weight adjustment, with a value range of 0≤α≤1;

[0071] B3: Normalization of weight matrix

[0072] During the dynamic adjustment process, in order to ensure the numerical stability of the weight matrix, the system normalizes R:

[0073]

[0074] Among them: max(R): the maximum value in the weight matrix R, used to convert R ij Normalized to the range [0, 1].

[0075] In this embodiment, in step S3, key features are extracted based on the unified table view, and the features are modeled using a multidimensional data analysis model; the model combines the importance weights of the fields and business rules to capture complex nonlinear relationships between data and adapt to business needs in different scenarios; during the modeling process, model parameters are dynamically adjusted based on real-time feedback and new data.

[0076] Specifically, in the unified table view T uni Based on the semantic association weight matrix R, the present invention proposes a dynamic multidimensional data analysis model M multi , used to capture complex nonlinear relationships between data and generate interpretable analysis results. The process of multidimensional feature modeling includes three core parts: feature extraction, weight adjustment and dynamic modeling;

[0077] C1: Feature extraction: Feature extraction aims to extract features from the unified table T uni Extract the key dimensions that can reflect the global and local characteristics of the data. The feature extraction formula is as follows:

[0078]

[0079] Where: T uni : kth eigenvalue; T uni [j]: j-th field value in the unified table view; w j : The weight of the j-th dimension, indicating the importance of the field; n: The total number of fields in the unified table view.

[0080] Weight w j The initial value of is calculated using the following formula:

[0081]

[0082] Where: Var(T uni [j]): The variance of field j, which is used to measure the degree of fluctuation of the field; w j : Initial weight value, meeting the normalization condition

[0083] The calculation of the initial weight is based on the variance of the field. Fields with larger fluctuations are given higher weights to emphasize their importance in data analysis.

[0084] C2: Dynamic weight adjustment: To adapt to real-time changes in data and business needs, the system adjusts the weight w j Perform dynamic adjustment. The dynamic adjustment formula is:

[0085]

[0086] in: The weight of the j-th dimension at time t+1; The weight of the j-th dimension at time t; Δw: the change in weight, calculated as follows:

[0087]

[0088] Where: Corr(T uni [j], Y): correlation coefficient between field j and target variable Y; Y: target variable of the analysis model;

[0089] β: adjustment coefficient, which controls the amplitude of weight adjustment and has a value range of β;

[0090] Dynamic adjustment introduces the correlation information of the target variable Y so that the weight distribution can reflect the real-time characteristics of the data and the needs of the analysis target;

[0091] C3: Dynamic modeling: Based on feature extraction and weight adjustment, a dynamic multidimensional data analysis model M is constructed. multi , its core formula is:

[0092] M multi =f(T uni , W, Q)

[0093] Where: M multi : Multidimensional data analysis model; f: Support vector machine; T uni : Unified table view; W = {w1, w2, ..., w n}: dimensional weight vector; Q = {q1, q2, ..., q k}: A set of business rules used to constrain the analysis logic;

[0094] In order to capture the nonlinear relationship between multidimensional data, the model f uses the nonlinear kernel function K(x, y) for feature mapping:

[0095]

[0096] Where: x, y: input features; ||xy||: Euclidean distance between features; σ: kernel width parameter, which controls the amplitude of nonlinear mapping;

[0097] Through nonlinear kernel functions, the model can map low-dimensional features to high-dimensional space, thereby capturing more complex data relationships.

[0098] During the modeling process, the objective function is designed as follows:

[0099]

[0100] in: Objective function; f(T uni [i], W): predicted value of the i-th data sample; Y[i]: true value of the i-th sample; Ω(f): model complexity penalty term, used to prevent overfitting; λ: regularization parameter, controlling the strength of complexity constraint;

[0101] C4: Model output and explanation: Model output feature φ k And the prediction result M multi After that, the system explains the importance of the feature. Through feature contribution calculation, the contribution of each field C j Defined as:

[0102]

[0103] Where: C j : Contribution of field j to the model; Var(T uni [j]): variance of field j,

[0104] Contribution results can help users intuitively understand the impact of each field on the analysis target and provide support for subsequent business decisions.

[0105] In this embodiment, in step S4, the interactive charts generated include the following types: drill-down analysis charts: allowing users to gradually drill down from high-dimensional overview data to specific dimensions, supporting the exploration of multi-level data; data slicing analysis charts: by filtering data of specific dimensions or ranges, generating corresponding subset analysis charts; linkage analysis charts: when the user selects a data point in a chart, other related charts will be updated synchronously to display the related information of the data point.

[0106] Specifically, the visualization module provides a multi-dimensional interactive display of analysis results, supporting data drilling down, slicing analysis, and linkage display; the visualization module can not only intuitively present the analysis results, but also help users quickly find the optimal analysis path through the path recommendation function, thereby significantly improving analysis efficiency and user experience.

[0107] D1: Path relevance score definition: Path relevance score Score (u, k) is the core parameter of the path recommendation function, which is used to measure the degree of adaptability between path k and user u's current analysis needs. Its calculation formula is:

[0108] Score(u,k)=α·Sim(D k , D u )+β·Feat(k,u)+γ·Hist(u,k)

[0109] Among them: path and scene similarity Sim(D k , D u ): Measure the data dimension D involved in path k k Dimension D of the user's current operation data u similarity;

[0110]

[0111] Where: D k : The set of data dimensions involved in path k; D u : The data dimension set of the user's current scenario; |·|: The number of elements in the set, indicating the size of the dimension set.

[0112] Path feature weight Feat(k,u): measures the impact of the fields involved in path k on the user's current analysis target Y u the importance of

[0113]

[0114] Among them: F k : The set of fields involved in path k; w j : The weight of field j, indicating the importance of the field to the analysis model; T uni [j]: Unify the data of field j in the table view; Y u : The target variable of the current analysis scenario of user u; Corr(T uni [j],Y u ): Field j and target variable Y u The correlation coefficient is used to measure the linear relationship between field data and target variables;

[0115] Historical path usage frequency Hist(u,k): measures the importance of path k in the historical operation records of user u.

[0116]

[0117] Where: n k : the number of times user u uses path k; P u : The historical path set of user u, that is, the set of all paths the user has ever used;

[0118] Weight coefficients α, β, and γ: respectively control the impact of path-scene similarity, path feature weight, and historical path frequency on the total score, satisfying the following conditions:

[0119] α+β+γ=1

[0120] D2: Route recommendation function

[0121] Based on the definition of the path relevance score Score(u,k), the system dynamically generates the optimal analysis path through the path recommendation formula:

[0122]

[0123] Where: P rec (u): the optimal path recommended for user u; k: represents a candidate analysis path; δ k : Time decay factor, used to reduce the weight of historical paths when they have not been used for a long time. Its calculation formula is:

[0124] δ k =exp(-λ·t k )

[0125] t k: The interval between the last time path k was used and the current time; λ: Time decay coefficient, which controls the decay speed. The larger the value, the faster the path weight decays over time.

[0126] By comprehensively considering the path relevance score and time decay weight, the system can dynamically recommend the optimal path to help users quickly locate the analysis direction;

[0127] D3: Dynamic interactive chart generation

[0128] The system dynamically generates the following types of interactive charts based on the output of the analysis model and the user's operational requirements: drill-down analysis charts, data slicing analysis charts, and linkage analysis charts.

[0129] In this embodiment, in step S5, the system collects the user's operation behavior and feedback data, and optimizes the analysis model in combination with real-time new data; by dynamically adjusting the model parameters, the system can continuously improve the accuracy and real-time performance of the analysis, while reducing the complexity of the model to ensure its stable performance in changing scenarios.

[0130] Specifically, F1: Optimization objective function definition; the core of the optimization objective function is to minimize the weighted sum between prediction error and model complexity, which is defined as follows:

[0131]

[0132] in: Optimize the objective function value; F(x t ): Model F for input data x at time t t The predicted value of y t : The true value (target value) at time t; || F(x t )-y t || 2 : The quadratic error between the predicted value and the true value; Ω(F): The complexity measurement function of the model F, which is used to limit the overfitting of the model; λ: The regularization parameter, which weighs the impact of the prediction error and model complexity on the objective function;

[0133] E2: Forecast error calculation; prediction error part || F(x t )-y t || 2 It represents the quadratic difference between the model's predicted value and the true value and is defined as follows:

[0134] ||F(x t )-y t || 2 =(F(x t )-y t ) 2

[0135] Where: F(x t ): The model is based on the input x t For example, x t Can be a sample data in a unified table view; t : True value (target variable value), usually derived from actual observation data or the target specified by business needs; quadratic difference (F(x t )-y t ) 2 Used to emphasize large forecast errors and avoid error cancellation;

[0136] E3: Model complexity metric. The model complexity metric function Ω(F) is a regularization term that limits the model complexity and is defined as follows:

[0137]

[0138] Where: Ω(F): model complexity value; w i : Model parameter w i (such as the i-th weight in the weight vector); n: the total number of model parameters.

[0139] By adjusting the model parameter w i Constraining the sum of squares can prevent the model parameter values ​​from being too large, thereby limiting the occurrence of overfitting;

[0140] E4: Dynamic feedback and optimization process: The system collects user operation behaviors (such as filtering fields, selecting charts, drilling up and down paths) and feedback data on the current analysis results in real time. These feedback data will be converted into target variables y t and input data x t , as the input of the optimization process; the system dynamically receives new data (such as real-time sales data or equipment operating status) and uses it as input data x t Incorporate model optimization calculations to ensure that the model can adapt to the latest data environment;

[0141] The optimization objective function is solved iteratively using the gradient descent method. The gradient update formula is:

[0142]

[0143] in: The value of the i-th weight after t+1 optimization iterations; The value of the i-th weight at the t-th optimization iteration; η: learning rate, used to control the step size of weight update;

[0144] The objective function is the weight w i The gradient of is calculated as:

[0145]

[0146] The first term is the gradient of the prediction error, and the second term is the gradient of the complexity regularization term. After the weight is updated, the model will be based on the latest weight value. Re-predict the input data to achieve dynamic adjustment.

[0147] Example 2:

[0148] like Figure 2 As shown, a multidimensional data intelligent analysis system based on Yibiaotong includes a data acquisition and preprocessing module, a semantic analysis module, a multidimensional feature modeling module, a visualization presentation module and a real-time feedback and optimization module; the data acquisition and preprocessing module is used to collect data from multiple data sources, and perform cleaning, standardization and semantic enhancement processing on the data to generate a unified table view; the semantic analysis module is used to calculate the semantic association weight matrix between data fields in the unified table view, and dynamically adjust the weights in combination with real-time new data; the multidimensional feature modeling module is used to extract key features from the unified table view and construct a multidimensional data analysis model based on field weights and business rules; the visualization presentation module is used to generate interactive charts, support up and down drilling, slice analysis and linkage display, and guide users to perform intelligent analysis through the path recommendation function; the real-time feedback and optimization module is used to collect user feedback data and real-time new data, and dynamically adjust the analysis model in combination with the optimization objective function.

[0149] In this embodiment, the data acquisition and preprocessing module includes a data interface unit, a field mapping unit and a data cleaning unit to realize multi-source data acquisition, cleaning and standardization, and finally generate a unified table; the semantic analysis module generates an association weight matrix through a semantic calculation unit, and updates the weight in real time through a dynamic adjustment unit to capture the semantic relationship between fields; the multidimensional feature modeling module includes a feature extraction unit and a model training unit, which builds a multidimensional analysis model and extracts key features based on a unified table view; the visualization module provides users with up and down drill-down analysis and intelligent path recommendation functions through a chart generation unit and a path recommendation unit, and intuitively displays data patterns; the feedback optimization module obtains user feedback data through a feedback acquisition unit, and dynamically adjusts model parameters in combination with a parameter optimization unit to ensure that the system continues to adapt to dynamically changing scenario requirements.

[0150] This solution addresses the technical bottlenecks of existing technologies in data integration, semantic association analysis, multidimensional modeling, and real-time optimization. The system implements a closed-loop process from efficient integration of multi-source data to output of analysis results, and features high efficiency, intelligence, real-time performance, and adaptability. Combined with intelligent visualization capabilities and continuous optimization mechanisms, the present invention can provide users with accurate analysis results and efficient decision-making support. This invention has a wide range of application scenarios and can significantly enhance the data processing and analysis capabilities of various industries, promoting the development of digital transformation and intelligent upgrades.

[0151] Finally, it should be noted that various modifications and variations of the present invention may be made by those skilled in the art without departing from the spirit and scope of the present invention. Thus, the present invention is intended to include such modifications and variations as fall within the scope of the claims and their equivalents.

Claims

1. A multi-dimensional data intelligent analysis method based on one table, characterized in that: The following steps are involved: S1: Collect data from multiple data sources, clean, standardize and semantically enhance the data, and generate a unified table view; S2: Based on the unified table view described in step S1, a semantic association weight matrix between data fields is calculated, and potential associations between data fields are captured through semantic labels and feature representations; S3: Extract key features from the unified table view, build a multidimensional data analysis model based on field weights and business rules, and capture complex nonlinear relationships between data; S4: Based on the output of the multidimensional data analysis model described in step S3, an interactive chart is generated, which supports drill-down, slice analysis, and linkage display, and recommends the optimal analysis path based on the user's historical operations and data characteristics; S5: Collect user feedback data and real-time new data, and dynamically optimize and adjust the multidimensional data analysis model.

2. The multidimensional data intelligent analysis method based on one table according to claim 1 is characterized in that: In step S1, when generating the unified table view, original data is collected from multiple data sources through a unified data interface, and the original data includes structured data, semi-structured data and unstructured data; During the preprocessing process, the system cleans and standardizes the collected data, removes outliers, fills in missing values, and ensures consistency in data format and range; at the same time, it uses field mapping and semantic labeling technology to semantically enhance the data and ultimately build a unified table view.

3. The multi-dimensional data intelligent analysis method based on one table according to claim 2 is characterized in that: In step S2, the calculation of the semantic association weight matrix includes: based on the semantic label set and feature vector of the field, using the semantic similarity algorithm to generate the semantic association relationship between the fields, and optimizing the weight matrix in real time through a dynamic adjustment mechanism.

4. The multi-dimensional data intelligent analysis method based on one table according to claim 3 is characterized in that: In step S3, key features are extracted according to the unified table view, and the features are modeled through a multidimensional data analysis model; The model combines the importance weights of fields and business rules to capture complex nonlinear relationships between data and adapt to business needs in different scenarios. During the modeling process, model parameters are dynamically adjusted based on real-time feedback and new data.

5. The multi-dimensional data intelligent analysis method based on one table according to claim 4 is characterized in that: In step S4, the generated interactive charts include the following types: drill-down analysis charts: allowing users to gradually drill down from high-dimensional overview data to specific dimensions, supporting multi-level data exploration; data slicing analysis charts: by filtering data of a specific dimension or range, generating corresponding subset analysis charts; linkage analysis charts: when a user selects a data point in a chart, other related charts will be updated synchronously to display the associated information of the data point.

6. The multi-dimensional data intelligent analysis method based on one table according to claim 5 is characterized in that: In step S5, the system collects the user's operation behavior and feedback data, and optimizes the analysis model in combination with real-time new data; by dynamically adjusting the model parameters, the system can continuously improve the accuracy and real-time performance of the analysis, while reducing the complexity of the model to ensure its stable performance in changing scenarios.

7. A multi-dimensional data intelligent analysis system based on one table, characterized in that: It includes data collection and preprocessing module, semantic analysis module, multi-dimensional feature modeling module, visualization module and real-time feedback and optimization module; The data collection and preprocessing module is used to collect data from multiple data sources, and perform cleaning, standardization and semantic enhancement processing on the data to generate a unified table view; The semantic analysis module is used to calculate the semantic association weight matrix between data fields in the unified table view, and dynamically adjust the weights in combination with real-time newly added data; The multidimensional feature modeling module is used to extract key features from the unified table view and build a multidimensional data analysis model based on field weights and business rules; The visualization module is used to generate interactive charts, support drill-down, slice analysis and linkage display, and guide users to perform intelligent analysis through the path recommendation function; The real-time feedback and optimization module is used to collect user feedback data and real-time new data, and dynamically adjust the analysis model in combination with the optimization objective function.

8. The method for intelligent analysis of multidimensional data based on one table according to claim 7, characterized in that: The data collection and preprocessing module includes a data interface unit, a field mapping unit and a data cleaning unit, which realizes multi-source data collection, cleaning and standardization, and finally generates a unified table; the semantic analysis module generates an association weight matrix through a semantic calculation unit, and updates the weight in real time through a dynamic adjustment unit to capture the semantic relationship between fields; The multidimensional feature modeling module includes a feature extraction unit and a model training unit, which builds a multidimensional analysis model and extracts key features based on a unified table view; The visualization module provides users with drill-down analysis and intelligent path recommendation functions through the chart generation unit and the path recommendation unit, and intuitively displays data patterns; the feedback optimization module obtains user feedback data through the feedback collection unit, and dynamically adjusts the model parameters in combination with the parameter optimization unit to ensure that the system continues to adapt to dynamically changing scenario requirements.

Citation Information

Cited By

  • Modeling method for multi-source data joint space

    CN121144287A

  • A Modeling Method for Joint Space of Multi-Source Data

    CN121144287B