A method and system for multidimensional dataset analysis and processing combined with AI

By combining AI with graph neural networks to construct a multidimensional dataset analysis method, the shortcomings of traditional methods in adaptability and implicit link capture are solved, and efficient and intelligent data analysis and display are achieved.

CN120744333BActive Publication Date: 2025-11-14CHENGDU BIG DATA GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511156548.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-14
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing technologies struggle to adapt to dynamic changes in business scenarios when analyzing multidimensional datasets. Manually maintaining association rules is costly, and they cannot capture implicit links, resulting in low data analysis efficiency and delays.

Method used

By combining AI with graph neural networks, we construct dimensional relationship graphs by cleaning data, generate dimensional feature representations using multi-class embedding technology, perform missing data imputation and data sorting, create a dimensional storage library, and optimize the display path through dimensional link sets.

Benefits of technology

It improves the accuracy and timeliness of data analysis, reduces human intervention, enhances the system's adaptability and intelligence, and optimizes data processing and display efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744333B_ABST
    Figure CN120744333B_ABST
Patent Text Reader

Abstract

This invention discloses a multidimensional dataset analysis and processing method and system that combines AI, belonging to the field of data analysis and processing technology. It includes: cleaning the multidimensional dataset, extracting dimensional features from the multidimensional dataset to obtain dimensional feature terms, which represent the number and type of dimensions corresponding to different data in the multidimensional dataset; integrating data based on the dimensional feature terms; imputing missing data in the multidimensional dataset to obtain a processed dataset. This invention constructs a dimensional link set and implements dynamic link adjustment. The system can intelligently identify key association paths, optimize the display method, and control the display complexity by setting display thresholds, ensuring that information is presented concisely and intuitively. These mechanisms work together to not only improve the automation level of data integration, storage, and display, but also significantly improve data processing efficiency and decision support capabilities, providing enterprises with a more efficient and intelligent analysis platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis and processing technology, specifically to a method and system for analyzing and processing multidimensional datasets using AI. Background Technology

[0002] Data analysis and processing is a combination of data processing and data analysis. It includes both the organization and processing of raw data and the in-depth mining and interpretation of data, ultimately serving to optimize decision-making. Data processing refers to the process of collecting, cleaning, transforming, and storing raw data, with the aim of removing invalid data such as missing values ​​and outliers, standardizing the format, and making the data meet the needs of analysis. Multidimensional dataset analysis and processing refers to the process of analyzing datasets composed of multiple dimensions. These datasets usually include information from multiple levels or perspectives, such as time, location, product category, user group, etc. The goal of multidimensional analysis is to observe and understand the structure, relationships, and trends of data from different perspectives simultaneously.

[0003] The data analysis and processing method and system disclosed in patent publication number CN105824974A allows access to functional nodes in a new data analysis and processing project. After importing data from target files, the system can call data calculation and processing scripts generated based on demand information to process the data. The system can execute data calculation and processing scripts to analyze and process complex data. All data processing is completed within the functional nodes, eliminating the need for data to flow between multiple nodes. This simplifies the data processing process and improves data processing efficiency.

[0004] When using the above and similar technical solutions, conventional methods rely on predefined dimensional levels and association rules, such as star or snowflake models, to analyze datasets. This makes it difficult to adapt to dynamic changes in business scenarios. Furthermore, the cost of manually maintaining association rules is high, and the system cannot capture the complex relationships hidden between dimensions, such as implicit links in the dataset. As a result, a high amount of computing power is wasted when analyzing and storing the dataset, and data adjustment speed is delayed. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for analyzing and processing multidimensional datasets that combines AI, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a multidimensional dataset analysis and processing method combining AI, comprising:

[0007] The multidimensional dataset is cleaned, and its dimensional features are extracted to obtain dimensional feature terms. These dimensional feature terms are used to represent the number of dimensions and the types of dimensions corresponding to different data in the multidimensional dataset.

[0008] Data integration is performed based on dimensional feature terms, and missing data is filled in the multidimensional dataset to obtain the processed dataset.

[0009] The dataset is sorted and analyzed, and the data is ordered according to its size to obtain a sorted dataset.

[0010] Based on the sorted dataset and the corresponding dimensional feature items, a dimensional storage library is created, and the sorted dataset is stored in categorized dimensions. The dimensional storage library includes at least two category directories corresponding to the dimensional feature items.

[0011] Obtain the link paths between dimension repositories to obtain a dimension link set, which includes at least the link paths between two dimension feature items.

[0012] The first display item is obtained by initially displaying links based on the dimensional link set;

[0013] Based on the continuous acquisition of dimensional feature terms, dimensional update terms are obtained. Based on the dimensional update terms, the classification dimensions of the dimensional storage library are adjusted to obtain link update terms.

[0014] The second display item is obtained by performing secondary link display based on the link update item. The link path of the classification dimension in the dimension storage library is optimized, so that the link path between classification dimensions with high weight is shorter, which facilitates the rapid storage of classification dimensions for the dimension update item. When displaying the stored dataset, the related dimension information is obtained based on the link optimization item and displayed synchronously, thereby improving the efficiency of analyzing, storing and displaying the dataset.

[0015] Furthermore, the method for obtaining the dimensional feature terms includes:

[0016] Clean the cube by removing duplicate data to obtain cleaned data items;

[0017] Based on cleaned data items, multi-class embedding technology is used to map categories to a continuous vector space. A relationship graph between dimensions is constructed using a dimensional hierarchy graph, transforming the hierarchical relationships in multidimensional data into a graph structure. A graph neural network is then used to generate dimensional feature representations, thereby obtaining dimensional feature items.

[0018] Furthermore, the method for obtaining the processed dataset includes:

[0019] Based on the dimensional feature terms corresponding to the multidimensional dataset, obtain the corresponding data for the missing dimensional feature terms in the multidimensional dataset to obtain the target data terms;

[0020] Based on the location of the missing dimensional features in the target data item, the target location item is obtained. The average value of the dimensional features corresponding to the target location item in the multidimensional dataset is obtained to obtain the target feature item. The target feature item is used as the missing filler to obtain the processed dataset.

[0021] Furthermore, the method for creating the dimensional repository includes:

[0022] Based on the feature types and number of features of the dimension feature items corresponding to the first sorted data information in the sorted dataset, a first storage library is created, which includes at least one storage directory for a dimension feature item.

[0023] Sequentially obtain the feature types and number of features of the dimension feature items corresponding to the sequentially sorted data information in the sorted dataset, and expand the feature types of the first storage to obtain the expanded feature items.

[0024] Based on the combination of the first storage library and the expanded category items, the dimension storage library is obtained.

[0025] Furthermore, the method for obtaining the dimensional link set includes:

[0026] Obtain the location information and list sorting information of the category directories corresponding to different dimension feature items in the dimension storage library to obtain the directory information items;

[0027] Based on the data acquisition method, the directory information items are linked to obtain the relationship link paths, a dimensional relationship graph is constructed, and the directory link items are obtained. The directory link items are used to represent the link paths between the target category directory and other category directories. The dimensional link set is obtained based on the comprehensive statistics of the directory link items.

[0028] Furthermore, the method for obtaining the first display item includes:

[0029] Obtain the corresponding category directory of the target display dimension in the dimension storage library to get the display target item;

[0030] Based on the dimensional link set, the target link path of the target item link and the link category directory corresponding to the link path are obtained and displayed to get the displayed link item;

[0031] The first item to be displayed is obtained by using the target item as the primary display target and the linked items as the secondary display targets.

[0032] Furthermore, the method for obtaining the link update item includes:

[0033] Based on the dimension update item, obtain the dimension storage quantity information of the classification directory corresponding to different dimension feature items in the dimension storage library, and obtain the dimension data item;

[0034] The dimensional data items are sorted by storage size. Based on the sorting result, the positions of the directory information items are adjusted, and the relationships are linked based on the adjusted positions to obtain the relationship link paths. The dimensional relationship graph is then reconstructed to obtain the updated link items, and thus the link update items are obtained.

[0035] Furthermore, the method for obtaining the second display item includes:

[0036] Set a display threshold, which is a threshold for the number of link paths. Based on the display threshold, the displayed link items are eliminated to obtain the filtered link items.

[0037] The primary display target is the item to be displayed, and the secondary display target is the filtered link item, resulting in the second display item.

[0038] Furthermore, a multidimensional dataset analysis and processing system incorporating AI utilizes the aforementioned multidimensional dataset analysis and processing method incorporating AI, including:

[0039] Data processing module: Cleans the multidimensional dataset, extracts the dimensional features of the multidimensional dataset to obtain dimensional feature terms, integrates the data based on the dimensional feature terms, imputes missing data in the multidimensional dataset, and obtains the processed dataset;

[0040] Data Analysis Module: Sorts and analyzes the processed dataset, sorts the data according to its size to obtain a sorted dataset, creates a dimension store based on the sorted dataset and its corresponding dimensional features, and stores the sorted dataset by category dimension.

[0041] Data visualization module: Obtain the link paths between dimension storage libraries to obtain a dimension link set. The dimension link set includes the link paths between at least two dimension feature items. Perform preliminary link display based on the dimension link set to obtain the first display item.

[0042] Updated display module: Based on the continuous acquisition of dimensional feature items, dimensional update items are obtained. Based on the dimensional update items, the link of the classification dimension in the dimensional storage library is adjusted to obtain link update items. Based on the link update items, secondary link display is performed to obtain the second display item. The link path of the classification dimension in the dimensional storage library is optimized to facilitate the rapid storage of dimensional update items for classification dimension. When displaying the stored dataset, the associated dimension information is obtained based on the link optimization items and displayed synchronously.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] A multidimensional dataset analysis and processing method and system combining AI is proposed. By cleaning the data and constructing a dimensional relationship graph, the hierarchical structure of multidimensional data is transformed into a graph structure. Then, a graph neural network is used to generate dimensional feature representations, enabling the system to automatically identify and adapt to new dimensions. This process not only reduces manual intervention and improves the generalization ability of the model, but also captures the evolution trend of the business environment in real time, thereby enhancing the accuracy and timeliness of data analysis. In addition, by continuously acquiring dimensional feature items and dynamically optimizing dimensional links, the system can automatically adjust the data classification and display path, making the relationship between high-weight dimensions more intuitive, and significantly improving the intelligence level of data processing and business response efficiency.

[0045] Meanwhile, the introduction of a missing item imputation strategy based on dimensional features during the data integration phase effectively improves data integrity and analysis quality. In the data sorting and classification storage stage, the system sorts data according to its size and builds a multi-dimensional storage library, providing structured support for subsequent link construction and display. In addition, by constructing a set of dimensional links and realizing dynamic link adjustment, the system can intelligently identify key association paths, optimize the display method, and control the display complexity by setting display thresholds to ensure that information is presented in a concise and intuitive manner. These mechanisms work together to not only improve the automation level of data integration, storage and display, but also significantly improve data processing efficiency and decision support capabilities, providing enterprises with a more efficient and intelligent analysis platform. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the overall process of the present invention;

[0047] Figure 2 This is a schematic diagram illustrating the sequential sorting of dataset A according to the present invention;

[0048] Figure 3 This is a schematic diagram of the dimension storage library acquisition process of the present invention;

[0049] Figure 4 This is a schematic diagram illustrating the expansion of the first storage feature types of the present invention;

[0050] Figure 5 This is a schematic diagram of the directory link item relationship of the present invention. Detailed Implementation

[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] Traditional data analysis methods, such as star or snowflake schemas based on predefined hierarchical dimensions and association rules, reveal a series of limitations when facing complex and ever-changing business scenarios, directly impacting the efficiency and quality of data analysis. First, insufficient adaptability is a major drawback of traditional methods. These models rely on pre-defined dimensions and association rules, failing to flexibly capture dynamic changes in business scenarios. When business models shift or new data dimensions emerge, large-scale adjustments and reconstructions of the model are required. This not only consumes significant time and resources but can also lead to lag in data analysis, hindering timely support for decision-making. For example, a retail company might initially build a data warehouse based on geographic, product category, and time dimensions. However, with the expansion of online business and the rise of social media, new dimensions such as user behavior, reviews, and preferences become crucial. If traditional models fail to incorporate these new dimensions in a timely manner, they will struggle to accurately grasp market trends and user needs. Second, the cost of manually maintaining association rules remains high. In traditional models, association rules between dimensions typically require manual definition and maintenance. This not only demands in-depth involvement from domain experts but is also susceptible to subjective influences, leading to inconsistencies in the rules. The inaccuracies or omissions in the rules, coupled with the increasing volume of data and the complexity of business operations, lead to a surge in the number of association rules and an exponential increase in maintenance difficulty. This results in enterprises investing significant resources in data analysis without achieving desired results. Furthermore, traditional methods struggle to effectively identify and analyze implicit links within the dataset, such as potential bottlenecks in the supply chain network. This application provides a multidimensional dataset analysis and processing method that combines AI. By introducing graph neural networks and embedding technology, it significantly enhances the intelligence and adaptability of data analysis. By cleaning the data and constructing a dimensional relationship graph, the hierarchical structure of multidimensional data is transformed into a graph structure. Then, graph neural networks are used to generate dimensional feature representations, enabling the system to automatically identify and adapt to new dimensions. This process not only reduces manual intervention and improves the model's generalization ability but also captures real-time trends in the evolving business environment, thereby enhancing the accuracy and timeliness of data analysis. Furthermore, by continuously acquiring dimensional feature items and dynamically optimizing dimensional links, the system can automatically adjust data classification and display paths, making the relationships between high-weight dimensions more intuitive. This significantly improves the intelligence level of data processing and business response efficiency. Figure 1 As shown, it includes steps S100-S800.

[0053] Step S100: Clean the multidimensional dataset and extract its dimensional features.

[0054] It is important to note that after cleaning and extracting the dimensional features of the multidimensional dataset, dimensional feature terms are obtained. These dimensional feature terms represent the number and types of dimensions corresponding to different data in the multidimensional dataset. The methods for obtaining dimensional feature terms include: cleaning the multidimensional dataset to remove duplicate data and obtain cleaned data terms; based on the cleaned data terms, using multi-class embedding technology to map the categories to a continuous vector space, constructing a relationship graph between dimensions using a dimensional hierarchy graph, transforming the hierarchical relationship in the multidimensional data into a graph structure, and generating dimensional feature representations using a graph neural network, thereby obtaining dimensional feature terms.

[0055] Specifically, the categorical fields are first converted into integer indices to ensure that each category corresponds to a unique number. Then, during the data cleaning stage, missing values ​​and outliers are handled to ensure the rationality of the category coding. Next, a set of random low-dimensional vectors (e.g., 8, 16, or 32 dimensions) is assigned to each category through random initialization as the embedding starting point. Pre-trained models such as Word2Vec and FastText are used as the starting point for embedding. The category embedding vectors are then used as features input to subsequent models, such as classifiers and predictors. A loss function is defined according to the target task, and backpropagation is used to optimize the embedding vectors. Target tasks include prediction and clustering. Then, nodes are defined through the construction of a relation matrix or graph. Nodes represent categories or feature items. Directed or undirected edges are established between nodes according to the relation type, and edge weights are assigned to reflect the closeness of the relationship. Edge weights include relation strength, co-occurrence frequency, correlation coefficient, etc. Relationships are stored using a relational graph database or graph structure data structure. Finally, a graph neural network is used to generate dimensional feature representations, thus obtaining dimensional feature items.

[0056] Step S200: Integrate data based on dimensional feature terms, impute missing data in the multidimensional dataset, and obtain the processed dataset.

[0057] It is important to note that the methods for obtaining the processed dataset include: obtaining the corresponding data for missing dimensional features in the multidimensional dataset based on the dimensional feature items corresponding to the multidimensional dataset, thus obtaining the target data item; obtaining the target position item based on the location of the missing dimensional features in the target data item; obtaining the average value of the dimensional features corresponding to the target position item in the multidimensional dataset, thus obtaining the target feature item; using the target feature item as the missing filler, thus obtaining the processed dataset.

[0058] In the specific implementation process, there is currently a sales data table of a retail company, including the following fields: order number, region, product category, sales quantity, sales amount, and order date. These fields together form a multidimensional dataset, the contents of which are shown in Table 1.

[0059] Table 1

[0060] Order Number area Product Category Sales volume Sales Order date 1001 Region A cell phone 2 2000 2025-07-01 1002 flat 1 1000 2025-07-02 1003 Region B cell phone 2 2025-07-03 1004 Region A notebook 1 4000 2025-07-04 1005 Region A cell phone 2 2000 2025-07-05

[0061] After inspection, it was found that some "Region" and "Sales Amount" fields had missing values, which needed to be filled. The corresponding data for the missing dimensional features in the cube were obtained to get the target data items, namely "Region" and "Sales Amount". Based on the location of the missing dimensional features in the target data items, the target position items were obtained, namely "1002" and "1003". The average value of the dimensional features corresponding to the target position items in the cube was obtained to get the target feature items. These target feature items were used for filling the missing values. The average value of the dimensional features corresponding to "1002" in the cube was "Region A", and the average value of the dimensional features corresponding to "1003" was "2000". Therefore, "Region A" and "2000" were used as the missing values ​​respectively to obtain the processed dataset.

[0062] Step S300: Sort and analyze the data in the dataset.

[0063] It is important to note that, such as Figure 2 As shown, the data is sorted according to its size to obtain a sorted dataset. In a multidimensional dataset, since it is composed of data from different dimensions, there are differences in size between the data. For example, a dataset A includes video content a, text content b, and image content c. There is a difference in the size of these three contents: video content a has a size of 150Mb, text content b has a size of 10Mb, and image content c has a size of 40Mb. Sort the data according to its size to obtain a sorted dataset, i.e., video content a > image content c > text content b.

[0064] Step S400: Create a dimension store based on the sorted dataset and the corresponding dimensional feature items.

[0065] It is important to note that, such as Figure 3 As shown, the sorted dataset is stored in a categorical dimension library. The dimension library includes at least two categorical directories corresponding to categorical feature items. The method for creating the dimension library includes: creating a first library based on the feature types and feature quantities of the categorical feature items corresponding to the first-ranked data information in the sorted dataset. The first library includes a storage directory for at least one categorical feature item; sequentially obtaining the feature types and feature quantities of the categorical feature items corresponding to the sequentially ranked data information in the sorted dataset, and expanding the feature types of the first library to obtain expanded feature items; and obtaining the dimension library based on the combination result of the first library and the expanded feature items.

[0066] In the specific implementation process, such as Figure 4As shown, a multidimensional dataset is obtained, including data 1, data 2, data 3, data 4, and data 5. After sorting, the sorting result is data 5 > data 4 > data 3 > data 2 > data 1, resulting in a sorted dataset. The sorted dataset is then categorized and stored. The feature types and number of the first sorted data are obtained, which are the feature types and number of the feature items corresponding to the first sorted data. For example, the feature type of data 5 is time and location, and the feature number is 2. A first storage library is then created, containing two storage directories: time and location. Next, the feature types and number of the feature items corresponding to data 4 are obtained. The feature types of data 4 are time, location, and people. The first storage library is then expanded with additional feature types, resulting in a "people" storage directory. This process is repeated to obtain the final dimension storage library.

[0067] Step S500: Obtain the link paths between dimension storage libraries to obtain the dimension link set.

[0068] It is important to note that a dimensional link set includes at least two link paths between dimensional feature items. The method for obtaining a dimensional link set includes: obtaining the location information and list sorting information of the category directories corresponding to different dimensional feature items in the dimensional storage library to obtain the directory information items; based on the data acquisition method, performing relational linking on the directory information items to obtain the relational link paths, constructing a dimensional relationship graph to obtain the directory link items. The directory link items are used to represent the link paths between the target category directory and other category directories. The dimensional link set is obtained based on the comprehensive statistics of the directory link items.

[0069] In the specific implementation process, such as Figure 5 As shown, a dimensional storage library A has been created. The category directories corresponding to different dimensional feature items in the dimensional storage library are: time, season, location, patient, symptom, and medicine. The location information and list sorting information are time-season-location-patient-symptom-medicine, respectively, to obtain the directory information items. At this time, based on the data acquisition method, the directory information items are linked. There are relationship link paths between patient-symptom-medicine. A dimensional relationship graph is constructed to obtain the directory link items.

[0070] Step S600: Perform preliminary link display based on the dimensional link set to obtain the first display item.

[0071] It should be noted that the method for obtaining the first display item includes: obtaining the corresponding category directory of the target display dimension in the dimension storage library to obtain the display target item; obtaining the target link path and the link category directory corresponding to the link path based on the dimension link set to obtain the display link item; and using the display target item as the primary display target and the display link item as the secondary display target to obtain the first display item.

[0072] Specifically, when staff need to display a dataset, the target display dimension corresponds to different category directories in the dimension storage library. At this time, the corresponding category directory is used as the primary display target. Since there are also relationship link paths with this category directory, the link category directory corresponding to the relationship link path is obtained as the secondary display target.

[0073] In the specific implementation process, a medical staff member needs to obtain the symptoms related to a certain disease from a multidimensional dataset. Therefore, the symptoms are the primary display target. The patients and drugs that are related to the symptoms are also displayed together as secondary display targets.

[0074] Step S700: Adjust the classification dimensions of the dimension storage library based on the continuous acquisition of dimensional feature items.

[0075] It is important to note that the continuous acquisition of dimensional feature items yields dimensional update items. Based on these dimensional update items, the classification dimensions in the dimensional storage library are adjusted to obtain link update items. The method for obtaining link update items includes: based on the dimensional update items, obtaining the dimensional storage volume information of the classification directories corresponding to different dimensional feature items in the dimensional storage library to obtain dimensional data items; sorting the dimensional data items by storage volume; adjusting the position of the directory information items based on the sorting result; establishing relationship links based on the adjusted positions to obtain relationship link paths; reconstructing the dimensional relationship graph to obtain updated link items, and thus obtaining link update items.

[0076] In the specific implementation process, a dimension store B is created. The category directories corresponding to different dimension feature items in the dimension store are: time, season, location, patient, symptom, and medicine. The location information and list sorting information are time-season-location-patient-symptom-medicine, respectively, to obtain the directory information items. At this time, based on the continuous acquisition of dimension feature items, the dimension storage volume information of the category directories corresponding to different dimension feature items in the dimension store is obtained as 12%, 5%, 9%, 25%, 30%, and 19%, respectively. Then, the storage volume is sorted by size, and the sorting result is symptom-patient-medicine-time-location-season. The position of the directory information items is adjusted, and the relationship link is established based on the adjusted position to obtain the relationship link path. The dimension relationship graph is reconstructed to obtain the updated link item, and then the link update item is obtained.

[0077] Step S800: Perform secondary link display based on the link update item to obtain the second display item.

[0078] It is important to note that optimizing the link paths of categorical dimensions in the dimensional storage library shortens the link paths between categorical dimensions with high weights, facilitating the rapid storage of categorical dimensions for updated dimensional items. Furthermore, when displaying the stored dataset, the related dimensional information is retrieved based on the link optimization items and displayed synchronously, thereby improving the efficiency of dataset analysis, storage, and display.

[0079] It is important to note that the method for obtaining the second display item includes: setting a display threshold, which is a threshold for the number of link paths, with a display threshold of 2; eliminating display link items based on the display threshold to obtain filtered link items, meaning that the maximum number of display link items is 2; using the display target item as the primary display target and the filtered link items as the secondary display target to obtain the second display item, thus limiting the number of display link items and ensuring the simplicity of the display.

[0080] A multidimensional dataset analysis and processing system combining AI, employing the aforementioned AI-integrated multidimensional dataset analysis and processing method, includes: a data processing module: cleaning the multidimensional dataset, extracting its dimensional features to obtain dimensional feature terms, integrating data based on these dimensional feature terms, and imputing missing data to obtain a processed dataset; a data analysis module: sorting and analyzing the processed dataset, ranking it according to data size to obtain a sorted dataset, creating a dimensional storage library based on the sorted dataset and its corresponding dimensional feature terms, and storing the sorted dataset by category; and a data display module: acquiring… The link paths between dimensional storage libraries are used to obtain dimensional link sets. Each dimensional link set includes link paths between at least two dimensional feature items. Based on the dimensional link sets, a preliminary link display is performed to obtain the first display item. The update display module obtains dimensional update items based on the continuous acquisition of dimensional feature items. Based on the dimensional update items, the link of the classification dimension in the dimensional storage library is adjusted to obtain link update items. Based on the link update items, a secondary link display is performed to obtain the second display item. The link paths of the classification dimension in the dimensional storage library are optimized to facilitate the rapid storage of classification dimension update items. Furthermore, when displaying the stored dataset, the associated dimension information is obtained based on the link optimization items and displayed synchronously.

[0081] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended embodiments and their equivalents.

Claims

1. A method for analyzing and processing multidimensional datasets using AI, comprising: The multidimensional dataset is cleaned, and its dimensional features are extracted to obtain dimensional feature terms. These dimensional feature terms are used to represent the number of dimensions and the types of dimensions corresponding to different data in the multidimensional dataset. Data integration is performed based on dimensional feature terms, and missing data is filled in the multidimensional dataset to obtain the processed dataset. Its features are: The dataset is sorted and analyzed, and the data is ordered according to its size to obtain a sorted dataset. Based on the sorted dataset and the corresponding dimensional feature items, a dimensional storage library is created, and the sorted dataset is stored in categorized dimensions. The dimensional storage library includes at least two category directories corresponding to the dimensional feature items. The method for creating the dimensional repository includes: Based on the feature types and number of features of the dimension feature items corresponding to the first sorted data information in the sorted dataset, a first storage library is created, which includes at least one storage directory for a dimension feature item. Sequentially obtain the feature types and number of features of the dimension feature items corresponding to the sequentially sorted data information in the sorted dataset, and expand the feature types of the first storage to obtain the expanded feature items. Based on the combination of the first storage library and the expanded category items, the dimension storage library is obtained; Obtain the link paths between dimension repositories to obtain a dimension link set, which includes at least the link paths between two dimension feature items. The first display item is obtained by initially displaying links based on the dimensional link set; Based on the continuous acquisition of dimensional feature terms, dimensional update terms are obtained. Based on the dimensional update terms, the classification dimensions of the dimensional storage library are adjusted to obtain link update terms. The second display item is obtained by performing secondary link display based on the link update item. The link path of the classification dimension in the dimension storage library is optimized, so that the link path between classification dimensions with high weight is shorter, which facilitates the rapid storage of classification dimensions for the dimension update item. When displaying the stored dataset, the related dimension information is obtained based on the link optimization item and displayed synchronously, thereby improving the efficiency of analyzing, storing and displaying the dataset. The multidimensional dataset includes order number, region, product category, sales quantity, sales amount, and order date.

2. The multidimensional dataset analysis and processing method combining AI according to claim 1, characterized in that: The method for obtaining the dimensional feature terms includes: Clean the cube by removing duplicate data to obtain cleaned data items; Based on cleaned data items, multi-class embedding technology is used to map categories to a continuous vector space. A relationship graph between dimensions is constructed using a dimensional hierarchy graph, transforming the hierarchical relationships in multidimensional data into a graph structure. A graph neural network is then used to generate dimensional feature representations, thereby obtaining dimensional feature items.

3. The multidimensional dataset analysis and processing method combining AI according to claim 1, characterized in that: The method for obtaining the dataset includes: Based on the dimensional feature terms corresponding to the multidimensional dataset, obtain the corresponding data for the missing dimensional feature terms in the multidimensional dataset to obtain the target data terms; Based on the location of the missing dimensional features in the target data item, the target location item is obtained. The average value of the dimensional features corresponding to the target location item in the multidimensional dataset is obtained to obtain the target feature item. The target feature item is used as the missing filler to obtain the processed dataset.

4. The multidimensional dataset analysis and processing method combining AI according to claim 1, characterized in that: The method for obtaining the dimensional link set includes: Obtain the location information and list sorting information of the category directories corresponding to different dimension feature items in the dimension storage library to obtain the directory information items; Based on the data acquisition method, the directory information items are linked to obtain the relationship link paths, a dimensional relationship graph is constructed, and the directory link items are obtained. The directory link items are used to represent the link paths between the target category directory and other category directories. The dimensional link set is obtained based on the comprehensive statistics of the directory link items.

5. The multidimensional dataset analysis and processing method combining AI according to claim 1, characterized in that: The method for obtaining the first display item includes: Obtain the corresponding category directory of the target display dimension in the dimension storage library to get the display target item; Based on the dimensional link set, the target link path of the target item link and the link category directory corresponding to the link path are obtained and displayed to get the displayed link item; The first item to be displayed is obtained by using the target item as the primary display target and the linked items as the secondary display targets.

6. The multidimensional dataset analysis and processing method combining AI according to claim 4, characterized in that: The method for obtaining the link update item includes: Based on the dimension update item, obtain the dimension storage quantity information of the classification directory corresponding to different dimension feature items in the dimension storage library, and obtain the dimension data item; The dimensional data items are sorted by storage size. Based on the sorting result, the positions of the directory information items are adjusted, and the relationships are linked based on the adjusted positions to obtain the relationship link paths. The dimensional relationship graph is then reconstructed to obtain the updated link items, and thus the link update items are obtained.

7. The multidimensional dataset analysis and processing method combining AI according to claim 5, characterized in that: The methods for obtaining the second display item include: Set a display threshold, which is a threshold for the number of link paths. Based on the display threshold, the displayed link items are eliminated to obtain the filtered link items. The primary display target is the item to be displayed, and the secondary display target is the filtered link item, resulting in the second display item.

8. A multidimensional dataset analysis and processing system combining AI, characterized in that: The method for analyzing and processing multidimensional datasets using any one of claims 1-7, comprising: Data processing module: Cleans the multidimensional dataset, extracts the dimensional features of the multidimensional dataset to obtain dimensional feature terms, integrates the data based on the dimensional feature terms, imputes missing data in the multidimensional dataset, and obtains the processed dataset; Data Analysis Module: Sorts and analyzes the processed dataset, sorts the data according to its size to obtain a sorted dataset, creates a dimension store based on the sorted dataset and its corresponding dimensional features, and stores the sorted dataset by category dimension. Data visualization module: Obtain the link paths between dimension storage libraries to obtain a dimension link set. The dimension link set includes the link paths between at least two dimension feature items. Perform preliminary link display based on the dimension link set to obtain the first display item. Updated display module: Based on the continuous acquisition of dimensional feature items, dimensional update items are obtained. Based on the dimensional update items, the link of the classification dimension in the dimensional storage library is adjusted to obtain link update items. Based on the link update items, secondary link display is performed to obtain the second display item. The link path of the classification dimension in the dimensional storage library is optimized to facilitate the rapid storage of dimensional update items for classification dimension. When displaying the stored dataset, the associated dimension information is obtained based on the link optimization items and displayed synchronously.

Citation Information

Patent Citations

  • Method and system for analyzing and processing data

    CN105824974A

  • Database cascade operation intelligent analysis execution method and system

    CN120256103A

  • Integrated test scene prediction method and system based on multi-dimensional data

    CN120336190A