Database multi-dimensional data management system with high expandability

By designing a highly scalable database multidimensional data management system, the problems of data semantic inconsistency and privacy protection in enterprises were solved, and the data consistency, security and compliance were improved, ensuring the efficiency and accuracy of data analysis.

CN121765030APending Publication Date: 2026-03-31TIANJIN SHENZHOU GENERAL DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional database management systems suffer from semantic inconsistencies and interoperability issues when processing multidimensional data, and struggle to meet data privacy and compliance requirements. This is especially true when different departments within an enterprise use different terminology and data storage systems, making it difficult to guarantee data security and compliance.

Method used

A highly scalable database multidimensional data management system was designed, including a data integration module, a semantic management module, a privacy management module, and a bias prediction module. Through data integration, semantic unification, privacy protection, and transparency management, data consistency and security are ensured.

Benefits of technology

It achieves semantic consistency of data across different departments, improves data comparability and accuracy, ensures the security and compliance of privacy data, reduces the risk of privacy leaks, and improves the efficiency and accuracy of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765030A_ABST
    Figure CN121765030A_ABST
Patent Text Reader

Abstract

The invention discloses a database multi-dimensional data management system with high expandability, and relates to the technical field of enterprise data management, the database multi-dimensional data management system comprises a data integration module, the data integration module comprises a collection unit and a distributed storage unit, the collection unit is used for collecting and integrating data, and the distributed storage unit is used for storing the data; the distributed storage unit is used for storing the data of the collection unit; the semantic management module is used for defining unified data terms for the data stored in the data integration module; through the application of the semantic management module, an enterprise can unify the use of key terms among different departments, so that the consistency of data is improved, the comparability and accuracy of the data in the whole organization are ensured through the consistency, and wrong decisions caused by term confusion or misunderstanding are reduced; the method is helpful for enterprises to better identify and process personal privacy information, and ensures that data collection and processing activities meet the requirements of laws and regulations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise data management technology, and more specifically, to a highly scalable database multidimensional data management system. Background Technology

[0002] In today's information age, the explosive growth of enterprise data has increased the difficulty of database management. Traditional office software such as Excel has limited samples, single data, and low efficiency in exporting basic data statistics from management systems. Moreover, data analysis is often subjective and it is difficult to guarantee the quality of analysis. It can no longer meet the needs of modern management work. To improve the efficiency of database management and adapt to the continuous growth of data volume, highly scalable database multidimensional data management systems have also come into the public eye. Their main function is to process and analyze multidimensional data, and they are widely used in many fields. However, in actual use, due to the large number of departments in an enterprise, different terms and definitions may be used to describe the same data entities, which leads to semantic inconsistency of data. Moreover, as enterprises adopt multiple SaaS applications and cloud services, data may be stored in different systems that may use different data models and technology stacks, resulting in data interoperability issues. Furthermore, for medical companies or human resources departments within these companies, there is a lot of sensitive user information. With the strengthening of data privacy regulations, users do not have much control over their data, including the right to withdraw consent and the right to have their data forgotten, which may lead to privacy leaks. Summary of the Invention

[0003] To address the aforementioned problems, this invention provides a highly scalable database multidimensional data management system.

[0004] This invention provides a highly scalable database multidimensional data management system, comprising: A data integration module, comprising a collection unit and a distributed storage unit, wherein the collection unit is used to collect and integrate data, and the distributed storage unit is used to store the data collected by the collection unit; A semantic management module is used to define unified data terms for the data stored in the data integration module, ensuring that different departments and systems have a consistent understanding of the data; The privacy management module includes an identification unit and a management unit. The identification unit is used to identify privacy data in the data integration module, and the management unit is used to manage the data consent requests of users corresponding to the privacy data, including requests to withdraw consent and requests for data to be forgotten. A bias prediction module is provided to offer transparency to the semantic management module and the privacy management module, ensuring the fairness of the results.

[0005] Preferably, the specific operation of the collection unit is as follows: Identify the data sources to be integrated and access data from the identified data sources; Integrate data from different sources according to specific logic; Verify that the integrated data meets the expected quality. The matching data is transmitted to the distributed storage unit; The specific working principle of the distributed storage unit is as follows: Receive integrated data pushed from the collection unit; Based on the preset sharding strategy, the data is distributed across different nodes; The sharded and replicated data is stored in a distributed database.

[0006] Preferably, the data integration module further includes an expansion unit, which is used to increase the scalability of the collection unit and the distributed storage unit, specifically: Based on the current and expected load, determine the number of nodes required for the distributed storage unit to ensure that the distributed storage unit can handle the expected requests. The total resource requirement of a single node is calculated by superimposing the CPU, memory, and storage resource requirements of a single node in the distributed storage unit. The total resource requirement of a single node is then multiplied by the number of nodes required by the distributed storage unit to obtain the total resource requirement of the distributed storage unit. The number of nodes is automatically adjusted based on the actual load.

[0007] Preferably, the specific working method of automatically adjusting the number of nodes according to the actual load further includes the following: A load threshold for the distributed storage unit is set in advance. ; According to the formula The new number of nodes is calculated. ,in This represents the current actual load of the distributed storage unit.

[0008] Preferably, the semantic management module operates as follows: The information gain value P of each term in the text data within the data integration module is obtained; Based on the obtained term information gain value P, the frequency of use and context of terms in different data sources are analyzed to identify semantic differences and obtain semantic difference value F. Based on the semantic difference value F and the information gain value P of each term, the relevance between terms and potential substitute terms is evaluated. For each term, calculate the correlation coefficient of all its possible substitute terms and multiply it by the information gain value of the term to obtain the comprehensive coefficient of each substitute term. Then select the term with the highest comprehensive coefficient as the recommended harmonized term. The generated harmonization terms will replace the terms in the text data within the data integration module, ensuring that all relevant data fields and documents use the new terms.

[0009] Preferably, the identification unit operates as follows: The data in the data integration module is categorized, and data fields containing personal privacy information are marked; Use regular expressions and keyword matching to identify data in data fields that contain personal privacy information; Assess the sensitivity level of the identified data, and mark data with a sensitivity level higher than a pre-set threshold as private data.

[0010] Preferably, the management unit operates as follows: Create and maintain records of user consent data; Monitor the access and usage of monitoring data; Handling users' right to be forgotten requests, Allow users to withdraw their previously given consent to data processing.

[0011] Preferably, the bias prediction module operates as follows: According to the formula ,in It is the weight of the i-th data. It is the recognition ratio of the recognition unit in the i-th data. This is the actual proportion; Obtain the deviation value B of the recognition unit; Further quantify the degree of deviation; Apply bias reduction techniques to reduce identified biases; Generate a transparency report detailing the model's biases and the mitigation measures taken.

[0012] Preferably, the specific process of reducing the identified deviation using the applied deviation reduction technology is as follows: Resample to obtain the adjusted sample weight values. ; Adjusted sample weight values Substitute the bias prediction module to obtain the new bias value B1 of the identification unit; Until the deviation value B1 of the obtained identification unit meets the preset requirements.

[0013] Preferably, the bias prediction module further includes a feedback adjustment unit, which operates as follows: Obtain the information content adjustment factor for the current time period. ; The obtained confidence adjustment factor is based on feedback. ; According to the formula The updated confidence adjustment factor is calculated and obtained. ,in It is a smoothing factor; Output the updated confidence adjustment factor Confidence adjustment factor The higher confidence adjustment factor directly affects the magnitude of the recognition unit update. This means the model will be updated significantly when it receives new data or feedback to quickly adapt to changes; a lower confidence adjustment factor. This means that model updates are more conservative, avoiding instability caused by over-adjustment.

[0014] Beneficial effects: By applying the semantic management module, enterprises can standardize the use of key terms across different departments, thereby improving data consistency. This consistency ensures the comparability and accuracy of data throughout the organization, reducing erroneous decisions caused by terminology confusion or misunderstanding. It helps companies better identify and process personal privacy information, ensure that data collection and processing activities comply with legal and regulatory requirements, and reduce the risk of data breaches and protect users' privacy rights by identifying and classifying sensitive information and implementing appropriate protection measures. Attached Figure Description

[0015] Figure 1 This is a flowchart of the management system of the present invention; Figure 2 How the collection unit works; Figure 3 How distributed storage units work. Detailed Implementation

[0016] Application scenarios: Through the design of the above modules, this multidimensional data management system can effectively handle issues such as data interoperability, semantic consistency, privacy protection, and algorithm transparency in enterprises, ensuring data security and compliance, while improving the efficiency and accuracy of data analysis; However, in actual use, due to the large number of departments in an enterprise, different terms and definitions may be used to describe the same data entities, which leads to semantic inconsistency of data. Moreover, as enterprises adopt multiple SaaS applications and cloud services, data may be stored in different systems that may use different data models and technology stacks, resulting in data interoperability issues. Furthermore, for medical companies or human resources departments within these companies, there is a lot of sensitive user information. With the strengthening of data privacy regulations, users do not have much control over their data, including the right to withdraw consent and the right to have their data forgotten, which may lead to privacy leaks.

[0017] like Figure 1 As shown: A highly scalable database multidimensional data management system includes a data integration module, a semantic management module, a privacy management module, and a bias prediction module; By designing the distributed storage unit as a node, nodes can be added as needed, satisfying the scalability requirements during use. Moreover, resource allocation, load balancing, monitoring, and expansion are all automated, achieving high scalability. The data integration module includes a collection unit and a distributed storage unit. The collection unit is used to collect and integrate data, and the distributed storage unit is used to store the data collected by the collection unit. The semantic management module is used to define unified data terms for the data stored in the data integration module, ensuring that different departments and systems have a consistent understanding of the data. The privacy management module includes an identification unit and a management unit. The identification unit is used to identify privacy data in the data integration module, and the management unit is used to manage the data consent requests of users corresponding to the privacy data, including requests to withdraw consent and requests for data to be forgotten. The bias warning module is used to provide transparency to the semantic management module and the privacy management module, ensuring the fairness of the results.

[0018] As an optional embodiment: such as Figure 2 As shown, the specific working method of the collection unit is as follows: Identify the data sources to be integrated and access data from the identified data sources; the specific access methods can be to access data from the identified data sources, such as API calls, database queries, file imports, etc. Data from different sources can be integrated according to specific logic; integration can be based on primary key matching, timestamp sorting, etc. Verify whether the integrated data meets the expected quality; in this embodiment, the integrity, consistency, and accuracy of the data can be calculated to make the judgment. The matching data is transmitted to the distributed storage unit; like Figure 3 As shown, the specific working method of the distributed storage unit is as follows: Receive integrated data pushed from the collection unit; Data is distributed across different nodes according to a preset sharding strategy; preset sharding strategies include hash sharding or range sharding, etc. The sharded and replicated data is stored in a distributed database. It should be noted that this allows for more efficient and reliable data processing, ensuring data quality and system stability, while also improving system scalability and flexibility.

[0019] As an optional embodiment: the data integration module further includes an expansion unit, which is used to increase the scalability of the collection unit and the distributed storage unit, specifically: Based on the current and expected load, the number of nodes required by the distributed storage unit is determined to ensure that the distributed storage unit can handle the expected requests; in this embodiment, this can specifically be achieved by: obtaining the expected load. and the storage capacity of a single node in the distributed storage unit ; According to the formula The expected number of nodes needed to be added to the distributed storage unit is calculated. ; The total resource requirement of a single node is calculated by superimposing the CPU, memory, and storage resource requirements of a single node in the distributed storage unit. The total resource requirement of a single node is then multiplied by the number of nodes required by the distributed storage unit to obtain the total resource requirement of the distributed storage unit. It should be noted that after determining the required number of nodes, the next step is to allocate CPU, memory, and storage resources based on the role and expected load of each node. This ensures that each node has sufficient resources to handle its expected workload while also avoiding resource waste. The number of nodes is automatically adjusted based on the actual load. By evenly distributing data across all nodes, single-point overload can be avoided. By designing the distributed storage unit as a node, nodes can be added as needed, satisfying scalability during use. Moreover, resource allocation, load balancing, monitoring, and expansion are all automated, achieving high scalability.

[0020] As an optional embodiment, the specific working method of automatically adjusting the number of nodes according to the actual load also includes the following: A load threshold for the distributed storage unit is set in advance. It should be noted that the load threshold It can be pre-set by staff to maintain a reasonable range; According to the formula The new number of nodes is calculated. ,in This represents the current actual load of the distributed storage unit. The number of nodes can be automatically adjusted to adapt to changing loads through monitoring.

[0021] As an optional embodiment, the semantic management module works as follows: The information gain value P of each term in the text data within the data integration module is obtained. In this embodiment, the information gain value P is calculated using term frequency (TF) and inverse document frequency (IDF), where term frequency (TF) represents the number of times a term appears in a single document. The higher the frequency of a term, the more important the term is in that document; while inverse document frequency (IDF) represents the rarity of a term across all documents. If a term appears in multiple documents, its IDF value is low because it is not rare enough; conversely, if a term appears in only a few documents, its IDF value is high because it is relatively rare. The specific calculation method is as follows: First, for each term, we calculate the product of its term frequency and inverse document frequency in each document; Then, we sum the products of the term frequency and the inverse document frequency of the term in all documents to obtain the numerator; Next, we sum the products of the term frequency and the inverse document frequency of all terms in all documents to obtain the denominator. We then divide the numerator by the denominator to obtain the information gain value P for each term. The higher the information gain value P, the more important the term is in the document collection. It helps us identify which terms are key and which terms provide more information, and can be used for various applications such as text mining, information retrieval, and document classification. Based on the obtained terminology information gain value P, the frequency of terminology use and context in different data sources are analyzed to identify semantic differences, resulting in a semantic difference value F. This score is calculated by comparing the terminology information gain values ​​P in two data sources, measuring the degree of difference in terminology use between the two data sources. The higher the score, the greater the semantic difference in terminology use between the two data sources. Based on the semantic difference value F and the information gain value P of each term, the relevance between the term and the potential replacement term is evaluated; it should be noted that this can be achieved through co-occurrence analysis, semantic similarity measures (such as cosine similarity) or machine learning methods (such as word embedding); For each term, calculate the correlation coefficient of all its possible substitute terms and multiply it by the information gain value of the term to obtain the comprehensive coefficient of each substitute term. Then select the term with the highest comprehensive coefficient as the recommended harmonized term. The generated harmonized terms will replace the terms in the text data within the data integration module, ensuring that all relevant data fields and documents use the new terms. It should be noted that this approach can handle existing semantic inconsistencies. In the human resources department of an enterprise, due to the large number of departments in the enterprise, different terms and definitions may be used to describe the same data entity, which leads to semantic inconsistency of data; For example, in actual use, the human resources department in a company may call an employee's "performance evaluation" "performance review," while the sales department may call it "performance assessment." In this case, the semantic management module can find this type of terminology and unify it into a consistent terminology that is easy to understand.

[0022] As an optional embodiment, the specific operation of the identification unit is as follows: The data in the data integration module is categorized, and data fields containing personal privacy information are marked; Use regular expressions and keyword matching to identify data in data fields that contain personal privacy information; The sensitivity level of identified data is assessed, and data with a sensitivity level exceeding a pre-set threshold is marked as privacy data. This identification unit not only improves data security and compliance but also provides enterprises with deeper data insights, helping them fully leverage the value of data while protecting user privacy. For example, in medical companies or human resources departments, the privacy data identified by the identification unit typically includes, but is not limited to, the following categories: Personal identification information, such as name, ID number, contact information, and home address, can directly identify an individual. Health-related information: In healthcare companies, patients' medical records, diagnostic results, treatment processes, etc., are all sensitive and private data; Work-related sensitive information: In the human resources department, information such as employees' salaries, performance evaluations, health status, and work experience are also considered private data; Biometric information, such as fingerprints and facial recognition information, may be used in attendance systems in modern enterprises, but it also needs to be strictly protected.

[0023] As an optional embodiment, the specific operation of the management unit is as follows: Create and maintain records of user consent data; ensure that all data collection and processing activities are based on the user's explicit consent; Monitor data access and usage; ensure that only authorized users can access private data; Process users' right to be forgotten requests and delete users' personal data from the system; Allowing users to withdraw their previously given consent to data processing, and ensuring that the data is no longer used, ensures that privacy data is properly managed and that users' privacy rights are respected and protected. These steps help improve the security and compliance of data integration and reduce the risk of privacy breaches. For example, in a medical company or the human resources department of a company, the specific working methods of a management unit can be exemplified as follows: In healthcare companies, management needs to ensure that all patient data collection and processing is based on the patient’s explicit consent. This may involve creating electronic consent forms for patients to sign before treatment begins, specifying which types of data processing activities they consent to. In the human resources department, this may involve data protection agreements that employees sign upon joining the company, specifying how they agree to how the company handles their personal and payroll information. The management unit needs to ensure that only authorized medical personnel can access patients’ health records. This may involve implementing role-based access controls to ensure that employees can only access the data required for their duties. In the human resources department, the management unit may need to monitor employees’ access to human resources information systems to ensure that they can only access information related to their work. When patients or employees request the deletion of their personal information, the management unit needs to remove this data from the system. In healthcare companies, this may involve deleting personal data from electronic health records while ensuring that healthcare continuity is not affected. In human resources departments, this may involve deleting the personal information of former employees from employee databases. If a patient or employee decides to withdraw their consent, the management unit needs to ensure that the relevant data is no longer used. In a healthcare company, this may mean ceasing further processing of the patient's data and storing it in an isolated environment until it can be safely deleted. In the human resources department, this may involve deleting the employee's data from all systems or isolating the data when the employee requests restrictions on data processing. The benefit of this identification method is that it ensures that privacy data is properly managed and that users' privacy rights are respected and protected.

[0024] As an optional embodiment, the specific operation of the bias prediction module is as follows: According to the formula ,in It is the weight of the i-th data. It is the recognition ratio of the recognition unit in the i-th data. This is the actual proportion; Obtain the deviation value B of the recognition unit; calculate it by comparing the performance of the recognition unit on different data. The degree of bias can be further quantified; the Jensen-Shannon divergence (JSD) can be used to quantify the similarity between two distributions, thereby quantifying the bias. Apply bias reduction techniques to reduce identified biases; Generate a transparency report detailing model biases and mitigation measures taken. The bias prediction module effectively identifies and reduces biases in the model, improving the system's transparency and fairness. This approach helps build user trust in the system.

[0025] As an optional embodiment: the specific process of applying the deviation reduction technology to reduce the identified deviation is as follows: Resample to obtain the adjusted sample weight values. In this embodiment, it can be achieved through: Calculated, where I is the indicator function. A subset of the i-th data point is typically partitioned based on some feature or attribute. This represents the i-th sample in the data; it is a function, when... Belongs to a subset It returns 1 if the condition is met, and 0 otherwise. It is used to determine which samples belong to the subset. ; Adjusted sample weight values Substitute the bias prediction module to obtain the new bias value B1 of the identification unit; Until the deviation value B1 of the obtained identification unit meets the preset requirements.

[0026] As an optional embodiment: the bias warning module further includes a feedback adjustment unit, which operates as follows: Obtain the information content adjustment factor for the current time period. This represents the previous performance of the identification unit. The value can be determined based on the previous performance of the identification unit. The better the previous performance, the higher this value, indicating a better degree of terminology coordination. The baseline value is 1. The obtained confidence adjustment factor is based on feedback. This reflects the adjustments made by the identification unit based on feedback; it can be determined by analyzing user or system feedback. If the degree of terminology harmonization is accurate, the value should increase; if it is inaccurate or there is a deviation, it should decrease, with a baseline value of 1. According to the formula The updated confidence adjustment factor is calculated and obtained. ,in The smoothing factor controls the balance between the old confidence adjustment factor and the feedback-based confidence adjustment; it is a hyperparameter that controls the balance between the old confidence adjustment factor and the feedback-based confidence adjustment, and its value is typically between 0 and 1. The specific value can be set through cross-validation or based on experience. Output the updated confidence adjustment factor Confidence adjustment factor The higher confidence adjustment factor directly affects the magnitude of the recognition unit update. This means the model will be updated significantly when it receives new data or feedback to quickly adapt to changes; a lower confidence adjustment factor. This means that model updates are more conservative, avoiding instability caused by over-adjustment; It should be noted that a higher confidence level means that the recognition module is very certain about its results, and therefore may require minor adjustments; a lower confidence level means that the recognition module is less certain about its results, and therefore may require major adjustments.

[0027] Smoothing factor This is used to balance the old confidence adjustment factor and the feedback-based confidence adjustment, ensuring that the update of the confidence adjustment factor takes into account both the model's previous experience and the new feedback. By continuously updating the confidence adjustment factor, the system can continuously learn new terms and semantic changes, thereby improving its effectiveness in identifying privacy data.

[0028] Working principle By applying the semantic management module, enterprises can standardize the use of key terms across different departments, thereby improving data consistency. This consistency ensures the comparability and accuracy of data throughout the organization, reducing erroneous decisions caused by terminology confusion or misunderstanding. It helps companies better identify and process personal privacy information, ensure that data collection and processing activities comply with legal and regulatory requirements, and reduce the risk of data breaches and protect users' privacy rights by identifying and classifying sensitive information and implementing appropriate protection measures.

[0029] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of this template.

Claims

1. A highly scalable database multidimensional data management system, characterized in that, include: A data integration module, comprising a collection unit and a distributed storage unit, wherein the collection unit is used to collect and integrate data, and the distributed storage unit is used to store the data collected by the collection unit; The semantic management module defines unified data terminology for the data stored in the data integration module, ensuring consistent understanding of the data across different departments and systems. The specific operation of the semantic management module is as follows: The information gain value P of each term in the text data within the data integration module is obtained; Based on the obtained term information gain value P, the frequency of use and context of terms in different data sources are analyzed to identify semantic differences and obtain semantic difference value F. Based on the semantic difference value F and the information gain value P of each term, the relevance between terms and potential substitute terms is evaluated. For each term, calculate the correlation coefficient of all its alternative terms and multiply it by the information gain value of the term to obtain the comprehensive coefficient of each alternative term. Then select the term with the highest comprehensive coefficient as the recommended harmonized term. Replace the terms in the text data within the data integration module with the generated harmonization terms, ensuring that all relevant data fields and documents use the new terms; The privacy management module includes an identification unit and a management unit. The identification unit identifies privacy data in the data integration module, and the management unit manages the data consent requests of users corresponding to the privacy data, including requests to withdraw consent and for data to be forgotten. The specific operation of the identification unit is as follows: The data in the data integration module is categorized, and data fields containing personal privacy information are marked; Use regular expressions and keyword matching to identify data in data fields that contain personal privacy information; Assess the sensitivity level of the identified data, and mark data with a sensitivity level higher than a pre-set threshold as private data; The specific working method of the management unit is as follows: Create and maintain records of user consent data; Monitor the access and usage of monitoring data; Handling users' right to be forgotten requests, Allow users to withdraw their previously given consent to data processing; A bias prediction module is provided to offer transparency to the semantic management module and the privacy management module, ensuring the fairness of the results.

2. The highly scalable database multidimensional data management system according to claim 1, characterized in that, The specific working method of the collection unit is as follows: Identify the data sources to be integrated and access data from the identified data sources; Integrate data from different sources according to specific logic; Verify that the integrated data meets the expected quality. The matching data is transmitted to the distributed storage unit; The specific working principle of the distributed storage unit is as follows: Receive integrated data pushed from the collection unit; Based on the preset sharding strategy, the data is distributed across different nodes; The sharded and replicated data is stored in a distributed database.

3. A highly scalable database multidimensional data management system according to claim 2, characterized in that, The data integration module further includes an expansion unit, which is used to increase the scalability of the collection unit and the distributed storage unit, specifically: Based on the current and expected load, determine the number of nodes required for the distributed storage unit to ensure that the distributed storage unit can handle the expected requests. The total resource requirement of a single node is calculated by superimposing the CPU, memory, and storage resource requirements of a single node in the distributed storage unit. The total resource requirement of a single node is then multiplied by the number of nodes required by the distributed storage unit to obtain the total resource requirement of the distributed storage unit. The number of nodes is automatically adjusted based on the actual load.

4. A highly scalable database multidimensional data management system according to claim 3, characterized in that, The specific working method of automatically adjusting the number of nodes according to the actual load also includes the following: A load threshold for the distributed storage unit is set in advance. ; According to the formula The new number of nodes is calculated. ,in This represents the current actual load of the distributed storage unit. This is the expected number of nodes that the distributed storage unit will need to add. It is the storage capacity of a single node in the distributed storage unit.

5. A highly scalable database multidimensional data management system according to claim 4, characterized in that, The application of bias reduction technology reduces the identified bias in the following specific process: Resample to obtain the adjusted sample weight values. ; Adjusted sample weight values Substitute the bias prediction module to obtain the new bias value B1 of the identification unit; Until the deviation value B1 of the obtained identification unit meets the preset requirements.