Data management method, system and device based on DorisDB and storage medium

By adopting DorisDB-based data governance methods in small and medium-sized housing and construction institutions, the problem of high resource demand in the data governance process is solved, and the optimization of resource utilization and improvement of data governance efficiency is achieved.

CN120067086APending Publication Date: 2025-05-30安徽省住房和城乡建设信息中心 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145107.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Small and medium-sized housing and construction institutions face the problem of high resource demand in the process of data governance. The existing Hadoop-based solutions consume high resources and are difficult to reasonably match actual needs, resulting in waste of resources and difficulty in implementing them.

Method used

Using DorisDB-based data governance methods, we collect and build housing and construction databases, store business data in partitions, and connect them to DorisDB for data cleaning and fusion, and use DorisDB's distributed computing and storage capabilities to optimize resource utilization.

Benefits of technology

It effectively reduces resource consumption in the data governance process, improves data quality and governance efficiency, adapts to the business needs of small and medium-sized housing and construction institutions, reduces costs and improves system performance and availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067086A_ABST
    Figure CN120067086A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data governance, in particular to a DorisDB-based data governance method, system and device and a storage medium, and the method comprises the steps: collecting business data of a residence construction institution, and constructing a residence construction database based on the collected business data; performing partition storage on the business data based on the type of the business data in the residence database; accessing the residence and building database after the partition storage to a DorisDB (Data Business); performing data cleaning on the accessed service data according to a preset cleaning rule in each partition; and performing data fusion operation based on the DorisDB in response to a data fusion demand of the residence establishment mechanism. According to the invention, the problem of high resource demand of small and medium-sized residence and construction institutions in the data management process can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data governance technology, and in particular to a data governance method, system, device and storage medium based on DorisDB. Background Art

[0002] In recent years, with the rapid development of big data technology and its widespread application in the field of housing and construction, the importance of data resources has become increasingly prominent. Relevant guidelines and technical specifications clearly state that it is necessary to strengthen the planning and management of data resources, promote the deep integration and efficient use of data, and support the modernization and upgrading of governance capabilities. With technological progress, data-driven modern governance has gradually become the core direction of digital transformation in the field of housing and construction, and the transformation from traditional "document processing" to "data governance" has become an industry consensus.

[0003] At present, most of the data construction of housing and construction institutions in the field generally adopt solutions based on Hadoop architecture. This architecture can support the storage, processing and query of large-scale data through distributed storage and large-scale parallel computing technology, and has good scalability and stability. This technical architecture is widely used in large and medium-sized data governance projects. Its functions cover metadata management, data cleaning, data fusion and other links, and can meet various needs in complex scenarios.

[0004] However, for small and medium-sized housing and construction agencies with relatively small data volume requirements, the existing Hadoop-based solutions have high resource requirements and are difficult to reasonably match actual needs. The high cost of computing and storage resources required for architecture deployment has caused resource waste and difficulty in implementation. This high resource demand has limited the promotion and application of data governance solutions in small and medium-sized housing and construction agencies, and has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] This application provides a data governance method, system, device and storage medium based on DorisDB, which can effectively solve the high resource demand problem faced by small and medium-sized housing and construction institutions in the data governance process. This application provides the following technical solutions:

[0006] In a first aspect, the present application provides a data governance method based on DorisDB, the method comprising:

[0007] Collecting business data of housing and construction institutions, and building a housing and construction database based on the collected business data;

[0008] In the housing and construction database, the business data is partitioned and stored based on the type of the business data;

[0009] Connect the partitioned housing and construction database to DorisDB;

[0010] Clean the accessed service data according to the preset cleaning rules in each partition;

[0011] In response to the need for data fusion of the housing construction agency, perform data fusion operations based on the DorisDB.

[0012] In a specific feasible implementation, the collection of the service data of the housing construction agency and the construction of the housing construction database based on the collected service data include:

[0013] Configure the connection information of each data source. After the configuration is completed, access each business system in the housing construction agency and collect data;

[0014] During the data collection process, adopt the full-volume collection method to obtain data and format the collected data;

[0015] The collected data will be stored in a new database to complete the construction of the housing construction database, and the housing construction database will be updated regularly.

[0016] In a specific feasible implementation, the partition storage of the service data based on the type of the service data includes:

[0017] For each type of service data, a set of related words and a description text are correspondingly set;

[0018] Calculate the correlation degree between each service data and the description text corresponding to each type as the first correlation degree;

[0019] Calculate the correlation degree between each service data and the set of related words corresponding to each type as the second correlation degree;

[0020] Add the first correlation degree and the second correlation degree to obtain the final correlation degree, and store the service data in the partition where the data type with the largest final correlation degree is located.

[0021] In a specific feasible implementation, the calculation of the correlation degree between each service data and the description text corresponding to each type as the first correlation degree includes:

[0022] The i-th service data D i and the description text T of the j-th type j of the first correlation degree S 1 (D i ,T j ) is calculated as follows:

[0023]

[0024] Among them, Vec(D i ) and Vec(T jThey are business data D i and description text T j 's vector representations. CosSim(Vec(D i ), Vec(T j )) represents the cosine similarity between business data D i and description text T j . α is an adjustment coefficient. Var(D i , T j ) is a term measuring text difference. The calculation method of Var(D i , T j ) is as follows:

[0025]

[0026] Among them, w k in the first relevance calculation process represents a single word or term in the business data and the description text. m represents the total number of words in the business data and the description text. TF(w k , D i ) represents the frequency of the word w k appearing in the business data D i . TF(w k , T j ) represents the frequency of the word w k appearing in the description text T j .

[0027] In a specific feasible implementation, calculating the relevance between each business data and the set of associated words corresponding to each type as the second relevance includes:

[0028] The second relevance S i between the i-th business data D j and the set of associated words W 2 (D i , W j ) has the following calculation formula:

[0029]

[0030] Among them, w k in the second relevance calculation process represents a single word or term in the business data and the set of associated words. β is an adjustment factor. J(D i , W j ) represents the Jaccard similarity between the business data D i and the set of associated words W j . The calculation formula is as follows:

[0031]

[0032] Among them, f(w k , D i ) and f(w k , W j ) are the occurrence frequencies of the word w k in the business data D i and the associated word set W j respectively; n represents the number of words in the associated word set W j ; C(w k , D i , w j ) represents the co-occurrence frequency of the word w i in the business data D j and the associated word set W k , and the calculation formula is as follows:

[0033] C(w k , D i , W j ) = min(f(w k , D i ), f(w k , W j ))。

[0034] In a specific feasible implementation, the data cleaning of the accessed business data according to the preset cleaning rules in each partition includes:

[0035] During the cleaning process, the data that cannot be repaired or does not meet the predetermined standards and rules is regarded as abnormal data;

[0036] Move the abnormal data to the abnormal data table and indicate the reason for the abnormality in the table.

[0037] In a specific feasible implementation, the data fusion operation based on the DorisDB in response to the needs of the data fusion of the housing construction agency includes:

[0038] Clarify the business logic and specific fusion rules of the housing construction agency;

[0039] Perform data fusion through DorisDB, use the SQL query ability of DorisDB to associate data from different sources, and match the fields set based on the business logic for the tables of different data sources to construct a comprehensive data view;

[0040] After completing the data fusion, store the associated data as a data model through DorisDB, and generate regular reports and visualization charts based on the fused data model.

[0041] In the second aspect, the present application provides a data governance system based on DorisDB, adopting the following technical solutions:

[0042] A data governance system based on DorisDB, comprising:

[0043] A data collection module, configured to collect business data of a housing construction agency and build a housing construction database based on the collected business data;

[0044] A data classification module, configured to store the business data in partitions in the housing construction database based on the type of the business data;

[0045] A data access module, configured to access the partitioned housing construction database to DorisDB;

[0046] A data cleaning module, configured to clean the accessed business data according to preset cleaning rules in each partition;

[0047] A data fusion module, configured to perform a data fusion operation based on the DorisDB in response to the need for data fusion of the housing construction agency.

[0048] In a third aspect, the present application provides an electronic device, which includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a data governance method based on DorisDB as described in the first aspect.

[0049] In a fourth aspect, the present application provides a computer-readable storage medium, in which a program is stored, and the program is used to implement a data governance method based on DorisDB as described in the first aspect when executed by a processor.

[0050] In summary, the beneficial effects of the present application at least include:

[0051] 1) During the data cleaning process, dedicated cleaning rules are set for each data partition, enabling customized processing according to the business requirements of different data partitions. Through this partition cleaning mechanism, it is possible to flexibly optimize specific types of data without interfering with the data in other partitions, thereby improving the quality of the data. At the same time, for data that cannot be cleaned, the system will mark it as abnormal data and record relevant information for subsequent optimization of the cleaning rules to enhance the data quality control ability.

[0052] 2) By performing data fusion in DorisDB, for business data from different systems and departments, summary, aggregation, and association processing are carried out according to preset business logics and data association rules to generate a unified data view, supporting cross-departmental collaboration and efficient decision-making analysis. Through the distributed computing ability and efficient data storage mechanism of DorisDB, the processing speed and consistency of data can be guaranteed, and multi-dimensional analysis can be performed on the generated unified data model, improving the quality and efficiency of decision support.

[0053] 3) Through the data governance method based on DorisDB, the problem of high resource requirements faced by small and medium-sized housing construction institutions in the data governance process can be effectively solved. Specifically, in scenarios with a small amount of data, through efficient data partition storage, precise business data classification, and flexible data fusion, the optimized utilization of resources is achieved. At the same time, the excessive complexity and resource waste of the traditional Hadoop architecture are avoided, improving the data governance ability of small and medium-sized housing construction institutions.

[0054] Through partition storage and a fine-grained classification method based on similarity, the need for a large amount of resources is avoided, and the data classification accuracy in scenarios with a small amount of data is guaranteed. By adopting personalized cleaning rules for different data partitions, the data quality is further improved. At the same time, using the powerful computing and storage capabilities of DorisDB, data fusion and decision support are effectively carried out, meeting the business needs of small and medium-sized institutions, reducing costs, and improving system performance and availability.

[0055] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly and implement it according to the content of the specification, the following takes the preferred embodiments of this application and combines with the attached drawings to elaborate in detail as follows. Brief Description of the Drawings

[0056] Figure 1 is a schematic flowchart of the data governance method based on DorisDB in an embodiment of this application.

[0057] Figure 2 is a schematic overall flowchart of the data governance method based on DorisDB in an embodiment of this application.

[0058] Figure 3 is a block diagram of the structure of the data governance system based on DorisDB in an embodiment of this application.

[0059] Figure 4 is a block diagram of the electronic device for data governance based on DorisDB in an embodiment of this application. Detailed Description of the Embodiment

[0060] The following will further describe in detail the specific implementation manners of the present application in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but not to limit the scope of the present application.

[0061] Optionally, the data governance method based on DorisDB provided in each embodiment of the present application is taken as an example for illustration in an electronic device. The electronic device is a terminal or a server. The terminal can be a mobile phone, a computer, a tablet computer, etc. The type of the electronic device is not limited in this embodiment.

[0062] Referring to Figure 1 , which is a schematic flowchart of the data governance method based on DorisDB provided in an embodiment of the present application. The method at least includes the following steps:

[0063] Step S101: Collect the business data of the housing construction agency, and construct a housing construction database based on the collected business data.

[0064] In step S101, the purpose is to collect relevant business data from various business systems of the housing construction agency and construct a housing construction database based on these data. First, it is necessary to identify and confirm various business systems involved within the housing construction agency, including but not limited to project management systems, financial management systems, personnel management systems, approval systems, etc. Each business system contains a large amount of data, such as project basic information, budget expenditures, financial statements, employee information, approval processes, etc. These data are all important bases for daily management and decision-making.

[0065] In implementation, to achieve data collection, first, the connection information of each data source needs to be configured, which includes necessary information such as the type of the database, IP address, port number, username, password, etc. The data source can be a relational database or a non-relational database, or an external file system or API interface. After configuration, the data collection system can access each business system and extract the required data. During the data collection process, a full-volume collection method is adopted to ensure that complete data is obtained at one time. To ensure that the data can be successfully integrated into the housing construction database, the collected data needs to be formatted. Specifically, for structured data, such as database tables, CSV files, etc., it is necessary to ensure the standardization of field types, lengths, units, etc. For unstructured data, such as log files, PDF files, pictures, etc., they need to be processed first during the data collection process, such as OCR recognition, data parsing, etc. When collecting data, the field naming rules, data types, and units, etc. should be unified to ensure that business data from different sources can be integrated into a unified format. The collected data will be stored in a new database to complete the construction of the housing construction database. After data collection, the housing construction database needs to be updated regularly to ensure the timeliness and accuracy of the data.

[0066] In addition, to ensure data persistence and reliability, the database also needs to be backed up regularly to prevent data loss. The full backup method can be adopted to ensure data recovery in any emergency and maintain data integrity and availability.

[0067] Step S102: In the housing construction database, partition and store business data based on the type of business data.

[0068] In step S102, the business data in the housing construction database is partitioned and stored according to the data type. For each type of business data, a set of related words and a description text are correspondingly set. The set of related words contains keywords or phrases related to this data type. These keywords are extracted from actual operations and can help identify data categories. The vocabulary in the set of related words can include specific business terms, domain-specific vocabulary, and common operation or status descriptions. The description text is usually a short summary or definition of this type of data, which can accurately describe the content and characteristics of this type of data. First, calculate the correlation degree between each business data and the description text corresponding to each type as the first correlation degree, and then calculate the correlation degree between each business data and the set of related words corresponding to each type as the second correlation degree. Add the first correlation degree and the second correlation degree to obtain the final correlation degree. Store the business data in the partition where the data type with the largest final correlation degree is located.

[0069] First, the i-th business data D i and the description text T of the j-th type j The first correlation degree S 1 (D i , T j ) is calculated as follows:

[0070]

[0071] Among them, Vec(D i ) and Vec(T j ) are the vector representations of the business data D i and the description text T j , respectively, calculated through the pre-trained model Bert. CosSim(Vec(D i ), Vec(T j )) represents the cosine similarity between the business data D i and the description text T j . α is an adjustment coefficient used to balance the influence of the semantic similarity part and the text difference part. This coefficient can be adjusted according to the actual application scenario. Var(D i , T j) is a term for measuring text difference, which is achieved by calculating the variance of semantic distribution in two texts. By analyzing the differences between texts, a moderate penalty can be imposed on overly similar texts to avoid the interference of high-frequency words. Var(D i ,T j ) is calculated as follows:

[0072]

[0073] Among them, w in the calculation process of the first correlation degree k represents a single word or phrase in business data and descriptive text. During the process of calculating the first correlation degree, it is necessary to segment or split the business data and descriptive text. w k is the word or phrase after segmentation. m represents the total number of words in business data and descriptive text. TF(w k ,D i ) represents the frequency of the word w k appearing in the business data D i . TF(w k ,T j ) represents the frequency of the word w k appearing in the descriptive text T j .

[0074] In the above process, when calculating the first correlation degree, a semantic similarity calculation formula based on the Bert model is adopted, and by introducing the text difference term, the interference of high-frequency words is reduced, and the matching accuracy between texts is improved. This can avoid the influence of excessive semantic similarity of texts (for example, in the same business field, the descriptive contents of multiple business data may be similar but actually have different semantics). Combining cosine similarity with semantic variance takes into account the context information of each data and descriptive text as well as the text differences between different data, and can more comprehensively measure the similarity between data. Combining cosine similarity with semantic variance takes into account the context information of each data and descriptive text as well as the text differences between different data, and can more comprehensively measure the similarity between data. By introducing text differences to adjust the calculation of similarity, the interference of high-frequency words caused by simple word frequency statistics is avoided. Especially in text similarity calculation, the repeated occurrence of some high-frequency words may affect the accuracy of business data classification.

[0075] Secondly, the second correlation degree S i of the i-th business data D j and the set of related words W 2 (D i ,W j ) is calculated as follows:

[0076]

[0077] Among them, w in the second correlation degree calculation process k represents a single word or term in the business data and the set of related words. In the process of calculating the second correlation degree, it is necessary to segment or split the business data and the set of related words. w k is the word or phrase after segmentation. J(D i ,W j ) represents the Jaccard similarity between the business data D i and the set of related words W j , which is used to calculate the similarity between the business data and the set of related words, and measures the proportion of common words in both to all words. The calculation formula is as follows:

[0078]

[0079] Among them, f(w k ,D i ) and f(w k ,W j ) are the occurrence times of the word w k in the business data D i and the set of related words W j respectively. n represents the number of words in the set of related words W j . C(w k ,D i ,W j ) represents the co-occurrence frequency of the word w i in the business data D j and the set of related words W k , which is used to represent the co-occurrence frequency of the same word in the business data and the set of related words, and can highlight the influence of those frequently co-occurring words on the final result. The calculation formula is as follows:

[0080] C(w k ,D i ,W j ) = min(f(w k ,D i ), f(w k ,W j ))

[0081] Among them, β is a regulation factor, which is used to adjust the weights of the Jaccard similarity and the co-occurrence frequency.

[0082] In the above process, when calculating the second correlation degree, through the combination of Jaccard similarity and co-occurrence word frequency, it is possible to deeply analyze the matching situation between each piece of business data and the set of associated words. The Jaccard similarity calculation considers the intersection and union of the data and the set of associated words, which can intuitively reflect the overlapping degree of the two; while the co-occurrence word frequency highlights the impact of commonly occurring words on classification by considering the frequency of the same word appearing in the business data and the set of associated words. The introduction of the adjustment factor enables the weights of the two (Jaccard similarity and co-occurrence word frequency) to be flexibly adjusted according to actual needs, ensuring that the algorithm effect can be adaptively adjusted in different business scenarios. Jaccard similarity and co-occurrence word frequency can effectively capture the matching degree of key words in business data and the set of associated words, and are suitable for efficiently finding the most relevant partitions in scenarios with a small amount of data.

[0083] It should be noted that in the case of a small amount of data, calculating the similarity between each piece of data and the type description text, as well as the matching degree with the set of associated words, can be quickly completed through efficient vector calculation and simple co-occurrence frequency statistics, without generating excessive computational burden. The vectorized representation after the pre-training of the Bert model can efficiently capture semantic information and is suitable for applications with a small and uncomplicated amount of data. At the same time, by storing and calculating similarities in partitions, the business data can be carefully divided and accurately classified. Especially when the amount of data is small, through the comprehensive consideration of the description text and the set of associated words, it is possible to effectively avoid misclassifying data into irrelevant partitions. However, when the amount of data is large, calculating the similarity between each piece of data and all description texts and the set of associated words will generate extremely high computational pressure. Especially the calculation of Jaccard similarity and co-occurrence word frequency, as the amount of data increases, the computational complexity rises sharply. In addition, although the Bert model is very effective for small-scale data, in the case of a very large amount of data, the vectorization process will also lead to performance bottlenecks. Moreover, as the amount of data increases, overly fine-grained partition calculations may lead to a decrease in the system's response speed and performance, making it unable to process efficiently.

[0084] To sum up, by combining the first correlation degree (semantic similarity of the description text) and the second correlation degree (matching degree of the set of associated words), it is possible to effectively store the business data in the housing construction database in partitions, enabling each piece of business data to be accurately classified into the corresponding data type partition. This method can accurately classify and integrate data in scenarios with a small amount of data, thus supporting decision-making by providing high-quality data views. For scenarios with a large amount of data, such fine-grained calculations may lead to performance problems, so it is more suitable for application scenarios with a small amount of data.

[0085] Step S103: Connect the housing construction database after partition storage to DorisDB.

[0086] In step S103, the data of each partition in the housing construction database (such as project data, financial data, approval data, etc.) has been stored according to the classification rules. On this basis, the data of each partition is imported into DorisDB using an ETL tool or a custom script. During the migration process, appropriate format conversion needs to be performed on the data to ensure that the fields and data formats of different data tables are compatible with the table structure in DorisDB. At the same time, the connection between the housing construction database and DorisDB is configured to ensure that the data can be transmitted efficiently and stably. For each data partition (such as project data, financial data, etc.), corresponding tables or data partitions are created in DorisDB. The table structure of each data partition will be consistent with the classification table in the housing construction database to ensure that no information loss or format inconsistency occurs when the data is accessed in DorisDB. After the data import is completed, data consistency and integrity verification are performed. By checking whether the data of each partition is complete and whether the fields match, it is ensured that no errors or data loss occur during the import process.

[0087] In implementation, after the data is accessed in DorisDB, necessary optimization and index construction are performed on each data partition table to ensure that the data can be accessed efficiently during subsequent data query, analysis, and report generation processes. According to the data query requirements, appropriate index strategies are selected, such as primary key index, range query index, composite index, etc., to improve the query efficiency. Partitioned storage can be accurately divided according to the type of business data, improving the query efficiency. For the case of a small amount of data, through refined partitioned storage and optimized query mechanisms, higher system response speed and user experience can be obtained.

[0088] Step S104: Clean the accessed business data according to the preset cleaning rules within each partition.

[0089] In step S104, the data cleaning process is carried out according to the specific cleaning rules of each data partition. These rules are tailored for different data partitions to ensure the quality and consistency of the data.

[0090] Specifically, each data partition (such as project data, financial data, approval data, etc.) has its own specific cleaning rules, which are usually designed by data administrators or business personnel according to actual business needs to ensure that the rules can meet the business logic and quality requirements of the data. During the cleaning process, if some data cannot be repaired or does not meet the predefined standards and rules, they will be regarded as abnormal data. The abnormal data will be moved to a dedicated abnormal data table, and the reasons for the anomalies will be indicated in this table. The abnormal data records can not only help optimize the cleaning rules subsequently but also serve as the basis for problem tracking and data quality control. Each data partition can set unique cleaning rules according to its own business characteristics and data quality requirements, flexibly meeting the needs of different data. Through partition cleaning, specific types of data can be optimized without affecting other data, improving data quality. The handling method of abnormal data ensures that data problems can be detected and recorded in a timely manner, thus supporting subsequent data analysis and improvement.

[0091] Step S105: In response to the need for data integration of the housing and construction agency, perform data integration operations based on DorisDB.

[0092] In step S105, during the operation of the housing and construction agency, since the data comes from different departments, systems, and sources, such as project data, financial data, approval data, etc. To provide a unified view and support efficient decision-making and business analysis, data integration operations are performed based on DorisDB, summarizing, aggregating, and correlating various types of data according to the business logic to ensure that the final data can comprehensively reflect the operation status of the housing and construction agency. The goal of data integration is to integrate data from different systems through predefined business logic to form a unified data view. This view supports subsequent data analysis, decision-making support, and cross-departmental collaboration.

[0093] Specifically, before data integration, it is first necessary to clarify the business logic of the housing and construction agency and specific integration rules. The business logic determines how different categories of data are combined, and clear field mapping and data merging rules need to be set according to business requirements. For example, keyword fields such as project ID and approval ID can be used as the basis for data association, and fields such as timestamps and status flags are used for data synchronization.

[0094] Secondly, data fusion is performed through DorisDB. Using the SQL query capabilities of DorisDB, data from different sources is correlated through methods such as JOIN statements and WHERE conditions. For tables from different data sources, matching is performed based on fields set according to business logic to construct a comprehensive data view. During the fusion process, data is summarized through aggregation functions in DorisDB (such as SUM(), AVG(), COUNT(), etc.) to calculate various key metrics. For example, multiple financial records are aggregated into the total expenditure for each project, or multiple steps in the project approval process are summarized.

[0095] Finally, after data fusion is completed, the summarized and correlated data is stored as a unified data model through DorisDB. This model can reflect data in multiple dimensions, including project progress, budget usage, approval status, etc. This data model can be used in subsequent data analysis, report generation, decision support, and other scenarios. By designing appropriate data dimensions and hierarchical structures, multi-angle analysis is supported. For example, the financial progress of projects can be presented according to different time dimensions (quarterly, monthly, etc.), or data slicing and aggregation can be performed through different dimensions such as project category and region. After the data fusion operation is completed, regular reports and visual charts can be generated based on the unified data model after fusion. DorisDB supports integration with external visualization tools and can display query results in the form of charts, dashboards, etc. to assist the housing construction agency in making decisions. By performing data fusion operations based on DorisDB, the housing construction agency can integrate data from different systems, form a unified data view, and support more accurate business analysis and decision-making.

[0096] In summary, combined with Figure 2 , this application proposes a data governance method based on DorisDB, aiming to solve the problems of high resource consumption and inefficient data processing faced by small and medium-sized housing construction agencies during the data governance process. This method optimizes the data governance process through the following steps: First, collect relevant data from each business system of the housing construction agency and construct a housing construction database based on the collected data; then, achieve precise partition storage by calculating the correlation degree between business data and a set of preset description texts and associated words, and subdivide the data according to business types; next, import the data after partition storage into DorisDB and perform optimization processing and index construction on DorisDB; subsequently, clean the incoming business data based on the specific rules of the business data to ensure data quality; finally, according to the data fusion requirements of the housing construction agency, perform summarization, aggregation, and correlation of multi-source data through DorisDB to provide a unified data view to support efficient decision-making and business analysis.

[0097] By combining semantic similarity and associated word set matching, precise partitioning and classification of small-scale data are achieved, thus avoiding resource waste in medium and small-sized housing construction institutions in solutions based on the Hadoop architecture. In addition, the efficient storage and computing capabilities of DorisDB ensure fast response and consistency in the data processing process, perform excellently in scenarios with small data volumes, effectively improve the flexibility and cost-effectiveness of data governance, and meet the actual needs of housing construction institutions.

[0098] Figure 3 FIG. is a structural block diagram of a data governance system based on DorisDB provided by an embodiment of the present application. The system at least includes the following modules:

[0099] A data collection module, configured to collect business data of a housing construction institution and construct a housing construction database based on the collected business data;

[0100] A data classification module, configured to perform partitioned storage on business data in the housing construction database based on the type of the business data;

[0101] A data access module, configured to access the partitioned and stored housing construction database to DorisDB;

[0102] A data cleaning module, configured to perform data cleaning on the accessed business data according to a preset cleaning rule in each partition;

[0103] A data fusion module, configured to perform a data fusion operation based on DorisDB in response to the need for data fusion of a housing construction institution.

[0104] For relevant details, refer to the above method embodiment.

[0105] Figure 4 FIG. is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.

[0106] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0107] The memory 402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 401 to implement the DorisDB-based data governance method provided in the method embodiments of the present application.

[0108] In some embodiments, the electronic device may optionally further include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.

[0109] Of course, the electronic device may also include fewer or more components, and this embodiment does not limit this.

[0110] Optionally, the present application further provides a computer-readable storage medium, and a program is stored in the computer-readable storage medium, and the program is loaded and executed by the processor to implement the DorisDB-based data governance method of the above method embodiments.

[0111] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium. A program is stored in the computer-readable storage medium and is loaded and executed by a processor to implement the data governance method based on DorisDB in the above method embodiment.

[0112] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0113] The above embodiments only express several implementation manners of the present application, and the description thereof is relatively specific and detailed. However, it should not be understood as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A data governance method based on DorisDB, characterized in that: The method comprises: Collecting business data of housing and construction institutions, and building a housing and construction database based on the collected business data; In the housing and construction database, the business data is partitioned and stored based on the type of the business data; Connecting the partitioned housing and construction database to DorisDB; Clean the incoming business data according to the preset cleaning rules in each partition; In response to the demand for data fusion of housing and construction agencies, data fusion operations are performed based on the DorisDB.

2. The data management method based on DorisDB according to claim 1 is characterized in that: The collecting of business data of the housing and construction agency and building a housing and construction database based on the collected business data include: Configure the connection information of each data source. After the configuration is completed, access various business systems in the housing and construction agency and collect data; In the data collection process, the full amount of data is collected and the collected data is formatted; The collected data will be stored in a new database to complete the construction of the housing and construction database, and the housing and construction database will be updated regularly.

3. The data management method based on DorisDB according to claim 1 is characterized in that: The partitioning and storing the business data based on the type of the business data includes: For each type of business data, a corresponding set of associated words and description text are set; Calculate the relevance between each business data and the description text corresponding to each type as a first relevance; Calculating the degree of association between each business data and the associated word set corresponding to each type as a second degree of association; The first relevance and the second relevance are added together to obtain a final relevance, and the business data is stored in the partition where the data type with the greatest final relevance is located.

4. The data management method based on DorisDB according to claim 3 is characterized in that: The calculating the relevance between each business data and the description text corresponding to each type as the first relevance comprises: The i-th business data D i and the description text T of the jth type j The first degree of association S1(D i ,T j ) is calculated as follows: Among them, Vec(D i ) and Vec(T j ) are business data D i and description text T j The vector representation of CosSim(Vec(D i ),Vec(T j )) indicates business data D i and description text T j The cosine similarity between them, α is the adjustment coefficient, Var(D i ,T j ) is a measure of text difference, Var(D i ,T j ) is calculated as follows: Among them, w in the first correlation calculation process k represents a single word or phrase in the business data and description text, m represents the total number of words in the business data and description text, TF(w k ,D i ) represents the word w k In business data D i The frequency of occurrence, TF(w k ,T j ) represents the word w k In the description text T j The frequency of occurrence in .

5. The data management method based on DorisDB according to claim 4 is characterized in that: The calculating the degree of association between each business data and the associated word set corresponding to each type as the second degree of association includes: The i-th business data D i and the j-th type of associated word set W j The second correlation degree S2(D i ,W j ) is calculated as follows: Among them, w in the second correlation calculation process k represents a single word or phrase in the business data and associated word set, β is the adjustment factor, J(D i ,W j ) represents business data D i and the associated word set W j The Jaccard similarity is calculated as follows: Among them, f(w k ,D i ) and f(w k ,W j ) are word w k In business data D i and the associated word set W j The number of occurrences in; n represents the associated word set W j The number of words in C(w k ,D i ,W j ) represents business data D i and the associated word set W j Chinese word w k The co-occurrence frequency of is calculated as follows: C(w k ,D i ,W j )=min(f(w k ,D i ),f(w k ,W j ))。 6. The data management method based on DorisDB according to claim 1 is characterized in that: The data cleaning of the accessed business data according to the preset cleaning rules in each partition includes: During the cleaning process, data that cannot be repaired or does not conform to the predetermined standards and rules are considered abnormal data; Move the abnormal data to the abnormal data table and indicate the abnormal reasons in the table.

7. The data management method based on DorisDB according to claim 1 is characterized in that: In response to the demand for data fusion of the housing and construction agency, the data fusion operation based on the DorisDB includes: Clarify the business logic and specific integration rules of housing and construction agencies; Through DorisDB, data is integrated and DorisDB's SQL query capability is used to associate data from different sources. For tables from different data sources, fields set based on business logic are matched to build a comprehensive data view. After completing the data fusion, the associated data is stored as a data model through DorisDB, and regular reports and visualization charts are generated based on the fused data model.

8. A data management system based on DorisDB, characterized in that: include: A data collection module, used to collect business data of housing and construction institutions, and build a housing and construction database based on the collected business data; A data classification module, used for partitioning and storing the business data in the housing and construction database based on the type of the business data; A data access module, used to access the housing and construction database after the partition storage into DorisDB; The data cleaning module is used to clean the accessed business data according to the preset cleaning rules in each partition; The data fusion module is used to respond to the data fusion needs of the housing and construction agency and perform data fusion operations based on the DorisDB.

9. An electronic device, characterized in that: The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a data governance method based on DorisDB as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores a program, which, when executed by a processor, is used to implement a data governance method based on DorisDB as described in any one of claims 1 to 7.