Big data platform construction method and device, equipment, medium and product
Patent Information
- Application Number
- CN202310031842.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-01-10
AI Technical Summary
[0003]目前,政务服务产生的数据主要存储在各部门内的单一数据库中,导致存在着大量的“信息孤岛”,绝大部分信息系统不能和外界实现数据共享,因此需要构建政务大数据平台,利用大数据技术对政务数据进行融合利用
[0022]本申请实施例中的大数据平台构建方法、装置、设备、介质及产品,通过至少一个源系统的系统信息、初始GSDM模型以及服务需求,确定目标GSDM模型,目标GSDM模型包括至少一个功能单元;基于至少一个功能单元的第一设计方案,得到目标GSDM模型的目标设计方案;根据目标设计方案,将各功能单元的第一设计方案对应的应用代码组合,得到目标大数据平台;对目标大数据平台进行测试,得到测试结果;在测试结果指示目标大数据平台测试通过的情况下,在生产环境中运行目标大数据平台,从而可以是大数据平台的构建标准化、规范化。这样,构建倒数据平台采用规范化的流程,可以提高大数据平台构建工作的质量,减少出错几率,提高效率。
Smart Images

Figure CN116049143B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of big data technology, and in particular relates to a method, apparatus, equipment, medium and product for building a big data platform. Background Technology
[0002] "Government services" refer to the administrative services provided by governments, relevant departments, and public institutions to social organizations, enterprises, and individuals, including licensing, confirmation, adjudication, rewards, and penalties, in accordance with laws and regulations. Currently, government service departments in various regions are required to input or maintain the implementation list elements of government service items in the government service item management system to sort out the government service items under their jurisdiction.
[0003] Currently, the data generated by government services is mainly stored in single databases within each department, resulting in a large number of "information silos." Most information systems cannot share data with the outside world. Therefore, it is necessary to build a government big data platform to integrate and utilize government data using big data technology.
[0004] However, existing big data platform construction methods have not formed a unified implementation process specification throughout the entire project development cycle, lack engineering implementation capabilities, and lack implementation process specifications for the entire life cycle of the big data project, which is detrimental to the efficiency of project development. Summary of the Invention
[0005] This application provides a method, apparatus, device, medium, and product for building a big data platform, which can improve the efficiency and quality of big data platform construction.
[0006] In a first aspect, embodiments of this application provide a method for constructing a big data platform, the method comprising:
[0007] Based on system information of at least one source system, an initial GSDM model, and service requirements, a target GSDM model is determined, wherein the target GSDM model includes at least one functional unit.
[0008] Based on the first design scheme of the at least one functional unit, the target design scheme of the target GSDM model is obtained;
[0009] Based on the target design scheme, the application code corresponding to the first design scheme of each functional unit is combined to obtain the target big data platform;
[0010] The target big data platform was tested, and the test results were obtained.
[0011] If the test results indicate that the target big data platform has passed the test, the target big data platform shall be run in the production environment.
[0012] Secondly, embodiments of this application provide a big data platform construction apparatus, the apparatus comprising:
[0013] The first determining module is used to determine the target GSDM model based on system information of at least one source system, an initial GSDM model, and service requirements. The target GSDM model includes at least one functional unit.
[0014] The second determining module is used to obtain the target design scheme of the target GSDM model based on the first design scheme of the at least one functional unit.
[0015] The third determining module is used to combine the application code corresponding to the first design scheme of each functional unit according to the target design scheme to obtain the target big data platform;
[0016] The testing module is used to test the target big data platform and obtain test results;
[0017] The runtime module is used to run the target big data platform in a production environment if the test results indicate that the target big data platform has passed the test.
[0018] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions;
[0019] When the processor executes the computer program instructions, it implements the steps of the big data platform construction method as described in any embodiment of the first aspect.
[0020] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the steps of the big data platform construction method as described in any embodiment of the first aspect.
[0021] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the steps of the big data platform construction method as described in any embodiment of the first aspect.
[0022] The big data platform construction method, apparatus, equipment, medium, and product in this application embodiment determine a target GSDM model based on system information of at least one source system, an initial GSDM model, and service requirements. The target GSDM model includes at least one functional unit. Based on a first design scheme of the at least one functional unit, a target design scheme for the target GSDM model is obtained. According to the target design scheme, the application code corresponding to the first design scheme of each functional unit is combined to obtain the target big data platform. The target big data platform is tested to obtain test results. If the test results indicate that the target big data platform has passed the test, the target big data platform is run in a production environment, thereby standardizing and normalizing the construction of the big data platform. In this way, the construction of the big data platform adopts a standardized process, which can improve the quality of the big data platform construction work, reduce the probability of errors, and improve efficiency. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating a method for constructing a big data platform according to an embodiment of this application;
[0025] Figure 2 This is a schematic diagram of the structure of a big data platform construction device provided in an embodiment of this application;
[0026] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0028] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0029] The acquisition, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.
[0030] Government Big Data: Data related to government affairs; the foundation and results of government departments in providing government services.
[0031] Implementation process: A set of technical specifications used to guide and standardize the construction, project implementation, and application development processes of data resource centers.
[0032] Source layer: Stores data that is consistent with the source system, including data generated by business systems and external data, and aggregates the collected data for downstream use.
[0033] Integration Layer: Data from various sources is integrated according to the theme model, making the data collected on the platform easier to manage and use.
[0034] Application layer: This is a data layer that is application-service oriented, processed according to application requirements, and directly supplied to applications.
[0035] GSDM: Government Service Data Model. The Government Service Data Model is a standardized logical data model designed during the process of participating in the construction of smart government. Based on government affairs and data requirements within the government data platform, and referencing national standards such as the "Basic Data Specifications for Government Service Platforms" and the standardization of government service interface data, it draws on data modeling achievements from the enterprise-level business model developed by China Construction Bank's new generation core system. This model covers the data content of government affairs handling matters across various government departments, establishing rules and regulations for the integration and sharing of government data.
[0036] As can be seen from the background section, existing big data platform construction methods do not have a unified implementation process specification throughout the entire project development cycle, lack engineering implementation capabilities, and lack implementation process specifications for the entire life cycle of the big data project, which is detrimental to the efficiency of project development.
[0037] To address the aforementioned issues, this application provides a method for constructing a big data platform. This method effectively aggregates and integrates multi-source information, such as business data and attribute data from various data sources, and effectively combines and applies data governance, big data mining, and application methods. This allows for the integrated application and presentation of multi-source business data. The method provides a complete plan and detailed step-by-step instructions for the entire lifecycle of the big data platform project, ensuring a well-reasoned and high-quality project delivery during implementation.
[0038] The following description, in conjunction with the accompanying drawings, details the big data platform construction method provided in this application through specific embodiments and application scenarios.
[0039] Figure 1 This is a flowchart illustrating a method for constructing a big data platform provided in an embodiment of this application.
[0040] like Figure 1 As shown, the method for building this big data platform may specifically include the following steps:
[0041] S101. Based on the system information of at least one source system, the initial GSDM model, and the service requirements, determine the target GSDM model, which includes at least one functional unit.
[0042] S102. Based on the first design scheme of at least one functional unit, the target design scheme of the target GSDM model is obtained;
[0043] S103. Based on the target design scheme, combine the application codes corresponding to the first design scheme of each functional unit to obtain the target big data platform;
[0044] S104. Test the target big data platform and obtain the test results;
[0045] S105. If the test results indicate that the target big data platform has passed the test, run the target big data platform in the production environment.
[0046] Therefore, by using system information from at least one source system, an initial GSDM model, and service requirements, a target GSDM model is determined. The target GSDM model includes at least one functional unit. Based on the first design scheme of at least one functional unit, a target design scheme for the target GSDM model is obtained. According to the target design scheme, the application code corresponding to the first design scheme of each functional unit is combined to obtain the target big data platform. The target big data platform is tested, and the test results are obtained. If the test results indicate that the target big data platform has passed the test, it is run in a production environment, thus standardizing and normalizing the construction of the big data platform. In this way, adopting a standardized process for building the big data platform can improve the quality of the big data platform construction work, reduce the probability of errors, and increase efficiency.
[0047] The specific implementation methods for each of the above steps are described below.
[0048] In some embodiments, in S101, a target GSDM model is determined based on system information of at least one source system, an initial GSDM model, and service requirements.
[0049] Specifically, building a big data platform requires first obtaining data information from the source system, determining the functional units of the initial GSDM model based on the system information of the source system, and then verifying the functional units of the initial GSDM model based on service requirements to obtain the final target GSDM model.
[0050] In some embodiments, S101 above may include the following steps:
[0051] Obtain basic system information, database information, and business information from each source system;
[0052] Based on system attribute information, database information, business information, and the initial GSDM model, the preset functional units and their attributes are mapped to the initial GSDM model to obtain the initial model mapping.
[0053] Based on service requirements, the initial model mapping is verified and adjusted to determine the mapping relationship between at least one functional unit and the initial GSDM model.
[0054] The target GSDM model is obtained based on the mapping relationship between at least one functional unit and the initial GSDM model.
[0055] This embodiment outlines the analytical steps for building a big data platform. It is demand-driven, based on upstream resources, and involves organizing the necessary elements to meet the requirements, confirming their feasibility, and guiding subsequent work. Specifically, the first step is to analyze the source system, gaining a basic understanding of its relevant aspects, collecting relevant materials, examining data, and analyzing key business processes to obtain the source system's fundamental information, database information, and business information.
[0056] In some embodiments, the above-mentioned system basic information includes at least one of the following: source system data dictionary, functional design specification, operation manual, system report list, data flow diagram and external data exchange relationship, table creation statement, ER diagram, etc.; database information includes at least one of the following: business table data information, field null value rate, code value, etc.; business information includes business information directory.
[0057] The source system basic information includes: data dictionary, functional design specification, operation manual, system report list, data flow diagram and external data exchange relationship, table creation statement, ER diagram, etc. Among them, the data dictionary is a must, as it will be needed for subsequent data collection and business analysis.
[0058] Data analysis: Analyze the database provided by the source system, including the data of all business tables, field null value rate, code value analysis, etc.
[0059] Source system business items sorting: Organize the business items of the source system, divide them into multiple levels, and compile an information catalog, which helps with business sorting and subsequent application analysis.
[0060] Functional units can include entity units in the GSDM model. As an example, based on system attribute information, database information, business information, and the initial GSDM model, the above-mentioned mapping of preset functional units and their attributes to the initial GSDM model yields the initial model mapping, which includes the following:
[0061] (1) GSDM mapping analysis:
[0062] Before data flows from the source layer to the integration layer, it needs to be mapped to the model based on the GSDM model, and stored according to themes. That is, the fields in the source system's data dictionary are mapped to the entities in the GSDM model.
[0063] Entity analysis: Analyze the standard layer physical tables and divide them according to the GSDM subject domains.
[0064] Attribute analysis: Analyze the attributes of the standard layer physical table and compare them with the attributes of GSDM.
[0065] GSDM Model Supplement: Entities or attributes that are not present in the model but are essential to the business will be supplemented.
[0066] GSDM Mapping Review: Submit the initial draft of the model mapping, confirm the accuracy of the model mapping relationship with the model designer, and finally form the final draft.
[0067] (2) Application service requirements analysis. Data is supplied to applications through the GSDM model.
[0068] Definition Analysis: For indicator-based requirements, it is necessary to analyze the meaning of the indicators in detail and confirm the measurement of the indicators.
[0069] Data service analysis: Starting from the data in the integration layer, analyze the data services that can be provided externally.
[0070] Data requirement scope analysis: Analyze the data scope of the requirement design and whether the production data can meet it.
[0071] Source system value analysis: Starting from the source system, analyze the data values and provide feedback on whether each requirement can be met. In other words, analyze whether the data provided by the source system meets the needs of the data application.
[0072] Integration layer value analysis: Analyze how to extract values from the integration layer. For parts that cannot be extracted, extract values from the source layer according to the actual situation, or analyze the integration scheme into the integration layer, and continuously improve the consistency between model and data matching.
[0073] In some embodiments, in S102 above, a target design scheme for the target GSDM model is obtained based on a first design scheme of at least one functional unit.
[0074] Specifically, by combining the design schemes of one functional unit within the target GSDM model, the target design scheme of the entire target GSDM model can be obtained.
[0075] In some embodiments, S102 may include:
[0076] Based on the data acquisition design scheme, source layer design scheme, standardization design scheme, integration layer physical table design scheme, data verification design scheme, and scheduling design scheme, the design scheme of the target GSDM model is obtained.
[0077] That is, the above-mentioned functional units can be data acquisition units, source layer units, standardization units, integration layer physical table units, data verification units, scheduling units, etc. Different units can be set according to actual needs. By first obtaining the design scheme of the functional units, the design scheme of the entire target GSDM model can be obtained based on the design scheme of each functional unit.
[0078] In one example, S102 may include the following steps:
[0079] 1. Data Acquisition Design
[0080] For different source databases, data supply methods, and data types, detailed design of the data acquisition process is required. Acquisition Scope: The scope of data acquisition is determined by analyzing the data items in the source system's business tables and the data situation in the source system's database. Acquisition Tools: Due to the wide variety of acquisition tools and the varying resources available in different production environments, this process flow only discusses the data acquisition function module of the data resource center.
[0081] 2. Source layer design
[0082] Data Integration Method: By analyzing specific tables, determine whether full integration is necessary or incremental integration can meet the requirements. Storage Strategy: Determine whether the source table is a transaction table or a status table, and adopt different storage strategies accordingly, such as slicing, snapshots, chaining, etc. Type Conversion: Design the storage field types and precision of the integrated data for corresponding data types in different databases. Data Deletion Strategy: Based on requirements, confirm the time span for retaining source layer data, and determine whether expired historical data should be physically deleted or archived.
[0083] 3. Standardized design
[0084] Code conversion: Converting fields in the source system that can be converted into standard code. Data cleaning: Removing dirty data from tables, such as null values and garbled characters.
[0085] 4. Integration Layer Physical Table Design
[0086] The design of physical entities is based on the GSDM model: Tables are designed according to granularity and information category, based on the analysis of the source layer tables and fields. Storage strategy design: Storage strategies are designed based on table characteristics, and corresponding technical fields are added according to specifications. Physical table naming: Physical tables are named according to topic division, root word translation, and storage strategy. Field naming: Attributes are named according to root words. Constraint design: Primary key, NOT NULL, unique, and foreign key constraints are designed for the tables. Integration layer mapping design: The mapping relationship from the standard layer to the integration layer is described in detail, including inter-table relationships, field processing logic, and multi-source data merging logic.
[0087] 5. Data validation design
[0088] Data verification and validation are required at every stage of the data flow. The fields and rules for data validation need to be designed according to actual needs.
[0089] 6. Scheduling Design
[0090] Based on the dependencies of the job flow, a reasonable scheduling design is carried out to ensure that each job in the platform runs normally every day.
[0091] In some embodiments, in S103 above, the application code corresponding to the first design scheme of each functional unit is combined according to the target design scheme to obtain the target big data platform.
[0092] Specifically, this step is the development step, meaning that after obtaining the design scheme, application code can be written based on the design scheme. Code can be written first for individual functional modules, and then the final big data platform code can be obtained from the application code of each functional unit, thereby improving development efficiency.
[0093] In some embodiments, the above-described S103 may include the following steps:
[0094] Based on the preset development specifications and the first design scheme corresponding to each functional unit, write the application code for each functional unit;
[0095] Based on the target design scheme, the application code corresponding to the first design scheme of each functional unit is combined to obtain the target big data platform.
[0096] This section sets up pre-defined development specifications, which can be development rules written by developers based on testing or historical development data. Following these development rules and combining them with the reviewed and verified design, standardized development is implemented to ensure accurate functional implementation and ease of version maintenance and iteration.
[0097] In one example, the source layer development specifications may include: unified storage of business data collected from various source systems, entering it into the database (STG area) and retaining recent daily data copies (creating data tables by date); merging and processing STG data to form full-scale detailed source data storage; data standardization processing, performing standardized mapping processing on necessary fields according to data standards, converting them into standardized code values; retaining historical state data (selecting important master data tables and retaining historical changes in a chained table format); registering data assets and incorporating them into the data asset system (the source layer is responsible for providing data to the application). The integration layer (model layer) development specifications may include: a model integrated based on the GSDM model. Data is stored according to subject domains to build a complete data asset system; detailed granular data is retained, while common aggregation processing and object wide table processing are performed based on OneID, and mounted under a unified model framework; data assets are registered and associated with the indicator system; data quality verification is strengthened (primary key uniqueness, value range check, correlation check, etc.). Application layer (topic library, marketplace library) development specifications may include: topic libraries designed by the application project to meet application needs, extracting and processing model layer data as needed; planning required technical components according to application scenarios, with each application topic independently applying for and allocating corresponding databases; data can be retrieved from the integration layer and the source layer (with priority given to the integration layer).
[0098] In some embodiments, in S104 above, the target big data platform is tested to obtain test results.
[0099] After the initial version of the target big data platform is developed, it needs to be tested to ensure that its functions and performance meet the expected requirements.
[0100] Optionally, the target big data platform may be tested, including the following steps:
[0101] Obtain test resources and test cases from the target big data platform;
[0102] Based on test resources and test cases, functional and non-functional tests are performed on the target big data platform, and a test report is generated, which includes the test results.
[0103] Through testing at each stage, we accurately identify various defects in the version, including both functional and non-functional ones, and generate various test reports to provide a basis for production deployment. Specifically, the testing process includes the following:
[0104] Environment preparation: Prepare the hardware and software resources required for this test to ensure that there are no differences between the subsequent testing and the production environment that could lead to production failure.
[0105] Non-functional testing: Non-functional testing aims to evaluate the readiness of an application and assess its performance under challenging conditions using various criteria such as load testing, scalability testing, stress testing, etc.
[0106] User testing: Testing is conducted based on user scenarios.
[0107] Unit testing: Checking and verifying the smallest testable unit.
[0108] Write test cases: Write test cases based on the system functions and usage scenarios.
[0109] Integration testing: Assemble all modules into a subsystem or system according to design requirements and perform integration testing.
[0110] In this embodiment, by conducting tests on the target big data platform at various stages, various defects in the version can be accurately reported, including functional and non-functional defects, and various test reports can be generated to provide a basis for production launch.
[0111] After testing the behind-the-scenes big data platform, the test report can be used to determine whether this version of the big data platform has passed the test.
[0112] In some embodiments, in S105, if the test results indicate that the target big data platform has passed the test, the target big data platform is run in the production environment.
[0113] Specifically, this may include the following: deploying the corresponding version to the production environment in accordance with standard operating procedures and detailed operating documents, including various operations before, during, and after deployment.
[0114] Resource List: List the hardware configuration and software environment required for the deployment version. Before deployment, confirm that the conditions in the list are met. After the deployment is completed, the deployment content on each machine needs to be maintained.
[0115] Deployment plan: Develop specific deployment timelines and plans to minimize the impact of new system or version launches on other systems.
[0116] Deployment and Execution: The specific deployment process, including the commands used at each step, the operations performed, and precautions.
[0117] Deployment verification and rollback: After deployment, it is necessary to verify whether all functions have been implemented. If deployment fails, it is necessary to roll back to the state before deployment.
[0118] This application provides an engineering implementation process for the development of a big data platform. This process supports all stages of data collection, integration, and application in the construction of a big data resource center. It provides specifications for the entire lifecycle of a big data platform project (analysis, design, development, testing, and deployment), supporting efficient project delivery.
[0119] The above embodiments are merely examples, and the embodiments can be combined and substituted with each other to ultimately form an embodiment of a data processing method.
[0120] It should be noted that the application scenarios described in the above embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0121] Based on the same inventive concept, this application also provides a big data platform construction apparatus. (Specifically combined with...) Figure 2 A detailed explanation will be provided.
[0122] Figure 2 This is a schematic diagram of the structure of a big data platform construction device provided in an embodiment of this application.
[0123] like Figure 2 As shown, the big data platform construction device 300 may include:
[0124] The first determining module 201 is used to determine the target GSDM model based on system information of at least one source system, an initial GSDM model and service requirements. The target GSDM model includes at least one functional unit.
[0125] The second determining module 202 is used to obtain the target design scheme of the target GSDM model based on the first design scheme of at least one functional unit.
[0126] The third determining module 203 is used to combine the application codes corresponding to the first design scheme of each functional unit according to the target design scheme to obtain the target big data platform;
[0127] Test module 204 is used to test the target big data platform and obtain test results;
[0128] The execution module 205 is used to run the target big data platform in the production environment when the test results indicate that the target big data platform has passed the test. The big data platform construction device 200 described above is described in detail below:
[0129] In some embodiments, the first determining module 201 described above is specifically used for:
[0130] Obtain basic system information, database information, and business information from each source system;
[0131] Based on system attribute information, database information, business information, and the initial GSDM model, the preset functional units and their attributes are mapped to the initial GSDM model to obtain the initial model mapping.
[0132] Based on service requirements, the initial model mapping is verified and adjusted to determine the mapping relationship between at least one functional unit and the initial GSDM model.
[0133] The target GSDM model is obtained based on the mapping relationship between at least one functional unit and the initial GSDM model.
[0134] In some embodiments, the system basic information includes at least one of the following: source system data dictionary, functional design specification, operation manual, system report list, data flow diagram and external data exchange relationship, table creation statement, ER diagram, etc.; database information includes at least one of the following: business table data information, field null value rate, code value, etc.; business information includes business information directory.
[0135] In some embodiments, the second determining module 202 is specifically used for:
[0136] Based on the data acquisition design scheme, source layer design scheme, standardization design scheme, integration layer physical table design scheme, data verification design scheme, and scheduling design scheme, the design scheme of the target GSDM model is obtained.
[0137] In some embodiments, the third determining module 203 is specifically used for:
[0138] Based on the preset development specifications and the first design scheme corresponding to each functional unit, write the application code for each functional unit;
[0139] Based on the target design scheme, the application code corresponding to the first design scheme of each functional unit is combined to obtain the target big data platform.
[0140] In some embodiments, the above-mentioned test module 204 is specifically used to: acquire test resources and test cases of the target big data platform;
[0141] Based on test resources and test cases, functional and non-functional tests are performed on the target big data platform, and a test report is generated, which includes the test results.
[0142] Therefore, by using system information from at least one source system, an initial GSDM model, and service requirements, a target GSDM model is determined. The target GSDM model includes at least one functional unit. Based on the first design scheme of at least one functional unit, a target design scheme for the target GSDM model is obtained. According to the target design scheme, the application code corresponding to the first design scheme of each functional unit is combined to obtain the target big data platform. The target big data platform is tested, and the test results are obtained. If the test results indicate that the target big data platform has passed the test, it is run in a production environment, thus standardizing and normalizing the construction of the big data platform. In this way, adopting a standardized process for building the big data platform can improve the quality of the big data platform construction work, reduce the probability of errors, and increase efficiency.
[0143] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0144] The electronic device 300 may include a processor 301 and a memory 302 storing computer program instructions.
[0145] Specifically, the processor 301 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0146] Memory 302 may include mass storage for data or instructions. For example, and not limitingly, memory 302 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 302 may include removable or non-removable (or fixed) media. Where appropriate, memory 302 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 302 is non-volatile solid-state memory.
[0147] In certain embodiments, the memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Thus, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this application.
[0148] The processor 301 reads and executes computer program instructions stored in the memory 302 to implement any of the big data platform construction methods in the above embodiments.
[0149] In some examples, the electronic device 300 may also include a communication interface 303 and a bus 310. For example, Figure 3 As shown, the processor 301, memory 302, and communication interface 303 are connected through bus 310 and complete communication with each other.
[0150] The communication interface 303 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0151] Bus 310 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not as a limitation, bus 310 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 310 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0152] For example, the electronic device 300 can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc.
[0153] The electronic device 300 can execute the big data platform construction method in the embodiments of this application, thereby achieving the combination of Figure 1 and Figure 2 The method and apparatus for building a big data platform are described.
[0154] Furthermore, in conjunction with the big data platform construction method in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the big data platform construction methods in the above embodiments. Examples of computer-readable storage media include non-transitory computer-readable storage media, such as portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, etc.
[0155] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0156] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0157] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0158] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0159] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for constructing a big data platform, characterized in that, include: Based on the system information of at least one source system, the initial GSDM model, and service requirements, a target GSDM model is determined. The target GSDM model includes at least one functional unit, which includes a data acquisition unit, a source-attaching layer unit, a standardization unit, an integration layer physical table unit, a data verification unit, and a scheduling unit. Based on the first design scheme corresponding to each of the at least one functional unit, the target design scheme of the target GSDM model is obtained. Based on the target design scheme, the application code corresponding to the first design scheme of each functional unit is combined to obtain the target big data platform; The target big data platform was tested, and the test results were obtained. If the test results indicate that the target big data platform has passed the test, the target big data platform shall be run in the production environment. The process of determining the target GSDM model based on system information from at least one source system, an initial GSDM model, and service requirements includes: Obtain the system basic information, database information, and business information of each of the source systems, wherein the system basic information includes the data dictionary of the source system; Before the data flows from the source layer to the integration layer, the fields in the data fields are mapped to the entities in the initial GSDM model; Analyze the standard layer physical table and divide it according to the subject domain of the initial GSDM model; The attributes of the standard layer physical table are analyzed and compared with the attributes of the initial GSDM model to obtain the initial model mapping draft; For entities or attributes that are not present in the initial GSDM model, if they are determined to be an essential part of the business based on the business information, they are supplemented to obtain a supplemented initial model mapping draft. The revised model mapping draft is reviewed to obtain the initial model mapping. The initial model mapping is verified and adjusted based on service requirements to determine the mapping relationship between at least one functional unit and the initial GSDM model. Based on the mapping relationship between the at least one functional unit and the initial GSDM model, the target GSDM model is obtained.
2. The method according to claim 1, characterized in that, The system basic information also includes at least one of the following: source system functional design specification, operation manual, system report list, data flow diagram and external data exchange relationship, table creation statement, ER diagram, etc.; the database information includes at least one of the following: business table data information, field null value rate, code value, etc.; the business information includes a business information directory.
3. The method according to claim 1, characterized in that, The design scheme for obtaining the target GSDM model based on the first design scheme corresponding to each of the at least one functional unit includes: Based on the data acquisition design scheme, source layer design scheme, standardization design scheme, integration layer physical table design scheme, data verification design scheme, and scheduling design scheme, the design scheme of the target GSDM model is obtained.
4. The method according to claim 1, characterized in that, The step of combining the application code corresponding to the first design scheme of each functional unit according to the target design scheme to obtain the target big data platform includes: Based on the preset development specifications and the first design scheme corresponding to each functional unit, write the application code for each functional unit; Based on the target design scheme, the application code corresponding to the first design scheme of each functional unit is combined to obtain the target big data platform.
5. The method according to claim 1, characterized in that, The test results obtained from testing the target big data platform include: Obtain the test resources and test cases of the target big data platform; Based on the test resources and test cases, functional and non-functional tests are performed on the target big data platform, and a test report is generated, which includes the test results.
6. A big data platform construction device, characterized in that, The device includes: The first determining module is used to determine the target GSDM model based on the system information of at least one source system, the initial GSDM model, and the service requirements. The target GSDM model includes at least one functional unit, which includes a data acquisition unit, a source-attached layer unit, a standardization unit, an integration layer physical table unit, a data verification unit, and a scheduling unit. The second determining module is used to obtain the target design scheme of the target GSDM model based on the first design scheme corresponding to the at least one functional unit; The third determining module is used to combine the application code corresponding to the first design scheme of each functional unit according to the target design scheme to obtain the target big data platform; The testing module is used to test the target big data platform and obtain test results; The runtime module is used to run the target big data platform in a production environment if the test results indicate that the target big data platform has passed the test. The first determining module is specifically used for: Obtain the system basic information, database information, and business information of each of the source systems, wherein the system basic information includes the data dictionary of the source system; Before the data flows from the source layer to the integration layer, the fields in the data fields are mapped to the entities in the initial GSDM model; Analyze the standard layer physical table and divide it according to the subject domain of the initial GSDM model; The attributes of the standard layer physical table are analyzed and compared with the attributes of the initial GSDM model to obtain the initial model mapping draft; For entities or attributes that are not present in the initial GSDM model, if they are determined to be an essential part of the business based on the business information, they are supplemented to obtain a supplemented initial model mapping draft. The revised model mapping draft is reviewed to obtain the initial model mapping. The initial model mapping is verified and adjusted based on service requirements to determine the mapping relationship between at least one functional unit and the initial GSDM model. Based on the mapping relationship between the at least one functional unit and the initial GSDM model, the target GSDM model is obtained.
7. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the steps of the big data platform construction method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the steps of the big data platform construction method as described in any one of claims 1-5.
9. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device causes the electronic device to perform the steps of the big data platform construction method as described in any one of claims 1-5.
Citation Information
Patent Citations
Data model construction method and device, electronic equipment and storage medium
CN114253939A