Statistical analysis system based on data elements
The data-driven statistical analysis system addresses the inefficiencies of manual data collection and analysis by providing a web-based, automated solution for data entry, cleaning, and visualization, enhancing data utilization and decision-making efficiency.
Patent Information
- Application Number
- CN202311586084.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
Existing data collection and analysis processes are cumbersome, time-consuming, and resource-intensive, involving extensive manual operations and paper-based documentation, which hinders efficient data utilization and analysis.
A data-driven statistical analysis system deployed on Windows servers, utilizing a web-based interface for data entry, incorporating data management, ETL tools, and data resource graph units for data collection, cleaning, and visualization, with modules for real-time optimization and data governance.
The system streamlines data collection and analysis, reducing manual effort and paper usage while enabling efficient, automated data processing and visualization, supporting data-driven decision-making for industrial development.
Smart Images

Figure CN120047086A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of statistical analysis systems and relates to a statistical analysis system based on data elements. Background Art
[0002] The project of the data element statistical analysis system is positioned as a think tank platform to assist the construction of the data element market in Shanghai. By building a digital business ecological resource library and establishing a digital communication bridge between the main bodies of various data element market activities, it comprehensively assists the construction of the data element markets in Shanghai and the whole country. The construction goal is to build market development monitoring and management services such as data element statistical analysis and data empowerment index for the government; provide services such as quasi-digital business recommendation, data product creation, and promotion of transaction implementation for the data exchange; and provide public services such as evaluation consultation and policy publicity for enterprises.
[0003] The existing data collection scheme needs to be carried out offline, with cumbersome steps, long time, and will generate a lot of paper documents, resulting in waste of resources, and a large amount of manual operations are required for data analysis, which is time-consuming and laborious.
[0004] Therefore, it is of great significance to study a statistical analysis system with simple steps based on data analysis rather than manual operations. Summary of the Invention
[0005] The purpose of the present invention is to solve the above problems existing in the prior art and provide a statistical analysis system based on data elements.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] A statistical analysis system based on data elements is deployed on a windows server and provided to enterprises for filling operations through web links, including a data management unit, an ETL tool, an accounting system, and a data element resource map unit;
[0008] The data management unit is used to detect the use of data resources and perform scheduling optimization in real time;
[0009] The ETL tool is used to support the collection and cleaning of various semi-structured and unstructured data in the statistical analysis system of data elements;
[0010] The accounting system is used to collect data element data that various enterprises need to fill through collection tasks for calculation;
[0011] The data element resource map unit is used to display the calculation results of the accounting system in the form of charts.
[0012] As a preferred technical solution:
[0013] A statistical analysis system based on data elements as described above, where the data management unit includes a data resource usage management module, a real-time scheduling optimization module, a data governance module, a data analysis module, and a platform management module.
[0014] Data resource usage management: Provide the ability to manage the entire life cycle of data resources from a multi-dimensional perspective, and provide services for the commercialization and valuation of data resources;
[0015] Real-time scheduling optimization module: Unified scheduling provides flexible, reliable, and multi-type integrated task scheduling capabilities, supporting a wide range of business scenarios;
[0016] Data governance module: Provide data governance capabilities driven by standards to ensure data quality;
[0017] Platform management module: Provide the capabilities of platform organization, user management, and operating environment management.
[0018] A statistical analysis system based on data elements as described above, where the data sources supported by the ETL tool include databases, file systems, Excel, and Xml.
[0019] A statistical analysis system based on data elements as described above, where the database is DB2, Oracle, Mysql, or SQLServer.
[0020] A statistical analysis system based on data elements of the present invention uses an ETL tool to perform various semi-structured and unstructured data collection and cleaning tasks; the accounting system unit is used to carry out data element statistical accounting work to promote the normalization of data element statistical accounting. The statistical analysis system based on data elements will comprehensively and multi-dimensionally collect data elements from key enterprises that master data element resources, exploring new paths for incorporating data production factors into the national economic statistical accounting system in the future. The data element resource map unit is used to construct various models such as a multi-dimensional data quotient portrait model, an intelligent recommendation model, a data element resource map model, and an intelligent prediction model, providing technical support for the intelligent service capabilities of the present invention, analyzing data element resource enterprises, and forming a visual data element resource distribution map and a data value distribution map, panoramically presenting the data element resources and value distribution in Shanghai.
[0021] In the design of the statistical analysis system architecture for data elements of the present invention, it covers aspects such as the user access layer, business application layer, data resource layer, platform support layer, infrastructure layer, standard specification system, security guarantee system, operation and maintenance system, etc. The system adopts a service-oriented architecture SOA, which connects different functional units (called services) of the statistical analysis system for data elements. The functional units include a data management unit, an ETL tool, an accounting system, and a data element resource map unit, through well-defined interfaces and contracts between these services; the interfaces are defined in a neutral manner, independent of the hardware platform, operating system, and programming language for implementing the services; enabling the constructed services to interact in a unified and general way.
[0022] In the technical system of the statistical analysis system based on data elements of the present invention, the Java architecture is selected, and the B / S / D three-layer model is adopted for the development of the statistical analysis system for data elements; Browser / WebServer / Database (i.e., the B / S / D three-layer model) is an application model most suitable for solving information services and interactive corresponding dynamic services, realizing a truly thin client, and greatly simplifying the distribution, configuration management, and version management work of the application system.
[0023] The development of this system supports the J2EE technical specification and can meet the requirements of rapid development and construction period.
[0024] The system architecture of the present invention is transformed based on the cloud to ensure the compatibility of the open-source interfaces of the system, and to achieve rapid iteration, flexible expansion, and efficient operation and maintenance.
[0025] The present invention adopts technologies such as parallel processing, distributed caching, cluster deployment, and load balancing to ensure the response speed of each service and the processing of massive data.
[0026] Beneficial effects:
[0027] (1) The present invention provides a statistical analysis system based on data elements, which can analyze the relevant factors affecting the development of the industry, summarize experience, discover laws, predict trends, and assist in decision-making; provide data support for the industry development, and make decisions based on data analysis rather than just expert experience.
[0028] (2) A statistical analysis system based on data elements of the present invention converts the collection method to online, reduces the number of steps and time, while reducing the generation of paper documents, and cleans and analyzes the data through built-in algorithm modules (such as ETL tools), reducing the labor input cost. Brief description of the drawings
[0029] Figure 1 It is the architecture diagram of the statistical analysis system based on data elements of the present invention. Detailed implementation manners
[0030] The present invention will be further described below in conjunction with specific implementation manners. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0031] A statistical analysis system based on data elements is deployed on a Windows server and includes a data management unit, an ETL tool, an accounting system, and a data element resource map unit;
[0032] The data management unit includes a data resource usage management module, a real-time scheduling optimization module, a data governance module, a data analysis module, and a platform management module; the data management unit is used to detect the usage of data resources and perform real-time scheduling optimization;
[0033] The ETL tool is used to support the data collection and cleaning work of various semi-structured and unstructured data in the statistical analysis system of data elements; the data sources supported by the ETL tool include databases, file systems, Excel, and Xml; the databases are DB2, Oracle, Mysql, or SQLServer; it supports data collection in the HADOOP big data environment;
[0034] The accounting system unit is used to perform operations on the data element data that various enterprises need to fill in through the collection tasks;
[0035] The data element resource map unit is used to display the operation results of the accounting system in the form of charts.
[0036] The present invention will be described below using specific embodiments, as follows:
[0037] A statistical analysis system based on data elements, the overall architecture design and overall technical route are as follows:
[0038] (1) Overall architecture design;
[0039] The architecture design of the statistical analysis system of data elements covers the design contents of the user access layer, business application layer, data resource layer, platform support layer, infrastructure layer, standard specification system, security guarantee system, operation and maintenance system, etc. The system architecture is as Figure 1 shown.
[0040] (1.1) User access layer;
[0041] The user access layer consists of multiple levels of users such as the government, data exchanges, and enterprises, and provides corresponding access content for users through system roles, user permissions, process permissions, menu permissions, and module permissions;
[0042] Users can view information such as to-do tasks and messages through the unified portal access entrance, and support the configuration of independent access addresses for access; it meets the business scenarios where users use the system through terminal browsers.
[0043] (1.2) Business application layer;
[0044] The business application layer includes an accounting system unit and a data element resource map unit, which are used to carry out statistical accounting work on data elements, analyze data element resource enterprises, and panoramically display the data element resources and value distribution in Shanghai.
[0045] (1.3) Data resource layer;
[0046] The data resource layer includes a data management unit, an ETL tool, and a database; the data resource layer provides data support services for the business application layer.
[0047] (1.4) System support layer;
[0048] Build multiple business application scenarios through microservice components such as unified process management, form management, report management, and log auditing, and realize low-code development and zero-code configuration of business applications.
[0049] (1.5) Infrastructure layer;
[0050] The infrastructure layer provides a basic operating environment for the system to run, and is built through a data management unit, an ETL tool, an accounting system, and a data element resource map unit, providing a centralized and unified basic operating environment and infrastructure services.
[0051] (1.6) Standard specification system;
[0052] The standard specification system provides a basis for the construction of the statistical analysis system project based on data elements of the present invention. At the same time, the standard specification system can be further improved during the construction of this project; the standard specification system of this project is formulated on the basis of following national e-government, combined with relevant policy guidelines on security and reliability and the characteristics of supervision services.
[0053] (1.7) Security guarantee system;
[0054] During the construction of this project, the content such as the project security level protection requirements for the scheme design, technical requirements, management specifications, and evaluation guidelines is strictly followed; after the project construction is completed, assist users to complete the corresponding security evaluation according to relevant requirements.
[0055] (1.8) Operation and maintenance system;
[0056] By establishing a standardized, normalized, and institutionalized operation and maintenance system, comprehensively monitor the operation status of this project, promptly handle problems that occur during the operation process, and ensure the safe, stable, efficient, and continuous operation of this project; at the same time, through the construction of an operation and maintenance platform and an operation and maintenance technical support team (including on-site operation and maintenance teams), realize the intelligentization and high efficiency of operation and maintenance work, and improve the operation and maintenance level of this project.
[0057] (2) Overall technical route;
[0058] (2.1) SOA architecture;
[0059] The system should adopt a service-oriented architecture SOA; it connects different functional units (called services) of the statistical analysis system of data elements, and the functional units include a data management unit, an ETL tool, an accounting system, and a data element resource map unit, through well-defined interfaces and contracts between these services; the interfaces are defined in a neutral manner, independent of the hardware platform, operating system, and programming language for implementing the services; enabling the constructed services to interact in a unified and general way; this feature with a neutral interface definition (not forcibly bound to a specific implementation) is called loose coupling between services.
[0060] SOA requires applications to be designed as a collection of services, requiring developers to think beyond the business application itself, consider the reuse of existing services, or think about how their services can be reused by other projects; a key benefit of "separate", "independent", and "well-encapsulated" services is that they can be combined into larger services in various different ways to achieve reuse.
[0061] (2.2) Domesticated platform adapted to the whole system;
[0062] Under the background of the full replacement of domestic products in the e-government information system, the construction of the informatization of this invention is urgently needed to implement national strategies and safeguard national information security. It is necessary to follow relevant security systems, standards, and specifications at the national and local levels, make full use of existing various safe and reliable facilities, comprehensively realize the selection and adaptation of domestic software and hardware, and achieve the purpose of technological independence, control, and effective security guarantee.
[0063] The construction of this invention supports domestic mainstream operating systems, middleware, databases, terminals, and virtual machines.
[0064] Operating system: Support domestic mainstream operating systems;
[0065] Database: Support domestic mainstream databases such as DM and Kingbase;
[0066] Middleware: Support domestic middleware such as TongTech, Kingdee, and Powerleader;
[0067] Terminal: Support domestic terminals such as Loongson and Phytium CPUs, desktop operating systems such as ZKFD, NeoKylin, and Galaxy Kylin, Firefox, 360 Browser, and WPS Office software;
[0068] Virtual Machine: Support CPUs such as Loongson, Zhaoxin, Phytium, Sunway, and Haiguang, virtual machine operating systems such as NeoKylin, ZKFD, and Galaxy Kylin, and domestic mainstream database and middleware products.
[0069] (2.3) Java technology system and B / S / D three-tier model;
[0070] The statistical analysis system based on data elements of the present invention selects the Java architecture in the technical system and adopts the B / S / D three-tier model to develop the statistical analysis system of data elements; Browser / WebServer / Database (i.e., the B / S / D three-tier model) is the most suitable application model for solving information services and interactive corresponding dynamic services, realizing a truly thin client, and greatly simplifying the distribution, configuration management, and version management of application systems.
[0071] (2.4) Microservices architecture (i.e., system support layer);
[0072] As an architecture composed of multiple microservice components, each microservice in the system can be independently deployed, and there is loose coupling between microservices. Each microservice only focuses on completing one task and doing it well; in all cases, each task represents a small business capability. Each service has its own process and uses lightweight mechanisms (usually HTTP source APIs) to achieve communication. These services are built around business functions and are independently deployed by virtue of an automated deployment mechanism. These services match a minimal set of centralized management mechanisms, and each service can be written in different programming languages and can use different data storage technologies.
[0073] The microservices architecture has the following characteristics:
[0074] (ⅰ) Componentization of applications through services;
[0075] In the microservices architecture, components are defined as software units that can be independently replaced and upgraded. In the application architecture design, componentization design is carried out by splitting the overall application into microservices that can be independently deployed and upgraded.
[0076] Organize services around business capabilities: The microservices architecture adopts a strategy of organizing services starting from business capabilities;
[0077] (ii) Intelligent endpoints and pipeline flattening;
[0078] In a microservices architecture, the relevant business logic / intelligence for communication between components can be placed on the component endpoint side rather than in the communication components. The communication mechanism or components should be as simple and loosely coupled as possible; the RESTful HTTP protocol and lightweight asynchronous mechanisms that only provide message routing functions are the most commonly used communication mechanisms in a microservices architecture;
[0079] (iii) "Decentralized" governance;
[0080] A microservices architecture can use appropriate tools to complete their respective tasks. Each microservice can consider choosing the best tools to complete (such as different programming languages). The technical standards of microservices tend to seek technologies that other developers have successfully verified to solve similar problems;
[0081] (iv) "Decentralized" data management;
[0082] A microservices architecture advocates the adoption of diverse persistence methods, allowing each microservice to manage its own database and permitting different microservices to adopt different data persistence technologies;
[0083] (v) Infrastructure automation;
[0084] Technologies such as cloudification and automated deployment have greatly reduced the difficulty of building, deploying, and operating microservices. By applying methods such as continuous integration and continuous delivery, it helps to achieve the goal of accelerating time to market;
[0085] (vi) Fault handling design;
[0086] A microservices architecture must consider the failure tolerance mechanism for each service. Therefore, microservices attach great importance to establishing real-time monitoring and logging mechanisms for architecture and business-related metrics;
[0087] (vii) Evolutionary design;
[0088] Microservices applications pay more attention to rapid updates, so the design of the system will change and evolve over time; the design of microservices is affected by factors such as the life cycle of business functions; for example, if an application is a monolithic application but gradually evolves towards a microservices application architecture, the monolithic application is still the core, but new functions will be built using the APIs provided by the application; another example is that in a microservices application, following the basic principle of replaceable modular design, if it is found that two microservices often need to be updated simultaneously after implementation, this may very likely mean that they should be merged into one microservice.
[0089] (2.5) SpringMVC framework;
[0090] Currently, the mainstream microservice frameworks include SpringCloud, Dubbo, DubboX, Hydra, etc. The application system uses the SpringMVC framework and builds a distributed microservice system based on SpringBoot. The front end of the application system supports frameworks such as VUE, VUEX, VUE-Router, iVew, Weex, etc., and uses front-end development technologies such as html, html5, CSS, less, Jquery, etc. It meets the requirements of using Tomcat and domestic middleware as WEB containers and accessing application services through http and https protocols. It supports load balancing technology, and performs traffic control, service circuit breaker and degradation, interface monitoring, and routing proxy on access at the security gateway layer. At the support service layer, the component management, three-person management, and log management of each business component are carried out through the management center. Each microservice component regularly calls the heartbeat detection interface to send heartbeat messages to verify whether each microservice component is alive.
[0091] SpringCloud is an ordered collection of a series of frameworks and a combination of a relatively complete set of microservice framework technical solutions currently. It provides each module required for building a distributed system, and the application of SpringCloud greatly simplifies the amount of code for developers. Using the SpringCloud framework, mainly its core components are used to complete various tasks. The core components include the Eureka component, Ribbon component, Fegin component, Hystrix component, Zuul component, etc.
[0092] (1) The Eureka component is equivalent to a service center, and its functions are service governance, registration, and discovery, responsible for the unified management of all services. It includes a server, a client, and a Server. Its main functions are to provide registration, offline, and renewal services for service providers; to provide acquisition, invocation, and offline services for service consumers. The existence of the service registration center enables the automatic registration and discovery of each microservice instance.
[0093] (2) The Ribbon component is a load balancer responsible for load balancing of the client. Based on the http and TCP protocols, it polls and accesses through the ribbonServerList server list configured in the client to achieve the effect of service balance. This component can be used in combination with the Eureka component.
[0094] (3) The key mechanism of the Fegin component is to use the dynamic proxy mechanism to complete remote service calls. Define the methods of other services to be called as abstract methods and encapsulate a set of simple interfaces, without the need for the component to build an HTTP request.
[0095] (4) The Hystrix component is a circuit breaker. It has powerful functions such as service degradation, request service fusing, thread and signal isolation (dependency isolation), request caching, request merging, and service monitoring.
[0096] (5) The Zuul component is a service gateway that mainly provides functions such as dynamic routing, monitoring, elasticity, response, and security control. Its core is a series of filters (pre - filters, post - filters, routing filters, and error filters), and it can cooperate with the above - mentioned components to complete specified functions.
[0097] (2.6) J2EE technical specifications;
[0098] The development of this system supports J2EE technical specifications and can meet the requirements of rapid development and project duration.
[0099] J2EE is a technical architecture completely different from traditional application development. It contains many components, mainly simplifies and standardizes the development and deployment of application systems, and then improves portability, security, and reuse value. J2EE simplifies the development of application programs, reduces the requirements for programming and trained programmers, and improves portability, security, and reuse value.
[0100] The core of J2EE is a set of technical specifications and guidelines. The various components, service architectures, and technical levels it contains have common standards and specifications, enabling good compatibility between different platforms that follow the J2EE architecture, and solving the dilemma that information products used in the enterprise backend in the past were incompatible with each other, making it difficult for internal or external communication within the enterprise.
[0101] Under the J2EE architecture, developers can develop enterprise - level applications based on the specification foundation. Different J2EE suppliers will support the standards formulated within different J2EE versions to ensure the compatibility between different J2EE platforms and products. In other words, application systems based on the J2EE architecture can basically be deployed on different Windows servers, and with little or only a small amount of code modification, the portability (Portability) of the application system can be greatly improved.
[0102] (2.7) Workflow technology;
[0103] The overall design idea of workflow technology is to schedule human - participated manual activities through participants obtained from the organizational model; to implement automatic activities that operate by calling external applications; to perform necessary routing judgments by accessing process - related data; and to record the running track of the process through process control data.
[0104] Workflow technology is responsible for the management of the entire life cycle of business processes. It defines processes for business application systems by providing services, realizes process flow through a process engine, and completes process adjustment, monitoring, and auditing using a browser-based management and monitoring interface. It can quickly customize business applications using a rich component library to achieve on-demand response.
[0105] Business application systems define processes through the modeling service in workflow technology. Workflow technology assigns and manages process tasks through a process scheduling engine and task management.
[0106] During operation, the process engine schedules human participation activities by obtaining appropriate participants from the organizational model; realizes automatic activities by calling business application systems; makes necessary routing judgments by accessing workflow-related data; records the process operation trajectory through workflow control data; the process scheduling engine and the task table manager are linked through the task table, and drive each other through the state transition of the task table.
[0107] (2.8) Distributed database technology;
[0108] Distributed database technology is a combination of database technology and distributed technology. Specifically, it refers to a database technology that combines databases that are geographically dispersed but logically belong to the same system in a computer system. It has database coordination and data distribution. The distributed database management system does not focus on centralized control of the system, but on the autonomy of each database node. In addition, to reduce the workload of programmers when writing programs and the possibility of system errors, generally, the data distribution situation is not considered at all, so that the data distribution situation of the system always remains transparent.
[0109] The concept of data independence is also a very important part in the distributed database management system. However, not only that, the distributed data management system also adds a new concept called distributed transparency. The role of this new concept is to ensure that the program correctness is not affected when data is transferred, just as if the data was not distributed when the program was written.
[0110] In a distributed database, data redundancy is a required feature, which is different from general centralized database systems. First, data is replicated at database nodes where it is needed to improve local application performance. Second, if a database node encounters a system error, before it is repaired, the system can continue to be used by operating on the replicated data in other database nodes, improving the system's availability.
[0111] (2.9) Distributed messaging technology;
[0112] Distributed Message Service (DMS) is a message middleware service based on highly available distributed cluster technology, providing a reliable and scalable managed message queue for sending, receiving, and storing messages.
[0113] Using DMS, a message queue can be created, which serves as a transfer station for messages, storing the messages passed between different components of the application. This enables message transmission between different components of the application without requiring all components to be available simultaneously.
[0114] In a distributed system, messages, as a means of communication between applications, are widely used. Messages can be stored in a queue until retrieved by the receiver. Since the message sender does not need to synchronously wait for the response from the message receiver, the asynchronous reception of messages reduces the coupling degree of system integration, improves the efficiency of distributed system collaboration, enables the system to respond to users faster, and provides higher throughput. When the system is under peak pressure, the distributed message queue can also act as a buffer to smooth out the peaks and valleys, relieve the pressure on the cluster, and prevent the entire system from collapsing.
[0115] (2.10) Distributed caching technology;
[0116] Caching technology generally refers to storing some frequently used data in a faster storage device for users to access quickly. Distributed caching means storing some popular data in a location close to the user and the application and preferably in a faster device in a distributed environment or system to reduce the latency of remote data transmission.
[0117] In a Redis distributed cache, each node is responsible for storing a part of the data. At the same time, each node also has a master-slave design to improve the reliability of Redis. The data structures supported by the Redis distributed cache include not only simple key / value types but also complex types such as List, Set, and Hash. Data is written from the volatile memory storage device to the disk to permanently save the data. Redis supports master-slave synchronization, and Redis can effectively ensure data consistency by configuring the two parameters min-replicas-to-write and min-replicas-max-lag.
[0118] (2.11) Distributed database query optimization technology;
[0119] In the research of database systems, it is very important to encapsulate database operations for users as much as possible, so as to enhance the generality of the database system. Another point is that the distributed database system also needs to hide the relevant details in the system from users, so as to enhance the security and convenience of the system in actual use. For relational databases, they can provide data interfaces to users, and all data can be transmitted through these interfaces. When querying data using SQL statements, only a simple description of the data to be queried is required, and there is no need to know how the data is obtained inside the system. When a user sends a request, the distributed database system will first check whether the accessed database exists locally. If it exists, the command needs to be executed. If it does not exist, the request needs to be broadcast to other databases, and the optimal query node is selected based on the query information. On this basis, the query can be executed, that is, query the relevant information of the database with the relevant query information, and it is necessary to ensure that the required query resources are minimized. After that, the query command can be sent to the database according to the address and returned to the database IP address. After the returned information is received by the client, a connection can be immediately established with the database. After the internal query of the database is completed, the query information can be sent back to the client.
[0120] First, the query speed is improved by reasonably setting indexes. In the current distributed database systems, data indexes are a very important data structure. To effectively improve the query speed, relevant application principles should be followed. Specifically, the principles to be followed in actual application mainly include the following points: Set indexes at positions that are not specified as foreign keys but need to be frequently joined. For some fields that are less frequently used for joining, the index can be automatically generated by the DBMS; Set indexes for columns related to frequent sorting and grouping operations; When there are many columns to be sorted, a composite index can be set. Generally speaking, in the default state, non-clustered indexes can be selected for index setting, but this type is not necessarily the most reasonable one. The selection of the appropriate index type should be based on the analysis of the query type. For example, when there are a large number of duplicate values, the construction of a clustered index should be considered. When multiple columns need to be frequently accessed and there are many duplicate values, a composite index can be constructed. When constructing a composite index, efforts should be made to form an index coverage in the key query.
[0121] Secondly, attention should be paid to avoiding sorting or streamlining sorting as much as possible. In the query optimization process of a distributed database system, for some large data tables, sorting operations should be avoided as much as possible. When the index can output according to a certain number of times, sorting operations can be effectively avoided, and the query speed can be improved in data query. At the same time, the effective reduction of sorting operations can also be achieved by appropriately increasing the index, and the appropriate merging of data tables can be realized. In addition, for some inevitable sorting situations, attempts should be made to simplify them. On the basis of continuously narrowing the sorting range, the sorting operation should be simplified as much as possible.
[0122] Thirdly, sequential access operations should be avoided as much as possible for large data tables. For distributed database queries, among the various factors affecting data query efficiency, the sequential access of nested queries has an important impact, which will significantly reduce the speed in the data query process. Therefore, for related columns with joins, the construction of an index method should be implemented to avoid sequential access operations on large data tables. At the same time, the query can also be processed by using the index path, and the union method can be selected to effectively avoid sequential access.
[0123] Fourthly, temporary tables can be used to accelerate data query speed. In the actual query process, for subsets of data tables, by sorting them and constructing relevant temporary data tables, the data query efficiency can be effectively improved. The number of rows in the temporary table should be relatively small compared to the main table, so the I / O cost can be reduced and the workload of query operations can be effectively reduced. In addition, on the basis of constructing the temporary table, repeated sorting operations can be effectively avoided, and the optimizer operation can be effectively reduced.
[0124] Fifthly, relatively difficult regular expressions and related subqueries should be avoided as much as possible, and the reduction of nested-level queries should be actively achieved to effectively improve query efficiency. To avoid re-querying the subquery when changing the values of the main query columns, the query time can be effectively saved. At the same time, non-starting substrings can also be effectively avoided to achieve satisfactory query results.
[0125] (2.12) High-concurrency processing technology;
[0126] Adopt technologies such as parallel processing, distributed caching, cluster deployment, and load balancing to ensure the response speed of each service and the processing of massive data.
[0127] Establishing a resource allocation system can fully utilize the idle resources of other virtual machines and client resources on the network to respond to service requests and undertake part of the services, which can greatly improve the system performance and save hardware investment at the same time. The system model provides the allocator with the resource information visible on all nodes in a timely manner. After obtaining the information, the allocator reasonably allocates the resources to tasks, thereby optimizing the system performance. By establishing a resource allocation system, load balancing and grid computing of services are achieved, so that when a large number of users operate concurrently on the network, the system efficiency will not be significantly reduced.
[0128] (2.13) Global view technology based on metadata;
[0129] By establishing the metadata management function and mechanism throughout the entire process of e-government data processing, not only the unified management of various types of metadata such as business metadata, technical metadata, and management metadata is realized, but also it runs through the entire life cycle of e-government data processing, including various links such as data source, ETL, storage, processing, analysis, presentation, use, and archiving.
[0130] Through a unified and standardized e-government metadata standard system, the self-description and self-interpretation of e-government data are realized, ensuring the standardization, consistency, and accuracy of e-government data query, analysis, and display.
[0131] (2.14) Visual development technology;
[0132] The management of the entire life cycle of the business process, customizing the process of the business application system through system services, relying on the process engine to achieve the automated transfer of the approval process, and based on the browser-based management and monitoring interface, completing the adjustment, monitoring, and auditing of the process, and using a rich component library to quickly customize business applications to achieve on-demand response.
[0133] The visual development technology has the following characteristics:
[0134] Graphical workflow definition tool;
[0135] Role-based task management and permission control;
[0136] Tracking and monitoring of workflow status;
[0137] It can judge the workflow path according to conditions;
[0138] Task reminder and exception handling;
[0139] It can statistically analyze all aspects of the workflow, such as the execution time of each task, etc., for the improvement and optimization of the process;
[0140] User logs;
[0141] It is easy to integrate with relational databases and other systems.
[0142] (2.15) Dynamic interactive visualization display technology;
[0143] The interactive chart analysis function provides users with a powerful and easy-to-use tool for data visualization analysis and presentation.
[0144] In terms of report production: Users can freely combine analysis indicators and analysis dimensions to dynamically form various query and summary reports, and support the production of complex reports such as multi-level nesting and cross-grouping.
[0145] In terms of statistical chart production: On the one hand, it provides users with a rich and beautiful graphics library. In addition to conventional graphics such as line charts, pie charts, histograms, bar charts, area charts, umbrella point charts, radar charts, and bubble charts, it also provides special graphics such as maps, heatmaps, and mosaic charts. On the other hand, it supports dynamic linked analysis of graphics. Users only need to perform simple associated attribute settings to obtain the effect of graphics linked analysis.
[0146] (2.16) Data unified storage model;
[0147] Data resource types mainly include two categories, namely structured data and unstructured data; semi-structured data is obtained as structured data after being loaded and cleaned by ETL tools;
[0148] (2.16.1) Structured data;
[0149] Data that can be displayed in one-dimensional or two-dimensional report forms mainly includes: basic data, summary data, comprehensive data, and macro data, etc.
[0150] (2.16.2) Unstructured data
[0151] Unstructured data mainly includes two major categories. One category is the finished products produced by text-based business processes (in formats such as PDF, WORD, XLS, XML, etc.), mainly including various text materials: various working documents and reports formed by government agencies at all levels. The other category is multimedia materials, including various pictures, videos and other materials.
[0152] (2.16.3) Advantages of data unified storage:
[0153] (i) Simple model;
[0154] Adopt an object-based integrated storage model to replace the original mode of one report corresponding to one data table, making the storage information structure simple and the number of data tables concise and controllable.
[0155] (ii) Facilitate business expansion;
[0156] Object-based integrated storage can ignore the data presentation form and the changes and adjustments of the report system, and support the arbitrary increase or decrease of indicators.
[0157] (ⅲ) Facilitate storage expansion;
[0158] The data unified storage model is an extended model of key-value pair storage based on objects, which conforms to the model definition concept of the current mainstream distributed storage architecture and can be well compatible with the mainstream distributed storage and high-performance in-memory databases.
[0159] (ⅳ) Strong consistency;
[0160] For any data indicator, there is and only one unique record corresponding to it in the data center, effectively avoiding the version chaos and information conflicts caused by storing data in multiple places.
[0161] (ⅴ) Simplify the processing logic and improve the processing performance;
[0162] The complexity of the traditional storage model leads to excessive association queries during the process of data processing and statistical analysis (such as cross-table queries, processing, summarization, and calculation of different report data), resulting in a geometric multiple increase in the complexity of the processing logic, directly affecting the ease and performance of processing. The data unified storage model avoids multi-table operations, simplifies the processing logic such as cross-table queries, processing, summarization, and calculation of different report data, and at the same time, relying on the distributed processing structure, can greatly improve the processing performance.
[0163] (2.17) Online analysis and mining technology;
[0164] The online analysis and mining technology integrates OLAP multi-dimensional analysis and data mining (DM) functions, which can not only meet the general multi-dimensional analysis application requirements of users, but also meet the advanced data processing and analysis scenarios based on the e-government analysis model.
[0165] (2.17.1) In addition to the drill-down function of mainstream BI products, the multi-dimensional analysis function can also allow users to directly drill down to the basic detailed data, facilitating users to conduct various exploratory data analyses.
[0166] (2.17.2) In addition to providing conventional visual big data analysis applications, the data mining function also integrates common programming languages, puts the custom tools for big data analysis in the hands of data analysts, supports users with data analysis skills to conduct custom and interactive data exploration, so as to obtain data insights and discover patterns and trends, and generate more value for customers.
[0167] The analysis model integration interface technically supports computer languages such as R, Python, Java, C / C++, etc. In particular, by introducing a large number of data mining algorithm libraries in R and Python, a large number of common machine learning algorithms can be implemented. For example, when used in combination with the algorithms in the R language, it can analyze the massive data in the existing platform at high speed.
[0168] (2.18) Knowledge graph technology;
[0169] Knowledge graph technology is an integral part of artificial intelligence technology. Its powerful semantic processing and interconnection organization capabilities provide a foundation for intelligent information applications. With the development of the Internet, the content of network data has shown an explosive growth trend. Due to the characteristics of large scale, heterogeneous diversity, and loose organizational structure of Internet content, it poses challenges to people's effective access to information and knowledge. Knowledge Graph, with its powerful semantic processing ability and open organizational ability, has laid a foundation for the knowledge-based organization and intelligent application in the Internet era. A knowledge graph aims to describe the entities existing in the real world and the relationships between entities. The original intention of knowledge graph technology is to improve the capabilities of search engines, enhance the search quality and experience of users. With the development and application of artificial intelligence technology, knowledge graph technology, as one of the key technologies, has been widely used in fields such as intelligent search, intelligent question answering, personalized recommendation, and content distribution.
[0170] The construction and application of large-scale knowledge bases require the support of multiple technologies. Through knowledge extraction technology, knowledge elements such as entities, relationships, and attributes can be extracted from the data in some publicly available semi-structured, unstructured, and third-party structured databases. Knowledge representation represents the knowledge elements by certain effective means for further processing and use. Then, through knowledge fusion, the ambiguity between the referential terms such as entities, relationships, and attributes and the factual objects can be eliminated to form a high-quality knowledge base. Knowledge reasoning further mines the implicit knowledge based on the existing knowledge base, thereby enriching and expanding the knowledge base. The comprehensive vectors formed by distributed knowledge representation are of great significance for the construction, reasoning, fusion, and application of knowledge bases. Next, this article will focus on knowledge extraction, knowledge representation, knowledge fusion, and knowledge reasoning technologies, select representative methods, and illustrate the relevant research progress and practical technical means.
[0171] (2.19) Cloud computing technology;
[0172] Cloud computing technology is a distributed storage and computing method. Through this method, shared software and hardware resources and information can be provided to computers and other devices on demand. The cloud platform consists of a configurable shared resource pool that provides various hardware and software resources such as networks, servers, storage, applications, and services. The resource pool has self-management capabilities, and users can conveniently and quickly obtain resources on demand with only a small amount of participation. Through the network "cloud", the huge data computing and processing capabilities are decomposed into countless resource units, and distributed processing and analysis are carried out based on the server cluster system to solve task distribution and achieve the merging and feedback of calculation results. Nowadays, cloud computing is the result of the hybrid evolution and leap of computer technologies such as distributed computing, utility computing, load balancing, parallel computing, network storage, hot backup redundancy, and virtualization. It has technical characteristics such as ultra-large scale, virtualization, high reliability, generality, high scalability, and on-demand service. Through this technology, the processing of tens of thousands of data can be completed in a very short time (a few seconds), thus achieving powerful data service benefits.
[0173] (2.20) Cloud-native containerized management technology;
[0174] After years of development, cloud computing has evolved from stages such as "application migration to the cloud" to the stage of "application construction for the cloud", that is, the cloud-native infrastructure stage that evolves from "resource-centered" to "application-centered". The application system with a cloud-native architecture is built using components, adapts to the Internet architecture and operation and maintenance capabilities, the system architecture is transformed based on the cloud to ensure the compatibility of the system's open-source interfaces, and realizes rapid iteration, flexible expansion, and efficient operation and maintenance.
[0175] The cloud-native architecture is born and grows in the cloud, maximizing the use of cloud capabilities. The IT architecture built relying on cloud products and cloud-native containerized management technology allows developers to focus on business rather than underlying technologies. From a technical perspective, the cloud-native architecture is a set of architectural principles and design patterns based on cloud-native containerized management technology, aiming to maximize the separation of non-business code parts in cloud applications, so that cloud facilities can take over a large number of original non-functional characteristics in the application (such as elasticity, resilience, security, observability, gray scale, etc.). While the business is no longer troubled by non-functional business interruptions, it has the characteristics of being lightweight, agile, and highly automated.
[0176] The original IaaS layer is upgraded to an agile infrastructure, while the PaaS layer and the SaaS layer are merged into a microservices architecture. Both the agile infrastructure and microservices belong to the architecture in the technical category. In the entire cloud-native architecture, there are also management measures such as automated continuous delivery and Devops.
[0177] The representative technologies of cloud native include containers, service meshes, microservices, immutable infrastructure, and declarative APIs. By encapsulating the entire software runtime environment through Docker container technology, a platform for building, releasing, and running distributed applications, Docker can quickly and automatically deploy applications inside containers, and provide resource isolation and security guarantees for containers through operating system kernel technology, achieving lightweight, second-level deployment, easy portability, and elastic scaling. The main operation objects of Docker include images, containers, networks, and storage. Images are used to store and transport application programs. You can use images alone to build containers, or customize them to add other elements to expand the current configuration.
[0178] The core design points that cloud native architecture focuses on include indicators such as serviceability, elasticity, serverless degree, observability, resilience, and automation capabilities.
[0179] The essence of serviceability is based on service traffic control strategies, such as traffic limiting and degradation, circuit breaking and isolation, grayscale, backpressure, and zero-trust security. From the perspective of serviceability, cloud native architecture emphasizes standardized traffic transmission, policy control, and governance between services, and uses microservices for management, where the service governance system is the key.
[0180] Elasticity means that the system deployment scale can be scaled with the change of business volume, which can save costs and effectively cope with business peaks. Its elasticity is determined by manual, semi-manual, automatic and other modes, and elasticity is also related to the system's expansion ability.
[0181] Use cloud services to make the design of applications as stateless as possible, and save the stateful part to cloud services, especially in the case of self-operating open-source software. Entrust stateless computing such as computing, networking, and big data, as well as stateful storage such as databases, files, and object storage, to the cloud, and run them in a serverless manner such as Serverless and FaaS.
[0182] Through means such as logs, link tracing, and APM metrics, the time-consuming, return values, and parameters of service calls behind business clicks can be clearly visible, and even drill down to each third-party software call, SQL request, node topology, network response, system resources, etc.
[0183] Resilience is the system's ability to resist anomalies, reflecting the software's ability to continuously provide business services. The core goal is to increase the mean time between failures. Analyzed from the architecture design, it includes standards such as high availability, disaster tolerance, and asynchronous capabilities, such as eliminating single points, service grading, traffic limiting and degradation, timeout retry, disaster tolerance units, master-slave / cluster, and multi-site active-active.
[0184] Through the practice of DevOps and a large number of automated delivery tools in the CI / CD pipeline, by standardizing the delivery process and the self-describing and end-state-oriented delivery process, enabling automated tools to understand delivery goals and environmental differences, and achieving the automation of the entire software delivery and operation and maintenance.
[0185] (2.21) Natural integration based on containerized microservices / distributed / containerized;
[0186] A complete microservices system integrates various microservices capabilities such as Dubbo and Cloud, combines K8S containerized management, integrates one-click containerized management of CI / CD applications, and naturally integrates distributed management capabilities. It provides a complete integration of microservices capabilities, including nacos / zookeeper, etc. Based on the management and technical shielding of the basic platform, there is no perception for new arrivals, and it is easier to cut into containerization.
[0187] The project depends on the hexagonal concept, integrates the characteristics of the current project engineering structure, and makes plans and divisions, which can quickly complete the construction of the code framework for new services. The purpose is to make all services created on the project platform be based on the same set of frameworks and code structures, with a unified style and technical route, avoiding difficulties in later upgrades, etc. The code generator provides the following capabilities;
[0188] The project code generated by the alinesno-init scaffolding can help developers prepare for the preliminary work of development;
[0189] Make development faster, without spending time on preliminary configuration, and reduce the framework learning cost;
[0190] The default generated code configuration and online code generation management enable developers to focus more on business logic development;
[0191] By integrating basic platform components and data support components, it provides base capabilities for the development of applications;
[0192] The low-code generator simplifies the R & D process by using an automatically generated platform and integrating various environment configurations;
[0193] Build a description of the overall microservices platform framework and automatically generate integrated container configurations;
[0194] Automatically generate containerized configurations, avoid cumbersome container configurations, and form a unified specification;
[0195] The standard engineering structure and code generation structure provide better specifications for later upgrade management, provide technical service management without perception, and low-cost technical switching and follow-up.
[0196] (2.22) A large number of user concurrent processing technologies;
[0197] By adopting technologies such as parallel processing and distributed caching, the response speed of each service and the processing of massive data are ensured. Establishing a resource allocation system to fully schedule the idle resources of other servers and clients on the network, respond to service requests, and undertake part of the services can greatly improve the system performance and save hardware investment at the same time. The system model provides the allocator with the resource information visible on all nodes in a timely manner. After obtaining the information, the allocator reasonably allocates the resources to tasks, thereby optimizing the system performance. By establishing a resource allocation system, load balancing of services and grid computing are realized, so that when a large number of users perform concurrent operations on the network, the efficiency of the system will not be significantly reduced.
[0198] (2.23) XML-based data representation;
[0199] Data exchange is an open function. If the data formats used in data exchange vary widely, complex data encoding and decoding work is required. Therefore, unifying the data encapsulation format used in data exchange is the primary task in the construction of e-government platforms.
[0200] XML (eXtensible Markup Language) is a currently popular data representation standard internationally. Due to its characteristics such as simplicity, openness, scalability, flexibility, and self-descriptiveness, XML has important uses in many fields such as data and information management, data exchange, Web applications, e-commerce, and application integration. It has received widespread support from the industrial community and is also the standard adopted in China's e-government.
[0201] Using the XML method to represent the data to be exchanged by the system can not only facilitate data exchange between systems but also be convenient for expansion.
[0202] (2.24) Data visualization technology;
[0203] Data visualization has been a hot topic in the big data field in recent years. It belongs to an interdisciplinary subject of multiple disciplines such as human-computer interaction, graphics, image science, statistical analysis, and geographic information. It combines various knowledge and skills such as data processing, algorithm design, software development, and human-computer interaction, and displays data through forms such as images, charts, and animations to interpret the relationships and trends between data and improve the efficiency of reading and understanding data.
[0204] (2.24.1) Multidimensional data visualization;
[0205] (ⅰ) Geometry-based visualization method;
[0206] Parallel coordinates: Use parallel vertical lines to represent different dimensions, depict the values of multidimensional data on the coordinate axes, and connect the coordinate points on the number axes, thereby displaying multidimensional data in a two-dimensional space.
[0207] Scatter plot matrix: It shows the relationship between variables through a set of points in a two-dimensional coordinate system. The data of each dimension are combined in pairs and arranged regularly to draw a scatter plot. By combining visualization methods with the scatter plot matrix, the display effect of multi-dimensional data is enhanced.
[0208] Andrews curve method: It shows the visualization effect through a coordinate system. The multi-dimensional data are reflected into the coordinate system curve through periodic functions, and users can perceive data clustering and other situations by observing the curve.
[0209] (ⅱ) Icon-based visualization method;
[0210] The icon-based visualization method mainly uses geometric figures as icons to depict multi-dimensional data. The characteristic attributes of the icons (hardness, shape, length, size, etc.) reflect the dimensions of the information, and the visualization effect is reflected by the relationship between the icons and the multi-dimensional data. The representative icon-based visualization methods include the star plot method and the Chernoff face method. The star plot method maps the information dimensions through the way from points to lines, and the length of the line segments reflects the numerical values of different dimensions; the Chernoff face method reflects the information dimensions by identifying the facial shapes, features, etc., and draws a facial diagram to visually observe the information data.
[0211] (ⅲ) Dimensionality reduction mapping-based visualization method;
[0212] The dimensionality reduction mapping visualization method regards multi-dimensional information data as points in a certain dimension, determines the coordinates of the points according to the dimension attributes, and maps the points to a visible low-dimensional space on the premise of keeping the relationship between the information data unchanged. When reducing the dimension, some information data are selectively omitted, and finally the data set is presented in a two- or three-dimensional space.
[0213] (2.24.2) Visualization of time series data;
[0214] Time series visualization collects information data as time develops and presents them using visualization techniques. There are mainly three types of visualization methods presented. One is the line graph, which shows the changes in information data at different time periods through the initial points. In the visualization process, the information data presents more time dimensions, and corresponding icons are established according to different dimensions for arrangement to observe the changes in the data; the second is the stacked graph, which mainly stacks all time series. When there are negative numbers, the stacked graph cannot handle all time series, greatly reducing the visualization presentation effect; the third is the horizon graph, which can clearly observe the change rate of information data as time changes, and the depth of color represents the positive and negative change effects.
[0215] (2.24.3) Visualization of network data;
[0216] The core of network data visualization technology is the automatic layout algorithm, which automatically lays out and calculates information data and draws them into a network structure. There are three types of applications that are widely used.
[0217] Force-directed layout: With the help of the concept of force, force-bearing nodes are connected to draw a network diagram. Due to the existence of mutual repulsion, the overlap between nodes can be reduced. It is suitable for describing the relationship between things, such as computer network relationships, social network relationships, and other relationship network scenarios.
[0218] Circular layout: All nodes are sorted in a custom order and arranged in order on a circle to quickly analyze the results. Due to the size of the screen, when there are a large number of nodes, the radius of the circle becomes larger and larger, making it difficult to intuitively display all the nodes. This is suitable for finding nodes with more associations. For example, in the circular layout diagram, you can clearly tell which nodes have more associations.
[0219] Grid layout: Use grid design to draw grid-shaped information data network diagram, which is suitable for hierarchical networks and is conducive to observing the overall hierarchy.
[0220] (2.24.4) Visualization of hierarchical information data;
[0221] Hierarchical structures are often used to describe objects with obvious hierarchical structures, including library labels, computer hierarchical systems, or inheritance relationships between object-oriented program classes. The methods used for visualizing hierarchical information data mainly include node connection, space filling, and hybrid methods.
[0222] Node connection: Point connection mainly draws nodes of different shapes to represent the content of information data, and the lines between nodes represent the relationship between data. Representative technologies of this type of hierarchy include space tree and cone tree.
[0223] Space filling: Space filling mainly uses bounding boxes to represent hierarchical information data. The bounding relationship between the upper node and the lower node represents the structural relationship between the information data. Such hierarchical representation technologies include tree diagrams, information cubes, etc.
[0224] Hybrid method: Hybrid method combines the advantages of multiple visualization technologies to make cognitive behavior more efficient. Representative technologies of this type of method include elastic hierarchy, hierarchical network, etc.
[0225] Based on the construction requirements of the statistical analysis system for data elements, combined with the current characteristics of statistical business, and comprehensively considering the performance bottleneck problems in the Xinchuang environment, the scalability and maintainability requirements of future business, a distributed architecture design is adopted, and the front-end and back-end separation technology is used. The Vue framework is used for the front-end, and the microservice technology is used for the back-end. Containers are used to manage microservices. The database is designed by combining relational databases and analytical databases. Relational databases are used in the data collection, data management, and data analysis and application stages to ensure the uniqueness of transactions. Analytical databases are used in the stages of summarization, analysis, and generation of data products to ensure the timeliness of data production. At the same time, big data technology is introduced to solve complex data calculation operations.
[0226] The overall technical roadmap of the statistical analysis system for data elements (including program development management) is mainly divided into four parts:
[0227] (1) Program development management components, including software such as gitlab, Jenkins, and maven that are dependent on program development. They are mainly deveops process components to facilitate continuous integration during the development process;
[0228] (2) Basic program operation components, including containers and container orchestration tools k8s and docker in the cloud environment, including unified authentication for the API routing gateway, service registry, service configuration center, and ELK logging system;
[0229] (3) Program-dependent middleware components, including data synchronization tools and system monitoring software. The data synchronization tools are mainly used for data synchronization between different storages; the system monitoring software is mainly used to monitor system performance and can be set to give warnings;
[0230] (4) Data storage components, including three types of databases, relational databases, and cache database redis; among them, relational databases are used for query summarization of each application system; redis is used to store frequently used data, such as user information, directory data, etc.
[0231] The overall application program technical roadmap adopts the front-end and back-end separation development mode. When selecting the technology stack, the adaptation problem in the Xinchuang environment is fully considered, and a technology stack with strong compatibility and high adaptability is selected. The Vue framework is used for the front-end, the Spring Cloud technology roadmap is used for the back-end, and the uni-app hybrid development framework is used for the mobile terminal. Distributed relational database DRDS technology or open-source sharing-jdbc technology is used for data storage. Technologies such as database sharding, smooth expansion, and read-write separation can be used, which can greatly improve efficiency in large-scale data summarization and data query, and is easy to expand horizontally in the future.
[0232] The front end adopts the mainstream Vue framework. The advantages of this framework are that it is easy to quickly build complex page applications, such as review and acceptance pages, summary table designers, data product design modules, etc. It has excellent compatibility and is compatible with all mainstream browsers on the market, such as Firefox, domestic Xinchuang browsers, 360 Security Browser, IE, and Edge, etc. And it can quickly adapt and adjust the layout for mobile phones, tablets, and PC devices.
[0233] The back end adopts a microservices architecture and is developed using the mainstream Spring Cloud. Spring Cloud features high development response efficiency, flexible deployment, high stability, and strong scalability, and can quickly respond to new requirements put forward by professionals. The microservices architecture can minimize the coupling relationship between various services, splitting data filling, network opening preparation, review and acceptance, data aggregation, data query, data products, and user management into independent services, which do not affect each other.
[0234] The statistical analysis system for data elements uses a relational database for data storage. For application links such as data query, aggregation, and analysis, an analytical database with higher query and analysis efficiency is used to separate reading and writing, thus solving the performance bottleneck and improving the user experience. When designing the storage, the query and aggregation efficiency issues of large data volume tables are fully considered. For example, both the data acceptance record table and the data review record table are large tables with a scale of millions. The database sharding and table partitioning design is carried out for large tables with a scale of millions.
[0235] The data storage adopts a read-write separation architecture. The basic principle is to let the master database handle transactional insert, update, and delete operations (INSERT, UPDATE, DELETE), while the slave database handles SELECT query operations. Database replication is used to synchronize the changes caused by transactional operations to the slave databases in the cluster. That is, the data of services such as data collection and data reporting is written into the "write database", and the data applications such as data retrieval and query of the corresponding support system read data from the "read database". If there is high concurrency and high read requests, the frequently used hot data can be stored in the cache cluster, and the client application service reads data from the "read database" or the "cache database".
[0236] According to the requirements of the data product business (data analysis, editing and publishing) in the statistical analysis system for data elements, the WPS component needs to be embedded in the business system to browse and edit data products in a Web form; in the data product template, content such as text, tables, and graphics can be added, and comments can also be added to the content. To meet the needs of statistical users to quickly compile reporting documents.
[0237] The technical stack and component list of the statistical analysis system for data elements of the present invention are shown in Table 1:
[0238] Table 1
[0239]
[0240]
[0241]
[0242]
[0243] The statistical analysis system of data elements is a large-scale information system based on the Internet. The following is the security design from the technical level:
[0244] (1) System architecture security;
[0245] Besides ensuring the stability of the system, different levels of business logic are encapsulated. The "black box" operation between various business components can effectively protect the concealment and independence of the system logic.
[0246] The multi-level architecture adopted in this solution encapsulates the business logic operations and controls that general users don't need to see and runs them on the servers in the system background. The user's client is just a simple browser and the final content display of the system.
[0247] Using this system architecture that isolates the business implementation logic from users, malicious visitors can't even see the business logic, let alone maliciously tamper with it, thus ensuring the security and correctness of the system logic.
[0248] (2) Transmission security;
[0249] Encryption protection of data objects. It supports encrypting the data objects themselves. During the upload and download of data, the data goes through an encryption process. Therefore, even if a malicious visitor intercepts the data content, what they get is just a bunch of meaningless data, thus protecting the security of the data objects.
[0250] Encryption protection of the transmission protocol supports the standard secure encryption transmission protocol provided by Https, further enhancing the security of data transmission.
[0251] (3) Unified security authentication;
[0252] The system adopts unified identity authentication. When a user performs an authorization operation, after the authentication server verifies successfully, it generates an authorization code. After the application obtains the authorization, it sends the authorization code together with its identity information (ID password) on the platform to the authentication server for another verification request, indicating that its identity is correct and it has obtained user authorization, so as to obtain the permission to access user resources.
[0253] (4) Traceable auditing;
[0254] The system has a built-in multi-granularity logging system that can record actions of various operation granularities as needed in the log for tracking and auditing users' historical operations.
[0255] The advantage of the multi-granularity logging system is that it can not only record in a flowing manner, but also record more detailed historical actions according to the requirements of security levels; or simplify and ignore some registrations with low security requirements, so as to simplify the tracking and auditing work and improve efficiency.
Claims
1. A statistical analysis system based on data elements, deployed on a Windows server, Characterized in that: It includes a data management unit, an ETL tool, an accounting system, and a data element resource map unit; The data management unit is used to detect the use of data resources and perform scheduling optimization in real time; The ETL tool is used to support the data collection and cleaning work of various semi-structured and unstructured data in the statistical analysis system of data elements; The accounting system is used to collect data element data that various enterprises need to fill in through collection tasks for calculation; The data element resource map unit is used to display the calculation results of the accounting system in the form of charts.
2. The statistical analysis system based on data elements according to claim 1, Characterized in that, The data management unit includes a data resource usage management module, a real-time scheduling optimization module, a data governance module, a data analysis module, and a platform management module.
3. The statistical analysis system based on data elements according to claim 1, Characterized in that, The data sources supported by the ETL tool include databases, file systems, Excel, and Xml.
4. The statistical analysis system based on data elements according to claim 3, Characterized in that, The database is DB2, Oracle, Mysql, or SQLServer.