A spatial science data sharing system and method integrating big data analysis functions
By designing a spatial scientific data sharing system that integrates big data analysis functions, the lack of online analysis functions in the existing technology has been solved, the integration of data resources and computing resources has been achieved, data processing efficiency and access speed have been improved, and resource requirements and costs have been reduced.
Patent Information
- Application Number
- CN202410552786.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-05-07
AI Technical Summary
In the prior art, the spatial science data sharing system lacks online analysis functions, resulting in local analysis after data download, inefficient efficiency, high demand for computing and storage resources, and slow access speed.
Design a spatial scientific data sharing system that integrates big data analysis functions, including user management module, data service module, data collection module and resource management module, combined with data integration module, analysis and computing module, permission management module and data cache module, to realize the integration of data resources and computing resources, and support distributed storage and online analysis.
By integrating data resources and computing resources, the cumbersome process of data download is avoided, the efficiency of data processing is improved, the demand for network and storage resources is reduced, the speed of users accessing data is improved, and the computing and storage costs are reduced.
Smart Images

Figure CN118445313B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of scientific data sharing and management, and particularly to a spatial science data sharing system and method integrating big data analysis functions. Background Art
[0002] Currently, the sharing and transmission of spatial science data mainly rely on traditional file transfer and data storage methods. Data is usually stored on local servers or cloud storage, and users can access and download the data through a web browser.
[0003] With the rapid growth of spatial science data, the traditional method of downloading data for local analysis later is not only inefficient but also requires a huge amount of network resources and storage resources. How to improve the convenience of data processing and use has become an urgent problem to be solved. The main defects and problems existing in the existing scientific data sharing technologies include:
[0004] (1) Low data download efficiency: Due to the huge amount of data in the scientific data center and the limited bandwidth of the current public network, the time for users to download data is too long. This not only affects the progress of scientific research work but also causes a huge pressure on network resources and storage resources.
[0005] (2) Lack of online analysis function: The existing scientific data sharing systems do not integrate data analysis modules, and scientific researchers need to download data for local analysis. This method not only requires a relatively high demand for computing resources and storage devices but also requires scientific researchers to deploy relevant basic software and scientific software by themselves, increasing the time cost.
[0006] (3) The access speed needs to be improved: Due to the huge amount of data and the relatively small size of each file, the speed is significantly slow when accessing massive data. This not only affects the user experience but also limits the data processing efficiency. Summary of the Invention
[0007] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a spatial science data sharing system and method integrating big data analysis functions. The present invention solves the problems of excessive data volume, low speed for each access, inability to perform online analysis for scientific data sharing, low data processing efficiency, and too high time cost in the prior art.
[0008] To achieve the above purpose, the present invention provides the following solutions:
[0009] A spatial science data sharing system integrating big data analysis functions, the spatial science data sharing system includes: a user management module, a data service module, a data aggregation module, and a resource management module, and further includes:
[0010] A data integration module, a personal data input module, a data caching module, an analysis and calculation module, and a permission management module;
[0011] The data integration module is connected to the resource management module, the personal data input module is connected to the data aggregation module, the analysis and calculation module is connected to the data service module, the permission management module is connected to the user management module, and the data caching module is respectively connected to the data integration module, the analysis and calculation module, and the permission management module;
[0012] The permission management module is used to analyze the identity information, download permission, and analysis permission of the user to be accessed, and obtain a first analysis result. The personal data input module is used to upload the user's personal data to the data aggregation module. The analysis and calculation module is used to analyze and calculate the model data, observation data, and user's personal data in the data aggregation module, obtain a second analysis result and process data, and record them. The data integration module is used to integrate the data resources and computing resources in the resource management module, obtain integrated data, and perform distributed storage, scheduling, and management on the integrated data using a network file system and a container orchestration system. The data caching module is used to obtain high-frequency data according to the first analysis result and cache and update the high-frequency data.
[0013] Preferably, the permission management module includes:
[0014] A user verification sub-module, a download permission verification sub-module, and an analysis permission verification sub-module;
[0015] The user verification sub-module is used to verify the identity information to be accessed and obtain a first verification result. The download permission verification sub-module is used to verify the download permission of the user to be accessed when the first verification result is passed, and obtain a second verification result. The analysis permission verification sub-module is used to verify the analysis permission of the user to be accessed when the first verification result is passed, and obtain a third verification result.
[0016] Preferably, the analysis and calculation module includes:
[0017] A first analysis sub-module, a second analysis sub-module, and a recording sub-module;
[0018] The first analysis sub-module is used to analyze and calculate the model data, observation data, and user's personal data to obtain process data. The second analysis sub-module is used to perform sensitive operation analysis on the process data to obtain a second analysis result. The recording sub-module is used to record the second analysis result and the process data.
[0019] Preferably, the data integration module includes:
[0020] A combination sub-module, a distributed storage sub-module, and a resource scheduling sub-module;
[0021] The combination sub-module is used to integrate the data resources and computing resources to obtain integrated data. The distributed storage sub-module uses a network file system to perform distributed storage on the integrated data. The resource scheduling sub-module is used to schedule and manage the integrated data using a container orchestration system.
[0022] Preferably, the resource scheduling sub-module includes:
[0023] A first environment setting unit and a second environment setting unit;
[0024] The first environment setting unit uses a container engine to set up an independent computing environment to achieve the management and scheduling of computing resources. The second environment setting unit uses Jupyternotebook to set up an independent analysis environment to achieve the management and scheduling of data resources.
[0025] Preferably, the data cache module includes:
[0026] A cache identification sub-module, a cache reception sub-module, and a cache update sub-module;
[0027] The cache identification sub-module is used to identify the data in the data service module according to the first analysis result to obtain high-frequency data. The cache reception sub-module is used to cache the high-frequency data. The cache update sub-module is used to update the current high-frequency data in the cache reception sub-module in real time according to the second analysis result and the integrated data.
[0028] A method for sharing space science data integrating big data analysis functions, the method including:
[0029] A user accesses the data service module and performs data retrieval, and determines whether to download or analyze data according to the retrieval result to obtain a determination result;
[0030] If the determination result is yes, the user accesses the user management module to submit a login request and uses the permission management module to verify and analyze the login request to obtain a first analysis result;
[0031] Identify the data in the data service module according to the first analysis result to obtain high-frequency data and cache it;
[0032] If the first analysis result is passed, the Web service reads the Jupyter Notebook service configuration information in the database service, creates a data storage area for the user according to the service configuration, mounts the scientific data storage area, creates a user container. The created container meets the CPU, memory, and storage resource limit amounts in the service configuration information, runs the Jupyter Notebook program, starts the Jupyter Notebook service, and the Web service forwards the URL path of the service to the user terminal through the reverse proxy service. The user terminal logs in to the analysis and calculation module using the obtained URL path to get the second analysis result;
[0033] Store the second analysis result in the resource management module;
[0034] Use the data integration module to integrate all the data in the resource management module and permanently retain it through NFS distributed storage and MySQL storage services. After the Jupyter Notebook service stops, the container resources where it is located are immediately deleted;
[0035] Use the second analysis result and the integrated data to update the high-frequency data in real time.
[0036] Preferably, it further includes:
[0037] The user stops the operation of the Jupyter Notebook service through the user terminal by using the "Stop Service" menu option in the page provided by the Jupyter Notebook service; or
[0038] The Jupyter Notebook service automatically stops after being idle for a fixed duration;
[0039] The container where the Jupyter Notebook service is located is terminated and deleted, and the occupied hardware resources are released.
[0040] The present invention discloses the following technical effects:
[0041] The present invention provides a space science data sharing system integrating big data analysis functions. The space science data sharing system includes: a user management module, a data service module, a data collection module, and a resource management module, and further includes: a data integration module, a personal data input module, a data cache module, an analysis and calculation module, and a permission management module; the data integration module is connected to the resource management module, the personal data input module is connected to the data collection module, the analysis and calculation module is connected to the data service module, the permission management module is connected to the user management module, and the data cache module is respectively connected to the data integration module, the analysis and calculation module, and the permission management module; the permission management module is used to analyze the identity information, download permission, and analysis permission of the user to be accessed to obtain a first analysis result, the personal data input module is used to upload the user's personal data to the data collection module, the analysis and calculation module is used to analyze and calculate the model data, observation data, and user's personal data in the data collection module to obtain a second analysis result and process data and record them, the data integration module is used to integrate the data resources and computing resources in the resource management module to obtain integrated data and perform distributed storage, scheduling, and management on the integrated data by using a network file system and a container orchestration system, and the data cache module is used to obtain high-frequency data according to the first analysis result and cache and update the high-frequency data. By using the data integration module to integrate data resources and computing resources, the present invention avoids the cumbersome process of data downloading; adding analysis modules to each module enables users to quickly obtain and use data, improving work efficiency. The cache module is used to identify and store high-frequency data, improving the speed of user data access and reducing the access pressure on the original data source. The present invention provides strong support for the research and application in the field of space science. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0043] Figure 1 FIG. is a schematic structural diagram of a space science data sharing system integrating big data analysis functions provided by an embodiment of the present invention;
[0044] Figure 2 FIG. is a schematic functional diagram of a space science data sharing system integrating big data analysis functions provided by an embodiment of the present invention;
[0045] Figure 3Schematic flow chart of a spatial science data sharing method integrating big data analysis function provided by an embodiment of the present invention.
[0046] Explanation of reference numerals:
[0047] 1 - User management module, 2 - Data service module, 3 - Data aggregation module, 4 - Resource management module, 5 - Data integration module, 6 - Personal data input module, 7 - Analysis and calculation module, 8 - Permission management module, 9 - Data cache module. Detailed implementation manners
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific implementation manners.
[0050] As Figure 1 shown, the present invention provides a spatial science data sharing system and method integrating big data analysis function. The spatial science data sharing system includes: user management module 1, data service module 2, data aggregation module 3, and resource management module 4. The system further includes:
[0051] Data integration module 5, personal data input module 6, data cache module 9, analysis and calculation module 7, and permission management module 8;
[0052] The data integration module 5 is connected to the resource management module 4, the personal data input module 6 is connected to the data aggregation module 3, the analysis and calculation module 7 is connected to the data service module 2, the permission management module 8 is connected to the user management module 1, and the data cache module 9 is respectively connected to the data integration module 5, the analysis and calculation module 7, and the permission management module 8;
[0053] The permission management module 8 is used to analyze the identity information, download permission, and analysis permission of the user to be accessed, and obtain a first analysis result. The personal data input module 6 is used to upload the user's personal data to the data aggregation module 3. The analysis and calculation module 7 is used to analyze and calculate the model data, observation data, and user's personal data in the data aggregation module 3, obtain a second analysis result and process data, and record them. The data integration module 5 is used to integrate the data resources and computing resources in the resource management module 4, obtain integrated data, and perform distributed storage, scheduling, and management on the integrated data using a network file system and a container orchestration system. The data caching module 9 is used to obtain high-frequency data according to the first analysis result and cache and update the high-frequency data.
[0054] Among them, the resource management module 4: mainly includes the management of computing resources and storage resources.
[0055] The data aggregation module 3: collects scientific data from multiple sources, mainly including model data, observation data, etc. These data have different formats and standards, so operations such as data cleaning, format conversion, and standardization need to be performed to ensure the accuracy and consistency of the data.
[0056] The data service module 2: needs to provide data sharing and distribution functions, mainly including modules such as classification navigation, data retrieval, and visualization display. Users can select data for download according to their needs. In order to improve the data access speed and efficiency, a data distribution network and a caching mechanism also need to be established.
[0057] The user management module 1: includes functions such as user registration approval and unified login.
[0058] Furthermore, as Figure 2 shown, an analysis function is added on the basis of the current mainstream functional modules, specifically including: (1) adding analysis permission management to user management; (2) adding analysis and calculation functions to data services; (3) allowing users to submit personal data in their personal space for comprehensive analysis in data aggregation; (4) adding the management of container resources to resource management.
[0059] Specifically, the present invention realizes the big data analysis function through technologies such as Kubernetes (K8s, an open-source container orchestration engine), Docker (containerization technology), JupyterHub, and NFS (Network File System). It combines data resources and computing resources to provide users with a service to directly analyze and obtain results in the system, avoiding the cumbersome process of data downloading. The NFS (Network File System) is used to provide file sharing services for users to achieve distributed storage of data. K8s is used for scheduling and management of computing resources. Through Docker containerization of applications, each analysis task has an independent computing environment without mutual influence. Jupyternotebook is used to provide an interactive programming analysis environment for users. Users can view and analyze all data resources with access permissions, and can also upload their own data to the system for joint analysis. The user analysis results are saved in the storage system, and users can view the analysis results in real time or export and save them.
[0060] Further, the permission management module 8 includes:
[0061] A user verification sub-module, a download permission verification sub-module, and an analysis permission verification sub-module;
[0062] The user verification sub-module is used to verify the identity information of the user to be accessed to obtain a first verification result. The download permission verification sub-module is used to verify the download permission of the user to be accessed when the first verification result is passed to obtain a second verification result. The analysis permission verification sub-module is used to verify the analysis permission of the user to be accessed when the first verification result is passed to obtain a third verification result.
[0063] Further, the analysis and calculation module 7 includes:
[0064] A first analysis sub-module, a second analysis sub-module, and a recording sub-module;
[0065] The first analysis sub-module is used to analyze and calculate model data, observation data, and user personal data to obtain process data. The second analysis sub-module is used to perform sensitive operation analysis on the process data to obtain a second analysis result. The recording sub-module is used to record the second analysis result and the process data.
[0066] Further, the data integration module 5 includes:
[0067] A combination sub-module, a distributed storage sub-module, and a resource scheduling sub-module;
[0068] The combination sub-module is used to integrate the data resources and computing resources to obtain integrated data. The distributed storage sub-module uses a network file system to perform distributed storage on the integrated data. The resource scheduling sub-module is used to schedule and manage the integrated data by using a container orchestration system.
[0069] Further, the resource scheduling sub-module includes:
[0070] A first environment setting unit and a second environment setting unit;
[0071] The first environment setting unit uses a container engine to set up an independent computing environment to achieve the management and scheduling of computing resources. The second environment setting unit uses Jupyternotebook to set up an independent analysis environment to achieve the management and scheduling of data resources.
[0072] Specifically, the system provides multiple operating environments and can be customized according to user needs, supports multiple programming languages such as Python and R, and supports deep learning frameworks such as Pytorch and TensorFlow. It can also integrate machine learning algorithm libraries such as XGBoost and SVM. The system supports multi-user simultaneous online analysis. Through container technology and dynamic resource scheduling, it ensures that each user can obtain sufficient computing resources. The system adopts a data access control mechanism to ensure the security and privacy of data. At the same time, sensitive operations during the user analysis process are audited and recorded to ensure that data is not misused.
[0073] In terms of system expansion and maintenance, through the container orchestration function of K8s, the system can be horizontally expanded to cope with large-scale data processing requirements. In terms of system maintenance, by using the image management function of Docker, system components can be quickly deployed and updated, improving the maintainability of the system.
[0074] The data cache module 9 includes:
[0075] A cache identification sub-module, a cache reception sub-module, and a cache update sub-module;
[0076] The cache identification sub-module is used to identify the data in the data service module according to the first analysis result to obtain high-frequency data. The cache reception sub-module is used to cache the high-frequency data. The cache update sub-module is used to update the current high-frequency data in the cache reception sub-module in real time according to the second analysis result and the integrated data.
[0077] Specifically, the data cache module will store high-frequency data that is frequently accessed to reduce the number of accesses to the original data source.
[0078] Specifically, the cache recognition sub-module includes: an interaction unit, a recognition unit, and an output unit;
[0079] The interaction unit is used to interact with the data in the data service module and the integration module to obtain interaction data, which includes: model data, user data, observation data, etc. When the recognition unit recognizes the interaction data, a preset high-frequency data standard is set, and the high-frequency data standard includes: the importance of the data, the access frequency of the data, the update frequency of the data, etc.; the output unit is used to output the recognized high-frequency data to the cache receiving sub-module.
[0080] More specifically, according to the access frequency of the data, the present invention adopts an algorithm based on the access frequency: by counting the access frequency of the data, the data that is frequently accessed can be identified, and thus it can be recognized as high-frequency data. The frequent pattern mining algorithm (such as the Apriori algorithm) can be used to discover the frequent access patterns of the data; or according to the data download permission and data analysis permission in the first analysis result, the download frequency and analysis frequency of the data are judged, and the high-frequency data is identified according to the download frequency and analysis frequency.
[0081] Furthermore, a log recording unit and a data access counting unit are set. After each user logs in to the system, the access to the data is recorded. When there is an access record for each data item in the log recording unit, the data access counting unit adds 1. The data in the data access counting unit is statistically analyzed for a period of time to calculate the data access frequency. A frequency threshold is set, and according to the comparison between the access frequency of each data item and the frequency threshold, the high-frequency data is determined. Among them, the frequency threshold can also be fine-tuned according to the access frequency of the data.
[0082] Specifically, the system can analyze the correlation between the data. By analyzing the correlation, dependence, and influence and other association relationships between the data, the importance of the data is judged according to the association between the data, so as to determine the high-frequency data.
[0083] Furthermore, the cache receiving sub-module can be a metadata cache (using memory) to improve the data retrieval speed; or a high-frequency data cache (local hard disk) to improve the access speed of entity data.
[0084] Furthermore, the data in the data service module, the data aggregation module, and the resource management module is monitored in real time. If the data is modified, the cache update sub-module is linked. Through the API of the data, the cache update sub-module obtains the data change items and updates the data.
[0085] Specifically, an effective cache update mechanism is established to ensure that the data in the cache is synchronized with the original data source. When the data changes, the data in the cache is updated in a timely manner.
[0086] As Figure 3 shown, this embodiment also provides a spatial science data sharing method integrated with big data analysis function. The method includes:
[0087] The user accesses the data service module and performs data retrieval, and judges whether to download or analyze the data according to the retrieval result to obtain a judgment result;
[0088] If the judgment result is yes, the user accesses the user management module to submit a login request, and the permission management module verifies and analyzes the login request to obtain a first analysis result;
[0089] Identify the data in the data service module according to the first analysis result, obtain high-frequency data and cache it;
[0090] If the first analysis result is passed, the Web service reads the Jupyter Notebook service configuration information in the database service, creates the user's data storage area, mounts the scientific data storage area, and creates a user container according to the service configuration. The created container meets the CPU, memory, and storage resource limit amounts in the service configuration information, runs the Jupyter Notebook program, starts the Jupyter Notebook service, and the Web service forwards the URL path of the service to the user terminal through the reverse proxy service. The user terminal logs in to the analysis and calculation module using the obtained URL path to obtain a second analysis result;
[0091] Store the second analysis result in the resource management module;
[0092] Use the data integration module to integrate all the data in the resource management module, and permanently retain it through the NFS distributed storage and MySQL storage service. After the Jupyter Notebook service stops, the container resources where it is located are immediately deleted;
[0093] Use the second analysis result and the integrated data to update the high-frequency data in real time.
[0094] Specifically, the system provides services for users through the portal website. Users can first retrieve data. If users wish to download or analyze data, they need to visit the login page and submit a login request. This request is passed to the Web service through the reverse proxy service. The Web service reads the user data in the database and performs a comparison to ensure that the verification is correct. For download tasks, the system guides users to the corresponding download page, and when users click the download link, the file download is initiated. For analysis tasks, the system calls the JupyterHub service. A storage area for the user data file is created, and the user configuration information is viewed from the database. Then, the system controls K8s to create containers according to the user configuration and mounts the storage area of the user data file into the containers. Finally, the system starts the JupyterNotebook service and forwards the page of the Jupyter Notebook service to the user through the reverse proxy service.
[0095] The specific description is as follows. When users use this system through the user terminal, they first visit the portal website and can retrieve data. If users wish to download or analyze data, they need to log in. After successful authentication, they will be automatically redirected to the download page or the JupyterNotebook programming page.
[0096] Among them, when users use the system for the first time, they need to register an account. After successful registration, the account information and the initialized JupyterNotebook service configuration information are stored in the MySQL database to complete the registration work. Default data access permissions and service resource configurations are granted during the initial registration. If special permissions are required, a separate application needs to be submitted. The download function process provided by the system is similar to the traditional method. The process of the analysis function is introduced in detail below.
[0097] When users log in for the first time and access the analysis function, the Web service reads the JupyterNotebook service configuration information in the database service, creates the data storage area for this user, mounts the scientific data storage area, creates user containers according to the service configuration. The created containers meet the resource limit amounts for CPU, memory, storage, etc. in the service configuration information, and the JupyterNotebook program is run to start the JupyterNotebook service.
[0098] After the JupyterNotebook service is started, the Web service forwards the URL path of this service to the user terminal through the reverse proxy service. The user terminal uses the obtained URL path to log in to the page provided by the Jupyter Notebook service for subsequent programming work.
[0099] Users can stop the operation of the JupyterNotebook service through the user terminal by using the "Stop Service" menu option in the page provided by the JupyterNotebook service. Alternatively, the JupyterNotebook service will automatically stop when it has been idle for a fixed period of time. The container where the JupyterNotebook service is located is terminated and deleted, and the hardware resources it occupies are released. However, the data storage area in the file sharing service and the JupyterNotebook service configuration file are permanently retained through the NFS distributed storage and the MySQL storage service respectively; the container resources where the JupyterNotebook service is located are immediately deleted after the service stops.
[0100] When the user logs in and accesses the analysis function again, the Web service reads the Jupyter Notebook service configuration data in the database and starts the JupyterNotebook service according to the configuration. The data storage area of this user created previously in the file sharing service remains unchanged.
[0101] The beneficial effects of this embodiment are as follows:
[0102] (1) Efficient data sharing: Through the integrated analysis function, the data sharing efficiency is significantly improved, thus accelerating the process of scientific research and innovation work. Users can quickly obtain and use the required data, reducing waiting and download time and improving work efficiency.
[0103] (2) Powerful data analysis capabilities: It supports multiple programming languages, providing users with rich choices. At the same time, it integrates a variety of machine learning frameworks and algorithm libraries, enabling users to select appropriate methods for data processing and analysis according to actual needs. This design meets the diverse needs of different users and makes the data processing process more intuitive and convenient.
[0104] (3) Reducing computing and storage costs: The system provides users with an integrated solution for data resources and computing resources. Users do not need to purchase a large number of storage devices or rent computing resources separately. This greatly reduces the costs of users and alleviates the economic burden.
[0105] (4) Strong stability and scalability: With the help of containerization technology, the stability and consistency of the analysis environment are ensured. Combining K8s and NFS technologies, the system can not only meet the current needs, but also has excellent scalability and can flexibly expand as the number of users and data volume grow.
[0106] (5) Flexible permission management mechanism: Permissions for the data that users can download and the data that can be analyzed can be assigned separately or managed uniformly. The computing resources and storage resources that users can use can be flexibly configured. Through Docker technology, in the case of multi-user operation, it meets the requirements of an operating environment where independent computing resources, data storage areas, program execution permissions, etc. are isolated from each other.
[0107] (6) Fast data access: The data cache module is used to cache high-frequency data, and high-frequency data in the data cache module can be directly accessed during each user access, improving the user's access speed.
[0108] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0109] In this article, specific examples are used to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A space science data sharing system integrating big data analysis functions, the space science data sharing system comprising: The user management module, the data service module, the data collection module and the resource management module are characterized by further comprising: Data integration module, personal data input module, data cache module, analysis and calculation module and authority management module; The data integration module is connected to the resource management module, the personal data input module is connected to the data collection module, the analysis and calculation module is connected to the data service module, the authority management module is connected to the user management module, and the data cache module is connected to the data integration module, the analysis and calculation module and the authority management module respectively; The permission management module is used to analyze the identity information, download permission and analysis permission of the access user to obtain a first analysis result. The personal data input module is used to upload the user's personal data to the data collection module. The analysis and calculation module is used to analyze and calculate the model data, observation data and user's personal data in the data collection module to obtain a second analysis result and process data and record them. The data integration module is used to integrate the data resources and computing resources in the resource management module to obtain integrated data and use the network file system and container orchestration system to perform distributed storage, scheduling and management on the integrated data. The data cache module is used to obtain high-frequency data according to the first analysis result and cache and update the high-frequency data; The data cache module comprises: Cache identification submodule, cache receiving submodule and cache updating submodule; The cache identification submodule is used to identify the data in the data service module according to the first analysis result to obtain high-frequency data, the cache receiving submodule is used to cache the high-frequency data, and the cache updating submodule is used to update the current high-frequency data in the cache receiving submodule in real time according to the second analysis result and the integrated data; The cache identification submodule includes: an interaction unit, an identification unit and an output unit; The interaction unit is used to interact with the data in the data service module and the integration module to obtain interaction data, and the interaction data includes: model data, user data and observation data. When the identification unit identifies the interaction data, a high-frequency data standard is preset, and the high-frequency data standard includes: data importance, data access frequency and data update frequency; the output unit is used to output the identified high-frequency data to the cache receiving submodule.
2. A spatial science data sharing system integrating big data analysis functions according to claim 1, characterized in that: The rights management module includes: User verification submodule, download permission verification submodule and analysis permission verification submodule; The user verification submodule is used to verify the identity information of the accessed user and obtain a first verification result; the download permission verification submodule is used to verify the download permission of the accessed user when the first verification result is passed and obtain a second verification result; the analysis permission verification submodule is used to verify the analysis permission of the accessed user when the first verification result is passed and obtain a third verification result.
3. The spatial science data sharing system with integrated big data analysis function according to claim 1 is characterized in that: The analysis and calculation module comprises: A first analysis submodule, a second analysis submodule and a recording submodule; The first analysis submodule is used to analyze and calculate the model data, observation data and user personal data to obtain process data. The second analysis submodule is used to perform sensitive operation analysis on the process data to obtain a second analysis result. The recording submodule is used to record the second analysis result and the process data.
4. The spatial science data sharing system integrating big data analysis function according to claim 1 is characterized in that: The data integration module includes: Combining submodule, distributed storage submodule and resource scheduling submodule; The combining submodule is used to integrate the data resources and computing resources to obtain integrated data. The distributed storage submodule uses a network file system to perform distributed storage on the integrated data. The resource scheduling submodule is used to schedule and manage the integrated data using a container orchestration system.
5. A spatial science data sharing system integrating big data analysis function according to claim 4, characterized in that: The resource scheduling submodule includes: a first environment setting unit and a second environment setting unit; The first environment setting unit uses the container engine to set up an independent computing environment to achieve management and scheduling of computing resources, and the second environment setting unit uses Jupyter notebook to set up an independent analysis environment to achieve management and scheduling of data resources.
6. A spatial science data sharing method integrating big data analysis function, applied to the system according to any one of claims 1 to 5, characterized in that: The method comprises: The user accesses the data service module and performs data retrieval, determines whether the data needs to be downloaded or analyzed based on the retrieval results, and obtains the judgment result; If the judgment result is yes, the user accesses the user management module to submit a login request and uses the authority management module to verify and analyze the login request to obtain a first analysis result; Identify the data in the data service module according to the first analysis result, obtain high-frequency data and cache it; If the first analysis result is passed, the Web service reads the Jupyter Notebook service configuration information in the database service, creates the user's data storage area, mounts the scientific data storage area, and creates a user container according to the service configuration. The created container meets the CPU, memory, and storage resource restrictions in the service configuration information. The Jupyter Notebook program is run to start the Jupyter Notebook service, and the Web service forwards the URL path of the service to the user terminal through the reverse proxy service. The user terminal uses the obtained URL path to log in to the analysis and calculation module to obtain the second analysis result; storing the second analysis result in a resource management module; All data in the resource management module are integrated by using the data integration module and permanently retained through NFS distributed storage and MySQL storage services, and the container resources where the JupyterNotebook service is located are immediately deleted after the JupyterNotebook service is stopped; The high-frequency data is updated in real time using the second analysis result and the integrated data.
7. A spatial science data sharing method integrating big data analysis function according to claim 6, characterized in that: Also includes: The user stops the JupyterNotebook service through the user terminal by using the "Stop Service" menu option on the page provided by the JupyterNotebook service; or the JupyterNotebook service automatically stops after being idle for a fixed period of time; The container where the JupyterNotebook service resides is terminated and deleted, and the occupied hardware resources are released.
Citation Information
Patent Citations
System for processing source network load multivariate data
CN112699162A
Large-scale reinforcement learning training task management platform based on Kubernetes cluster
CN115906999A