A numerical wave tank virtual test management system based on machine learning

Through an automated distributed storage system based on machine learning, the cumbersome problem of numerical pool virtual test data management in the existing technology is solved, efficient and automated data management and storage is achieved, experimental accuracy and reliability are improved, and costs are reduced.

CN116821238BActive Publication Date: 2025-07-11HARBIN ENG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310770494.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-28
Publication Date
2025-07-11
Estimated Expiration
2043-06-28

AI Technical Summary

Technical Problem

When managing and maintaining numerical pool virtual test data, existing distributed storage systems are difficult to adapt to different usage scenarios and workloads, resulting in cumbersome management and inability to achieve efficient and automated storage management.

Method used

An automated distributed storage system based on machine learning is adopted to realize automatic data collection, processing and analysis through data collection modules, central storage modules and automated storage modules. It combines distributed storage technology to allocate storage tasks between multiple computer nodes, and uses machine learning to identify and process outliers to improve experimental accuracy and reliability.

Benefits of technology

Improves test efficiency and accuracy, reduces test costs, realizes automated management and efficient and reliable data storage, and supports fault-tolerant processing and multi-copy backup.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821238B_ABST
    Figure CN116821238B_ABST
Patent Text Reader

Abstract

The present invention discloses a numerical tank virtual test management system based on machine learning, including a data collection module for obtaining initial data in various scenarios of digital virtual tank tests and randomly generating simulation data; the simulation data being data to be processed in various scenarios of hydrodynamic phenomena including water flow movement, water level change, and waves in an actual hydrological environment; a central storage module connected to the data collection module for, after receiving a storage task sent by the data collection module, processing the simulation data and performing underlying storage of the processed simulation data; and an automated storage module connected to the central storage module for automatically adjusting the storage resource configuration according to the prediction analysis result and maintaining data consistency of multiple copies. The present invention improves the storage capacity of the system for test data, as well as the accuracy and efficiency of the test, and realizes the automated management of numerical tank virtual tests.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence and distributed systems, and particularly relates to a numerical tank virtual test management system based on machine learning. Background Art

[0002] Numerical tank virtual test data refers to the data generated when using computer simulation methods to numerically simulate hydrodynamic phenomena. In hydrodynamic research, numerical tank virtual tests are an important means. Through computer simulation, hydrodynamic phenomena such as water flow movement, water level changes, and waves in the actual hydrological environment can be simulated. Numerical tank virtual test data usually includes simulated hydrological environment parameters, hydrodynamic parameters such as water flow velocity, water level height, and wave shape in the simulation results, as well as information such as relevant numerical calculation errors. These data can be used for the establishment, verification, and optimization of hydrodynamic models, and can also be used in applications in fields such as engineering design, water resource management, and disaster warning. The acquisition and processing of numerical tank virtual test data involve knowledge and technologies in multiple fields such as numerical calculation, computer science, and mathematics, and have a certain degree of difficulty and complexity.

[0003] Driven by technologies such as cloud computing, big data, and artificial intelligence, the current data storage demand for numerical tanks shows a trend of linear growth. To meet this demand, distributed storage systems have gradually become the mainstream, but their management and maintenance work is very cumbersome and complex. To address this problem, existing distributed storage systems usually require manual intervention for management and maintenance, cannot adaptively adapt to different usage scenarios and workloads, and are difficult to achieve efficient and automated storage management. Summary of the Invention

[0004] The purpose of the present invention is to establish a numerical tank virtual test management system based on machine learning automated distributed storage to manage numerical tank virtual tests more efficiently and accurately. Through machine learning technology, the command system can automatically collect, process, and analyze a large amount of data, thereby realizing the automated management of numerical tank virtual tests. In this way, researchers can focus more on analyzing test results without spending a lot of time and energy on data processing. At the same time, to improve the storage capacity and the accuracy and efficiency of the tests, the system adopts distributed storage technology, which can distribute data storage tasks among multiple computer nodes, thereby increasing the storage capacity of the system for test data. At the same time, the system can also automatically identify and process outliers and noise in the data through machine learning technology, thereby improving the accuracy and reliability of the tests.

[0005] To achieve the above object, the present invention provides the following solution: A numerical tank virtual test management system based on machine learning, comprising:

[0006] A data collection module, configured to obtain initial data in various scenarios during digital virtual pool tests and randomly generate simulation data; the simulation data is data to be processed for various scenarios of hydrodynamic phenomena including water flow movement, water level changes, and waves in an actual hydrological environment.

[0007] A central storage module, connected to the data collection module, configured to process the simulation data and perform underlying storage of the processed simulation data when receiving a storage task sent by the data collection module.

[0008] An automated storage module, connected to the central storage module, configured to automatically adjust storage resource allocation according to prediction analysis results and maintain data consistency of multiple copies.

[0009] Preferably, the data collection module includes a data transmission unit, an intelligent agent unit, a data collection unit, a mode conversion unit, a data chunking unit, and a communication unit that are connected in sequence.

[0010] The data transmission unit is configured to transmit initial data in various scenarios during digital virtual pool tests to the intelligent agent unit, and is also configured to transmit the chunked simulation data to the central storage module through the communication unit.

[0011] The intelligent agent unit is configured to automatically identify data categories in various scenarios; the data collection unit is configured to collect initial data in various scenarios.

[0012] The mode conversion unit is configured to determine a data storage model according to different data categories and provide an ordered traversal method.

[0013] The data chunking unit is configured to chunk according to keys, divide the entire key-value space into many segments, each segment is a series of consecutive keys, a certain segment of consecutive keys is stored on a storage node, and each segment is divided into a block area, which is described by a left-closed right-open interval like [start key, end key).

[0014] Preferably, the central storage module includes a machine learning unit, a data distribution unit, and an underlying storage unit MicaDB.

[0015] The machine learning unit is configured to predict and optimize the operation state and performance of the pool according to historical data and real-time data.

[0016] The data distribution unit is configured to distribute data storage tasks among multiple computer nodes through distributed storage technology, and integrate and manage data of multiple pools.

[0017] The underlying storage unit MicaDB is configured to perform underlying data storage on the block area through the Star protocol.

[0018] Preferably, the machine learning unit includes a communication interface unit, a data preprocessing unit, a model training unit, and a prediction unit;

[0019] The communication interface unit is connected to the communication unit and is used to receive the analog data after data chunking;

[0020] The data preprocessing unit is used to perform data cleaning, feature selection, feature extraction, and feature transformation on the analog data after data chunking;

[0021] The model training unit is used to train a machine learning model with the data processed by the data preprocessing unit to obtain a target prediction model;

[0022] The prediction unit is used to input real-time data into the target prediction model to perform predictive analysis on future storage requirements and obtain a prediction result.

[0023] Preferably, the central storage module includes storage nodes and a central management node. Among them, the storage nodes are deployed in the machine learning unit, and an intelligent agent is installed on each storage node to automatically identify and process anomalies of the storage nodes and data and report them to the central management node; the intelligent agent is also used to summarize and report the usage information of the storage nodes and data to the central management node;

[0024] The central management node is used to automatically analyze the usage patterns of storage nodes and data to discover and predict possible faults, and automatically start corresponding fault handling procedures when faults occur.

[0025] Preferably, the central management node includes a machine learning engine and a controller;

[0026] The machine learning engine is used to train with historical data and existing data to automatically analyze the usage patterns of storage nodes and data and predict future usage;

[0027] The controller is used to formulate a storage strategy based on the prediction results of the machine learning engine and issue corresponding commands to the storage node intelligent agent for execution.

[0028] Preferably, the storage nodes can dynamically join or leave the system, and the communication between the intelligent agents of the storage nodes and the central management node is carried out in a secure encryption manner.

[0029] Preferably, the automated storage module adopts the Star protocol, and the Star protocol includes leader selection, log replication, and security measures;

[0030] The leader selection requires that a new leader must be selected if the existing leader fails;

[0031] The log replication means that when the leader accepts commands from the client, records them as logs, and replicates them to other servers in the cluster, and forces the logs of other nodes to be consistent with the leader.

[0032] The security measures are used to ensure the security of the system through some measures.

[0033] Preferably, the node states in the Star protocol include leader, follower, and candidate.

[0034] The leader is used to receive all requests from the client, including processing write requests, managing log replication, and continuously sending heartbeat messages to notify other nodes not to initiate a new election. The Star protocol ensures that there is only one leader at any time.

[0035] The follower is used to passively receive and process messages from the leader. When the leader's heartbeat times out, it will actively stand up and recommend itself as a candidate.

[0036] The candidate is used to elect a new leader. By sending vote RPC messages to other nodes to notify them to vote, it can upgrade to a leader by winning the majority of votes.

[0037] Compared with the prior art, the present invention has the following advantages and technical effects:

[0038] (1) Improved test efficiency: Traditional numerical tank tests require a large amount of time and human resources, while the virtual test management system based on machine learning can automate the test process, greatly improving the test efficiency.

[0039] (2) Improved test accuracy: Virtual tests can simulate numerical tank tests under different parameter combinations, providing more accurate test data, thereby improving the test accuracy.

[0040] (3) Reduced test cost: Traditional numerical tank tests require the establishment of actual test equipment and human resources, with high costs. While virtual tests do not require actual equipment and human resources, greatly reducing the test cost.

[0041] (4) Improved data management efficiency: The virtual test management system based on machine learning can automatically store test data, facilitating management and analysis. In addition, the distributed storage system can provide more efficient and reliable data storage services.

[0042] (5) Fault tolerance processing: If abnormal situations such as system downtime or data corruption occur, the data and logs can be retrieved from other nodes for recovery and replay through the fault tolerance mechanism of the Star protocol. At the same time, this system also supports multi-copy backup to achieve fault tolerance and high availability through data replication on multiple nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0044] Figure 1 is a schematic diagram of the system structure of an embodiment of the present invention;

[0045] Figure 2 is a system structure diagram of the intermediate storage module of an embodiment of the present invention;

[0046] Figure 3 is a structure diagram of the Star protocol module of an embodiment of the present invention;

[0047] Figure 4 is a flowchart of the execution of the storage function of an embodiment of the present invention;

[0048] Figure 5 is a flowchart of the execution of the machine learning module of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0050] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0051] As Figure 1 shown, a numerical tank virtual test management system based on machine learning provided by the present invention mainly consists of three modules: a data collection module, a central storage module, and an automated storage module.

[0052] (1) Data collection module

[0053] In the data collection module, it mainly includes an intelligent agent module, a data collection module, and a system performance module. As a system for storing data, the first thing to decide is the data storage model, that is, in what form the data is to be saved. The choice of this system is the key-value model, and an ordered traversal method is provided. Two key points of data storage: First, this is a huge Map, that is, key-value pairs are stored; Second, the key-value pairs in this Map are ordered according to the binary order of the keys, that is, the position of a certain key can be found, and then the successive traversal method is continuously called to obtain the key-values larger than this key in ascending order. The system can collect various stored data and fragment them, and these data will be transmitted to the central storage module for analysis and processing, making the system more scalable.

[0054] The scalability of the system refers to an expansion ability provided by the system to cope with the changes in the storage requirements of future digital pool virtual tests. When new storage requirements appear, the distributed system can support them with little or no change, without the need to reconstruct or rebuild the entire system. The intelligent agent of the data collection module will automatically identify the categories of data in various scenarios and pave the way for transmission to the central storage module. The function of the data collection module is to process the raw data in the digital virtual pool test. Correspondingly, the data acquisition module needs to comprehensively process the data in various scenarios; by randomly generating simulated data, it simulates the water flow movement, water level changes, waves and other hydrodynamic phenomena in the actual hydrological environment and other scenarios of the data to be processed, and then it can efficiently and flexibly test whether the data acquisition of the digital pool test system can correctly, completely and stably process the data to be processed in different scenarios, that is, the scalability, in a targeted manner.

[0055] (2) Central storage module

[0056] The central storage module mainly includes a machine learning module and a bottom-layer storage module MicaDB. The machine learning module is mainly composed of five modules, including a communication interface, a data preprocessing module, a model training module, a prediction module, and a storage node selection module. The central storage module is located between the data collection module and the automated storage module. After receiving the storage task sent by the data collection module, it processes the data and stores it in the bottom layer.

[0057] Data preprocessing is a very important step in the machine learning process, usually including multiple links such as data cleaning, feature selection, feature extraction, and feature transformation. The purpose of data preprocessing is to process the original data so that the data can be better applied to the training of machine learning models.

[0058] The model training module uses historical data to train a machine learning model so that future storage requirements can be predicted. Model training is a core step of the machine learning module. Its role is to adjust the model parameters by inputting a labeled dataset, enabling the model to better fit the data, thereby achieving prediction or classification tasks.

[0059] The prediction module can predict future storage requirements based on the prediction analysis results. After the model is loaded, the prediction module can receive input data, input it into the model for prediction, and obtain the prediction results.

[0060] After passing through the machine learning module, the data is sent to MicaDB for storage. For any persistent storage engine, the data ultimately has to be stored on disk, and this system is no exception. However, in this embodiment, instead of directly writing data to the disk, the data is stored in MicaDB, and the specific data landing is the responsibility of MicaDB. Here, MicaDB can be simply regarded as a stand-alone key-value Map.

[0061] (3) Automated storage module

[0062] In the automated management module, the system can automatically adjust the storage resource configuration according to the prediction analysis results, such as automatically increasing the storage capacity, adjusting the data distribution strategy, etc. This module is also responsible for one of the core issues of the distributed storage system: maintaining data consistency of multiple replicas.

[0063] To maintain consistency, this system adopts the Star consistency protocol. The algorithm is divided into several parts, including: leader election, log replication, and security measures. Leader election requires that a new leader must be elected when the existing leader fails; log replication means that when the leader receives a command from the client, it records it as a log and replicates it to other servers in the cluster, and forces the logs of other nodes to be consistent with the leader. Security measures: Ensure the security of the system through some measures, such as measures to ensure that all state machines execute the same commands in the same order.

[0064] In Star, nodes have the following three states: (1) Leader: Receives all requests from clients. The normal work of the leader includes three parts: processing write requests, managing log replication, and continuously sending heartbeat messages to notify other nodes that "I am the leader, I am still alive, and you should not initiate a new election now." The Star protocol ensures that there is only one leader at any given time. (2) Follower: Passively receives and processes messages from the leader. When the leader's heartbeat times out, it will actively stand up and recommend itself as a candidate. (3) Candidate: Used to elect a new leader. The candidate sends a vote RPC message to other nodes, notifying other nodes to vote. If it wins a majority of the votes, it will be promoted to a leader.

[0065] Embodiment 1

[0066] Figure 1 In this system, a key-value database based on machine learning distributed automatic storage is provided. It provides a distributed transaction interface that meets ACID constraints, and ensures multi-copy data consistency and high availability through the Star protocol. MicaDB, as the underlying storage layer of the system, provides persistence and read-write services for data written by users to the system, and also stores relevant statistical information data. Data can be automatically stored in the MicaDB underlying storage through the machine learning module.

[0067] As mentioned above, this system can be regarded as a huge ordered key-value table. In order to achieve horizontal expansion of storage, data will be distributed on multiple machines. For a key-value system, there are two typical solutions for distributing data on multiple machines: (1) Hashing: Hash according to the key and select the corresponding storage node according to the hash value. (2) Blocking: Block according to the key, a certain continuous key is stored on a storage node. This embodiment selects the second method in the data collection module, dividing the entire key value space into many segments, each of which is a series of continuous keys. Each segment is called a block area, which can be described by a left-closed and right-open interval such as [start key, end key). The amount of data stored in each block area is maintained at about 96 MiB by default (can be modified through configuration). Data is then exchanged with the central storage module through the communication module.

[0068] In the machine learning module, this embodiment first performs data collection and processing, which involves collecting data from various data sources, including real-time sensor data in the numerical water pool, historical data, meteorological data, hydrological data, etc. These data need to be preprocessed and cleaned to ensure the quality and integrity of the data and to prepare for subsequent analysis.

[0069] The data is then fed into the feature engineering module, a process that involves converting raw data into more meaningful features to improve the accuracy and performance of the machine learning model. For example, time series data can be converted into frequency domain features to improve the accuracy of time series forecasting.

[0070] After extracting the features of the data, this embodiment performs model selection and training, which involves selecting an appropriate machine learning algorithm and training the model using feature-engineered data. For example, a supervised learning-based regression model can be used to predict water level changes. The process may also involve hyperparameter adjustment and model evaluation to ensure the generalization performance of the model.

[0071] Finally, in this embodiment, model deployment and inference are performed. This process involves deploying the trained model to an actual numerical water tank and using the model for real-time inference and prediction. For example, the model can be used to predict future water level changes and make corresponding adjustments based on the prediction results.

[0072] Meanwhile, the model designed in this embodiment also needs to be monitored and maintained. This process involves monitoring the performance and health status of the model and promptly detecting and handling any abnormal situations. For example, if the prediction performance of the model deteriorates, it may be necessary to retrain the model or reselect the algorithm.

[0073] Different from the traditional full-node backup method, MicaDB, this storage system designs a replica mechanism. The data is divided into roughly equal slices (collectively referred to as block areas below) according to the key range. Each slice will have multiple replicas (usually 3), and these replicas are kept consistent through the Star protocol.

[0074] As Figure 2 shown, in this embodiment, multiple storage nodes are deployed in the machine learning module, and intelligent agents are installed on each storage node. The intelligent agent can automatically identify and handle abnormal situations of the storage node and data and report them to the central management node. In addition, the intelligent agent is also responsible for summarizing and reporting the usage information of the storage node and data to the central management node.

[0075] Deploy the central management node and install the machine learning engine and controller in it. The machine learning engine can use historical data and existing data for training, thereby automatically analyzing the usage patterns of storage nodes and data and predicting future usage situations. The controller can formulate storage policies based on the prediction results of the machine learning engine and issue corresponding commands to the storage node intelligent agent for execution.

[0076] The communication between the storage node intelligent agent and the central management node is carried out in a secure encryption manner. The encryption method can ensure the security and reliability of data during transmission. At the same time, the storage node intelligent agent also needs to perform identity verification to ensure that only authorized intelligent agents can communicate with the central management node.

[0077] The machine learning engine predicts future usage patterns by automatically analyzing the usage patterns of storage nodes and data. The machine learning engine can analyze based on multiple factors such as data volume, data type, access pattern, and access frequency to predict future usage situations.

[0078] The controller formulates a storage policy based on the prediction results of the machine learning engine and issues corresponding commands to the storage node intelligent agent for execution. The storage policy can include operations such as data migration, backup, and recovery. For example, when the machine learning engine predicts that the future usage rate of a certain storage node will increase sharply, the controller can formulate a corresponding data migration policy to migrate some data to an idle storage node, thus avoiding the situation of storage node overload.

[0079] The intelligent agent can automatically identify and handle abnormal situations of storage nodes and data, and report to the central management node. For example, when a hardware failure or data corruption occurs in a certain storage node, the intelligent agent can promptly report to the central management node and perform data recovery or re-allocate storage according to the instructions of the controller.

[0080] As Figure 3 shown in the operation process of maintaining data consistency of replicas in the Star protocol. In the Star protocol, nodes have the following three states: (1) Leader: Receives all requests from clients. The normal work of the leader includes three parts: processing write requests, managing log replication, and continuously sending heartbeat messages to notify other nodes that "I am the leader, I am still alive, and you should not initiate a new election now". The Star protocol ensures that there is only one leader at any time. (2) Follower: Passively receives and processes messages from the leader. When the leader's heartbeat times out, it will actively stand up and recommend itself as a candidate. (3) Candidate: Used to elect a new leader. The candidate sends a vote RPC message to other nodes, notifying them to vote, and if it wins a majority of the votes, it will be upgraded to a leader.

[0081] Next, it will be explained that in the three states, the follower and the candidate are passive roles, and the leader is an active role. The role of the node may change over time and network conditions. The following is a detailed description of the node role changes:

[0082] In the initial state, all nodes are set to the follower state.

[0083] Follower -> Candidate: If a follower does not receive a heartbeat message from the leader within a certain period of time, it will consider the leader to have failed and become a candidate. The candidate sends a message requesting votes to other nodes, asking them to vote to support itself as the new leader.

[0084] Candidate -> Leader: If the candidate receives votes from a majority of nodes in the cluster, it will become the new leader. The leader will send heartbeat messages to other nodes to stay synchronized and start processing requests from clients.

[0085] Leader -> Follower: If the leader finds that it can no longer act as the leader, for example, due to network partitioning or downtime, it will revert to a follower. The leader will send a demotion message to notify other nodes that it is no longer the leader and start following the new leader.

[0086] Candidate -> Follower: If the candidate does not receive votes from a majority of nodes in the cluster, it will revert to a follower.

[0087] In the Star protocol, the role changes of nodes are achieved through communication protocols and election algorithms. Through these mechanisms, nodes can switch to different roles at different times to ensure the correctness and fault tolerance of the cluster. The role changes of nodes usually occur automatically without manual intervention.

[0088] As Figure 4 shown in the overall flowchart, it is mainly the user registration and login process: Users need to register and log in to the system to use the system functions. This process includes steps such as user input of information, verification of user information, creation of user accounts, and generation of login tokens. Trial management process: This process includes steps such as creating new trials, viewing existing trials, modifying trial parameters, running trials, and deleting trials. Data preprocessing process: This process includes steps such as collecting data from sensors or other data sources, preprocessing and filtering the data, and storing the data in a distributed storage system. Data storage and query process: This process includes steps such as storing data in a distributed storage system, querying and retrieving data from the storage system, and aggregating and processing the data. Machine learning process: This process includes steps such as obtaining data from the storage system, preprocessing the data, selecting appropriate machine learning algorithms, training machine learning models, using the models for prediction and inference. Visualization display process: This process includes steps such as obtaining data from the storage system, processing the data, generating charts and visualization tools, and presenting the visualization results to users.

[0089] As Figure 5 shown in the flowchart of the machine learning module execution. First is the data preprocessing stage: In this stage, operations such as data cleaning, transformation, and feature extraction need to be performed on the data to better prepare for model training. Data preprocessing includes the following steps: a. Data cleaning: Check whether there are invalid values, missing values, or outliers in the data and remove them from the dataset or fill in appropriate values. b. Data transformation: Convert the data into a format that the model can handle, such as converting categorical variables to numerical encodings, discretizing continuous variables, performing normalization or standardization, etc. c. Feature extraction: Extract features from the data to help the model learn. This may involve using techniques such as statistical analysis, natural language processing, and image processing to extract useful information from the data.

[0090] The second stage of model training: In this stage, a model suitable for the dataset and the problem needs to be selected, the dataset is used to train the model, and the model parameters are adjusted to obtain better performance. The model training includes the following steps: a. Model selection: Select a model suitable for the problem and the dataset, such as linear regression, decision tree, neural network, etc. b. Dataset division: Divide the dataset into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to adjust the model parameters, and the test set is used to evaluate the model performance. c. Model training: Use the training set to train the model and use the validation set to adjust the model parameters to obtain better performance. d. Model saving: Save the trained model for subsequent prediction and inference.

[0091] The third stage of model evaluation: In this stage, the performance of the model needs to be evaluated to determine whether the model can solve the problem well. The model evaluation includes the following steps: a. Model prediction: Use the test set to make model predictions and compare the prediction results with the true results. b. Model performance evaluation: Use various evaluation metrics (such as accuracy, recall rate, F1 value, etc.) to evaluate the performance of the model and compare the performance between different models. c. Model improvement: According to the model performance evaluation results, adjust the model parameters and retrain the model.

[0092] The above is only a preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the technical field of the present application within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A numerical wave tank virtual test management system based on machine learning, characterized in that Comprising: A data collection module, configured to obtain initial data in various scenarios of a digital virtual pool test and randomly generate simulation data; the simulation data being data to be processed for various scenarios of hydrodynamic phenomena including water flow movement, water level change, and waves in an actual hydrological environment; A central storage module, connected to the data collection module, configured to process the simulation data and perform underlying storage on the processed simulation data when receiving a storage task sent by the data collection module; An automated storage module, connected to the central storage module, configured to automatically adjust storage resource configuration according to a prediction analysis result and maintain data consistency of multiple copies; The data collection module includes a data transmission unit, an intelligent agent unit, a data collection unit, a mode conversion unit, a data chunking unit, and a communication unit connected in sequence; The data transmission unit is configured to transmit initial data in various scenarios of a digital virtual pool test to the intelligent agent unit, and is further configured to transmit the chunked simulation data to the central storage module through the communication unit; The intelligent agent unit is configured to automatically identify data categories in various scenarios; the data collection unit is configured to collect initial data in various scenarios; The mode conversion unit is configured to determine a data storage model according to different data categories and provide an ordered traversal method; The data chunking unit is configured to chunk according to keys, divide the entire key-value space into many segments, each segment being a series of continuous keys, and a certain segment of continuous keys is stored on a storage node, and each segment is divided into a chunk area, which is described by a left-closed and right-open interval [start key, end key); The central storage module includes a machine learning unit, a data distribution unit, and an underlying storage unit MicaDB; The machine learning unit is configured to predict and optimize the operation state and performance of the pool according to historical data and real-time data; The data distribution unit is configured to distribute data storage tasks among multiple computer nodes through a distributed storage technology and integrate and manage data of multiple pools; The underlying storage unit MicaDB is configured to perform underlying data storage on the chunk area through the Star protocol.

2. The machine learning-based numerical pool virtual test management system according to claim 1, wherein The machine learning unit includes a communication interface unit, a data preprocessing unit, a model training unit, and a prediction unit; The communication interface unit is connected to the communication unit and is configured to receive the chunked simulation data; The data preprocessing unit is configured to perform data cleaning, feature selection, feature extraction, and feature transformation on the chunked simulation data; The model training unit is configured to train a machine learning model with the data processed by the data preprocessing unit to obtain a target prediction model; The prediction unit is configured to input real-time data into the target prediction model to perform prediction analysis on future storage requirements and obtain a prediction result.

3. The machine learning-based numerical pool virtual test management system according to claim 1, wherein The central storage module includes storage nodes and a central management node. Among them, the storage nodes are deployed in the machine learning unit, and an intelligent agent is installed on each storage node. The intelligent agent automatically identifies and processes anomalies of the storage nodes and data, and reports them to the central management node; the intelligent agent is also used to summarize and report the usage information of the storage nodes and data to the central management node; The central management node is used to automatically analyze the usage patterns of storage nodes and data to discover and predict possible faults, and automatically start the corresponding fault handling program when a fault occurs.

4. The virtual test management system of a numerical wave tank based on machine learning according to claim 3, wherein, The central management node includes a machine learning engine and a controller; The machine learning engine is used to train using historical data and existing data, so as to automatically analyze the usage patterns of storage nodes and data and predict future usage; The controller is used to formulate a storage strategy according to the prediction results of the machine learning engine, and issue corresponding commands to the intelligent agent of the storage node for execution.

5. The virtual test management system of a numerical wave tank based on machine learning according to claim 3, wherein, The storage nodes can dynamically join or leave the system, and the communication between the intelligent agent of the storage node and the central management node is carried out in a secure encryption manner.

6. The virtual test management system of a numerical wave tank based on machine learning according to claim 3, wherein, The automated storage module adopts the Star protocol, and the Star protocol includes leader selection, log replication, and security measures; The leader selection requires that if the existing leader fails, a new leader must be elected; The log replication means that when the leader receives a command from the client, it records it as a log and replicates it to other servers in the cluster, and forces the logs of other nodes to be consistent with the leader; The security measures are used to ensure the security of the system through some measures.

7. The virtual test management system of a numerical wave tank based on machine learning according to claim 6, wherein, The node states in the Star protocol include leader, follower, and candidate; The leader is used to receive all requests from the client; including processing write requests, managing log replication, and continuously sending heartbeat messages to notify other nodes not to initiate a new election. The Star protocol ensures that there is only one leader at any time; The follower is used to passively receive and process messages from the leader. When the leader's heartbeat times out, it will actively stand up and recommend itself as a candidate; The candidate is used to elect a new leader; by sending a vote RPC message to other nodes to notify them to vote, and upgrading to a leader if it wins the majority of votes.

Citation Information

Patent Citations

  • Numerical pool application characteristic performance acquisition and monitoring system and operation method thereof

    CN110990227A

  • KR20220141170A