Data-driven intelligent material design research and development platform system and method

By building a data-driven intelligent material design and R&D platform system, combining multi-scale simulation, AI models and high-throughput computing, the shortcomings of cloud computing platforms in resource scheduling and data management are solved, and efficient scientific research materials database management and scientific research efficiency are achieved.

CN120409307AInactive Publication Date: 2025-08-01TIANMUSHAN LABORATORY

Patent Information

Application Number
CN202510919983.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In terms of resource scheduling, existing cloud computing platforms have problems such as low resource utilization, slow response speed, difficult to guarantee data security and privacy, insufficient elastic scaling capabilities, and high operation and maintenance costs, which cannot meet the special needs of the field of materials science.

Method used

Build a data-driven intelligent material design and R&D platform system, including multi-scale simulation modules, high-throughput modules, AI model modules and material databases. Through multi-scale simulation operations, process engine-driven high-throughput calculations and AI model prediction, combined with resource scheduling modules and identity authentication modules, cross-scale simulation and efficient data management are achieved.

Benefits of technology

It improves resource utilization, improves the accumulation and management efficiency of scientific research material databases, simplifies complex business processes, reduces operation and maintenance costs, ensures data security and privacy, and enhances the data processing and management capabilities of scientific researchers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409307A_ABST
    Figure CN120409307A_ABST
Patent Text Reader

Abstract

The invention discloses a data-driven intelligent material design research and development platform system and method, and the system comprises a multi-scale simulation module which comprises nano-scale simulation software, micro-scale simulation software, mesoscale simulation software and macro-scale simulation software, matches the simulation software based on material data inputted by a user, configures simulation information, and carries out simulation operation through the simulation software obtained through matching; the high-throughput module is used for performing node operation based on all nodes in a preset process template corresponding to the task file to obtain a task operation result; the AI model module comprises an AI model deployed on nodes in a containerization manner, generates input parameters of a current task based on a parameter template of the AI model, performs model prediction, and is used for online deployment of the AI model of a user; the material database stores material data and is used for a user to inquire the material data online and display the material data in a templated mode. Aiming at special requirements in the field of material science, the limitation of an existing cloud computing platform is overcome, and the resource utilization rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of information technology, and particularly relates to a data-driven intelligent design and R & D platform system and method for materials. Background Art

[0002] With the rapid development of information technology, cloud computing, as a computing resource sharing mode of on-demand allocation, has become an important support for informatization construction and digital transformation. Most existing cloud computing platforms provide basic services such as virtualized resources, data storage, and application deployment, and have certain resource scheduling and management capabilities. However, in actual applications, there are still several problems with existing cloud computing platforms, which limit the improvement of their performance and efficiency. Summary of the Invention

[0003] The technical problem to be solved by this application is to overcome the limitations of existing cloud computing platforms and improve resource utilization according to the special needs of the materials science field. On the one hand, this application provides a data-driven intelligent design and R & D platform system for materials. On the other hand, this application provides a data-driven intelligent design and R & D method for materials. On the third hand, this application provides a computer-readable storage medium. On the fourth hand, this application provides a computer device.

[0004] To solve the above technical problems, this application discloses a data-driven intelligent design and R & D platform system for materials, including a multi-scale simulation module, a high-throughput module, an AI model module, and a materials database; The multi-scale simulation module includes simulation software at the nano-scale, micro-scale, meso-scale, and macro-scale. Based on the material data input by the user, the simulation software is matched, and after configuring the simulation information, the simulation operation is performed through the matched simulation software; The high-throughput module is used to perform node operations on all nodes in the preset process template corresponding to the task file to obtain the task operation result; The AI model module includes an AI model containerized and deployed on the node, generates input parameters for the current task based on the parameter template of the AI model, performs model prediction, and is used for the online deployment of the user's AI model; The materials database stores material data, and is used for users to query material data online and display material data in a templated manner.

[0005] Optionally, it further includes a resource scheduling module for creating user supercomputer accounts, allocating resource storage space, allocating computing power, and counting computing power consumption, monitoring the usage of user computing power and space, automatically freezing and thawing accounts. When the user's computing power is insufficient, the supercomputer account will be automatically downgraded or frozen. After the background adds computing power again, the account will be automatically promoted or thawed. When the user's storage space is insufficient, the account will be automatically frozen and prohibited from use. When the administrator adds storage allocation or clears the space storage independently in the background, the system will automatically unfreeze the account again.

[0006] Optionally, the high-throughput module is used to decompose the task file into the smallest unit processes corresponding to the nodes according to the preset process template, and execute all the nodes corresponding to the task file in a directed graph traversal manner.

[0007] Optionally, the processes of the preset process template include ordinary processes, judgment processes, loop processes, and repeated parallel processes.

[0008] Optionally, it further includes an identity authentication module for verifying the user's identity through the user information verification method each time the user logs in to use the application program of the module. If the verification is passed, a security token will be generated. When the user uses other application programs, if there is a security token and the verification is passed, the user is allowed to use the corresponding application program.

[0009] Optionally, the user's computing power includes shared resources and exclusive resources. When using computing power resources, the exclusive resources are preferentially used. If the exclusive resources are exhausted, the shared resources are used. If the occupancy of the shared resources reaches the preset ratio, wait for the exclusive resources.

[0010] Optionally, it further includes a task configuration module for generating the task file based on the user input parameters and the preset script template.

[0011] This application also discloses a data-driven intelligent design and research and development method for materials, including: Matching the simulation software of the multi-scale simulation module based on the material data input by the user, configuring the simulation information, and performing simulation operations through the obtained simulation software. The multi-scale simulation module includes simulation software at the nano-scale, micro-scale, meso-scale, and macro-scale; Performing node operations on all nodes in the preset process template corresponding to the task file through the high-throughput module to obtain the task operation result. The AI model of the AI model module is containerized deployed on the node, generating the input parameters of the current task based on the parameter template of the AI model and performing model prediction, and the user can deploy the AI model online through the AI model module; Online querying material data through the material database and template-displaying the material data stored in the material database.

[0012] The present application also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the data-driven method for the intelligent design and R & D platform of materials as described above.

[0013] The present application also discloses a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the data-driven method for the intelligent design and R & D platform of materials as described above.

[0014] Beneficial effects: The present application constructs a data-driven intelligent design and R & D platform system for materials, which includes a multi-scale simulation module, a high-throughput module, an AI model module, and a materials database. The present application provides cross-scale material simulation operations through the multi-scale simulation module, and performs simulation operations by automatically matching and configuring simulation information, improving the simulation operation efficiency of users. The high-throughput module of the present application is used to perform node operations on all nodes in the preset process template corresponding to the task file to obtain the task operation result, and builds a high-throughput process templated calculation based on the process engine. The process engine is a tool used to drive the business to flow according to the set fixed process. In complex and changeable business situations, using the established process can greatly reduce the cost of design business and ensure the accuracy of business execution. In summary, the platform system of the present application will realize the accumulation and management of the scientific research materials database, expand the platform in the digital R & D direction through high-throughput, AI models, multi-scale simulation software, and integrated private supercomputer clusters. The platform aims to solve the data processing, management, and efficiency of scientific research personnel, simplify the original repetitive and cumbersome work process into a process template, and empower through the platform and AI, greatly improving scientific research efficiency. Description of the Drawings

[0015] The following further specifically describes the present application in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present application will become clearer.

[0016] Figure 1 It is a structural diagram of an embodiment of the data-driven intelligent design and R & D platform system for materials of the present application; Figure 2 It is a schematic diagram of the materials database of an embodiment of the data-driven intelligent design and R & D platform system for materials of the present application; Figure 3 It is a flowchart of an embodiment of the data-driven intelligent design and R & D method for materials of the present application; Figure 4 It shows a schematic structural diagram of a computer device suitable for implementing the embodiment of the present invention. Detailed Embodiments

[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0018] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0019] In the prior art, in terms of resource scheduling in the cloud computing platform, static or simple dynamic scheduling strategies are usually adopted, which cannot adjust resource allocation in real time according to changes in business load, resulting in problems such as low resource utilization and slow response speed. In a cloud computing environment, multiple tenants share the same physical resources. How to effectively isolate tenants and ensure data security and privacy is another challenge faced by the prior art. With the continuous change of business requirements, the cloud computing platform needs to have flexible elastic scaling capabilities to cope with sudden business peaks. However, existing platforms often have problems such as response latency and uneven resource allocation in terms of elastic scaling. The operation and maintenance costs and management complexity of the cloud computing platform increase with the expansion of scale. How to reduce operation and maintenance costs and simplify management processes has become an urgent problem to be solved in the development of the cloud computing platform.

[0020] To solve at least one of the problems existing in the prior art, according to one aspect of the present application, an embodiment discloses a data-driven intelligent design and R & D platform system for materials. As Figure 1 shown, the system includes a multi-scale simulation module 11, a high-throughput module 12, an AI model module 13, and a materials database 14.

[0021] The multi-scale simulation module 11 includes simulation software at the nano-scale, micro-scale, meso-scale, and macro-scale. Based on the material data input by the user, the simulation software is matched, and after configuring the simulation information, the simulation operation is performed through the obtained simulation software.

[0022] Among them, multiscale simulation is a computer simulation technology that covers the simulation process of coupling and correlation at different scales from the atomic (molecular) level to the macroscopic scale. This technology aims to establish a multiscale simulation model through mathematical and physical descriptions to study and predict the behaviors and characteristics of complex systems. The application fields of multiscale simulation are extensive, including but not limited to material science, semiconductor device simulation, and composite material forming process simulation, etc. By comprehensively considering factors at different scales in the system, it aims to understand and predict phenomena and behaviors across multiple spatial, temporal, and organizational levels, and solve practical problems by combining microscopic details with macroscopic phenomena. The development history and current situation of multiscale modeling and simulation technology, the classification and characteristics of common methods, the construction of key technologies and challenges, as well as its future development trends have jointly promoted the application of multiscale technology in various fields. This application is based on a supercomputer cluster environment to build a multi-element, multi-state, and multiscale computational simulation software. It is built online in the form of web and shell, eliminating the cumbersome software building and usage steps for users, refining the simplest usage links, and using the web online form to carry the simplest usage methods. It has functions such as basic network communication, control strategy configuration, and human-machine interface software platform, etc.

[0023] The user inputs material data such as material structures or files. The multiscale simulation module 11 selects a simulation software according to the material data, configures and adjusts simulation information such as environmental parameters and final operation instructions to execute the operation process of the simulation software. If it is necessary to adjust the job name, it can also be directly modified and adjusted. It is also possible to confirm the supercomputer, select the nodes to be used, and the number of computing cores, etc. for simulation calculations.

[0024] The high-throughput module 12 is used to perform node operations on all nodes in the preset process template corresponding to the task file to obtain the task operation result. Among them, the task operation result can be the performance analysis result of material design or the ratio of material design, etc., which promotes the efficiency and accuracy of material design through intelligent analysis.

[0025] The AI model module 13 includes an AI model containerized and deployed on the nodes, generates input parameters for the current task based on the parameter template of the AI model and conducts model prediction, and is used for the online deployment of the user's AI model. The AI model can adopt a deep learning model trained with a large amount of text data, and can generate natural language text or understand the meaning of language text. The large language model can handle various natural language tasks, such as text classification, question answering, dialogue, etc., and is an important way to artificial intelligence. Currently, the large language model adopts a similar Transformer architecture and pre-training objectives (such as LanguageModeling) as the small model. The difference from the small model is to increase the model size, training data, and computing resources.

[0026] The material database 14 stores material data for users to query material data online and display material data in a templated manner.

[0027] This application constructs a data-driven intelligent design and R & D platform system for materials, which includes a multi-scale simulation module 11, a high-throughput module 12, an AI model module 13, and a material database 14. This application provides cross-scale material simulation operations through the multi-scale simulation module 11, and performs simulation operations by automatically matching and configuring simulation information, improving the simulation operation efficiency of users. The high-throughput module 12 of this application is used to perform node operations on all nodes in the preset process template corresponding to the task file to obtain the task operation result, and build a high-throughput process templated calculation based on the process engine. The process engine is a tool used to drive the business to flow according to the set fixed process. In complex and changeable business situations, using the established process can greatly reduce the cost of design business and ensure the accuracy of business execution. In summary, the platform system of this application will realize the accumulation and management of the scientific research material database 14, and expand the platform in the direction of digital R & D through high-throughput, AI models, multi-scale simulation software, and an integrated private supercomputing cluster. The platform aims to solve the data processing, management, and efficiency of scientific research personnel, simplify the original repetitive and cumbersome work process into a process template, and empower through the platform and AI, greatly improving scientific research efficiency.

[0028] This application can be implemented through the IDM (Inter-scale Data Mat Explorer Lab) platform. The IDM platform realizes the acquisition, storage, and sharing of databases in the material R & D process, as well as the controllability of cross-scale simulation and high-throughput calculation through information means. The material database 14 mainly realizes the acquisition of online public material data, the acquisition and storage of the calculation result data of the IDM platform, and the maintenance of its own private data set. The cross-scale simulation calculation realizes the use of a series of online simulation software, and the visualization of high-throughput calculation realizes the controllability of calculation, which is mainly reflected in that the calculation process is driven by the process engine, and the relevant work is disassembled into the most basic calculation units, which are uniformly called by the process engine and submitted to the computing power unit for processing.

[0029] The material database 14 is a database specifically used for storing and organizing various material property parameters. These databases systematically collect and organize the performance data of materials so that engineers, designers, and researchers can easily search for and compare the performance parameters of different materials. The content of the material database 14 is extensive, including but not limited to mechanical properties, thermal properties, electrical properties, optical properties, etc.

[0030] The integration function of the materials database 14 primarily involves logically or physically integrating data from diverse sources, formats, and characteristics, thereby providing comprehensive data sharing for enterprises. This integration not only includes standardized and intelligent data collection, inspection, consolidation, and storage, but also involves data quality control to ensure the quality of attribute and spatial databases. By establishing a powerful database system, it is possible to achieve three-dimensional data presentation, display results, and provide application services, providing a basis for decision support.

[0031] The integrated functionality of Materials Database 14 improves data efficiency and value by: Data sharing and integration platform: Through the development of data integration middleware and real-time database synchronization functions, we can achieve logical or physical centralization of data from different sources and support the comprehensive data sharing needs of enterprises.

[0032] Database design and implementation: formulate unified survey database standards and specifications, collect identified and potential data, standardize database construction content, database structure, database construction methods, results data quality inspection and results submission requirements, etc., to ensure the quality and consistency of the database.

[0033] Material properties database: such as the TotalMateria global material properties database, which includes metal and non-metallic material data of more than 74 countries and international standards, provides real material properties data display and retrieval method updates, supports multi-dimensional search conditions, and quickly locates target data.

[0034] Data integration verification and application: In the construction of the composite material processing technology database, the materialized and chemical integration method based on the middle layer is adopted for data integration. Verification is carried out from the aspects of mechanical processing system integration, typical process package integration, knowledge base integration, etc. to ensure the accuracy and availability of the data.

[0035] Functional Materials Crystal Facet Database: By developing new strategies that integrate crystal facet databases, machine learning, mathematical models, and ionic liquid experimental synthesis, we can achieve the rational design and controllable synthesis of specific crystal facets, accelerating the discovery and application of functional materials.

[0036] The material database 14 may include a public database and an IDM database. The public database integrates the online public database into the IDM platform through the background global dictionary configuration to facilitate all platform users to view. The IDM database can realize crystal structure retrieval. The retrieval bottom layer uses elasticsearch support. Figure 2As shown, it mainly provides four retrieval methods: crystal structure elements, chemical formula, ID identifier, and full-text retrieval. Among them, the retrieval of ID is hp- plus the ID number. The full-text retrieval is the index matching mode of elasticsearch. On the left are some property retrieval and screening conditions of the crystal structure. Symmetry, band gap, formation energy, envelope energy, density, number of atoms, and calculated properties. The retrieval of these properties is generally directly through keyword rather than the index of elasticsearch. Below the periodic table is the retrieval result, including structure diagram, ID, density, crystal system, point group, space group, crystal volume, and lattice constant. The structure diagram is a static picture saved on the server.

[0037] The IDM database can also display the three-dimensional structure components of the crystal structure and show some basic information fields. It includes the display of 4 properties: energy band, density of states, cluster expansion, and X-ray diffraction pattern. Among them, the energy band, density of states, and cluster expansion can be docked with the high-throughput process for calculation through the process.

[0038] The material database 14 can also include a private database. Users can import their own material data sets. The system will automatically generate an online database according to the data imported by the users. Users can add attributes such as name, classification, and associated items to this data, and can also share it with other users so that the shared users can also view, copy, and download this data set for their own use and reference.

[0039] Thus, the data set can include "My Data Set" and "Shared with Me". My Data Set: What is displayed is the data set imported or copied by the user himself; Shared with Me: What is displayed is the data set shared by other users to oneself.

[0040] The private database preferably supports at least one of the following functions: Create data set: You can import data set files by yourself or copy from existing data sets; Create classification: Create classifications with different names to facilitate the management of data set classifications; Manage classification: You can manage the created classifications; View data set: View the details of the current data set; Edit data set: Edit the current data set; Share data set: Share the current data set with other users; Adjust classification: Adjust the classification to which the current data set belongs; Export to CSV: Export the current data set in the form of a csv file; Copy data set: Directly copy the current data set as a new data set; Delete dataset: Directly delete the current dataset, and it cannot be recovered after deletion.

[0041] In an optional embodiment, the system further includes a resource scheduling module. The resource scheduling module is used for user supercomputer account creation, resource storage space allocation, computing power allocation, and computing power consumption statistics. It monitors the usage of user computing power and space, and automatically freezes and unfreezes accounts. When the user's computing power is insufficient, the supercomputer account will be automatically downgraded or frozen. After the background adds computing power again, the account will be automatically upgraded or unfrozen. When the user's storage space is insufficient, the account will be automatically frozen and prohibited from use. The administrator can add storage allocation in the background or clean up the space storage independently, and the system will automatically unfreeze the account again.

[0042] Among them, the resource scheduling module realizes flexible scheduling of resources by configuring and adjusting computing power and resources for users. For valid users who have passed the review, the supercomputer accounts of the users can be opened, the storage space and computing power can be allocated, and the supercomputer capacity quota of the users can be adjusted. For example, to enable supercomputing: the supercomputer account of this user can be enabled; to configure resources: storage space and supercomputing quota can be allocated to this user's supercomputer account; to allocate computing power: the available supercomputing power can be allocated to this account. In a specific example, the resource scheduling module can implement the creation of supercomputer accounts for each user, resource storage space allocation, computing power allocation, job pulling computing power consumption statistics, and writing of account software UAS, etc. And monitor the usage of user computing power and space, automatically freeze and unfreeze accounts. When the user's computing power is insufficient, the supercomputer account will be automatically downgraded or frozen. After the background adds computing power again, the account will be automatically upgraded or unfrozen. When the user's storage space is insufficient, the account will be automatically frozen and prohibited from use. The administrator can be contacted to add storage allocation in the background or clean up the space storage independently, and the system will automatically unfreeze the account again. Automatically pull the simulation jobs submitted by the supercomputer for management on the system platform.

[0043] Optionally, the functions of this module, except for creating supercomputer accounts, allocating resources and computing power which belong to "User Management - Resource Allocation Management" in the background management, all others belong to system automation tasks.

[0044] The following is the design of supercomputer automation tasks: Software account authorization UAS task: Execute once per minute, write the newly opened supercomputer account into the software authorization file that requires a license for users to use.

[0045] Synchronize simulation job task: Execute once per minute, pull the jobs of users using cross-scale simulation software into the system database to facilitate users to manage and view the jobs by themselves.

[0046] Monitor user computing power consumption task: Execute once every 30 seconds. When the remaining computing power of the user is less than the monitoring threshold, the account will be downgraded or frozen. When the computing power reallocated to the user is greater than the threshold, it will be automatically upgraded or unfrozen.

[0047] Disable and enable account task: Execute once every 10 seconds. When the administrator disables and enables the user account in the background, the system will automatically freeze and unfreeze the corresponding supercomputer account.

[0048] Synchronize user storage task: Execute once every 100 seconds. Synchronize the entire storage usage of the user's supercomputer account back to the system to facilitate the system's monitoring of user storage usage.

[0049] In an alternative embodiment, the high-throughput module 12 is used to decompose the task file into the smallest unit processes corresponding to the nodes according to a preset process template, and execute all the nodes corresponding to the task file in a directed graph traversal manner.

[0050] The high-throughput module 12 runs supported by a process engine. The design of the start and end of the process is mainly to determine the scope of loops and repeated parallelism. Supporting loops and repeated parallelism is the core function of the process engine. Repeated parallelism allows the process engine to support unified processing of multiple material structure inputs or other inputs, which reflects high-throughput. And loops allow the process engine to support repeated calculation of a certain process, solving the computational problems that need to be continuously optimized on supercomputers. A directed graph itself does not have a loop structure and cannot have a loop structure, so a set of symmetric start and end is used to handle the scope problem.

[0051] The main data structure of the process is a directed graph Digraph. The directed graph should include the following methods: add node, add connection, query the connections with the incoming direction of the connected nodes, query the connection lines with the outgoing direction of the connected nodes, query the previous nodes, query the next nodes, query the node information, query the number of nodes, query the number of connection lines, query the in-degree of the nodes, query the out-degree of the nodes, query the in-degree of the judgment nodes, use the depth-first traversal algorithm, query the start node, query the end node, query all nodes, query all edges.

[0052] The start and end of the smallest group can be regarded as a process, and the process of the smallest unit can be regarded as a node. The process of transforming a process into a node is called subprocess reuse. The input and local attributes of the node are taken from the start node, and the output attributes are taken from the end node. Therefore, both the start node and the end node of the subprocess are value references. The nodes inside the subprocess are all address references. If the nodes inside the subprocess are edited, all processes will be updated synchronously, which is different from the reuse of task templates. At the same time, in data persistence, we will put the process and the subprocess in the same collection because they are essentially the same thing, except for some differences in details. Inside the collection, I distinguish between subprocesses and processes through the field type.

[0053] The execution of a process is actually the traversal of a directed graph, starting from the head node and traversing to the tail node. However, several special cases need to be handled. First, for any node, we can find the nearest start and end nodes. Therefore, we can divide the processes into four categories: ordinary processes, judgment processes, loop processes, and repeated parallel processes. Then, the execution processes of ordinary processes, judgment processes, loop processes, and repeated parallel processes are provided. At the same time, some data belonging to the current process may need to be cached in memory during execution, so each process has its own context.

[0054] Explanation of the four basic types of processes: Ordinary process: Wrapped by an ordinary start node and an end node, it can be regarded as a simple directed graph and can be executed from the head node to the tail node.

[0055] Judgment process: Wrapped by a judgment start node and a judgment end node, it has a few branches that are not taken compared to an ordinary process. Therefore, a judgment process can be split into ordinary processes corresponding to the number of judgment conditions. At the judgment start, a condition judgment is made to select the ordinary process to be executed.

[0056] Loop process: Wrapped by a loop start and a loop end node. The characteristic of a loop process is that it is a ring structure, that is, the head and tail are connected. The programming method is while(condition){dosomething}, that is, the condition is judged first and then the execution is carried out. Therefore, we make a condition judgment in the executor of the loop start node. If the loop condition is met, loop processing is carried out. If the loop condition is not met, it jumps to the loop end node and continues to execute the directed graph.

[0057] Repeated parallel process: Wrapped by a repeated parallel start node and a repeated parallel end node. There is a parameter called parallelism, which determines the number of parallel task lines. Each task line can be regarded as an ordinary process. For example, if the input is a set of material structures, the input of each task line is the material with the task line number in the array. And the process of executing the repeated parallel process becomes the process of executing the ordinary process for the number of times of parallelism.

[0058] For sub-processes, they are also quite special. The concept of a return point is required because after the end node of the sub-process is executed, it should return to execute the next set of nodes of the sub-process. It is necessary to return from the end node to the sub-process node to continue executing the directed graph. Note that here we do not use the method of expanding all sub-processes and then executing a huge directed graph. On the one hand, because the expansion logic will be very complex and the connection lines need to be changed. On the other hand, the attributes of the sub-process nodes may also be referenced by the following nodes, and they need to be split into two nodes, the start and the end, and the access identifier changes from one to two. And the design we adopt now is to execute the process of the smallest unit. Therefore, the execution of the sub-process and the execution of the process have no difference for the executor. The sub-process needs to find the return point and thus has its own context.

[0059] For ordinary tasks, they are execution units that cannot be further split and can be directly executed. Their belonging context can be any one of the four processes.

[0060] A process is composed of ordinary start / end nodes. Therefore, the input attributes of the ordinary start node are the inputs required at the process level. When executing each time, we only need to assign the attributes without default values to start executing the process. The output attributes defined by the end node are the outputs at the process level and also the final result of the process. Of course, if the result has been stored in the middle, such as the warehousing logic, there may not be a need for an output. A process is like a method. It can have parameters and returns, or it can have no parameters and no results, depending on the designer.

[0061] In an alternative embodiment, the system further includes an identity authentication module. The identity authentication module is used to verify the user's identity by means of user information verification when the user logs in to use the application program of the module each time. If the verification is passed, a security token is generated. When the user uses other application programs, if there is a security token and the verification is passed, the user is allowed to use the corresponding application program.

[0062] Specifically, in a specific example, when a user accesses the first application, they need to authenticate and log in to the system. This process can use any conventional authentication method, such as username and password, two-factor authentication, etc. After successful authentication, the system generates a security token (Token), stores it in the user's browser, and simultaneously stores the token information in the SSO server. When the user accesses other applications, the application sends a token verification request to the SSO server. The SSO server checks the token information in the browser and confirms the user's identity. If the token is valid and the user has been authenticated, the SSO server returns an authorization token to the application, authorizing the user to access the application. The application uses the authorization token to verify the user's identity and allows the user to access the application's resources.

[0063] In addition, to ensure security, some measures need to be taken, such as using the HTTPS protocol for communication, encrypting the token, setting the expiration time of the token, etc.

[0064] In an alternative embodiment, the user computing power includes shared resources and exclusive resources. When using computing power resources, the exclusive resources are preferentially used. If the exclusive resources are exhausted, the shared resources are used. If the occupancy of the shared resources reaches a preset ratio, wait for the exclusive resources.

[0065] It can be understood that, in order to improve resource utilization, this application introduces shared resources and exclusive resources to achieve resource isolation. If multiple downstream services of a service are blocked simultaneously and a single downstream interface never reaches the circuit breaker standard (for example, the abnormal ratio and slow request ratio do not reach the threshold), then it will lead to a decrease in the throughput of the entire service and more thread occupancy. In extreme cases, it may even cause the thread pool to be exhausted. After introducing resource isolation, the maximum thread resources available to a single downstream interface can be restricted to ensure that the throughput of the entire service is affected as little as possible before the circuit breaker is triggered.

[0066] Since the traffic and RT of each interface are different, it is difficult to set a reasonable maximum available thread count, and it is also difficult to maintain this threshold as the business iterates. Here, sharing plus exclusivity can be used to solve this problem. Each interface has its own exclusive thread resources. When the exclusive resources are fully occupied, shared resources are used. When the shared pool reaches a certain water level, exclusive resources are forcibly used and queued for waiting. The advantage of this mechanism is obvious, that is, it can ensure isolation while maximizing resource utilization. The thread count here is just one type of resource, and resources can also be connection counts, memory, etc.

[0067] In an alternative embodiment, the system further includes a task configuration module for generating the task file based on user input parameters and a preset script template.

[0068] In this application, by pre - defining a script template, pre - defined commands, events, notification programs, and handlers for performing certain operations are set in advance, eliminating the need for users to write scripts in real - time. The setting and management of AI model scripts are also crucial for improving work efficiency. By reasonably setting parameters for generating scripts, such as defining project requirements and goals, setting code styles and specifications, specifying the scope and granularity of code generation, etc., efficient automated programming and task management can be achieved. This includes steps such as creating automated task flows, using generated scripts for task scheduling, monitoring, and optimizing task execution, thereby reducing development costs and enhancing work efficiency.

[0069] Preferably, modifiable parameters in the script template can be replaced with user - input parameters to generate a task file. In one or more embodiments, the preset script template can be used to call at least one of the multi - scale simulation module 11, high - throughput module 12, and AI model module 13, so as to circularize the process by linking the material database 14, AI model, high - throughput, and multi - scale simulation software through the script, and obtain accurate experimental instructions and characterization preparations through continuous iterative calculations.

[0070] When building the platform system, the platform back - end service is developed using java + go + python, the front - end uses the VUE framework, the network load uses nginx reverse proxy, and the single - node deployment uses single - machine docker for containerization. The system can be split. The purpose of splitting is not to reduce the unavailable time, but to reduce the impact area of failures. Because a large system is split into several small independent modules, a problem in one module will not affect other modules, thus reducing the impact area of failures. System splitting also includes access - layer splitting, service splitting, and database splitting. The access layer and service layer are split according to dimensions such as business modules, importance, and change frequency. The data layer is generally split according to business first, and if necessary, vertical splitting can also be done, such as data sharding, read - write separation, and data cold - hot separation. After the system is split, it will be divided into multiple modules. The dependencies between modules are of different strengths. If it is a strong dependency, then if the dependent party has a problem, it will also be affected. Sort out the call relationships of the core module processes of the entire IDM platform to make weak - dependency calls. Use the REDIS queue method to achieve decoupling. Even if there is a problem downstream, it will not affect the current module.

[0071] After a full - scale evaluation in multiple aspects such as applicability, advantages and disadvantages, product reputation, community activity, practical cases, and scalability, through comparison, testing, and research, mysql is used for simple metadata and mongodb is used for complex structured data to provide overall data support.

[0072] Before the system goes live, capacity assessments need to be carried out on the machines, DB, and cache used by the entire service. The machine capacity is evaluated in the following way: 1) Define the expected traffic metric - QPS; 2) Define the acceptable latency and safety water level metrics (e.g., CPU% ≤ 40%, core link RT ≤ 50ms); 3) Through stress testing, evaluate the maximum QPS that a single machine can support below the safety water level (it is recommended to verify through a mixed scenario, such as stress testing multiple core interfaces simultaneously according to the estimated traffic ratio); 4) Finally, the specific number of machines can be estimated.

[0073] For the DB and cache evaluations, in addition to QPS, the data volume also needs to be evaluated. The method is roughly the same. After the system goes live, scaling can be performed according to the monitoring metrics.

[0074] Systematic protection is an undifferentiated traffic limit. In one sentence, the concept is to perform undifferentiated traffic limiting on all traffic entrances before the system is about to collapse, and stop the traffic limiting when the system returns to a healthy water level. More specifically, it combines monitoring metrics in several dimensions such as the application's Load, overall average RT, entrance QPS, and the number of threads, so that the entrance traffic of the system and the system's load reach a balance, allowing the system to run at the maximum throughput as much as possible while ensuring the overall stability of the system.

[0075] When a system failure occurs, we first need to find the cause of the failure, then solve the problem, and finally restore the system. The speed of troubleshooting largely determines the duration of the entire failure recovery, and the greatest value of observability lies in rapid troubleshooting. Secondly, based on the three pillars of Metrics, Traces, and Logs, alarm rules can be configured to detect potential risks and problems in the system in advance and avoid the occurrence of failures.

[0076] Based on the same principle, this application also discloses a data-driven intelligent design and R & D method for materials. As Figure 3 shown, the method includes: S100: Match the simulation software of the multi-scale simulation module 11 based on the material data input by the user. After configuring the simulation information, perform simulation operations through the matched simulation software. The multi-scale simulation module 11 includes simulation software at the nano-scale, micro-scale, meso-scale, and macro-scale.

[0077] S200: The high-throughput module 12 performs node operations based on all nodes in the preset process template corresponding to the task file to obtain the task operation result. The AI model of the AI model module 13 is containerized and deployed on the node. The input parameters of the current task are generated based on the parameter template of the AI model, and model prediction is performed. In addition, the user can deploy the AI model online through the AI model module 13.

[0078] S300: Online query the material data through the material database 14 and template-display the material data stored in the material database 14.

[0079] Since the principle of this method for solving problems is similar to the above system, the implementation of this method can refer to the implementation of the system and will not be elaborated here.

[0080] The system, device, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with a certain function. A typical implementation device is a computer device. Specifically, the computer device can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0081] In a typical example, the computer device specifically includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method executed by the client as described above, or when the processor executes the program, it implements the method executed by the server as described above.

[0082] Next, refer to Figure 4 , which shows a schematic structural diagram of a computer device 600 suitable for implementing the embodiments of the present application.

[0083] As Figure 4 shown, the computer device 600 includes a central processing unit (CPU) 601, which can perform various appropriate operations and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer device 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0084] The following components are connected to the I / O interface 605: an input part 606 including a keyboard, a mouse, etc.; an output part 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 608 including a hard disk, etc.; and a communication part 609 including a network interface card such as a LAN card, a modem, etc. The communication part 609 performs communication processing via a network such as the Internet. The drive 610 is also connected to the I / O interface 605 as required. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 610 as required so that a computer program read therefrom is installed in the storage part 608 as required.

[0085] Specifically, according to an embodiment of the present invention, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 609, and / or installed from the removable medium 611.

[0086] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0087] For convenience of description, when describing the above device, it is divided into various units according to functions and described separately. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.

[0088] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0089] The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0090] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0091] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A data-driven intelligent design and R & D platform system for materials, characterized in that, It includes a multi-scale simulation module, a high-throughput module, an AI model module, and a material database; The multi-scale simulation module includes simulation software at the nano-scale, micro-scale, meso-scale, and macro-scale. It matches the simulation software based on the material data input by the user, configures the simulation information, and then performs simulation operations through the matched simulation software; The high-throughput module is used to perform node operations on all nodes in the preset process template corresponding to the task file to obtain the task operation result; The AI model module includes an AI model containerized and deployed on nodes, generates input parameters for the current task based on the parameter template of the AI model and performs model prediction, and is used for the online deployment of the user's AI model; The material database stores material data and is used for the user to query the material data online and display the material data in a templated manner.

2. The data-driven intelligent material design and R & D platform system according to claim 1, characterized in that It further includes a resource scheduling module, which is used for creating user supercomputer accounts, allocating resource storage space, allocating computing power, and counting computing power consumption. It monitors the usage of the user's computing power and space, automatically freezes and unfreezes the account. When the user's computing power is insufficient, the supercomputer account will be automatically downgraded or frozen. After the background adds computing power again, the account will be automatically promoted or unfrozen. When the user's storage space is insufficient, the account will be automatically frozen and prohibited from use. When the administrator adds storage allocation in the background or clears the space storage independently, the system will automatically unfreeze the account again.

3. The data-driven intelligent design and R & D platform system for materials according to claim 1, characterized in that The high-throughput module is used to decompose the task file into the smallest unit processes corresponding to the nodes according to the preset process template, and execute all nodes corresponding to the task file in a directed graph traversal manner.

4. The data-driven intelligent material design and R & D platform system according to claim 3, characterized in that The process of the preset process template includes a normal process, a judgment process, a loop process, and a repeated parallel process.

5. The data-driven intelligent design and R & D platform system for materials according to claim 1, wherein It further includes an identity authentication module, which is used to verify the user's identity through the user information verification method when the user logs in to the application program of the module each time. If the verification is passed, a security token is generated. When the user uses other application programs, if there is a security token and the verification is passed, the user is allowed to use the corresponding application program.

6. The data-driven intelligent material design and R & D platform system according to claim 2, characterized in that, The user's computing power includes shared resources and exclusive resources. When using computing power resources, the exclusive resources are preferentially used. If the exclusive resources are exhausted, the shared resources are used. If the occupancy of the shared resources reaches the preset ratio, wait for the exclusive resources.

7. The data-driven intelligent material design and R & D platform system according to claim 1, characterized in that It further includes a task configuration module, which is used to generate the task file based on the user input parameters and the preset script template.

8. A data-driven research and development method for intelligent material design, characterized in that, It includes: Match the simulation software of the multi-scale simulation module based on the material data input by the user, configure the simulation information, and then perform simulation operations through the matched simulation software. The multi-scale simulation module includes simulation software at the nano-scale, micro-scale, meso-scale, and macro-scale; Through the high-throughput module, perform node operations on all nodes in the preset process template corresponding to the task file to obtain the task operation result. The nodes are containerized and deployed with the AI model of the AI model module, generate input parameters for the current task based on the parameter template of the AI model and perform model prediction, and the user can deploy the AI model online through the AI model module; Online query the material data through the material database and display the material data stored in the material database in a templated manner.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the data-driven intelligent design and research and development method for materials as described in claim 8.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the data-driven intelligent design and research and development method for materials as described in claim 8.

Citation Information

Patent Citations

  • Universal cross-scale material calculation simulation system and method thereof

    CN110299191A

  • Unified access method of multiple application programs and related equipment

    CN110417730A

  • Material computing framework, method and system and computer equipment

    CN112685911A

  • High-throughput task processing method based on super computer

    CN112882810A

  • Multi-scale material intelligent computing platform combined with artificial intelligence

    CN116861736A

Cited By

  • Method for automatically generating simulation calculation task on supercomputing platform

    CN120892057A