A multi-working-condition parallel scheduling method and system based on a marine wind power simulation platform
By adopting a multi-condition parallel scheduling method in the offshore wind power simulation platform, and utilizing the vernier-based amortized scheduling index and cross-stage pipeline incremental ready semantics, the problems of low resource utilization and computation cycle exceeding limits are solved, achieving efficient computation management and shortening the simulation cycle.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI INVESTIGATION DESIGN & RES INST CO LTD
- Filing Date
- 2026-07-06
- Publication Date
- 2026-08-04
AI Technical Summary
In offshore wind power simulation platforms, resource utilization is low, calculation cycles exceed limits, and there is a lack of a unified scheduling mechanism. Computational management capabilities have become a key weakness in engineering efficiency.
A multi-condition parallel scheduling method based on an offshore wind power simulation platform is adopted. The method realizes the condition assignment of the ordered task queue through a cursor-based amortized scheduling index, uses configurable regular expressions to identify transient I/O errors and perform exponential backoff retries, and combines cross-stage pipeline incremental readiness semantics for parallel solution.
It improved resource utilization, shortened the simulation calculation cycle, solved the problems of idle resources and exceeding the time limit, and achieved efficient calculation management.
Smart Images

Figure CN122507525A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of engineering simulation technology, specifically to a multi-condition parallel scheduling method and system based on an offshore wind power simulation platform. Background Technology
[0002] For a long time, the design of offshore wind power projects has relied on various self-developed, commercial, or open-source single-physics-domain solvers to complete various simulation analyses. Each solver typically runs and is managed independently, and hardware calls are manually triggered sequentially by engineers in different software. As integrated simulation platforms integrate multiple solution modules such as time-domain coupled loads, frequency-domain hydrodynamics, structural finite element analysis, and fatigue analysis into a unified platform, each solution stage within the platform performs batch calculations for tens of thousands of design load cases (DLC) lifecycle conditions. The computational management capabilities within the platform have become a key bottleneck affecting engineering efficiency. Summary of the Invention
[0003] This invention provides a multi-condition parallel scheduling method and system based on an offshore wind power simulation platform to solve the problems of low average resource utilization and excessive calculation cycle in offshore wind power simulation platforms.
[0004] In a first aspect, the present invention provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform, the method comprising: Obtain the list of working conditions in a single solution stage during the integrated simulation of offshore wind power, initialize the state of the list of working conditions in a single solution stage, and obtain an ordered task queue. The work conditions in the ordered task queue are distributed to the solver backend using a cursor-based amortized scheduling index. Obtain the solver backend output stream, use configurable regular expressions to identify transient I / O error conditions in the solver backend output stream, perform exponential backoff retries for transient I / O error conditions, and schedule loops for conditions after exponential backoff retries. When the working condition in a single solution stage reaches the termination state, the working condition in the next solution stage is solved in parallel using cross-stage pipeline incremental ready semantics until the integrated simulation of offshore wind power is completed.
[0005] This invention provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform. It achieves balanced allocation of conditions in an ordered task queue through a cursor-based amortized scheduling index, eliminating the performance bottleneck of traversal scheduling and allowing the bounded process pool to continuously operate near full load, thus improving resource utilization. An exponential backoff retry mechanism for transient I / O errors avoids condition failures and reruns caused by temporary storage or lock conflicts, reducing unnecessary computational overhead. Simultaneously, based on cross-stage pipeline incremental ready semantics, dependent subsequent stage conditions are immediately reached after the completion of the preceding stage condition, eliminating stage bubbles of full-batch waiting and significantly shortening the overall computation cycle. This significantly improves the resource utilization of the offshore wind power simulation platform, greatly reduces the simulation computation cycle, and effectively solves the core problems of resource idleness and project timeline exceeding limits.
[0006] Secondly, this invention provides a multi-condition parallel scheduling system based on an offshore wind power simulation platform, the system comprising: The state initialization module is used to obtain the list of working conditions in a single solution stage during the integrated simulation of offshore wind power, and to initialize the state of the list of working conditions in a single solution stage to obtain an ordered task queue. The assignment module is used to assign the working conditions in the ordered task queue to the solver backend using a cursor-based amortized scheduling index. The identification module is used to obtain the solver backend output stream, identify transient I / O error conditions in the solver backend output stream using configurable regular expressions, perform exponential backoff retries for transient I / O error conditions, and schedule loops for conditions after exponential backoff retries. The solver module is used to solve the working conditions in the next solver stage in parallel by using cross-stage pipeline incremental ready semantics when the working conditions in a single solver stage reach the termination state, until the integrated simulation of offshore wind power is completed.
[0007] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the multi-condition parallel scheduling method based on an offshore wind power simulation platform as described in the first aspect or any corresponding embodiment.
[0008] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the multi-condition parallel scheduling method based on an offshore wind power simulation platform as described in the first aspect or any corresponding embodiment.
[0009] Fifthly, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the multi-condition parallel scheduling method based on an offshore wind power simulation platform according to the first aspect or any corresponding embodiment described above. Attached Figure Description
[0010] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the first step of a multi-condition parallel scheduling method based on an offshore wind power simulation platform according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the state switching process of a five-state working condition state machine according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the scheduler startup detection checkpoint, incremental recovery, and atomic write process according to an embodiment of the present invention; Figure 5 This is a second flowchart illustrating a multi-condition parallel scheduling method based on an offshore wind power simulation platform according to an embodiment of the present invention. Figure 6 This is a flowchart illustrating the vernier-based O(1) amortized scheduling scheme according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the periodic sampling and dynamic adjustment decision-making process of the resource monitor according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the third process of a multi-condition parallel scheduling method based on an offshore wind power simulation platform according to an embodiment of the present invention. Figure 9 This is a schematic diagram of the fourth process of a multi-condition parallel scheduling method based on an offshore wind power simulation platform according to an embodiment of the present invention. Figure 10 This is a schematic diagram of the selective parallel scheduling and three-level progress aggregation process with group awareness according to an embodiment of the present invention; Figure 11 This is a structural block diagram of a multi-condition parallel scheduling system based on an offshore wind power simulation platform according to an embodiment of the present invention; Figure 12 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0014] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0015] The integrated simulation platforms generally suffer from the following problems in multi-condition computation management: multiple conditions are still executed sequentially in a loop within a single solution phase, resulting in low average utilization of multi-core CPU (Central Processing Unit) resources; each solution module maintains its own independent batch processing logic, lacking a unified scheduling abstraction across modules; failure of a single condition may block the entire batch of tasks, and completed results cannot be incrementally recovered; the progress of conditions is not observable, and there is a lack of interruption and recovery mechanisms, which cannot effectively support engineers' interactive working methods for large-scale batch computations.
[0016] Specifically, the relevant integrated simulation platform has the following shortcomings: 1) The scale of operating conditions makes multi-core parallelization a rigid engineering requirement, while a unified cross-module scheduling mechanism is still immature: According to standards such as IEC 61400-1 / 61400-3, a complete offshore wind power design load analysis includes at least DLC1.x to DLC9.x scenarios. In floating offshore wind power engineering practice, after each type of DLC is expanded by Cartesian combination according to dimensions such as wind speed range, turbulence seed number, wind and wave direction, yaw error, and ocean current / water level, the complete DLC case set of a single project can reach more than 28,000 independent calculation cases (statistics of actual floating cases). In other words, if the cases are called sequentially, the total time is as high as tens of days, which seriously exceeds the engineering design iteration cycle. Therefore, the scale of the cases itself determines that process-level parallelization is a rigid engineering requirement for offshore wind power simulation. However, a complete integrated simulation platform covers multiple independently encapsulated solution modules such as load, hydrodynamics, structural strength, and fatigue. If each module implements its own parallel calling logic, common behaviors such as task grouping, task queuing, concurrency control, and progress acquisition will be scattered across the modules. This not only makes it impossible to uniformly allocate cross-stage concurrent resources at the workflow level, but also results in a large amount of repetitive implementation. There is still a lack of systematic solutions at the engineering implementation level for designing a unified scheduling kernel that is transparent and applicable to all solution modules.
[0017] 2) The solvers for the simulation platform come from diverse sources, lacking a unified access and scheduling standard: The solvers used in different solution stages differ fundamentally in their source, execution form, and interface specifications: some solvers are released as standalone executable programs (EXE), some are embedded in the platform as dynamic link libraries (DLLs), and some need to be submitted to supercomputing clusters or cloud computing nodes for remote execution. These differences are not due to human choice, but rather determined by the evolution of solver technology and the realities of the deployment environment. Hardcoding the calling methods for different types of solvers separately leads to a lack of uniformity in workflow-level scheduling. Common logic such as concurrency control, state management, and progress acquisition at each stage is forced to be implemented repeatedly, resulting in high platform maintenance costs and limited cross-stage scheduling and allocation capabilities.
[0018] 3) Concurrent execution of external processes introduces various reliability issues, affecting the success rate of batch computation: In large-scale, multi-condition concurrent environments, the following documented reliability flaws are commonly observed in industry practice: When multiple processes concurrently write database result files to the same storage path, the database's internal file locking mechanism can trigger transient errors. If the scheduling system fails to recognize these errors and misjudges them as solution failures, normal processes will be forced to restart. User-initiated interruptions (such as discovering incorrect input parameters) and solver process crashes exhibit similar behaviors. If the scheduling system does not distinguish between the two termination reasons, completed results cannot be safely retained, and incremental continuation of computation is impossible. If a single abnormal condition is not isolated, it may propagate in a chain reaction within the batch and halt the entire batch computation. In engineering scenarios where a single complete batch computation takes several hours to several days, the losses caused by these reliability flaws are particularly significant.
[0019] 4) Concurrent deployment of third-party EXE solvers causes file conflicts and startup I / O (Input / Output) bottlenecks: Third-party solvers running as EXE processes typically write intermediate files (logs, lock files, etc.) to their own directory. If multiple workstations share the same EXE path during concurrent execution, handle conflicts will occur. Allocating an independent EXE copy to each workstation introduces disk copying overhead (approximately 434ms / 8 workstations end-to-end), a surge in peak disk usage, and I / O serialization waits before workstation startup, creating a significant throughput bottleneck in large-scale batch scenarios.
[0020] 5) Large-scale operational scenarios lack group management and on-demand execution selection mechanisms: In large-scale simulation scenarios involving tens of millions of load conditions, engineers typically do not need to execute all load conditions at once. For example, in the early stages of parameter iteration, it is often only necessary to run specific subsets of load conditions such as DLC1.1 and DLC2.1 to quickly verify the results; after a specific design change, only the affected DLC groups need to be recalculated. However, relevant simulation platforms generally lack the ability to manage load condition groups, only supporting full execution or manual filtering. This cannot meet the actual engineering needs of large-scale simulation projects for on-demand selection and step-by-step iterative calculation, resulting in a large amount of repetitive calculations and low engineering efficiency.
[0021] 6) Long-duration batch computations lack sufficient observability and interruption recovery mechanisms: When batch calculations last for hours or even days, engineers need to monitor the running status of each working condition in real time, be able to interrupt and correct parameters for subsequent calculations without losing completed results, and recover from the most recent valid state after an unexpected system crash. However, the relevant simulation platforms output progress frequently (tens of times per second per working condition) from each solver under concurrent multi-working-condition scenarios. If the interface is refreshed directly without processing, the progress signal density can reach hundreds of times per second, causing UI (User Interface Design) rendering lag and reducing platform availability. At the same time, the loss of scheduling state after the host process crashes will force a full rerun, resulting in a serious loss of engineering efficiency.
[0022] In summary, the aforementioned shortcomings stem from the inherent characteristics of offshore wind power simulation scenarios, including large scale, objectively heterogeneous solver architecture, uncontrollable external process output behavior, and high reliability requirements for long-term batch computation. Related parallel batch processing solutions (such as general workflow engines and process pool frameworks) are all designed for homogeneous computational tasks or internally controllable algorithm kernels, and do not offer a systematic solution to the aforementioned combined characteristics.
[0023] This invention provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform. It shields the differences in solver execution forms through a unified solver backend interface, supporting various solvers (including integrated load, frequency domain hydrodynamic, structural modal, finite element strength, and structural fatigue solvers) running as independent executable programs (EXE) or dynamic link libraries (DLLs). At the architecture level, it implements standard extension interfaces for backend forms such as supercomputing clusters, containerization, and heterogeneous computing. The scheduling system controls the parallel scale of multiple conditions with a bounded number of concurrent units, dynamically adjusts the concurrency limit through a resource monitor to prevent memory overflow-type faults caused by excessive concurrency, implements amortized O(1) task assignment indexes using monotonically increasing cursors, identifies transient I / O errors such as database file locks using configurable regular expression patterns, and automatically retryes them with an exponential backoff strategy, prefetches the working directory of the next batch of conditions in a pipelined manner in a background thread, and aggregates high-frequency progress signals and pushes them to the UI (User Interface) at a controlled frequency. Interface (user interface) layer; five-state working condition state machine supports incremental continuation of calculations after interruption; session state persistence and checkpoint mechanism support cross-process breakpoint recovery after host process crash, and the results of completed working conditions remain valid; for both fixed and floating engineering workflows, cross-solution stage dependencies are modeled through directed acyclic graphs, and fine-grained cross-stage incremental ready semantics are supported, eliminating stage waiting bubbles in multi-stage scheduling. Applicable to large-scale offshore wind power engineering design and calculation scenarios involving hundreds to tens of thousands of design working conditions.
[0024] As an optional application scenario of this invention, such as Figure 1As shown, the terminal equipment is equipped with a multi-condition parallel scheduling system based on an offshore wind power simulation platform. The system has a five-layer architecture, including an access abstraction layer L1, a scheduling decision layer L2, an operation support layer L3, a group management layer L4, and an observation interaction layer L5. The overall scheme is organized in a five-layer closed-loop manner. The five layers form a closed operation loop through task assignment, condition grouping, operation feedback, state convergence, and persistent recovery.
[0025] The access abstraction layer L1 unifies the backend interface, shields the execution differences between EXE / DLL / extension interfaces, and supports remote clusters and heterogeneous extension semantics; the scheduling decision layer L2 includes bounded process pools, cursor indexes, five-state state machines (i.e., five-state FSM), retry and priority policy components; the operation support layer L3 includes directory prefetching, progress throttling, log and checkpoint persistence components; the group management layer L4 includes work condition group identifier parsing, tag filtering, selective scheduling, progress aggregation, and group-level interrupt / recovery control components; and the observation interaction layer L5 includes session lifecycle event streams, UI-level state trees, and progress panel components.
[0026] The access abstraction layer L1 unifies the semantics of launch (starting the solver task, supporting EXE / DLL / extension interface forms), cancel (canceling the running solver task), status (querying the current running status of the task), and output (returning the solver output stream). The scheduling decision layer L2 schedules and calls launch, sending the task to the access abstraction layer L1. The access abstraction layer L1 starts the solver process / instance and returns the task handle and status. The scheduling decision layer L2 receives the solver's progress, logs, and error output through output. When the scheduling decision layer L2 issues a cancel command, the access abstraction layer L1 performs process termination / resource reclamation operations. The group management layer L4 issues group-level control commands, which are executed by the scheduling decision layer L2, which then provides status feedback. The group management layer L4 also pushes group status and progress data to the observation interaction layer L5. The scheduling decision layer L2 initiates scheduling execution and status write-back requests, and the operation support layer L3 writes back the recovery point and running data to the scheduling decision layer L2. The scheduling decision layer L2 pushes session events and single-condition progress data to the observation interaction layer L5.
[0027] According to an embodiment of the present invention, a multi-condition parallel scheduling method based on an offshore wind power simulation platform is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0028] This embodiment provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform, which can be used in the aforementioned terminal equipment. Figure 2 This is a flowchart of a multi-condition parallel scheduling method based on an offshore wind power simulation platform according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain the list of working conditions in a single solution stage during the integrated simulation of offshore wind power, initialize the state of the list of working conditions in a single solution stage, and obtain an ordered task queue.
[0029] Specifically, the N working conditions to be calculated are arranged into an ordered task queue, and each working condition is initialized to a Pending state. The working condition belongs to any of the solution stages of integrated time-domain load, frequency-domain hydrodynamic analysis, structural modal analysis, structural finite element strength analysis, or structural fatigue analysis.
[0030] Step S202: Use the cursor-based amortized scheduling index to distribute the working conditions in the ordered task queue to the solver backend.
[0031] Specifically, a monotonically increasing cursor is maintained in the task graph corresponding to the ordered task queue, pointing to the next candidate position in the queue to be assigned.
[0032] Furthermore, a unified solver backend interface is used to assign work conditions. This backend interface masks the differences in solver execution forms, supporting various forms such as independent executable programs (EXE process backend), dynamic link libraries (DLL plugin backend), remote clusters / supercomputing nodes (remote backend), and containerized / heterogeneous computing (extended backend). When the solver backend is a DLL plugin backend, solver dynamic link library instances are loaded independently in multiple dedicated worker processes or threads, with memory spaces isolated between instances. The scheduler manages the work condition execution state of the DLL backend through the same state machine interface as the EXE process backend, and the upper-level scheduling logic does not distinguish between backend forms. When the solver backend is a remote cluster backend, the scheduler submits the work condition description to the supercomputing cluster or cloud computing environment through remote procedure calls or standard job scheduling protocols (including SLURM, PBS, etc.), obtains the job status through polling or event callback mechanisms, and maps it to a five-state machine. The upper-level scheduling logic remains consistent with the local backend, enabling on-demand switching between local and cluster computing scenarios.
[0033] A unified SolverBackend interface is defined (used to abstract a unified programming interface for different solver backend implementations). Regardless of the backend type, it exposes the following four semantic interfaces: Startup interface, used to start a single task and return an execution handle for subsequent status management; Cancel interface, used to cancel the execution of a specified task; Status query interface, used to query the current running status of a specified task; Output acquisition interface, used to acquire the output data stream of a specified task. The scheduler only depends on the SolverBackend interface and is unaware of the specific backend implementation. Currently, four types of backends are supported, as shown in Table 1 below.
[0034] Table 1. Backend Types:
[0035] In Table 1 above, SLURM (Simple Linux Utility for Resource Management) / PBS (Portable Batch System) jobs represent a highly scalable and fault-tolerant cluster manager and job scheduling system / portable batch processing system job that can be used in large compute node clusters; Docker / K8sJob represents a containerized platform / a controller resource.
[0036] The technical necessity of the aforementioned unified interface is that without defining a unified backend interface, core scheduling mechanisms such as five-state machine, cursor scheduling index, database lock-aware retry, and session state persistence will be scattered with the business code of each form, making it impossible to reuse across forms, and the interruption recovery semantics of SLURM jobs will not be consistent with the local EXE.
[0037] Furthermore, during the execution process of the solver backend, each working condition maintains a finite state machine containing five states: Pending, Running, Completed, Failed, and Interrupted. The Interrupted state is triggered by user-initiated interruption and is strictly distinguished from the Failed state.
[0038] For each working condition, a five-state finite state machine (FSM) is defined, with the following semantics for each state: Pending, meaning pending execution and waiting for scheduling, its entry condition is initial creation or Interrupted reset; Running, meaning in execution and dispatched to the backend, its entry condition is the scheduler triggering launch; Completed, meaning completed and the result is valid, its entry condition is the backend exiting normally (exitCode=0); Failed, meaning failed and all retries have been exhausted, its entry condition is the backend exiting abnormally and the allowed number of retries has been exhausted; Interrupted, meaning interrupted and recoverable, its entry condition is user or system termination (strictly distinguished from Failed).
[0039] Among them, such as Figure 3 As shown, the state update process of the five-state finite state machine is as follows: A task is created, initially in the Pending state; the scheduler calls `tryStartPendingTasks` to start the task, and the task enters the Running state; if the task completes normally (i.e., the solver process has an `exitCode == 0`), the task enters the Complete state; if the task fails (i.e., a non-zero exit code, process crash, or startup failure), the task enters the Failed state; if actively interrupted (i.e., the user calls `stopTask` or the system calls `stopAll`), the task enters the Interrupted state; when the task is in the Failed or Interrupted state, and the user initiates a recovery operation, it is reset to Pending via `retryTask`, allowing for rescheduling and execution, directly entering the session end node, and the task's lifecycle terminates.
[0040] Standard workflow engines typically maintain four states (without interrupted). This embodiment explicitly introduces the interrupted state, supporting user-initiated ordered interruptions and incremental recovery. The completed state remains unchanged during incremental recovery and is not recalculated.
[0041] Furthermore, once a batch of work cases starts, the scheduler submits a directory prefetching task to the system background thread pool with the lowest priority. The prefetching task creates a working directory for the next batch of work cases and writes it into the solver startup script. The creation of the working directory required for the next batch of work cases and the writing of the solver startup script / configuration file are completed in parallel in the background thread (if the target file already exists, it returns idempotently). The prefetching task is executed in parallel with the solution calculation of the current batch, eliminating the I / O waiting time before the start of the work case (in actual tests, it was compressed from about 150ms to about 2ms).
[0042] Furthermore, the runtime support layer L3 performs zero-copy EXE distribution and adaptive degradation for each work case. The zero-copy EXE distribution scheme is as follows: before building the work case task list, the shared source EXE path is pre-parsed and directly assigned to each LoadCaseTask.executablePath (the shared source path of the solver backend executable program), without performing any file copying or hard linking operations. The batch script generated by the process launcher switches the current working directory (CWD) to the work case-specific directory using cd / d "%~dp0", so that the solver backend JSON configuration reading and writing and result output are all completed within the isolated CWD. The DLL is loaded by the operating system from the directory where the EXE is located, and the processes of each work case do not interfere with each other. The adaptive degradation scheme is as follows: during the batch work case task construction phase, the existence of the shared source EXE is checked. If it does not exist, it automatically degrades to the hard link / copy strategy and outputs a [WARNING] log, ensuring that parallel computing is available throughout without manual intervention.
[0043] The comparison table between the full copy scheme and the aforementioned zero-copy EXE distribution and adaptive degradation scheme is shown in Table 2 below.
[0044] Table 2. Comparison of the full copy scheme and the above-mentioned zero-copy EXE distribution and adaptive degradation scheme:
[0045] As shown in Table 2 above, compared to the full copy solution, which independently copies the EXE for each working condition, the end-to-end time is approximately 434.4ms in the 8 working conditions × 103.6MB scenario, with a peak disk usage of approximately 829MB. The aforementioned zero-copy end-to-end speedup is approximately 63.9 times.
[0046] Step S203: Obtain the solver backend output stream, use configurable regular expressions to identify transient I / O error conditions in the solver backend output stream, perform exponential backoff retries for transient I / O error conditions, and schedule loops for conditions after exponential backoff retries.
[0047] Specifically, when any operating condition ends, its exit status and output content are read, and the output stream is matched with a set of configurable regular expressions to determine whether it is a transient I / O error. The matching should at least cover database file lock conflict error patterns.
[0048] Furthermore, as the number of parallel working conditions increases, excessive push frequency can cause UI lag. Therefore, each solver writes progress data to the output stream at a higher frequency (approximately 20Hz / working condition). After receiving the backend output, the scheduler stores the latest progress data in a progress buffer indexed by the working condition identifier. In other words, the progress data parsed from the working condition output is cached with the working condition identifier as the index, and new data overwrites old values. A fixed-period timer (default 50ms) triggers batch pushes to push all the progress data to be sent in the buffer to the UI observers at once. The timer period determines the maximum progress update frequency received by the upper-level observers (such as the UI rendering thread). The periodic aggregation and throttling of the progress signal is shown in Table 3 below.
[0049] Table 3. Periodic Aggregation Throttling Table of Progress Signals under Parallel Operating Conditions:
[0050] Step S204: When the working condition in a single solution stage reaches the termination state, the working condition in the next solution stage is solved in parallel using cross-stage pipeline incremental ready semantics until the integrated simulation of offshore wind power is completed.
[0051] Specifically, a Directed Acyclic Graph (DAG) model is used to model the data dependencies between different solution stages in a multi-stage engineering simulation workflow. The output of the previous stage is used as the data dependency predecessor of the next stage's work case. The scheduler only includes the work cases of the subsequent stage in the Pending candidate set and triggers the assignment after all the work cases of the predecessor stage have entered the Completed state. Work cases without dependencies within the same stage are executed in parallel.
[0052] Furthermore, in the full-batch waiting scheme, subsequent stages must wait for all cases in the preceding stage to complete. The waiting of the last case in a large-scale case set forms a significant stage bubble. Therefore, this embodiment sets incremental ready semantics on the cross-stage dependency edges of the DAG. After a single case i in the preceding stage enters Completed, if there is a case j in the subsequent stage that only depends on the output of case i, then case j immediately enters the Pending candidate set and triggers the scheduling loop, without waiting for the completion of other cases in the preceding stage. The results show that in an 8-core machine with 4320 cases × 2 stages, the pipeline mode eliminates about 16 minutes of stage bubbles in the full-batch waiting scheme, and the total completion time is shortened by about 28%.
[0053] In this system, concurrent execution units across stages share the same bounded process pool, enabling soft pipeline parallelism across multiple solution stages and eliminating stage wait bubbles in the full batch wait strategy. This involves arranging all work cases within the same solution stage into an ordered task queue and setting a configurable maximum number of concurrent execution units P (the default value is recommended based on the number of system logical cores). The scheduling loop is triggered when each work case is completed, dynamically detecting the current number of running units and replenishing and dispatching new work cases from the queue to the backend for execution, keeping the concurrency close to the maximum number of concurrent execution units P. Each work case runs in an execution isolation unit (independent process or dedicated worker thread) provided by the solver backend, and the failure of a single work case does not affect other work cases in the same batch.
[0054] Furthermore, the session ends when all work conditions reach the terminated state (Completed / Failed) or receive an interrupt command; when an interrupt command is received, work conditions in the Running state are converted to Interrupted, while work conditions in the Completed state remain unchanged; when scheduling is resumed later, all Interrupted work conditions are idempotently reset to Pending, and the scheduling loop is retried after the cursor is positioned to the earliest Pending position to achieve incremental continuation of computation.
[0055] In some optional implementations, after completing a preset number of work conditions, the full state snapshot of the current session is atomically written to the session state checkpoint file according to the adaptive checkpoint interval and time fallback condition; when the work condition scheduling process exits abnormally and restarts, the session state checkpoint file is automatically loaded, and incremental recovery is performed based on the session state checkpoint file until the work condition execution is completed.
[0056] Specifically, after completing a preset number of task conditions, the scheduler atomically writes a full state snapshot of the current session (session identifier, cursor position, state of each task condition, and retry history) in JSON format to a checkpoint file. It also serializes the scheduling state of the current session into a persistent checkpoint file. The specific steps are: first, write a temporary file tmp.ck, then atomically rename it to the official file session-{id}.ck.json to ensure checkpoint consistency. The scheduler checks the checkpoint file upon startup; if a valid checkpoint is detected, it automatically loads the session state and performs incremental recovery, ensuring the results of completed tasks remain valid even after an unexpected crash of the host process. If the checkpoint file is missing or corrupted, a full initialization is performed. The writing process takes approximately 1.2ms, which is negligible in terms of amortized scheduling overhead.
[0057] Furthermore, when the time elapsed since the last checkpoint write exceeds a preset time threshold (default 5 minutes) and a new work condition is completed, a checkpoint write is triggered regardless of whether the work condition counting conditions are met, ensuring that no checkpoints are missed when a single work condition takes an extremely long time. The adaptive checkpoint interval I is calculated adaptively based on the total number of work conditions N, taking into account both I / O frequency and crash penalties. The formula for calculating the checkpoint interval I is as follows:
[0058] The time safety net condition is: if more than [time period] since the last checkpoint. (Default 5 minutes) and a new working condition is completed, regardless of whether the count has reached 1, a write operation will be triggered.
[0059] Furthermore, when the host process crashes and restarts unexpectedly (i.e., the task scheduling process exits and restarts abnormally), the scheduler automatically detects valid checkpoints, resets the Interrupted / Running task to Pending, keeps the Completed task unchanged, and resumes calculation directly after the cursor is positioned at the next Pending position. Completed tasks are not rerun. In a scenario where there are 4500 tasks × 10 minutes / task and a power outage occurs when running to the 3060th task, the checkpoint mechanism avoids the complete waste of 64 hours of completed computing resources.
[0060] Furthermore, such as Figure 4As shown, the specific steps of the scheduler's start-up checkpoint detection, incremental recovery, and atomic write are as follows: After the scheduling process starts, it checks whether the checkpoint file session-{id}.ck.json exists in the working directory; if the checkpoint exists, it loads the checkpoint file and reads key information: session_id, cursor position, tasks[] status, and retry history; it determines whether the checkpoint is valid. If the session_id matches and the checkpoint file is complete, the checkpoint is valid, and incremental recovery is performed. That is, the interrupted / running state tasks are reset to Pending, the completed state tasks remain unchanged, and no rescheduling is performed. The cursor is positioned at the earliest Pending task position in the queue, and calculation continues from that position. Completed tasks are not rerun to avoid wasting computing resources. All completed tasks before the crash will not be rerun; if the checkpoint does not exist / is invalid / is damaged, all tasks will be reset to Pending and the cursor will be reset to 0; the scheduling loop will be executed, and all tasks will be run in batches normally; if the host process crashes unexpectedly or the user terminates the session, the current session state has been persisted to the checkpoint file. When the scheduler starts again, it will re-enter the checkpoint loading process, continue the calculation based on the previously saved state, and re-execute the scheduling loop; if the task state changes, the checkpoint will be updated incrementally (triggered every N tasks completed); atomic write: first write to the temporary file tmp.ck; after writing, the temporary file will be atomically renamed to the official file session-{id}.ck.json to ensure consistency, where the checkpoint path is: working directory / session-{id}.ck.json.
[0061] This embodiment provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform. It achieves balanced allocation of conditions across an ordered task queue using a cursor-based amortized scheduling index, eliminating the performance bottleneck of traversal scheduling and allowing the bounded process pool to continuously operate near full load, thus improving resource utilization. An exponential backoff retry mechanism for transient I / O errors avoids condition failures and reruns caused by temporary storage or lock conflicts, reducing unnecessary computational overhead. Simultaneously, based on cross-stage pipeline incremental ready semantics, dependent subsequent stage conditions are immediately reached after the completion of the preceding stage condition, eliminating stage bubbles of full-batch waiting and significantly shortening the overall computation cycle. This significantly improves the resource utilization of the offshore wind power simulation platform, greatly reduces the simulation computation cycle, and effectively solves the core problems of resource idleness and project timeline exceeding limits.
[0062] This embodiment provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform, which can be used in the aforementioned terminal equipment. Figure 5 This is a flowchart of a multi-condition parallel scheduling method based on an offshore wind power simulation platform according to an embodiment of the present invention, such as... Figure 5As shown, the process includes the following steps: Step S501: Obtain the list of working conditions within a single solution stage during the integrated offshore wind power simulation. Initialize the state of the working condition list within a single solution stage to obtain an ordered task queue. For details, please refer to [link to relevant documentation]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.
[0063] Step S502: Use the cursor-based amortized scheduling index to distribute the working conditions in the ordered task queue to the solver backend.
[0064] In some optional implementations, step S502 above includes: Step S5021: Obtain the current position of the cursor, scan the ordered task queue backward from the current position of the cursor, and retrieve the first task that is in the pending execution state.
[0065] Specifically, such as Figure 6 As shown, the scheduler scans linearly forward from the current position of the cursor each time. After finding the first task in the Pending state, it moves the cursor to that position and assigns it. If no task in the Pending state is found, it continues to scan forward until the traversal is complete. If there are still tasks in the Pending state in the ordered task queue, it waits for the next scheduling trigger.
[0066] Step S5022: Assign the first work case in the pending state to the solver backend and advance the cursor to the queue index position corresponding to the first work case in the pending state.
[0067] Specifically, such as Figure 6 As shown, under the normal path without retries, the total number of scheduling accesses for the entire batch of N work conditions is O(N), and the amortized number of dispatch operations per operation is O(1). Compared with the scheme that traverses the entire task array from the beginning each time, the total number of accesses for the entire batch is O(N²), and the scheduling overhead of this implementation is significantly reduced at the scale of tens of millions of work conditions.
[0068] Step S5023: Use the cursor position after advancement to assign the remaining tasks in the ordered task queue until the current number of concurrent execution units reaches the maximum number of concurrent execution units, or the ordered task queue is exhausted.
[0069] Specifically, a maximum number of concurrent execution units P is set, triggering a scheduling loop. The loop linearly scans forward from the cursor's position, finds the first pending condition, updates the cursor to that position, and starts the process through the backend interface until the current number of concurrent execution units reaches P or the queue is exhausted.
[0070] In some alternative implementations, the system resource status of the local computing node is periodically sampled, and the maximum number of concurrent execution units is dynamically adjusted based on the system resource status of the local computing node.
[0071] Specifically, in situations involving multiple calculations and long processing times, to avoid impacting users' ability to conduct other production projects, a resource monitor is deployed at the scheduling decision layer (L2) / operation support layer (L3). The resource monitor runs as an independent background thread, periodically sampling the system resource status of local computing nodes and sending events suggesting adjustments to the concurrency count to the scheduler module. Upon receiving the adjustment suggestion, the scheduler module atomically updates the maximum number of concurrent execution units (P) without interrupting the current operating conditions and records the adjustment event in the session log.
[0072] Furthermore, the resource monitor samples the available memory, average CPU utilization, and disk I / O wait queue depth of the local computing node at fixed intervals (default 2 seconds). When resource pressure exceeds a preset threshold and lasts for more than a preset window period (e.g., available memory < 2GB or CPU > 95% for 30 seconds), the maximum number of concurrent execution units P is reduced by one unit. When resource pressure is detected to be below a slack threshold and lasts for more than a preset window period (e.g., sufficient memory and CPU < 70% for 30 seconds), i.e., when resources are slack, the maximum number of concurrent execution units P is increased by one unit. The adjustment range of P is constrained within [Pmin, Pmax], and any adjustment operation does not interrupt the current operating condition. All adjustment operations are accompanied by event log records for engineers to audit afterward, realizing dynamic adaptation of computing resources and preventing memory overflow failures caused by excessive concurrency. In particular, on a server with 32GB of memory and a peak memory of approximately 3.5GB for single-condition FEM (Finite Element Method) solution, this mechanism will prevent OOM (Out of Memory) errors. The process crash rate (due to memory overflow) dropped from approximately 12% to approximately 0%.
[0073] Furthermore, such as Figure 7As shown, the steps for resource-aware adaptive concurrent dynamic adjustment are as follows: System startup initializes the maximum concurrent execution unit count P to equal the number of logical cores; the resource monitor samples the available memory, average CPU utilization, and disk I / O wait queue depth of the local computing node at fixed intervals (default 2 seconds); periodically performs resource health assessments to obtain the stress level; if memory < 2GB or CPU > 95% for 30 seconds, a downgrade operation is performed, P = P - 1, without interrupting the currently running condition, but pausing the addition of new conditions; if P after adjustment is less than P_min, it is forcibly set to P_min, and a WARNING message is output. The process pool scheduler records the adjustment event to session.log, executes concurrently according to the current P, and continues resource monitoring in the next sampling period. If memory is sufficient and CPU usage is <70% for 30 seconds, an upgrade operation is performed, increasing the concurrency P to P+1, immediately triggering the scheduling loop to supplement and assign new tasks. If the adjusted P is greater than P_max, it is forcibly set to P_max, and a DEBUG log is output. The adjustment event is recorded to session.log, and the scheduler executes according to the new P. If resources are normal, the current P is maintained, and normal scheduling continues. The process pool scheduler waits for the next sampling period and re-enters the resource evaluation process.
[0074] Furthermore, the source-aware adaptive concurrency dynamic adjustment mechanism, the cross-stage fine-grained incremental ready pipeline mechanism, and the session state persistence and cross-process breakpoint recovery mechanism work together under the constraint of shared process pool resources. That is, the resource monitor dynamically adjusts the P value to affect the resource allocation ratio of the pipeline stage, and each P value adjustment event is synchronously written to the checkpoint to ensure that the P setting is correct after recovery, forming a complete closed loop of runtime dynamic optimization and power failure recovery.
[0075] Step S503: Obtain the solver backend output stream, use configurable regular expressions to identify transient I / O error conditions in the solver backend output stream, perform exponential backoff retries for transient I / O error conditions, and schedule loops for the conditions after exponential backoff retries. For details, please refer to [link to relevant documentation]. Figure 2 Step S203 of the illustrated embodiment will not be described again here.
[0076] Step S504: When the working condition within a single solution stage reaches the termination state, the working condition within the next solution stage is solved in parallel using cross-stage pipeline incremental ready semantics, until the integrated offshore wind power simulation is completed. For details, please refer to [link to relevant documentation]. Figure 2 Step S204 of the illustrated embodiment will not be described again here.
[0077] This embodiment provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform. By maintaining a monotonically increasing cursor and scanning linearly forward from its current position, the first condition to be executed is found, and the cursor is assigned and advanced. Under normal path conditions, the total number of scheduling accesses for all batches of conditions is O(N), and the amortized time complexity of a single assignment is O(1). This completely eliminates the O(N²) overhead caused by traversing the queue from the beginning each time, allowing the bounded process pool to continuously replenish new conditions with extremely low scheduling latency, always maintaining operation close to the maximum concurrency, avoiding idle execution units due to scheduling blockage, significantly improving resource utilization and scheduling efficiency under large-scale conditions, and effectively solving the problem of excessive overhead and slowing down the overall parallel efficiency of traditional traversal scheduling in scenarios with tens of millions of conditions.
[0078] This embodiment provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform, which can be used in the aforementioned terminal equipment. Figure 8 This is a flowchart of a multi-condition parallel scheduling method based on an offshore wind power simulation platform according to an embodiment of the present invention, such as... Figure 8 As shown, the process includes the following steps: Step S801: Obtain the list of working conditions within a single solution stage during the integrated offshore wind power simulation. Initialize the state of the working condition list within a single solution stage to obtain an ordered task queue. For details, please refer to [link to relevant documentation]. Figure 5 Step S501 of the illustrated embodiment will not be described again here.
[0079] Step S802: The work cases in the ordered task queue are distributed to the solver backend using a cursor-based amortized scheduling index. For details, please refer to [link to relevant documentation]. Figure 5 Step S502 of the illustrated embodiment will not be described again here.
[0080] Step S803: Obtain the solver backend output stream, use configurable regular expressions to identify transient I / O error conditions in the solver backend output stream, perform exponential backoff retries for transient I / O error conditions, and schedule loops for conditions after exponential backoff retries.
[0081] In some optional implementations, step S803 above includes: Step S8031: Merge the standard error stream in the solver back-end output stream into a standard output stream.
[0082] Specifically, when the backend is an EXE process backend, the standard error stream of the child process is merged into the standard output stream at the operating system level when the process is created. The merged stream is then processed by a single output callback, eliminating the race condition problem of the same line of content being processed repeatedly under the dual-channel separation architecture.
[0083] Furthermore, when the EXE process backend creates child processes, it inherits the host process's environment variables and overlays the runtime configuration key values specific to that working condition (including the solver authorization path, database file lock disable flag, and working condition parameter path). The overlay operation only applies to the current working condition's child processes and does not affect the running environment of other child processes in the same batch.
[0084] Step S8032: The standard output stream is processed uniformly through a single callback, and a session lifecycle event stream is generated in sequence.
[0085] Specifically, the scheduler registers a single output callback for the merged standard output stream. Through this callback, it uniformly receives and parses the output data of all operating conditions, eliminating the race condition problem of the same line of content being processed repeatedly under the dual-channel separation architecture. At the same time, the output lines of each operating condition are written to the output buffer indexed by the operating condition identifier in the order of arrival, for use by the regular expression semantic analysis in the subsequent step S8033.
[0086] Furthermore, the scheduler maintains a complete session lifecycle event stream, sending them sequentially in time: session start event, queue events for each work condition (N times in total), start events for each work condition (N times in total), progress update events for each work condition (multiple times), output events for each work condition (multiple times), completion events for each work condition (N times in total), and session end event; this event stream provides observable interfaces to the interface layer and external data acquisition components, supporting real-time progress panels, work condition status trees, and log scrolling views.
[0087] Step S8033: Identify transient I / O error types in the session lifecycle event stream using configurable regular expressions.
[0088] Specifically, the scheduler performs semantic analysis on the backend output stream, identifying transient I / O error types (strictly distinguishing them from actual solution failures) through a configurable set of regular expression patterns. Transient I / O error types include: database file lock conflicts, outputting the pattern database.unabletolockfile; shared memory write timeouts, outputting the configurable storage I / O timeout error pattern; and other recoverable I / O errors, which are defined by the user in an extended pattern list.
[0089] Step S8034: Obtain the number of retries for the working condition. If the working condition is a transient I / O error type and the number of retries for the working condition has not reached the maximum number of retries, then reset the working condition status to the pending execution status and return the cursor to the queue index position corresponding to the transient I / O error type working condition.
[0090] Specifically, if the operating condition is a transient I / O error and the retry count has not reached the limit, the operating condition status is reset to Pending. Figure 6As shown, the retry path is executed, that is, the cursor rolls back to the index of the current condition queue. The cursor rollback and the condition state reset atom are executed in pairs to ensure that the subsequent scan does not skip the current condition.
[0091] Step S8035: Obtain the waiting time of the current retry round, delay it according to the waiting time of the current retry round, and reschedule the transient I / O error type of the working condition until the working condition in a single solution stage reaches the termination state.
[0092] Specifically, for operating conditions identified as transient I / O error types, an exponential backoff retry strategy is adopted. The operating condition is rescheduled after a delay based on the waiting time of the current retry round; otherwise, the operating condition is marked as Completed or Failed based on the exit status.
[0093] Among them, the waiting time for the current retry round. The calculation formula is as follows:
[0094] in, Indicates the initial delay reference value. This indicates the current number of retries (counting from the first retrieval). The maximum number of retries and the maximum backoff delay can be configured independently according to the solver type. Before retrying, the cursor rollback and state reset atomic pairing operations are performed, and the retry history is appended to the working condition-specific log.
[0095] The scheduler maintains an independent work condition tracking log for each work condition, recording the work condition start time, solution stage, backend type, output summary, exit status, and complete retry history in an append-only manner. In multiple retry scenarios, it retains cross-retry log aggregation records, supporting post-event analysis and evidence preservation.
[0096] Furthermore, after the transient I / O error type is processed, the scheduling loop in step S803 is triggered again to supplement and assign new working conditions, so as to keep the number of concurrent execution units close to the upper limit P.
[0097] Step S804: When the working condition in a single solution stage reaches the termination state, the working condition in the next solution stage is solved in parallel using cross-stage pipeline incremental ready semantics, until the integrated simulation of offshore wind power is completed. For details, please refer to... Figure 5 Step S504 of the illustrated embodiment will not be described again here.
[0098] This embodiment provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform. By merging the standard error stream from the solver backend into the standard output stream, and processing the merged stream with a single callback to generate a time-series session event stream, it eliminates the race condition problem under the dual-channel separate architecture and provides a unified data source for subsequent error identification. Furthermore, it accurately identifies transient I / O error conditions through configurable regular expressions, and combines exponential backoff delay strategy, atomic cursor backoff, and state reset operation to achieve automatic retry of recoverable errors. This avoids the interruption of conditions and waste of recalculation caused by misjudging the solution failure, improves the success rate and robustness of the simulation task, and ensures the correctness of the scheduling logic and the integrity of the session event stream. It effectively solves the problems of chaotic output stream processing, misjudgment of transient errors, and non-closed-loop retry scheduling logic.
[0099] This embodiment provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform, which can be used in the aforementioned terminal equipment. Figure 9 This is a flowchart of a multi-condition parallel scheduling method based on an offshore wind power simulation platform according to an embodiment of the present invention, such as... Figure 9 As shown, the process includes the following steps: Step S901: Obtain the list of working conditions in a single solution stage during the integrated simulation of offshore wind power, initialize the state of the list of working conditions in a single solution stage, and obtain an ordered task queue.
[0100] In some optional implementations, step S901 above includes: Step S9011: Configure metadata fields for each condition in the condition list within a single solution phase.
[0101] Specifically, the DLC specification will be continuously released and updated. If the simulation platform hardcodes the specification classification rules, it will cause compatibility risks when the specification version is changed. Therefore, the group management layer L4 adds the following general metadata fields to the LoadCaseTask model, i.e., the load list. The platform does not interpret any semantics. The metadata fields are shown in Table 4 below.
[0102] Table 4. Metadata Fields:
[0103] Ultimately, the platform does not need to track IEC (International Electrotechnical Commission) standard versions; users can organize themselves according to the latest specifications. The same operating condition can belong to multiple tag dimensions simultaneously, supporting flexible cross-queries. Group identifiers are transmitted as metadata without affecting the correctness of scheduling logic.
[0104] Furthermore, the mechanism standardizes the coordinated operation of unrelated group identification and selective batch execution, group-level interruption and selective incremental recovery, resource-aware adaptive concurrent dynamic adjustment, and session state persistence and cross-process breakpoint recovery. That is, group filtering determines the task queue size, resource monitor dynamically adjusts the P value, P value changes are synchronously written to the checkpoint, and group-level interrupt events are also written to the checkpoint, ensuring that the group status, priority settings, and resource configuration are correctly aligned after crash recovery at any time, forming a complete closed loop for large-scale work condition set management.
[0105] Step S9012: Filter the working conditions after configuring the metadata fields to obtain a subset of working conditions to be scheduled.
[0106] Specifically, at the scale of tens of millions of work cases, users typically do not need to calculate all DLCs at once (e.g., only DLC1.1 and DLC2.1 need to be quickly verified in the early stages of iteration). The relevant scheduling system lacks a mechanism to select the execution subset according to the specification-independent group identifier. Therefore, the group management layer L4 selects the work case subset with matching group identifiers from the full list of work cases based on the set of specification-independent group identifiers selected by the user and injects it into the scheduler. The platform does not interpret the semantics of the group identifiers. The user defines the group hierarchy himself. All scheduling logic (cursor dispatch, process pool, state machine, retry, prefetch, etc.) can be executed transparently without modification. Multiple selective executions are supported, and completed work cases are not rescheduled.
[0107] Furthermore, each work case carries a specification-independent grouping identifier field and a multi-dimensional label list field. This field is user-defined, and the simulation platform does not interpret its mapping relationship with any specific DLC specification entry. Work cases of any specification version can be used directly without platform specification adaptation. Before the scheduler starts, it filters the entire work case list based on the user-selected specification-independent grouping identifier set and only includes the matching subset of work cases into the task queue. During subsequent selective executions, work cases that are already in a completed state are not re-executed, realizing incremental scheduling across batches.
[0108] Step S9013: Initialize the working state of the subset of working states to be scheduled to the state to be scheduled, and obtain an ordered task queue.
[0109] Specifically, the status of the work conditions in the subset of work conditions to be scheduled is initialized to the Pending state.
[0110] Furthermore, based on the aforementioned schedule aggregation and throttling mechanism, the group management layer L4 performs group-level schedule aggregation and three-level visualization; it maintains a schedule aggregation cache with the group identifier as the key, and updates the weighted completion progress of the group whenever the working condition changes; the UI presents a three-level hierarchical view of workflow overview, group level and single working condition level, and the group aggregation update frequency is shared with the existing schedule throttling timer without introducing additional signal channels.
[0111] Among them, grouping Overall progress The calculation formula is as follows:
[0112] in, Indicates grouping Internal working conditions The current progress value (percentage or 0-1). This indicates a user-defined operating condition group (e.g., the DLC1.1 group). Indicates working conditions The weight of the estimated execution time.
[0113] Furthermore, the UI presents a three-level hierarchical view: workflow overview, DLC group level, and single-case level. Each level supports expand / collapse interaction, and the group aggregation update frequency is shared with the progress throttling timer, without introducing additional UI signal channels.
[0114] Among them, such as Figure 10As shown, the steps of selective parallel scheduling and three-level progress aggregation data flow with group awareness are as follows: The user imports the working condition parameter table; the group panel parses the parameter table and generates a structured working condition task list LoadCaseTask[]; the group manager (loop processing) groups the working conditions according to the user-defined DLC specification number (such as DLC1.1, DLC2.1), and sets metadata such as groupId and tag for each working condition; the group list (DLC1.1:2850, DLC1.2:6336...) is displayed through the group panel; the user selects the groups to be executed in the group panel according to the simulation requirements (such as selecting only DLC1.1 and DLC2.1 and skipping other groups); the group panel filters out the matching working condition subset from the full set of working conditions according to the group identifier selected by the user, outputs the tasks filteredTasks, and the group manager... The processor calls the `setTasks(filteredTasks)` interface to pass the filtered subset of task cases to the scheduler; the user clicks "Start Calculation" in the grouping panel, and the grouping panel sends the `start()` command to the scheduler to start the scheduling session; the scheduler calls `launch(task)` to send the task case to the solver backend for execution; the solver backend provides feedback on the progress / completion status; the scheduler updates the task case status and sends it back to the group manager; the group manager refreshes the group aggregation progress: the group manager calculates the weighted progress by group dimension (e.g., DLC1.1:45%|DLC2.1:12%) and pushes it to the grouping panel for display; when all selected group task cases have been executed, the scheduler triggers the `sessionFinished` event; the grouping panel displays the group completion matrix, clearly showing which groups have been completed, which have failed, or which have been interrupted.
[0115] Step S902: Use the cursor-based amortized scheduling index to distribute the working conditions in the ordered task queue to the solver backend.
[0116] Specifically, the group management layer extends the interruption control interface. The group identifier is user-defined, and the platform does not interpret its semantics. Specifically: Group Stop: Cancels and marks the running state of the specified specification-independent group identifier as interrupted, and the working states of other groups in the same session are not affected; Group Resume: Resets the working states in the specified group that are interrupted to the pending state, triggering the scheduling loop to resume execution; Repeated calls do not produce side effects; Global Stop: Performs a stop operation on all groups, which has the same final effect as performing a stop operation on each group individually in sequence, maintaining backward compatibility.
[0117] Furthermore, a cancellation command is sent to the running condition in a specific group according to the group identifier and marked as interrupted, and the running conditions of other groups in the same session are not affected; during recovery, the interrupted running condition of the specified group can be idempotently reset to Pending according to the group identifier and the scheduling loop is triggered to achieve selective incremental recovery at the group level.
[0118] Step S903: Obtain the solver backend output stream, use configurable regular expressions to identify transient I / O error conditions in the solver backend output stream, perform exponential backoff retries for transient I / O error conditions, and schedule loops for the conditions after exponential backoff retries. For details, please refer to [link to relevant documentation]. Figure 8 Step S803 of the illustrated embodiment will not be described again here.
[0119] Step S904: When the operating condition within a single solution stage reaches the termination state, the operating condition in the next solution stage is solved in parallel using cross-stage pipeline incremental ready semantics, until the integrated offshore wind power simulation is completed. For details, please refer to [link to relevant documentation]. Figure 8 Step S804 of the illustrated embodiment will not be described again here.
[0120] This embodiment provides a multi-condition parallel scheduling method based on an offshore wind power simulation platform. By configuring specification-independent metadata fields for the conditions, it achieves compatibility with different DLC specification versions without requiring modifications to the platform logic with standard iterations. Furthermore, it filters conditions based on user-defined group identifiers, injecting only the user-selected subset of conditions to be scheduled into the task queue, avoiding invalid scheduling and resource waste of all conditions. At the same time, it provides basic support for subsequent group-level scheduling, progress aggregation, and interruption control, effectively solving the problems of difficult specification adaptation, poor execution flexibility, and low resource utilization under large-scale conditions.
[0121] The following specific embodiments illustrate the steps of a multi-condition parallel scheduling method based on an offshore wind power simulation platform.
[0122] Example 1: The related multi-condition parallel scheduling methods have the following problems: How to make full use of multi-core CPU computing resources through process-level parallel scheduling, fundamentally solving the problem of multi-core resource idleness and serious overrun of computing cycle caused by serial execution (the average utilization rate of multi-core CPU under the related serial method is only about 12%); and how to support the non-discriminatory collaborative scheduling of three types of backends, namely local independent executable programs (EXE), dynamic link libraries (DLL) and supercomputing clusters, through a unified solver access interface, so as to ensure that the scheduling logic does not need to be rewritten when the platform is elastically expanded from a single engineering workstation to a high-performance computing environment, and eliminate the problem of repeated implementation of the scattered maintenance of scheduling logic in each solution stage; how to efficiently allocate a large-scale set of conditions with amortized O(1) time complexity, and ensure that the scheduling throughput does not degrade under tens of thousands of conditions; how to accurately distinguish database files To address transient I / O errors such as locks and actual solution failures, automatic retries and case-level breakpoint continuation are implemented; how to eliminate I / O serialization overhead during each startup of third-party EXE solvers while ensuring file conflict safety; how to push high-frequency progress to window display without causing rendering stutters using aggregation throttling mechanisms; how to model multi-stage serial dependency workflows across stages, supporting case-level parallelism within stages and incremental pipeline readiness scheduling across stages; how to support on-demand subset execution of tens of thousands of case sets with group-aware selective scheduling, supporting controllable interruptions and incremental recovery at the group granularity; and how to automatically restore to the most recent valid checkpoint after an unexpected crash of the host process through session state persistence mechanisms, avoiding full reruns of long-term batch calculations due to unexpected termination, and solving the problem of insufficient fault tolerance in large-scale case set management.
[0123] To address the problems existing in the aforementioned multi-condition parallel scheduling methods, a specific step-by-step method for multi-condition parallel scheduling based on an offshore wind power simulation platform is as follows: In real-world engineering projects, multiple stages are linked together to form a workflow through DAG dependency edges. This embodiment demonstrates how fixed and floating workflows can utilize the aforementioned single-stage parallel scheduling capability for full-process scheduling.
[0124] Taking a fixed workflow as an example, the complete parallel scheduling steps for a multi-stage workflow are as follows: Integrated time-domain load stage: 4320 DLC cases are scheduled in parallel as EXE process backends. Parallel scheduling of multiple cases in a single solution stage is the core scenario of this embodiment. Taking integrated time-domain load solution as an example: the system initializes the list of 4320 DLC cases to the Pending state, the scheduler reads the number of system logical cores (taking 8 as an example), sets the maximum number of concurrent processes to P=8, and selects the EXE process backend.
[0125] Among them, the single-solution-stage multi-condition parallel scheduling-DLL plugin backend is integrated with the platform in the form of a DLL. The scheduler selects the DLL plugin backend and loads the solver DLL into multiple dedicated working processes. Each working unit independently loads the DLL instance and calls the solution logic through agreed function symbols. The memory states of each condition are completely isolated. The state machine, cursor scheduling, retry logic, progress aggregation and event flow of the condition task are completely consistent with the EXE process backend of Example 1. The upper-level scheduler does not need to distinguish the backend form, which demonstrates the cross-form transparency of the unified backend interface.
[0126] Initial scheduling loop: Starting from position 0, the cursor continuously assigns 8 cases to the backend of the EXE process. Each case starts the integrated time-domain simulation solver as an independent subprocess. The memory between processes is completely isolated, and the runtime environment variables required by the solver (including the database file lock disable flag) are injected.
[0127] Whenever any child process ends, the scheduler reads the exit code and output stream to determine the final state of the process, and then triggers the scheduling loop. The cursor continues to advance to supplement and start new processes, keeping 8 parallel processes in the pool until all 4320 processes are completed and the sessionFinished event is triggered.
[0128] Comparison of time estimates for the 8-core solution: serial processing takes about 6 days → parallel processing takes about 18 hours, and CPU utilization increases from ~12% to ~85%.
[0129] Structural strength stage: Establishing DAG dependency: After all 4320 load cases in the load stage have been completed, the scheduler automatically triggers the structural finite element load case queue and begins parallel solution.
[0130] Structural fatigue stage: Similarly, establish DAG dependency edges and establish upward dependency on the structural strength stage.
[0131] The workflow is compatible with the same scheduling kernel, allowing users to achieve consistent parallel efficiency improvements, progress observability, and interrupt / recovery experience at any stage.
[0132] The cursor-based O(1) scheduling steps are as follows: In a batch with multiple work conditions, the naive scheme traverses the entire task array from the beginning each time the scheduling loop is triggered, and the total number of accesses for the entire batch is O(N²); the cursor scheme monotonically increases as the scheduling progresses, and when the k-th supplementary assignment is performed, it directly starts scanning from the cursor position, and the total number of accesses for the entire batch is O(N); when work condition i needs to be retried due to a transient error, the cursor rolls back to i and is atomically paired with the work condition state reset to ensure that subsequent scans do not skip this work condition.
[0133] The database file lock-aware retry steps are as follows: During batch processing with a concurrency of 8, when writing the database result file in the 37th working condition, a file lock is encountered. The child process exits with a non-zero exit code, and the output includes the database error: unable to lock file. The scheduler identifies this as a transient I / O error. The first retry waits for 2 seconds and then resets the working condition state, and the cursor synchronously rolls back. If the error is the same for 5 consecutive times, it is marked as Failed (all retries have been exhausted) and the complete retry history is recorded. No further retries are made.
[0134] The pipeline directory prefetching steps are as follows: After the start of processes 1-8, the scheduler immediately submits the directory prefetching task to the background thread pool. In the background, the working directory is created for processes 9-24 and written into the startup script (if it already exists, it returns idempotently). Prefetching is executed in parallel with the current batch solution. When the scheduler enters process 9, its working directory is ready, and the startup wait is compressed from about 150ms to about 2ms.
[0135] The progress signal aggregation and throttling steps are as follows: each of the 8 parallel working conditions outputs a progress line at a frequency of about 20Hz. The scheduler may receive about 160 outputs per second. The scheduler stores the latest progress in the progress buffer (overwriting the old value). The latest progress of the 8 working conditions is sent to the interface layer in batches by a 50ms timer. The actual UI refresh frequency is about 20Hz, which makes the visual experience smooth and the CPU overhead controllable.
[0136] The interruption and incremental recovery steps are as follows: When a user executes a task in the tens of millions of tasks, the task is interrupted when it reaches the 1728th task (approximately 40% complete): The scheduler prevents the assignment of new tasks, sends cancellation instructions to all running backends, marks the Running task as Interrupted, and the Completed task remains unchanged; Subsequent recovery from the point of interruption: All Interrupted tasks are idempotently reset to Pending, the cursor is positioned at the earliest Pending position, the scheduling loop is retried, and completed tasks 1–1727 are not rerun.
[0137] The zero-copy EXE distribution and adaptive degradation steps are as follows: In an 8-condition parallel scenario, the zero-copy solution directly points task.executablePath to the shared source EXE ({application directory} / solver / integratedloads / XMarine.exe), without performing any file copying operations. The distribution time is 1.4ms, and the end-to-end time is 6.8ms, which is about 63.9 times faster than the full copy solution. The peak disk usage is reduced from 829MB to 103.6MB. If the shared source EXE does not exist, it automatically degrades to solution B (hard link / copy), outputs a [WARNING] log, and the scheduling logic is not affected.
[0138] The steps for connecting to the supercomputing cluster backend are as follows: The system switches to the remote / cluster backend and submits the job status via the SLURM protocol: Startup: The job status parameters are serialized into a job submission script, the job-specific environment variables are injected, the cluster job submission command is called to submit the job and record the job number; Status query: The cluster job status query is initiated, and the cluster status (queued / running / completed / failed / cancelled) is mapped to a five-state machine; Output acquisition: The cluster shared storage job log is parsed, the progress is extracted and pushed to the interface through the throttling module; Cancellation: The cluster job cancellation command is called.
[0139] The aforementioned backend switching does not require any code modification to the core mechanisms such as the scheduler's five-state FSM, cursor index, retry mechanism, pipeline prefetching and persistence; in a 128 logical core (4 nodes × 32 cores) HPC environment, the total time for the 4320 working condition is approximately 135 minutes, which is about 87.5% shorter than the local 8-core solution.
[0140] Example 2: The specific steps for large-scale group management in DLC are as follows: In the analysis input file, engineers assign group identifiers (such as “DLC1.1”, “DLC2.1”, “DLC6.1”, etc.) and multi-dimensional labels (such as “turbulent wind”, “extreme strength condition”, etc.) to 28,000 operating conditions according to the IEC61400-3-1 standard. The group identifiers are standard-independent group identifiers, and the platform does not interpret their mapping relationship with IEC standard entries. Engineers define the classification hierarchy themselves.
[0141] The steps for setting specification-independent group identifiers and selective batch execution are as follows: In the analysis input file, the engineer assigns groupId (such as "DLC1.1", "DLC2.1", "DLC6.1", etc.) and multidimensional labels (such as ["turbulent-wind", "ULS"]) to 28,000 operating conditions according to the IEC61400-3-1 standard. The groupId is a specification-independent group identifier. The platform does not interpret its mapping relationship with IEC specification entries. The engineer defines the classification hierarchy himself.
[0142] Iteration Phase 1 (Early Parameter Verification): Engineers select two groups, DLC1.1 and DLC2.1 (approximately 480 cases). The scheduler filters by group and injects only these 480 cases into the task queue. Parameter verification is performed quickly, and all scheduling logic (cursor dispatch, state machine, retries, prefetching) runs transparently without modification, completing in approximately 90 minutes.
[0143] Iteration Phase 2 (Incremental Supplementary Calculation): After correcting the parameters, the engineer adds the remaining 13 DLC groups. The scheduler automatically skips the 480 working conditions that have been completed and only performs scheduling on the newly added subset of working conditions, realizing incremental supplementary calculation across batches.
[0144] The group-level progress aggregation and three-level visualization steps are as follows: The UI interface simultaneously presents: Workflow overview layer: total number of work cases 28,000, completion rate 17% (4,760); DLC group layer: DLC1.1:100%, DLC2.1:100%, DLC6.1:32%,...; Single work case layer: can be expanded to view the status, progress percentage and log summary of each specific work case.
[0145] The three-level progress updates share the same 50ms throttling timer, and the UI refresh rate is stable at about 20Hz without introducing additional signal channels.
[0146] The steps for group-level interruption and selective incremental recovery are as follows: When the batch execution reaches the 12,000th case, the engineer finds that the input parameters of the DLC3.x group are incorrect. The engineer triggers stopGroup("DLC3.x"): only the cases in the Running state in DLC3.x are canceled and marked as Interrupted. The execution of the other 17 groups is unaffected and continues to run.
[0147] After the parameters are corrected, the engineer triggers resumeGroup("DLC3.x"): the Interrupted condition of DLC3.x is idempotently reset to Pending, the scheduling loop is triggered, and it proceeds concurrently with other groups that are still executing, and completed conditions are not rerun.
[0148] The resource-aware adaptive concurrency steps for large-scale scenarios are as follows: In a long-duration batch computation containing 28,000 scenarios, the system activates a resource monitor (with a 2-second sampling period). Midway through the batch run, multiple FEM solvers concurrently write results to the database file. The disk I / O wait queue depth continuously exceeds the threshold, triggering a downgrading of P from 8 to 7 (the currently running scenarios are unaffected). Once the I / O queue returns to normal, an upgrading is triggered again. All downgrading / upgrading is recorded in session.log.
[0149] The resource-aware adaptive mechanism for large-scale scenarios reduced the OOM crash rate from about 12% to about 0% on servers with 32GB of memory and a single-condition FEM peak memory of about 3.5GB.
[0150] The incremental readiness step in the cross-stage pipeline is as follows: In the batch computation of the workflow, incremental readiness semantics are set between the load calculation stage and the structural calculation stage. Whenever load condition j enters Completed, the corresponding fatigue condition j is immediately promoted to Pending to trigger dispatch, and the process pool concurrency P is shared between the two stages. This eliminates the stage wait bubble of approximately 16 minutes to several hours in the traditional full-batch waiting scheme, reducing the total completion time by approximately 28%.
[0151] The session persistence and ultra-large-scale breakpoint continuation steps are as follows: In an ultra-large-scale batch computation containing 28,000 work cases (approximately 10 minutes per work case, 8 cores in parallel, with an expected total time of approximately 583 hours), the checkpoint interval is 280. Every 280 work cases completed by the scheduler, the checkpoint is atomically written to the checkpoint file (taking approximately 1.2ms). After approximately 300 hours of computation, the compute node experiences an unexpected power outage. Upon restart, the scheduler automatically loads the last checkpoint, and the first 24,000+ Completed work cases remain unchanged. The cursor is positioned at the correct location to directly resume computation. The recalculation rate after the crash is approximately 0%, avoiding hundreds of hours of wasted computing resources.
[0152] The multi-condition parallel scheduling method based on the offshore wind power simulation platform is compared with the serial scheme. The key technical indicators are compared in Table 5 below.
[0153] Table 5. Comparison of Key Technical Indicators:
[0154] In summary, the differences between the multi-condition parallel scheduling scheme based on the offshore wind power simulation platform and related multi-condition parallel scheduling schemes are shown in Table 6 below.
[0155] Table 6. Differences between this embodiment and related multi-condition parallel scheduling schemes:
[0156] The above embodiments have the following advantages: Solver independence: A unified backend interface eliminates the need for the scheduling system to be aware of the solver's form, allowing transparent access to three types of backends: EXE, DLL, and remote supercomputing. The only cost of expansion is implementing the backend interface. Significantly improved CPU utilization: Bounded process pools increase the average utilization of multi-core CPUs from approximately 12% in serial solutions to over 80%. In a 128-core cluster environment, this can be further increased to approximately 90%. Significantly reduced overall latency: In typical scenarios where dependency-free conditions dominate, the total latency is approximately 1 / min(P,N) of the serial solution. An 8-core machine can achieve a total latency reduction of approximately 75–87%. Condition-level fault isolation: A single condition failure does not affect other conditions in the same batch. Multi-stage dependencies are precisely constrained through DAG edges, making fault propagation controllable. Transient error self-healing: Transient I / O errors such as database file locks are automatically recovered through exponential backoff retries, requiring no manual intervention or modification of the storage architecture. The workload loss rate has been reduced from approximately 8% to approximately 0%; consistent experience across workflows: multiple solution stages (load, hydrodynamic, intensity, fatigue, etc.) share the same scheduling kernel, providing users with consistent progress visualization and interrupt / resumption interaction at any solution stage; elimination of EXE distribution overhead: the zero-copy strategy reduces EXE distribution time from 376.6ms to 1.4ms, an end-to-end speedup of approximately 63.9 times, and a peak disk usage reduction of 87.5%; adaptive degradation strategy ensures that parallel computing can run normally under any deployment conditions; negligible scheduling overhead: cursor optimization reduces the total number of scheduling accesses in the entire batch to O(N), and the scheduler's own CPU consumption is less than 0.1% in batches of thousands of workloads; fine-grained control at the group level: support for selective execution of interrupt recovery at the group level by DLC grouping at the scale of 28,000+ workloads, significantly improving the interaction efficiency of engineers in large-scale batch computing.
[0157] This embodiment also provides a multi-condition parallel scheduling system based on an offshore wind power simulation platform. This system is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.
[0158] This embodiment provides a multi-condition parallel scheduling system based on an offshore wind power simulation platform, such as... Figure 11 As shown, it includes: The state initialization module 1101 is used to obtain the list of working conditions in a single solution stage during the integrated simulation of offshore wind power, and to initialize the state of the list of working conditions in a single solution stage to obtain an ordered task queue. The assignment module 1102 is used to assign the working conditions in the ordered task queue to the solver backend using a cursor-style amortized scheduling index. The identification module 1103 is used to obtain the solver backend output stream, identify transient I / O error conditions in the solver backend output stream using configurable regular expressions, perform exponential backoff retries for transient I / O error conditions, and perform scheduling loops for conditions after exponential backoff retries. The solver module 1104 is used to solve the working conditions in the next solver stage in parallel by using cross-stage pipeline incremental ready semantics when the working conditions in a single solver stage reach the termination state, until the integrated simulation of offshore wind power is completed.
[0159] The multi-condition parallel scheduling system based on an offshore wind power simulation platform provided in this embodiment of the invention can execute the multi-condition parallel scheduling method based on an offshore wind power simulation platform provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments above, and will not be repeated here.
[0160] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0161] The following is a detailed reference. Figure 12 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 1201, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1202 or a program loaded from memory 1208 into random access memory (RAM) 1203. The RAM 1203 also stores various programs and data required for the operation of the electronic device. The processor 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0162] Typically, the following devices can be connected to I / O interface 1205: input devices 1206 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1207 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 1208 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1209. Communication device 1209 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 12 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0163] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 1209, or installed from a memory 1208, or installed from a ROM 1202. When the computer program is executed by the processor 1201, it performs the functions defined in the multi-condition parallel scheduling method based on an offshore wind power simulation platform according to embodiments of the present invention.
[0164] Figure 12 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0165] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, it implements the multi-condition parallel scheduling method based on an offshore wind power simulation platform shown in the above embodiments.
[0166] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0167] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A multi-condition parallel scheduling method based on an offshore wind power simulation platform, characterized in that, The method includes: Obtain the list of working conditions in a single solution stage during the integrated simulation of offshore wind power, initialize the state of the list of working conditions in the single solution stage, and obtain an ordered task queue. The working conditions in the ordered task queue are distributed to the solver backend using a cursor-based amortized scheduling index. Obtain the solver backend output stream, use configurable regular expressions to identify transient I / O error conditions in the solver backend output stream, perform exponential backoff retries for the transient I / O error conditions, and schedule loops for the conditions after exponential backoff retries. When the working condition in a single solution stage reaches the termination state, the working condition in the next solution stage is solved in parallel using cross-stage pipeline incremental ready semantics until the integrated simulation of offshore wind power is completed.
2. The method according to claim 1, characterized in that, The step of using a cursor-based amortized scheduling index to distribute the task conditions in the ordered task queue to the solver backend includes: Obtain the current position of the cursor, scan the ordered task queue backward from the current position of the cursor, and retrieve the first task that is in the pending execution state; The first work case in the pending state is assigned to the solver backend, and the cursor is advanced to the queue index position corresponding to the first work case in the pending state; The remaining tasks in the ordered task queue are assigned using the cursor position after the advancement, until the current number of concurrent execution units reaches the maximum number of concurrent execution units, or the ordered task queue is exhausted.
3. The method according to claim 2, characterized in that, The method of using a cursor-based amortized scheduling index to distribute the work conditions in the ordered task queue to the solver backend also includes: The system resource status of the local computing node is periodically sampled, and the maximum number of concurrent execution units is dynamically adjusted based on the system resource status of the local computing node.
4. The method according to claim 1, characterized in that, The process of identifying transient I / O errors in the solver's backend output stream using configurable regular expressions, performing exponential backoff retries on these transient I / O error cases, and scheduling loops for the cases after exponential backoff retries includes: The standard error stream in the solver back-end output stream is merged into a standard output stream; The standard output stream is processed uniformly through a single callback, and a session lifecycle event stream is generated in sequence. The transient I / O error type in the session lifecycle event stream can be identified using configurable regular expressions; Get the number of retries for the working condition. If the working condition is a transient I / O error type and the number of retries for the working condition has not reached the maximum number of retries, then reset the working condition status to the pending execution state and return the cursor to the queue index position corresponding to the transient I / O error type working condition. Obtain the waiting time of the current retry round, and after delaying according to the waiting time of the current retry round, reschedule the working condition of the transient I / O error type in a loop until the working condition in a single solution stage reaches the termination state.
5. The method according to claim 1, characterized in that, The process of obtaining the working condition list within a single solution stage during the integrated offshore wind power simulation, and initializing the state of the working condition list within the single solution stage to obtain an ordered task queue, includes: Configure metadata fields for each condition in the condition list within the single solution phase; Filter the working conditions after configuring the metadata fields to obtain a subset of working conditions to be scheduled; The working condition status of the subset of working conditions to be scheduled is initialized to the state to be scheduled, thus obtaining the ordered task queue.
6. The method according to claim 1, characterized in that, Also includes: After completing a preset number of working conditions, according to the adaptive checkpoint interval and time fallback condition, the full state snapshot of the current session is atomically written to the session state checkpoint file; When the work condition scheduling process exits abnormally and restarts, the session state checkpoint file is automatically loaded, and incremental recovery is performed based on the session state checkpoint file until the work condition execution is completed.
7. A multi-condition parallel scheduling system based on an offshore wind power simulation platform, characterized in that, The system includes: The state initialization module is used to obtain the list of working conditions in a single solution stage during the integrated simulation of offshore wind power, and to initialize the state of the list of working conditions in the single solution stage to obtain an ordered task queue. The assignment module is used to assign the working conditions in the ordered task queue to the solver backend using a cursor-based amortized scheduling index. The identification module is used to acquire the solver backend output stream, identify transient I / O error conditions in the solver backend output stream using configurable regular expressions, perform exponential backoff retries on the transient I / O error conditions, and schedule the conditions after exponential backoff retries in a loop. The solver module is used to solve the working conditions in the next solver stage in parallel by using cross-stage pipeline incremental ready semantics when the working conditions in a single solver stage reach the termination state, until the integrated simulation of offshore wind power is completed.
8. An electronic device, characterized in that, include: The system includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes the computer instructions to perform the multi-condition parallel scheduling method based on an offshore wind power simulation platform as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the multi-condition parallel scheduling method based on the offshore wind power simulation platform as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, It includes computer instructions for causing a computer to execute the multi-condition parallel scheduling method based on an offshore wind power simulation platform as described in any one of claims 1 to 6.