Non-intrusive concurrent orchestration method, system and device for heterogeneous black-box programs

CN122526657BActive Publication Date: 2026-09-25湖南省气象信息中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611016715.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-25
Estimated Expiration
2046-07-09

AI Technical Summary

Technical Problem

[0003]将上述传统的物理计算约束模型作为“黑盒节点”接入现代DAG编排系统时,面临着与通用互联网业务截然不同的极端工业约束,主要体现在以下三个维度:(1)物理发散导致的“静默错误”;(2)遗留架构导致的同机并发踩踏与资源枯竭;(3)多源异步观测带来的时窗阻塞与空间边界异常跳变

Benefits of technology

本申请通过解析气象业务流水线中各任务节点的时空依赖关系,并对时空依赖关系中上游数据源中的数据标记关键依赖标识与非关键依赖标识;获取气象业务流水线中各任务流转过程中产生的全局变量、数据指针以及跨节点状态信号,跨节点状态信号包括业务状态标识或降级模式标识;在严格时间窗约束下,根据关键依赖标识和非关键依赖标识,执行任务触发或降级决策;根据任务触发或降级决策触发任务流转,创建当前任务节点的临时工作目录,并基于全局变量、数据指针以及跨节点状态信号,生成临时配置文件;通过重定向临时工作目录将临时配置文件的路径注入黑盒程序的启动环境中,以启动黑盒程序中的目标子进程;将目标子进程的日志流以非阻塞队列方式引流至异步日志泵,通过异步日志泵对抽取的日志流提取目标数据指针和目标跨节点状态信号;日志流为包含目标子进程的标准输出和标准错误的数据流,异步日志泵为用于读取标准输出和标准错误的数据流的独立异步读取协程;基于目标跨节点状态信号和目标子进程退出时的状态信号,判断执行完当前任务节点后的成败状态;通过上报成败状态和目标数据指针驱动下游任务执行,同时销毁当前任务节点的临时工作目录和临时配置文件,以完成黑盒程序的非侵入式并发编排。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122526657B_ABST
    Figure CN122526657B_ABST
Patent Text Reader

Abstract

The application discloses a non-intrusive concurrent arrangement method, system and device for a heterogeneous black box program. The method executes a task trigger or a degradation decision according to a key dependency identifier and a non-key dependency identifier under strict time window constraints. A task flow is triggered, a temporary working directory of a current task node is created, and a temporary configuration file is generated. A path of the temporary configuration file is injected into a starting environment of the black box program through redirection of the temporary working directory, a target sub-process in the black box program is started, a log stream of the target sub-process is guided to an asynchronous log pump in a non-blocking queue mode, a target data pointer and a target cross-node state signal are extracted from the extracted log stream through the asynchronous log pump, a success or failure state after execution of the current task node is judged, and downstream task execution is driven through reporting of the success or failure state and the target data pointer. The application improves the non-intrusive concurrent arrangement efficiency of the heterogeneous black box program, and improves the internal performance and computing throughput of a computer system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a non-intrusive concurrent orchestration method, system and device for heterogeneous black-box programs. Background Technology

[0002] With the rapid integration of meteorological observation technology and artificial intelligence, building multimodal meteorological data lakes to support high-frequency rolling forecasts (such as short-term forecasts) has become a core evolutionary direction for meteorological information systems. In modern high-precision meteorological operational pipelines, it is typically necessary to aggregate and process massive amounts of heterogeneous multi-source data (such as weather radar, satellite cloud images, and automatic weather station observations) in real time, and then entrust them to a Directed Acyclic Graph (DAG) scheduling system for unified computational orchestration. In this hybrid computing architecture, task nodes exhibit high heterogeneity: downstream processes often rely on advanced deep learning prediction models based on graph attention networks (GAT) or Transformers, while upstream key front-end operations such as data quality control (QC), polar-to-Cartesian coordinate interpolation, objective precipitation correction (such as based on inverse probability weighted wave (IPW), and multi-station radar mosaicking heavily depend on traditional physical computation constraint models that have been validated over many years.

[0003] When the aforementioned traditional physical computation constraint model is used as a "black box node" to connect to a modern DAG orchestration system, it faces extreme industrial constraints that are completely different from those of general Internet services. These constraints are mainly reflected in the following three dimensions: (1) "silent errors" caused by physical divergence; (2) concurrent system crashes and resource exhaustion caused by legacy architecture; and (3) time window blocking and abnormal spatial boundary jumps caused by multi-source asynchronous observation. Existing related technologies cannot solve the above problems without modifying the black box source code. Therefore, existing related technologies have problems such as greatly limiting the internal performance and computational throughput of the system and destroying the spatial physical continuity of meteorological products. Summary of the Invention

[0004] This application aims to propose a non-intrusive concurrent orchestration method, system, and device for heterogeneous black-box programs, which can intercept and block "silent errors" caused by physical divergence, solve the problems of concurrent overload and resource exhaustion caused by legacy architecture, ensure the spatial physical continuity of meteorological products, thereby improving the efficiency of non-intrusive concurrent orchestration of heterogeneous black-box programs and improving the internal performance and computing throughput of computer systems.

[0005] In a first aspect, embodiments of this application provide a non-intrusive concurrent orchestration method for heterogeneous black-box programs, the method comprising: The spatiotemporal dependencies of each task node in the meteorological operational pipeline are analyzed, and the data in the upstream data source in the spatiotemporal dependencies are marked with critical dependency identifiers and non-critical dependency identifiers. Acquire global variables, data pointers, and cross-node status signals generated during the workflow of each task in the meteorological business pipeline. The cross-node status signals include business status identifiers or degradation mode identifiers. Under strict time window constraints, task triggering or degradation decisions are executed based on the critical dependency identifier and the non-critical dependency identifier. Based on the task triggering or degradation decision, the task flow is triggered, a temporary working directory for the current task node is created, and a temporary configuration file is generated based on the global variables, the data pointer, and the cross-node status signal. By redirecting the temporary working directory, the path of the temporary configuration file is injected into the startup environment of the black-box program to start the target subprocess in the black-box program. The log stream of the target subprocess is diverted to an asynchronous log pump in a non-blocking queue. The asynchronous log pump extracts the target data pointer and the target cross-node status signal from the extracted log stream. The log stream is a data stream containing the standard output and standard error of the target subprocess. The asynchronous log pump is an independent asynchronous reading coroutine used to read the standard output and standard error data stream. Based on the target cross-node status signal and the target subprocess exit status signal, determine the success or failure status after completing the current task node; By reporting the success or failure status and the target data pointer, the downstream task is driven to execute, while the temporary working directory and temporary configuration file of the current task node are destroyed, so as to complete the non-intrusive concurrent orchestration of the black-box program.

[0006] In some implementations, the step of triggering task flow based on the task triggering or degradation decision, creating a temporary working directory for the current task node, and generating a temporary configuration file based on the global variables, the data pointer, and the cross-node status signal includes: If no local degradation occurs during the task triggering or degradation decision, a globally unique universal identifier is generated for the current task node; a temporary working directory for the current task node is created based on the globally unique universal identifier; the global variables, the data pointer, and the cross-node status signals are used to replace the original configuration placeholders in the static configuration file preset template to generate a temporary configuration file. If a local degradation occurs during task triggering or degradation decision-making, it indicates that a meteorological observation station is missing. The geographical coverage of the missing meteorological observation station is obtained; a globally unique universal identifier is generated for the current task node; a temporary working directory for the current task node is created based on the globally unique universal identifier; an effective coverage mask and boundary transition weight file are generated based on the geographical coverage; a degradation control parameter package containing the effective coverage mask and the boundary transition weight file is constructed; the degradation control parameter package, the global variables, the data pointer, and the cross-node status signal replace the original configuration placeholders in the static configuration file preset template to generate a temporary configuration file.

[0007] In some implementations, generating an effective overlay mask and boundary transition weight file based on the geographic coverage area includes: Set the grid values ​​within the geographic coverage area to invalid values, and set the other grid values ​​except the invalid values ​​to valid values ​​to generate an effective coverage mask; Calculate the minimum vertical distance between any grid point in the planar coordinate system and the effective coverage boundary of the missing meteorological observation station; The physical boundary of the geographic coverage area is extended outward by the minimum vertical distance to generate a boundary transition weight file.

[0008] In some implementations, after generating the temporary configuration file, the method further includes: Read the temporary configuration file; calculate the fusion weight using a smooth transition function based on the degradation control parameter package; the degradation control parameter package also includes available site identifiers and degradation alternative data sources, the available site identifiers are used to mark meteorological observation sites as available, and the degradation alternative data sources are data generated in the previous time segment or neighboring interpolated data; Data on available meteorological observation stations can be obtained through the available station identifiers; The data from the available meteorological observation stations and the data from the downgraded alternative data sources are weighted and fused using the fusion weights to obtain continuous physical field products.

[0009] In some implementations, after starting the target subprocess in the black-box program, the method further includes: Construct an active task mapping table containing the unique identifier of the target subprocess, the temporary working directory, and the startup timestamp; The active task mapping table is traversed at a fixed period. If the target subprocess entity corresponding to the unique identifier of the target subprocess in the active task mapping table does not exist, and the lifespan of the temporary working directory associated with the target subprocess entity exceeds a preset aging time threshold, then the temporary working directory is determined to be an abnormal residual and is deleted.

[0010] In some implementations, the step of diverting the log stream of the target subprocess to an asynchronous log pump in a non-blocking queue, and extracting the target data pointer and target cross-node status signal from the extracted log stream through the asynchronous log pump, includes: Bind non-blocking read descriptors to the standard output and standard error data streams in the log stream; Based on the non-blocking read descriptor, data in the log stream is read with the highest priority through the asynchronous log pump and pushed into memory in a first-in-first-out queue; When the length of the first-in-first-out queue reaches a preset threshold, a backpressure strategy is triggered; the backpressure strategy discards ordinary computation process logs that do not contain a preset control prefix. Perform single-line length validation on the log stream that is not discarded, and retain the remaining log stream with a character length less than or equal to the preset length; The remaining log stream is subjected to fixed prefix matching to obtain candidate lines; Extract the target data pointer and the target cross-node status signal from the candidate rows.

[0011] In some implementations, determining the success or failure status after completing the current task node based on the target cross-node status signal and the target subprocess exit status signal includes: Construct a multi-level state arbitration tree that includes priority adjudication of business semantics, verification of system exceptions and interruptions, and a fallback system exit code; Based on the target cross-node status signal and the target subprocess exit status signal, the success or failure status after the current task node is completed is determined by the multi-level status arbitration tree.

[0012] Secondly, embodiments of this application also provide a non-intrusive concurrent orchestration system for heterogeneous black-box programs, the system comprising: The data parsing module is used to parse the spatiotemporal dependencies of each task node in the meteorological business pipeline, and to mark the data in the upstream data source in the spatiotemporal dependencies as critical dependencies and non-critical dependencies. The data acquisition module is used to acquire global variables, data pointers and cross-node status signals generated during the workflow of each task in the meteorological business pipeline. The cross-node status signals include business status identifiers or degradation mode identifiers. The decision execution module is used to execute task triggering or degradation decisions based on the critical dependency identifier and the non-critical dependency identifier under strict time window constraints. The file generation module is used to trigger task flow based on the task triggering or degradation decision, create a temporary working directory for the current task node, and generate a temporary configuration file based on the global variables, the data pointer, and the cross-node status signal. The subprocess startup module is used to inject the path of the temporary configuration file into the startup environment of the black-box program by redirecting the temporary working directory, so as to start the target subprocess in the black-box program. The data extraction module is used to divert the log stream of the target subprocess to the asynchronous log pump in a non-blocking queue, and extract the target data pointer and target cross-node status signal from the extracted log stream through the asynchronous log pump; the log stream is a data stream containing the standard output and standard error of the target subprocess, and the asynchronous log pump is an independent asynchronous reading coroutine for reading the standard output and standard error data stream; The status judgment module is used to determine the success or failure status after the current task node is completed based on the target cross-node status signal and the target subprocess exit status signal. The orchestration completion module is used to drive the execution of downstream tasks by reporting the success or failure status and the target data pointer, while destroying the temporary working directory and temporary configuration file of the current task node, so as to complete the non-intrusive concurrent orchestration of the black-box program.

[0013] Thirdly, embodiments of this application also provide an electronic device, including at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to execute a non-intrusive concurrent orchestration method for heterogeneous black-box programs as described above.

[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a non-intrusive concurrent orchestration method for heterogeneous black-box programs as described above.

[0015] Compared with the prior art, this application has the following beneficial effects: This application analyzes the spatiotemporal dependencies of each task node in the meteorological operational pipeline and marks the data in the upstream data source with critical and non-critical dependency identifiers. It acquires global variables, data pointers, and cross-node status signals generated during the task flow in the meteorological operational pipeline, including operational status identifiers or degradation mode identifiers. Under strict time window constraints, it executes task triggering or degradation decisions based on critical and non-critical dependency identifiers. Based on the task triggering or degradation decision, it triggers task flow, creates a temporary working directory for the current task node, and generates a temporary configuration file based on global variables, data pointers, and cross-node status signals. Finally, it injects the path of the temporary configuration file into the blacklist by redirecting the temporary working directory. In the startup environment of the black-box program, the target subprocess in the black-box program is started; the log stream of the target subprocess is diverted to the asynchronous log pump in a non-blocking queue, and the target data pointer and target cross-node status signal are extracted from the log stream by the asynchronous log pump; the log stream is a data stream containing the standard output and standard error of the target subprocess, and the asynchronous log pump is an independent asynchronous reading coroutine used to read the standard output and standard error data stream; based on the target cross-node status signal and the status signal when the target subprocess exits, the success or failure status after the current task node is completed is determined; the success or failure status and the target data pointer are reported to drive the execution of downstream tasks, and at the same time the temporary working directory and temporary configuration file of the current task node are destroyed to complete the non-intrusive concurrent orchestration of the black-box program.

[0016] Thus, by executing task triggering or degradation decisions based on critical and non-critical dependency identifiers under strict time window constraints, instead of directly determining task failure, degradation control parameter packages are generated upon degradation. This allows for gradual weighted fusion based on spatial smooth transition functions at the physical influence boundaries of missing observation stations, ensuring the spatial physical continuity of meteorological products. Based on the target cross-node state signals and the state signals when the target subprocess exits, the success or failure status after completing the current task node is determined. This breaks the single assumption of strong dependence on exit codes in traditional scheduling systems, fundamentally preventing erroneous and dirty data from penetrating downstream tasks, thereby intercepting and blocking "silent errors" caused by physical divergence. By redirecting the temporary working directory, the path of the temporary configuration file is injected into the startup environment of the black-box program to start the target child process in the black-box program. The log stream of the target child process is diverted to the asynchronous log pump in a non-blocking queue. After the current task node is completed, the temporary working directory and temporary configuration file of the current task node are destroyed, providing sufficient memory resources for downstream tasks. This solves the problem of concurrent overload and resource exhaustion caused by legacy architecture, and can improve the non-intrusive concurrent orchestration efficiency of heterogeneous black-box programs, and improve the internal performance and computing throughput of the computer system. Attached Figure Description

[0017] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart illustrating an embodiment of the non-intrusive concurrent orchestration method for heterogeneous black-box programs provided in this application; Figure 2 This is a schematic diagram of a multi-level state decision tree process in the best embodiment of the non-intrusive concurrent orchestration method for heterogeneous black-box programs provided in this application; Figure 3 This is a flowchart illustrating the generation of readiness determination and functional degradation control parameters based on time windows in the best embodiment of the non-intrusive concurrent orchestration method for heterogeneous black-box programs provided in this application. Figure 4 This is a schematic diagram of an embodiment of the non-intrusive concurrent orchestration system for heterogeneous black-box programs provided in this application; Figure 5 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application. Detailed Implementation

[0018] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0019] In the description of this application, the use of terms such as "first," "second," etc., is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated.

[0020] In the description of this application, it should be understood that the orientation descriptions, such as up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0021] In the description of this application, it should be noted that, unless otherwise explicitly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.

[0022] To address the problems in related technologies that severely limit the internal performance and computational throughput of the system and disrupt the spatial physical continuity of meteorological products, this application proposes a non-intrusive concurrent orchestration method, system, and device for heterogeneous black-box programs.

[0023] First, let's analyze some of the terms used in this application: Directed Acyclic Graph (DAG) parser: controls the unidirectional flow of task nodes according to dependencies, without generating loops.

[0024] Black-box programs: These specifically refer to meteorological physics calculation executables that were developed a long time ago, have no source code maintenance, and cannot be modified by introducing modern software development kits (SDKs) or application programming interfaces (APIs) (such as radar quality control or jigsaw puzzle algorithms compiled in Fortran / C++). Their input and output formats are extremely fixed (usually only recognizing local configuration files with absolute paths and printing process information through standard output).

[0025] Instance-level configuration sandbox: An isolation mechanism designed to address the "single-machine exclusivity" defect of black-box programs. By generating a temporary working directory and temporary configuration file carrying a globally unique identifier (UUID) for each concurrent grid task, combined with the operating system's redirection technology, it achieves safe concurrency of multiple processes on the same machine and prevents configuration crashes.

[0026] Functional data refers to data structures that are not intended for human reading but rather "have a predetermined technical purpose and automatically control subsequent processing within a technical system." In this embodiment, the degradation control parameter package containing "effective overlay mask file path" and "boundary transition weight" is a typical example of functional data.

[0027] Silent errors: These refer to situations where a serious business error has occurred in the underlying meteorological algorithm (such as matrix calculation divergence or empty input), but it has not triggered a segmentation fault or crash in the operating system, causing the operating system to still return to a normal exit state (ExitCode=0). This embodiment uses an arbitration mechanism that prioritizes business semantics to block such errors.

[0028] OOMKiller and SIGKILL: Self-protection mechanisms of the Linux operating system kernel. When a server runs out of memory, the kernel sends a SIGKILL (-9) signal to instantly kill processes that are consuming excessive memory. Because this signal cannot be caught by the program, it causes conventional "process termination cleanup code" to fail.

[0029] Orphan directory: A concurrent temporary working directory left on the host machine due to OOM forced termination or unexpected crash of the agent process, which prevented the normal execution of cleanup logic.

[0030] Cleaner: A daemon thread within the agent node, independent of the main control flow. It periodically inspects the system process tree and active task mapping table, and is specifically responsible for forcibly reclaiming orphan directories to prevent host innode and disk resource exhaustion.

[0031] Asynchronous Log Pump: An independent asynchronous read coroutine in the agent node used to extract the standard output (stdout) of child processes, preventing read blocking.

[0032] Backpressure mechanism: a system flow control strategy. When the instantaneous meteorological calculation log surge is about to fill the bounded queue in memory, the system actively discards non-critical debug logs to ensure that the operating system's default pipe buffer is not filled, thereby preventing child processes from being suspended by the kernel (apparent death / deadlock).

[0033] Regular expression backtracking denial-of-service attack: When using complex regular expressions to match extremely long or specific abnormal text, the regular expression engine may get stuck in exponential backtracking calculations, causing CPU resources to be exhausted instantly. This embodiment uses " Length checksum The funnel-shaped probe strategy of "fixed prefix initial screening restricted regularization" completely blocks this risk.

[0034] Objective precipitation correction (such as the IPW algorithm): A classic physical constraint algorithm used in traditional meteorological operations to correct errors in precipitation grid points from radar or numerical weather prediction. It is extremely sensitive to input extreme values ​​and is prone to "silent errors".

[0035] Boundary Gradual Fusion: In radar mosaic or space weather field processing, when local data is missing, a spatial smoothing transition function is used to perform weighted calculations at the boundary between the missing area and the valid area, eliminating "cliff-like" physical anomaly jumps.

[0036] Multimodal fusion model: This can be a cutting-edge architecture in the field of meteorological artificial intelligence (AI). Downstream nodes use deep learning models such as graph attention networks (GAT) to absorb the output of upstream physical processes (i.e., the result data output by upstream physical processes). This embodiment enables this type of AI model to automatically read masks and complete the dynamic alignment and assembly of physical features by transmitting functional data.

[0037] Reference Figure 1 This application provides a flowchart illustrating a non-intrusive concurrent orchestration method for heterogeneous black-box programs. This method is applied to electronic devices, such as servers or mobile terminals. Figure 1 As shown, the non-intrusive concurrent orchestration method for heterogeneous black-box programs may include the following steps S101 to S108.

[0038] Step S101: Analyze the spatiotemporal dependencies of each task node in the meteorological business pipeline, and mark the data in the upstream data source in the spatiotemporal dependencies as critical dependencies and non-critical dependencies.

[0039] In this step, the spatiotemporal dependencies of each task node in the meteorological operational pipeline are analyzed using a directed acyclic graph parser. The meteorological operational pipeline can range from radar-based data quality control to multimodal fusion forecasting, and also includes a series of data stream processing procedures. In other words, the entire meteorological data processing process can be viewed as a pipeline for handling meteorological operations.

[0040] When resolving spatiotemporal dependencies, the directed acyclic graph parser declares that the data in the upstream data source in the spatiotemporal dependency relationship will be specifically divided into a critical dependency set (corresponding to the critical dependency identifier) ​​and an optional dependency set (corresponding to the non-critical dependency identifier).

[0041] Step S102: Obtain global variables, data pointers, and cross-node status signals generated during the workflow of each task in the meteorological business pipeline. The cross-node status signals include business status identifiers or degradation mode identifiers.

[0042] In this step, the context relation library can be used to store global variables, data pointers, and cross-node status signals generated during the workflow of each task in the meteorological operational pipeline. Global variables may include globally unique universal identifiers (UUIDs) representing task execution instances, timestamp parameters of forecast rolling time windows, etc. Data pointers may include absolute file paths of meteorological products output by upstream task nodes, multimodal tensor dimension information for downstream tasks to read, etc. Cross-node status signals may include operational status identifiers (PHYSICS_DIVERGE) of the representation matrix extracted by regularization probes, or degradation mode identifiers (PARTIAL_MISSING) representing local data loss.

[0043] Step S103: Under strict time window constraints, execute task triggering or degradation decisions based on critical dependency identifiers and non-critical dependency identifiers.

[0044] In this step, in an asynchronous observation scenario with strict time window constraints, flexible task triggering or degradation decisions are made based on the arrival status of critical dependency identifiers and non-critical dependency identifiers.

[0045] The aforementioned key dependency identifier (i.e., the identifier corresponding to the core dependency set) can refer to the radar station number (i.e., the meteorological observation station number) covering the core forecast area. If this part of the underlying data (i.e., the missing core radar station number) does not arrive within the operational time window (i.e., the strict time window), it will result in the loss of core physical field features, and the ready fence controller will forcibly determine that the current grid mosaicking task has failed.

[0046] The aforementioned strict time window constraint can be a pre-set time, such as a 6-minute strict time window for rolling updates.

[0047] The aforementioned non-critical dependency identifiers (i.e., identifiers corresponding to optional dependency sets) can refer to radar station numbers located in the edge region of the grid mosaic or those that play an auxiliary smoothing role. For the missing data corresponding to such identifiers, the system will not trigger a blockage in the overall meteorological operational pipeline. Instead, it will transform the status of the missing radar station number corresponding to this type of identifier into a functional degradation control parameter package (i.e., issuing the effective overlay mask file path and the boundary transition weight file path), driving the downstream legacy mosaic black-box algorithm or multimodal artificial intelligence (AI) model to perform spatial gradient fusion processing.

[0048] Step S104: Trigger task flow based on task triggering or degradation decision, create a temporary working directory for the current task node, and generate a temporary configuration file based on global variables, data pointers, and cross-node status signals.

[0049] In this step, if no local degradation occurs during task triggering or degradation decision, a globally unique universal identifier is generated for the current task node; a temporary working directory for the current task node is created based on the globally unique universal identifier; and the original configuration placeholders in the static configuration file preset template are replaced with global variables, data pointers, and cross-node status signals to generate a temporary configuration file. If a local degradation occurs during task triggering or degradation decision-making, it indicates that a meteorological observation station is missing. The process involves: obtaining the geographical coverage of the missing meteorological observation station; generating a globally unique universal identifier for the current task node; creating a temporary working directory for the current task node based on the globally unique universal identifier; generating an effective coverage mask and boundary transition weight file based on the geographical coverage; constructing a degradation control parameter package containing the effective coverage mask and boundary transition weight file; and replacing the original configuration placeholders in the static configuration file's preset template with the degradation control parameter package, global variables, data pointers, and cross-node status signals to generate a temporary configuration file.

[0050] The aforementioned globally unique universal identifier serves as the unique identifier for the current task node.

[0051] The above describes creating a temporary working directory for the current task node based on a globally unique universal identifier (UUID). This could be something like ` / tmp / caelus_run_`. <uuid>The temporary working directory under / directory.

[0052] The aforementioned static configuration file preset template can be a pre-set static configuration file template, such as a template in .ini or .nml format.

[0053] After generating the temporary configuration file, the above may also include: Read the temporary configuration file; calculate the fusion weights using a smooth transition function based on the degradation control parameter package; the degradation control parameter package also includes available site identifiers and degradation alternative data sources, where available site identifiers are used to mark available meteorological observation sites, and degradation alternative data sources are data generated in the previous time segment or neighboring interpolated data; obtain available meteorological observation site data through available site identifiers; and use fusion weights to perform weighted fusion of the available meteorological observation site data and the data in the degradation alternative data sources to obtain the continuous physical field product.

[0054] The above calculation of fusion weights using a smooth transition function includes: ; in, Indicates the fusion weight. Represents any grid point in a planar coordinate system The minimum vertical distance between the effective coverage boundary of the missing meteorological observation station and the station. Indicates the width of the boundary transition zone. It represents the sine square function.

[0055] The above generates effective overlay masks and boundary transition weight files based on geographical coverage, including: Set the grid values ​​within the geographic coverage area to invalid values ​​and set other grid values ​​to valid values ​​to generate an effective coverage mask; calculate the minimum vertical distance between any grid point in the planar coordinate system and the effective coverage boundary of the missing meteorological observation station; extend the physical boundary of the geographic coverage area outward by the minimum vertical distance to generate a boundary transition weight file.

[0056] The minimum vertical distance between any grid point in the above-mentioned planar coordinate system and the effective coverage boundary of the missing meteorological observation station can be calculated using a method known to those skilled in the art. This embodiment will not describe this in detail.

[0057] Step S105: Inject the path of the temporary configuration file into the startup environment of the black-box program by redirecting the temporary working directory, so as to start the target child process in the black-box program.

[0058] In this step, a redirection technique is used to redirect the temporary working directory and inject the path of the temporary configuration file into the startup environment of the black-box program in order to start the target child process in the black-box program.

[0059] The aforementioned redirection techniques include: 1. Configure input redirection: By modifying the target child process's environment variables (such as redirecting the temporary working directory) or by using operating system-level symbolic links / mounting technology, the temporary configuration file is disguised as a fixed absolute path required by the target black-box program and injected into its startup environment.

[0060] 2. Output stream redirection: Deprive the target child process of the permission to directly write to the standard console or default local files, and forcibly redirect its standard output (stdout) and standard error (stderr) data streams (i.e. log streams or output streams) to the asynchronous log pump inside the agent node in the form of non-blocking pipes (PIPE).

[0061] After starting the target subprocess in the black-box program, the above may also include: Construct an active task mapping table containing the unique identifier of the target subprocess, the temporary working directory, and the startup timestamp; traverse the active task mapping table at fixed intervals. If the target subprocess entity corresponding to the unique identifier of the target subprocess in the active task mapping table does not exist, and the lifespan of the temporary working directory associated with the target subprocess entity exceeds the preset aging time threshold, then the temporary working directory is determined to be an abnormal remnant and is deleted.

[0062] Step S106: The log stream of the target subprocess is diverted to the asynchronous log pump in a non-blocking queue. The target data pointer and target cross-node status signal are extracted from the log stream by the asynchronous log pump. The log stream is a data stream containing the standard output and standard error of the target subprocess. The asynchronous log pump is an independent asynchronous reading coroutine used to read the standard output and standard error data stream.

[0063] In this step, non-blocking read descriptors are bound to the standard output and standard error data streams in the log stream; based on the non-blocking read descriptors, data in the log stream is read with the highest priority via an asynchronous log pump and pushed into memory in a first-in-first-out queue; when the length of the first-in-first-out queue reaches a preset threshold, a backpressure strategy is triggered; ordinary computation process logs that do not contain a preset control prefix are discarded using the backpressure strategy; single-line length verification is performed on the log streams that are not discarded, and the remaining log streams with character lengths less than or equal to the preset length are retained; fixed prefix matching is performed on the remaining log streams to obtain candidate lines; target data pointers and target cross-node status signals are extracted from the candidate lines.

[0064] The aforementioned preset control prefix and fixed prefix can be strings that begin with "!!CAELUS_CTRL::".

[0065] Step S107: Based on the target cross-node status signal and the target subprocess exit status signal, determine the success or failure status after completing the current task node.

[0066] In this step, a multi-level state arbitration tree is constructed, which includes priority adjudication of business semantics, verification of system exception interruption, and fallback of system exit code. Based on the target cross-node state signal and the state signal when the target subprocess exits, the success or failure status after the current task node is completed is determined through the multi-level state arbitration tree.

[0067] The aforementioned business semantic priority decision can be achieved by the proxy node querying whether the current instance context records a specific business status identifier (i.e., a target cross-node status signal) extracted from the log stream. For example, does it detect failure identifiers such as abnormal business physical calculations / empty grids? If the extracted identifier representing internal matrix iteration divergence is !!STATUS=PHYSICS_DIVERGE!!, or the identifier indicating all input data is missing values ​​is !!STATUS=EMPTY_GRID!!, then regardless of the final exit code returned by the operating system (even if ExitCode=0), the proxy node forcibly marks the instance as a business failure (i.e., task execution failure) and immediately melts down the branch in the directed acyclic graph (DAG), completely blocking the flow of dirty data into the downstream multimodal meteorological data lake (i.e., downstream tasks). If a clear business success flag is extracted (e.g., !!STATUS=BUSINESS_SUCCESS!!), then it is marked as a successful task execution.

[0068] The aforementioned system exception interruption verification can involve the agent node checking the termination signal of the child process if no business status identifier is extracted from the log stream. Under the Portable Operating System Interface (POSIX) system specification, if a child process is forcibly terminated by the system kernel (e.g., due to timeout or being forcibly killed by OOMKiller due to host memory exhaustion), the return code is usually a specific negative value (e.g., ReturnCode < 0). When such an abnormal interruption signal is hit, the agent node determines that the task execution has failed and records the underlying hardware or system-level exception cause.

[0069] The aforementioned system exit code fallback mechanism allows the proxy node to revert to the traditional judgment method—reading the operating system's ExiCode—only when the first-level business semantics are missing and there are no system anomaly signals at the second level. If the ExitCode is 0, the task is considered successful, triggering subsequent DAG nodes; if the ExitCode is non-zero, the task is considered to have failed.

[0070] Step S108: Drive the execution of downstream tasks by reporting success or failure status and target data pointer, and destroy the temporary working directory and temporary configuration file of the current task node to complete the non-intrusive concurrent orchestration of the black-box program.

[0071] In this step, the proxy node reports the final success or failure status and the extracted target data pointer to the orchestration center context library to drive downstream tasks. Simultaneously, the proxy node proactively destroys the temporary working directory and temporary configuration files created for the current task node, releasing the host machine's storage and file index node resources. After determining the final success or failure status, the non-intrusive concurrent orchestration of the black-box program corresponding to the current task node is completed, and the success or failure status and target data pointer are reported to drive downstream tasks. Once all task nodes have completed, the non-intrusive concurrent orchestration of the entire black-box program is finished.

[0072] To facilitate understanding by those skilled in the art, a set of preferred embodiments is provided below: When traditional physical computing models are integrated into modern DAG orchestration systems as "black box nodes," they face extreme industrial constraints that are drastically different from those of general internet IT operations. These constraints are mainly reflected in the following three dimensions: 1. "Silent error" caused by physical divergence.

[0073] Traditional scientific computing programs typically follow standard operating system process exit semantics (i.e., returning ExitCode=0 upon successful completion). However, when processing meteorological grid data, extreme value perturbations in local input data can easily cause matrix iterations within the underlying computing program to diverge or produce empty grid outputs of all zeros. In such cases, even though the operational results severely violate physical principles, the operating system still determines that the process has successfully exited. If the scheduling system blindly relies on the system exit code to advance the Directed Acyclic Graph (DAG), dirty data will directly penetrate into the downstream multimodal data lake (i.e., downstream tasks), causing widespread forecast distortion.

[0074] 2. Legacy architecture leads to concurrent server overload and resource depletion.

[0075] To meet the timeliness requirements of large-scale (e.g., cross-provincial or nationwide) gridded processing, computational tasks typically need to be divided into hundreds of sub-grids and executed concurrently on modern multi-core server clusters. However, legacy black-box binary programs, designed primarily for single-machine serial scenarios, often have their input parameters hard-coded to local configuration files in fixed paths (such as config.ini in the current working directory). When the system attempts to launch multiple black-box instances concurrently on the same compute node, configuration file read / write operations can easily overwrite and disrupt the process, preventing the effective reuse of modern cloud-native computing power and leading to resource exhaustion.

[0076] 3. Time window blocking and spatial boundary jumps caused by multi-source asynchronous observation.

[0077] Taking high-frequency radar mosaicking as an example, due to network jitter or mechanical failures at various observation stations, radar base data often exhibits spatiotemporal misalignment when reaching the central control plane. Traditional DAG scheduling employs a "Wait-All" mechanism, where a single station's delay can block the entire global mosaic's 6-minute rolling cycle. If a simple "ignore missing meteorological observation stations" strategy is adopted, the downstream black-box program (i.e., downstream task) will experience drastic spatial boundary anomalies at the boundary between the influence of missing and normal meteorological observation stations during mosaic generation, severely disrupting the continuity of the physical field.

[0078] Existing general workflow orchestration techniques and traditional radar meteorological data processing systems objectively suffer from the following technical shortcomings when dealing with the distributed scheduling of legacy black-box physical computing programs: 1. Physical calculation "silent errors" lack business-level state blocking capabilities, which can easily lead to dirty data pollution downstream.

[0079] Existing general-purpose workflow orchestration systems (such as CI / CD agents or traditional DAG schedulers) rely heavily on operating system-level process exit codes (ExitCodes) when determining the execution status of external nodes. However, legacy scientific computing black boxes (such as objective precipitation correction models) can cause internal matrix divergence after encountering local grid extremum perturbations, or when upstream base data is empty. In these cases, the operating system will still terminate the process in a normal state (e.g., ExitCode=0). Current technologies cannot deeply detect such business-level failures without intruding into the source code, leading to abnormal physical computation products flowing into downstream multimodal data lakes as normal results, causing widespread cascading computational distortions.

[0080] 2. The fixed local configuration reading logic of legacy programs leads to high concurrency overload and idle computing power on the same machine.

[0081] To improve the processing efficiency of large-scale grid data, modern clusters typically need to concurrently schedule a large number of identical algorithm instances on the same compute node. However, traditional closed-source physical programs often hardcode fixed local configuration input paths (such as forcibly reading fixed .ini files in the current working directory). Existing orchestration tools fail to provide matching instance-level dynamic isolation mechanisms, leading to file read / write overwriting and crosstalk when multiple task instances are executed concurrently. This degrades modern multi-core servers to single-machine, single-threaded use, severely limiting the internal performance and computational throughput of the computer system.

[0082] 3. In the face of missing asynchronous observation data, there is a lack of automated degradation control surfaces that can ensure the continuity of the physical field.

[0083] In the real-time aggregation of multi-source high-frequency meteorological data (such as radar mosaics), delayed arrival or missing data from a single station is a frequent occurrence. Existing related technologies typically employ a "wait-all" approach, which can lead to prolonged pipeline blockages, or a simple "ignore missing meteorological observation stations" strategy. The latter only skips the data at the operational level and fails to translate the missing status into technical parameters that can control the downstream black-box execution logic. This results in severe abrupt changes in the final physical field (such as radar mosaic products) at the boundary between the coverage of missing and normal meteorological observation stations, disrupting the spatial physical continuity of meteorological products.

[0084] 4. For high-throughput industrial log streams, there is a lack of underlying pipeline deadlock prevention and computing resource protection mechanisms.

[0085] When using the standard output (stdout) stream as a non-intrusive control channel, meteorological algorithms can spew out massive amounts of grid processing logs in a short period of time. Existing wrapper techniques typically employ simple synchronous line-by-line reading without considering the capacity limit of the operating system's pipe buffer, making them highly susceptible to causing child process blocking and system deadlocks under log flooding. Furthermore, they lack safety filtering strategies to prevent backtracking disasters caused by regular expressions, and complex text input can easily exhaust control plane CPU resources.

[0086] To address the technical problems existing in the aforementioned related technologies, the purpose of this embodiment is to provide a non-intrusive concurrent orchestration method for heterogeneous black-box programs. The specific objectives of this embodiment are as follows: 1. Construct a multi-level state arbitration mechanism that takes precedence over system exit codes to achieve precise interception of silent errors. This embodiment aims to intercept the output stream of black-box processes in real time through proxy nodes, extract business status signals, and establish a multi-level state arbitration tree with "business semantics first, system anomalies second, and exit codes as a fallback." Even if the operating system successfully determines the error, workflow circuit breaking can be implemented based on the extracted business anomaly identifier, thereby blocking physical computation divergence or empty set penetration and ensuring the data quality of the multimodal data lake.

[0087] 2. By leveraging instance-level temporary configuration and a fallback cleanup mechanism, the single-machine concurrency bottleneck of legacy programs is overcome. This embodiment aims to dynamically generate a temporary configuration file with a unique identifier for each execution instance through the cooperation of the orchestration center and proxy nodes. It also utilizes directory isolation or path mapping to replace the startup parameters of the black-box program, allowing it to run "unnoticeably" in the concurrent sandbox. Simultaneously, combined with the Janitor anomaly cleanup mechanism, it effectively prevents concurrent file corruption and host machine residual contamination, significantly improving the resource utilization and processing throughput within the computer system.

[0088] 3. Introducing time-window control based on functional degradation instruction packages (i.e., functional degradation control parameter packages) to ensure the physical continuity of the spatial field under local missing data. This embodiment aims to automatically generate functional degradation control parameter packages containing effective overlay masks and boundary transition weights by the control surface for missing non-critical observation data (i.e., data with non-critical dependency identifiers) under strict time window constraints, and use these packages to drive downstream black-box programs. This allows downstream algorithms to automatically perform gradual fusion processing at the missing boundaries based on these parameter packages, ensuring business timeliness while completely eliminating spatial field jumps caused by traditional hard missing data.

[0089] 4. Design a log pump with backpressure control and a low-cost probe anti-backtracking strategy to ensure the robustness of the high-frequency scheduling control surface. This embodiment aims to ensure that, by optimizing the independent asynchronous read queue and pipeline buffer, combined with a defensive parsing syntax of "fixed prefix initial screening and controlled regular expression matching," the proxy node will not experience pipeline blockage deadlock or fall into CPU backtracking disaster when faced with massive untrusted weather operation logs, thus ensuring the high availability of the overall cluster scheduling system.

[0090] This embodiment's technical solution primarily relies on a distributed scheduling system comprising an orchestration center and proxy nodes. In actual meteorological data processing operations (i.e., meteorological operational pipelines, such as objective precipitation correction based on global datasets, low-altitude visibility prediction including physical constraints, etc.), the orchestration center is responsible for parsing and distributing the topology of directed acyclic graph tasks; proxy nodes are deployed on each Linux server worker node, responsible for packaging and executing legacy physics computation programs, real-time parsing of output streams, multi-level state arbitration, and dynamic construction of concurrent sandbox environments without modifying the original C / C++ / Fortran code. This embodiment's technical solution specifically includes: 1. Overall system architecture and core control flow.

[0091] (1) Arrange the central layer.

[0092] The orchestration center layer is deployed on the main control server and is responsible for scheduling and maintaining the status of global meteorological data processing tasks. Its main components include: Directed Acyclic Graph Parser: Used to resolve the spatiotemporal dependencies of each task node in meteorological operational pipelines (such as radar-based data quality control to multimodal fusion forecasting).

[0093] Context Relationship Library: Used to store global variables (e.g., globally unique universal identifier UUID representing task execution instance, timestamp parameter of forecast rolling time window), data pointers (e.g., absolute file path of meteorological products output by upstream physical nodes (i.e. upstream task nodes), multimodal tensor dimension information for downstream deep learning models (i.e. downstream tasks) and cross-node status signals (e.g., business status identifier (PHYSICS_DIVERGE) of the representation matrix extracted by regularization probe, or degradation mode identifier (PARTIAL_MISSING) representing local data missingness).

[0094] Ready Fence Controller: Used in asynchronous observation scenarios with strict time window constraints to execute flexible task triggering or degradation decisions based on the arrival status of critical and non-critical dependency identifiers. Critical and non-critical dependency identifiers are typically statically declared in the task configuration file during the DAG task orchestration phase (i.e., during the resolution process of the directed acyclic graph resolver), or obtained based on the geospatial topology of meteorological observation stations. Taking a high-frequency radar intelligent mosaic system as an example: Key dependency identifiers (i.e., identifiers corresponding to the core dependency set): These typically refer to the radar station numbers (i.e., meteorological observation station numbers) covering the core forecast area. If this underlying data does not arrive within the operational time window, it will result in the loss of core physical field features, and the ready fence controller will force the current grid mosaicking task to fail.

[0095] Non-critical dependency identifiers (i.e., identifiers corresponding to optional dependency sets): These typically refer to radar station numbers located in the edge region of the jigsaw puzzle or those that play an auxiliary smoothing role. For missing data corresponding to such identifiers, the system will not trigger overall pipeline blocking. Instead, it will transform this missing state into a functional degradation instruction package (i.e., a degradation control parameter package containing the paths to the valid overlay mask file and the boundary transition weight file), driving downstream legacy jigsaw puzzle black-box algorithms or multimodal artificial intelligence (AI) models to perform spatial gradient fusion processing.

[0096] (2) Proxy node layer.

[0097] The proxy node layer, deployed as a resident daemon process on various distributed worker nodes with computing power (such as Linux server worker nodes), is the core control plane for implementing the non-intrusive orchestration in this embodiment. Its main components include: Instance-level configuration sandbox: Before task execution, it is responsible for dynamically rendering and creating an independent temporary working directory and temporary configuration file for the currently executing task instance based on information such as global variables, data pointers and cross-node status signals issued by the orchestration center layer; and it has an embedded anomaly cleaner to reclaim residual resources on the host machine.

[0098] Process wrapper and starter: Responsible for launching the target black-box program as a child process (i.e., the target child process) and taking over its environment space and standard output / error stream through redirection techniques. These redirection techniques include: 1) Configure input redirection: By modifying the target child process environment variables (such as redirecting the temporary working directory) or by using operating system-level symbolic links / mounting technology, the temporary configuration file is disguised as a fixed absolute path required by the target black box program and injected into its startup environment.

[0099] 2) Output stream redirection: Deprive the target child process of the permission to directly write to the standard console or default local files, and forcibly redirect its standard output (stdout) and standard error (stderr) data streams (i.e. log streams or output streams) to the asynchronous log pump inside the agent node in the form of non-blocking pipes (PIPE).

[0100] Deadlock-proof asynchronous log pump: It adopts a non-blocking queue combined with a backpressure mechanism to safely extract the output stream of the target child process under the flood of massive operation logs, and prevents the operating system pipe buffer from being exhausted (i.e., prevents the pipe from locking up).

[0101] Low-cost probe and parsing engine (i.e., low-cost probe anti-backtracking strategy): Responsible for performing "funnel-style" prefix screening and controlled regular expression matching on the extracted log stream to extract target data pointers and target cross-node status signals (e.g., business status identifiers) to prevent backtracking disasters. The specific matching process is as follows: First, execute... The complexity involves single-line length validation, discarding abnormally long text; then, execution... The complexity is fixed-prefix exact matching (i.e., only log lines containing a specific leading identifier are allowed); finally, only for candidate lines that pass the initial screening, the restricted mode linear time regularization engine is called to extract the target cross-node status signal or target data pointer.

[0102] Multi-level state arbitration tree: Responsible for executing failure judgments based on business state priority when a child process terminates. Its fusion arbitration method is as follows: interception is performed using a three-layer decision tree with decreasing priority. The first level prioritizes checking the target cross-node state signal of the business layer. If a failure flag such as business physical calculation divergence is detected, the circuit breaker is forcibly broken and the task execution is judged as failed, regardless of the system state. The second level checks system-level interrupt signals (such as timeout termination). Only the third level rolls back to using the normal operating system exit code as a fallback decision, thus executing a strict task node success or failure judgment.

[0103] (3) Black-box execution layer (containing the target subprocess).

[0104] The black-box execution layer resides at the bottom of the system architecture, initiated and non-intrusively managed by the proxy node layer through a target subprocess startup / control mechanism. This layer contains uncontrolled, closed-source legacy executable programs (including radar-based data quality control programs or multimodal fusion models, such as objective precipitation correction binaries compiled in Fortran or C++, or encapsulated jigsaw puzzle algorithm modules) that actually perform meteorological physics calculations or deep learning inferences. These programs constitute the specific physical constraints of the implementation environment in this embodiment. 1) Closed source and non-intrusive: It is impossible to make it actively access the software development kit (SDK) or application programming interface (API) of modern scheduling frameworks by modifying the source code.

[0105] 2) Static configuration dependency: Input parameters can only be obtained by reading static configuration files (such as ini or nml format) in a fixed path on the local host machine.

[0106] 3) Unreliable process semantics: It is susceptible to business-level errors such as calculation divergence caused by local data extrema, but the operating system can still terminate the process in a normal state.

[0107] 2. Specific process of non-intrusive orchestration method for heterogeneous black-box programs.

[0108] Step S110: Task parsing and time window interception. The spatiotemporal dependencies of the meteorological business pipeline are parsed by the directed acyclic graph parser inside the orchestration center; when the rolling time window is reached, the ready fence controller executes task triggering or local degradation decisions based on the arrival status of critical dependency identifiers and non-critical dependency identifiers.

[0109] Step S120: Instance Sandbox Construction and Rendering. The orchestration center triggers task flow based on the aforementioned task triggering or partial degradation decision, distributing the global variables, data pointers, and cross-node status signals required by the current task node to the corresponding proxy node. If there is no partial degradation, the proxy node generates a globally unique universal identifier (UUID) for this task execution as the instance identifier; subsequently, the proxy node stores this identifier in the host file system (e.g., / tmp / caelus_run_). <uuid>A separate temporary working directory is created under the ` / ` directory. Based on the distributed global variables and data pointers, the original configuration placeholders in the target black-box program's required static configuration file preset template (i.e., a pre-set static configuration file template, such as a .ini or .nml format template) are replaced, dynamically generating a temporary configuration file specific to this task instance. If there is a local degradation, the proxy node needs to generate a functional degradation control parameter package. Based on the degradation control parameter package, global variables, data pointers, and cross-node status signals, the original configuration placeholders in the target black-box program's required static configuration file preset template (such as a .ini or .nml format template) are replaced, dynamically generating a temporary configuration file specific to this task instance.

[0110] It should be noted that the globally unique universal identifier (UUID) generated in this embodiment adopts a method known to those skilled in the art. For example, a combination of random numbers and timestamps can be used to generate the globally unique universal identifier. This embodiment does not describe or limit this in detail.

[0111] Step S130: Process Packaging and Input Redirection. The agent node does not modify the source code of the target black-box program (such as a physics computation binary compiled by Fortran or C++). The agent node injects the temporary configuration file path generated in step S120 into the startup environment of the black-box program by modifying environment variables (such as redirecting the temporary working directory) or by using operating system-level mounts / symbolic links. The agent node starts the target black-box program (i.e., the target child process) and completely takes over the standard output (stdout) and standard error (stderr) data streams of the target child process in a non-blocking pipe manner (i.e., a non-blocking queue), diverting them to an asynchronous log pump with backpressure control.

[0112] Step S140: Output Interception and Feature Extraction. During the execution of the target subprocess, the asynchronous log pump inside the agent node extracts pipeline data (i.e., data in the log stream) in real time. For the extracted pipeline data, the agent node executes a preset low-cost probe anti-backtracking strategy to extract two types of key structured semantic entities from the massive unstructured computation log stream: one type is the target data pointer mapped to the subsequent workflow (such as output product path, tensor dimension, etc.); the other type is the target cross-node status signal mapped to the actual execution status of the business (such as anomaly control indicators representing business-level warnings such as physical model divergence and empty input mesh).

[0113] Step S150: Multi-level state arbitration and success / failure determination. After the black-box process (i.e., the target sub-process) terminates (whether normally or abnormally), the proxy node, based on the structured semantic entities extracted in step S140 and combined with the system-level state information at the time of the target sub-process's exit, enters the multi-level state arbitration tree module. The proxy node determines the final success or failure of the task instance at the business flow level according to the principle that business state takes precedence over exit code (a multi-level state arbitration tree can be used for determination).

[0114] Step S160: Sandbox Resource Destruction and Context Reporting. The proxy node reports the final success or failure status determined in step S150 and the target data pointer extracted in step S140 to the orchestration center context library to drive downstream tasks. Simultaneously, the proxy node actively destroys the temporary working directory and temporary configuration files created in step S120, releasing the host machine's storage and file index node resources. After determining the final success or failure of the task instance at the business flow level, the system completes the non-intrusive concurrent orchestration and security control of the black-box program corresponding to the current target subprocess, and outputs the extracted result data pointer or degradation status to drive downstream black-box programs (i.e., downstream tasks) to perform subsequent nodes such as multimodal fusion prediction. Once all executions are complete, the non-intrusive concurrent orchestration of the black-box program is finished.

[0115] 3. Asynchronous log pump with backpressure control and multi-level state arbitration tree.

[0116] In large-scale scheduling of meteorological algorithms, this embodiment faces two major underlying engineering challenges: first, the instantaneous log surge can easily fill the limited pipeline buffer of the operating system, leading to deadlock of child processes; second, simply relying on the operating system exit code cannot intercept "silent errors" in complex physical calculations.

[0117] (1) Anti-deadlock asynchronous log pump and security filtering mechanism.

[0118] Asynchronous Log Pump Pumping and Backpressure Control: The agent node uses an asynchronous log pump thread / coroutine independent of the main control flow. When the black-box program subprocess (i.e., the target subprocess) starts, the agent node binds a non-blocking read descriptor to stdout / stderr=PIPE. The asynchronous log pump reads pipeline data (i.e., data in the log stream) with the highest priority and pushes it into a bounded first-in-first-out queue (i.e., a FIFO queue) in memory. When the FIFO queue length (i.e., the queue backlog) reaches a preset high water level (i.e., a preset threshold, such as 85% of the capacity), the system triggers a backpressure strategy. When the backpressure strategy is triggered, the asynchronous log pump does not pause or block the reading of underlying pipeline data (to ensure that the 64KB pipeline buffer of the underlying operating system is continuously cleared to maintain the operation of the black-box program). Instead, it initiates an active discard mechanism for newly read data: the system will directly discard ordinary meteorological calculation process logs (such as Debug / Info logs) that do not contain preset control prefixes (such as !!CAELUS_CTRL::), and only push critical control commands into the queue to ensure that special log lines containing control commands are never lost. This mechanism fundamentally avoids the target child process 'feigning death' or deadlock caused by the Linux system's default pipe buffer being full.

[0119] Low-cost probe backtracking prevention strategy: To prevent regular expression backtracking denial-of-service attacks from overwhelming the proxy node's CPU, the proxy node implements a "low-cost probe priority" strategy. The first step is to execute on the log streams that have not been discarded, selected through the backpressure strategy. The complexity of single-line length validation (e.g., if the single-line length validation character length is greater than 1024 characters, discard the exception stack output exceeding 1024 characters (i.e., discard the long exception text to prevent memory overflow); if the single-line length validation character length is less than or equal to 1024 characters, proceed to the next step) yields the remaining log stream; the second step executes on the remaining log stream. The first step involves a fixed-prefix exact match (e.g., only allowing strings starting with "!!CAELUS_CTRL::" and proceeding to the next step; strings not starting with "!!CAELUS_CTRL::" are ignored as ordinary meteorological calculation process logs) to obtain candidate rows. The second step triggers restricted pattern matching, which only calls a linear time regularization engine with no risk of backtracking disaster for candidate rows that pass the prefix exact match, to extract target cross-node status signals (e.g., business status identifiers) or target data pointers.

[0120] The aforementioned linear time regularization engine, free from the risk of backtracking disasters, can be an engine that employs a non-backtracking algorithm and guarantees that the matching time is proportional to the length of the input text. This fundamentally eliminates the possibility of CPU overload caused by malicious input.

[0121] The aforementioned non-backtracking algorithm is a method known to those skilled in the art, and will not be described in detail in this embodiment.

[0122] This embodiment can prevent the target child process from "freezing" or deadlocking due to the instantaneous filling of the operating system's default Linux pipe buffer by massive amounts of unstructured computation logs, block regular expression backtracking (ReDoS) attacks that may be caused by complex malicious text, and ensure the high availability of the proxy node's CPU computing resources.

[0123] (2) Intercepting the multi-level state arbitration tree of silent errors.

[0124] To address the issue of the traditional assumption that "ExitCode=0 means success" failing, refer to... Figure 2 This embodiment constructs a three-layer decision logic with decreasing priority: Level 1 Arbitration (Business Semantic Priority Adjudication): Business semantic priority adjudication can involve the proxy node querying the context of the current instance (i.e., the task instance) to see if a specific business status identifier (i.e., the target cross-node status signal) extracted from the log stream is recorded. For example, does it hit a failure identifier such as a business physical calculation anomaly / empty grid? If the extracted identifier representing the divergence of the internal matrix iteration is: !!STATUS=PHYSICS_DIVERGE!!, or the identifier indicating that all input data are missing values ​​is: !!STATUS=EMPTY_GRID!!, then regardless of the final exit code returned by the operating system (even if ExitCode=0), the proxy node will forcibly mark the instance as a business failure (i.e., the task execution failed) and immediately melt down the branch in the directed acyclic graph (DAG), completely blocking the flow of dirty data into the downstream multimodal meteorological data lake (i.e., the downstream task). If a clear business success identifier is extracted (e.g., !!STATUS=BUSINESS_SUCCESS!!), then it will be marked as a successful task execution.

[0125] Second-level arbitration (system exception interruption verification): If no business status identifier is extracted from the log stream, the agent node checks the termination signal of the child process. Under the Portable Operating System Interface (POSIX) system specification, if the child process is forcibly terminated by the system kernel (i.e., a system-level exception interrupt signal is detected, such as a timeout, or it is forcibly killed by OOMKiller due to host memory exhaustion), the return code is usually a specific negative value (e.g., ReturnCode < 0). When such an abnormal interrupt signal is hit, the agent node determines that the task execution has failed and records the underlying system or hardware-level exception cause. If no system-level exception interrupt signal is detected, the process terminates naturally and enters the next level of arbitration.

[0126] Third-level arbitration (system exit code fallback): The agent node will only revert to the traditional judgment method, i.e., reading the operating system's process exit code (ExiCode), if and only if the first-level business semantics are missing and there is no system exception signal at the second level. If ExiCode is 0, the task is considered to have succeeded, triggering subsequent DAG nodes (i.e., driving downstream tasks); if ExiCode is not 0, the task is considered to have failed.

[0127] This embodiment employs a multi-level state arbitration tree to accurately intercept "silent errors" generated by the black-box physical model. This mechanism breaks the traditional scheduling system's strong reliance on the single assumption of exit codes, fundamentally preventing erroneous and dirty data from penetrating into the downstream multimodal data lake.

[0128] 4. Instance-level configuration of sandbox and sweeper exception fallback mechanism.

[0129] In the context of gridded meteorological big data processing, to improve computational efficiency, operating systems typically need to divide large-scale data from across the country or province into dozens or even hundreds of subgrids, and perform high-concurrency scheduling on modern multi-core cloud servers. However, legacy physical meteorological calculation programs (such as black-box programs compiled in C / C++ or Fortran) often have hard-coded reads of static configuration files (e.g., forcibly reading / opt / weather_app / input.nml) from fixed absolute paths on the host machine or in the current working directory due to historical architectural limitations. This "single-machine exclusive" design inevitably leads to configuration file read / write overwriting and cross-contamination when multiple black-box instances are launched concurrently on the same machine. To overcome this concurrency bottleneck, this embodiment proposes a dynamic isolation and "use-and-burn" sandbox mechanism. The specific implementation scheme is as follows: (1) Dynamic Concurrent Configuration Sandbox Construction: Instance Identification and Directory Isolation: The proxy node receives concurrent subtasks and context parameters issued by the orchestration center; the proxy node assigns a globally unique universal identifier (UUID, such as run_202605260800_UUID) to each execution instance. The proxy node creates an independent temporary working directory that is strictly bound to the UUID in the host machine's high-speed storage medium (such as the / tmp directory or the memory file system tmpfs). Parameter Rendering and Seamless Injection: The proxy node reads the preset static configuration file template, replaces the context parameters of the current subtask (including global variables, data pointers, such as the latitude and longitude boundaries and specific input data source paths) with the original configuration placeholders in the preset static configuration file template, and generates a temporary configuration file exclusive to the instance. When starting the black-box program, the proxy node disguises the temporary configuration file as the fixed path format required by the original black-box program by modifying the environment variables of the subprocess, redirecting the temporary working directory, or establishing symbolic links. This mechanism enables legacy black-box programs to run concurrently in a "seamless" state, effectively reusing the multi-core computing power of modern servers and constituting a substantial improvement in the internal performance of the computer system. Normal lifecycle termination: Under normal circumstances, when the child process ends (triggers step S160 above), the agent node actively executes the deletion command, destroys the temporary working directory and temporary configuration file, and deregisters the record from the active task mapping table.

[0130] (2) Janitor Abnormal Recovery Mechanism: Under the POSIX operating system specification, if the agent process crashes, or the black-box process (i.e., the target child process) is instantly killed by a kernel-level process killer (SIGKILL, -9) due to host machine memory overflow, its normal sandbox cleanup code will not be executed, causing the temporary working directory to become an orphan residue occupying disk space and file inodes (i.e., an orphan temporary directory). To prevent concurrent residual files from causing host machine resource exhaustion and cross-contamination during long-term rolling scheduling, an independently running Janitor daemon thread is designed inside the agent node. Each time the agent node launches the target child process, it maintains an active task mapping table in memory. The active task mapping table records [Execution_ID, target child process PID, temporary working directory path, startup timestamp]. Execution_ID can represent the unique identifier (UUID) of the current task node instance, and the target child process PID represents the unique identifier of the target child process. The Janitor traverses the active task mapping table at a set fixed period (e.g., every 30 seconds). If the cleaner checks the operating system process table and finds that the target child process entity corresponding to a certain PID no longer exists, and the lifecycle of its associated temporary working directory (i.e., the temporary working directory's lifespan) has exceeded the preset aging time threshold, then the temporary working directory is determined to be an abnormal remnant (i.e., an orphan temporary directory). The cleaner then triggers fallback garbage collection, forcibly performs garbage collection, deletes the temporary working directory, cleans up the abnormal remnants in the active task mapping table, and records the temporary working directory in the orphan collection log for auditing purposes.

[0131] This embodiment overcomes the single-machine concurrency bottleneck caused by hard-coded absolute path configuration in legacy scientific computing programs. It solves the problems of file cross-over and exhaustion of host machine (Inode) resources under high-concurrency scheduling.

[0132] 5. Time window-based readiness determination and generation of functional degradation control parameters.

[0133] In short-term forecasting and high-frequency radar mosaic operations, rolling updates are typically required within a strict 6-minute time window. Due to network jitter and equipment differences among hundreds of radars nationwide, meteorological observation data arrives at the central control area with a highly asynchronous characteristic. If a traditional wait-all mechanism is adopted, a single station's delay will cause the entire mosaic pipeline to stall. If only simple alerts are issued at the operational level and missing stations are skipped, the radar mosaic product or downstream deep learning model will experience severe physical spatial field anomalies (abrupt changes) at the boundary between missing and normal meteorological observation stations.

[0134] To address this issue, this embodiment proposes dynamically generating a structured functional degradation control parameter package from the DAG control surface, directly taking over the boundary fusion processing of the downstream black-box program. (Refer to...) Figure 3 The specific implementation steps are as follows: (1) Time window interception and local degradation status determination: The aggregation node of the orchestration center sets a strict time window. The list of upstream radar station numbers that the current jigsaw puzzle task depends on is divided into "critical dependency set" and "optional dependency set". When the time window expires, it is determined whether the critical dependency set is ready and whether the optional dependency set is missing. If the critical dependency set data is complete, but some data in the optional dependency set is missing, the system does not determine that the aggregation task has failed, but enters the local degradation state (record degrade_mode=PARTIAL_MISSING). If the local degradation is not met, it is processed according to the normal path, and the downstream task is triggered normally or the global task is determined to have failed.

[0135] (2) Generating a functional degradation control parameter package: After the agent node intercepts the degradation state, it queries and obtains the geographical coverage of the missing meteorological observation stations based on the pre-set station metadata database (containing the center latitude and longitude and fixed scanning radius of each meteorological observation station). It then uses this geographical coverage as the invalid area to generate an effective coverage mask and calculates the minimum vertical distance calculated outward from the physical boundary of this geographical coverage, thereby generating a boundary transition weight file (i.e., a file containing boundary transition weights). For example, the effective coverage mask can directly set the grid values ​​within the geographical coverage to invalid (e.g., 0) and the rest to valid (e.g., 1); while the boundary transition weight can be obtained by calculating the minimum vertical distance calculated outward from the physical boundary of the geographical coverage. The agent node assembles the above information (including the effective coverage mask and the boundary transition weight file) into a degradation control parameter package bound to the current execution instance. This parameter package is not a regular business alarm log, but rather functional data with a predetermined technical purpose for automatically controlling the downstream computational processing method. The minimum set of fields in the degradation control parameter package includes at least: missing site identifier, available site identifier, degradation alternative data source (such as using the previous time product or neighbor interpolation data), effective overlay mask file path, boundary transition weight file path, and boundary transition band width.

[0136] (3) Replacement Injection and Downstream Automatic Fusion Driver: If a local degradation state is determined, the proxy node writes the degradation control parameter package into a temporary configuration file (e.g., generates a mosaic_runtime.ini with a [degrade_context] segment), and replaces the original configuration placeholder of the downstream program in the startup parameters (e.g., rewrites the original startup command mosaic.out / conf / mosaic.ini to mosaic.out / tmp / run_ under proxy control). <uuid> / mosaic_runtime.ini). Downstream legacy puzzle black-box algorithms or multimodal artificial intelligence (AI) models (i.e., downstream tasks) read this temporary configuration file and automatically assemble the input tensors or compute the boundaries based on the mask file and transition band width within it.

[0137] Mathematical Model for Gradual Fusion of Physical Field Boundaries: To eliminate spatial jumps, the proxy node controls the downstream to perform gradual fusion at the influence boundaries of missing sites through a boundary transition weight file. Let any grid point in the planar coordinate system be... The minimum vertical distance from any grid point to the effective coverage boundary of the missing radar station (i.e., meteorological observation station) is denoted as . The width of the boundary transition band issued in the instruction packet is The downstream black-box program calculates the fusion weights based on the implicit smooth transition function in the configuration file: ; in, Indicates the fusion weight. Represents any grid point in a planar coordinate system The minimum vertical distance between the effective coverage boundary of the missing meteorological observation station and the station. Indicates the width of the boundary transition zone. It represents the sine square function.

[0138] The downstream black-box program (i.e., the downstream task) automatically uses the fusion weight matrix to perform weighted fusion of the extracted available site observation data (obtained directly through available site identifiers) and the data from the downgraded alternative data source, outputting a continuous physical field product with no boundary spatial jumps. By elevating the "missing state" from a simple business prompt to a technical parameter controlling the behavior of the underlying black-box algorithm input organization, this embodiment completely eliminates the problem of abrupt changes in the physical boundary of the product caused by the lack of spatiotemporal data at the orchestration level without modifying the original program kernel, ensuring the robust operation of the meteorological business pipeline.

[0139] It should be noted that the effective coverage boundary and boundary transition zone width in this embodiment can be implicit parameters in the configuration file, which are directly obtainable parameters, and will not be specifically described in this embodiment. The previous time product can be the output result generated by the upstream task in the previous 6 minutes. When the previous time product is missing, it can be replaced by neighbor interpolated data (obtained by interpolation). The missing station identifier is used to mark the meteorological observation station as missing, and the available station identifier is used to mark the meteorological observation station as available, and will not be specifically described in this embodiment.

[0140] This embodiment ensures the timeliness of the high-frequency pipeline while eliminating the abrupt change in the spatial boundary of meteorological products caused by the rigid removal of missing sources (i.e., missing meteorological observation stations).

[0141] Compared with existing related technologies, the technical solution of this embodiment has the following advantages: Compared to existing general-purpose workflow orchestration tools (such as traditional DAG schedulers or CI / CD agents), this embodiment overcomes the limitations of relying solely on operating system exit codes for success or failure, as well as the concurrency bottleneck of legacy programs running exclusively on a single machine. This embodiment prioritizes intercepting specific business identifiers (i.e., cross-node status signals) in the output stream through a multi-level state arbitration mechanism. This allows for precise interception and blocking of "silent errors" when the underlying black-box program exits normally at the system level (ExitCode=0) but internal physical computation diverges, ensuring the underlying data quality of the multimodal data lake. Simultaneously, this embodiment, through dynamically generating instance-level configuration sandboxes, redirecting temporary working directories, and introducing a built-in exception cleaner (Janitor), successfully and safely transforms legacy algorithms that originally only supported single-threaded execution on fixed paths into components supporting multi-core high-concurrency scheduling, without modifying the old closed-source program source code. This significantly improves the computational throughput and resource reuse rate within the computer system.

[0142] Secondly, compared with existing parallel processing systems for radar meteorological data, the outstanding advantage of this embodiment lies in providing a decoupled, automated degradation mechanism for the underlying control surface, which completely solves the problem of physical field anomalies caused by local missing data from multi-source asynchronous observations. Existing operational systems often rely on hard-coded timeout skipping when faced with missing radar site data, easily causing product boundary abrupt changes. This embodiment innovatively transforms the spatiotemporal missing state into a functional degradation control parameter package (including effective coverage masks and boundary transition weights) at the orchestration level and injects it into the downstream black-box program. This mechanism allows the downstream model to autonomously perform spatially gradual weighted fusion at the influence boundary of the missing region without manual intervention, ensuring the stringent timeliness of the high-frequency short-term operational pipeline while maximizing the spatial-physical continuity of meteorological products.

[0143] Reference Figure 4 This application also provides a non-intrusive concurrent orchestration system for heterogeneous black-box programs, the system comprising: The data parsing module 100 is used to parse the spatiotemporal dependencies of each task node in the meteorological business pipeline, and to mark the data in the upstream data source in the spatiotemporal dependencies as critical dependencies and non-critical dependencies. The data acquisition module 200 is used to acquire global variables, data pointers and cross-node status signals generated during the workflow of each task in the meteorological business pipeline. The cross-node status signals include business status identifiers or degradation mode identifiers. The decision execution module 300 is used to execute task triggering or degradation decisions based on critical dependency identifiers and non-critical dependency identifiers under strict time window constraints. The file generation module 400 is used to trigger task flow based on task triggering or degradation decisions, create a temporary working directory for the current task node, and generate a temporary configuration file based on global variables, data pointers, and cross-node status signals. The subprocess startup module 500 is used to inject the path of the temporary configuration file into the startup environment of the black-box program by redirecting the temporary working directory, so as to start the target subprocess in the black-box program. The data extraction module 600 is used to divert the log stream of the target subprocess to the asynchronous log pump in a non-blocking queue. The asynchronous log pump extracts the target data pointer and the target cross-node status signal from the extracted log stream. The log stream is a data stream containing the standard output and standard error of the target subprocess. The asynchronous log pump is an independent asynchronous reading coroutine used to read the standard output and standard error data stream. The status judgment module 700 is used to determine the success or failure status after the current task node is completed based on the target cross-node status signal and the target child process exit status signal. The orchestration completion module 800 is used to drive the execution of downstream tasks by reporting success or failure status and target data pointers, while destroying the temporary working directory and temporary configuration file of the current task node, so as to complete the non-intrusive concurrent orchestration of the black-box program.

[0144] It should be noted that since the non-intrusive concurrent orchestration system for heterogeneous black-box programs in this embodiment is based on the same inventive concept as the non-intrusive concurrent orchestration method for heterogeneous black-box programs described above, the corresponding content in the method embodiment is also applicable to this system embodiment, and will not be described in detail here.

[0145] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant regulations. The acquisition, storage, use and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations.

[0146] Reference Figure 5 This application also provides an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in memory, and the processor executes at least one program to implement the non-intrusive concurrent orchestration method for heterogeneous black-box programs described above in this disclosure.

[0147] This electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0148] The electronic devices according to embodiments of this application will now be described in detail.

[0149] The processor 1600 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure. The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700, and the processor 1600 calls and executes the non-intrusive concurrent orchestration method for heterogeneous black-box programs according to the embodiments of this disclosure.

[0150] The input / output interface 1800 is used to implement information input and output. The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 2000 transmits information between various components of the device (e.g., processor 1600, memory 1700, input / output interface 1800, and communication interface 1900); The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.

[0151] This disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described non-intrusive concurrent orchestration method for heterogeneous black-box programs.

[0152] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0153] The embodiments described in this disclosure are for the purpose of more clearly illustrating the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided by this disclosure. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by this disclosure are also applicable to similar technical problems.

[0154] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this disclosure, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0155] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0156] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0157] The terms "first," "second," "third," "fourth," etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0158] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0159] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0160] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0161] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0162] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. The embodiments of this application have been described in detail above with reference to the accompanying drawings, but this application is not limited to the above embodiments. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of this application.

[0163] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.< / uuid> < / uuid> < / uuid>

Claims

1. A non-intrusive concurrent orchestration method for heterogeneous black-box programs, characterized in that, The method includes: The spatiotemporal dependencies of each task node in the meteorological operational pipeline are analyzed, and the data in the upstream data source in the spatiotemporal dependencies are marked with critical dependency identifiers and non-critical dependency identifiers. Acquire global variables, data pointers, and cross-node status signals generated during the workflow of each task in the meteorological business pipeline. The cross-node status signals include business status identifiers or degradation mode identifiers. Under strict time window constraints, task triggering or degradation decisions are executed based on the critical dependency identifier and the non-critical dependency identifier. Based on the task triggering or degradation decision, the task flow is triggered, a temporary working directory for the current task node is created, and a temporary configuration file is generated based on the global variables, the data pointer, and the cross-node status signal. By redirecting the temporary working directory, the path of the temporary configuration file is injected into the startup environment of the black-box program to start the target subprocess in the black-box program. The log stream of the target subprocess is diverted to an asynchronous log pump in a non-blocking queue. The asynchronous log pump extracts the target data pointer and the target cross-node status signal from the extracted log stream. The log stream is a data stream containing the standard output and standard error of the target subprocess. The asynchronous log pump is an independent asynchronous reading coroutine used to read the standard output and standard error data stream. Based on the target cross-node status signal and the target subprocess exit status signal, determine the success or failure status after completing the current task node; By reporting the success or failure status and the target data pointer, the downstream task is driven to execute, while the temporary working directory and temporary configuration file of the current task node are destroyed, so as to complete the non-intrusive concurrent orchestration of the black-box program.

2. The non-intrusive concurrent orchestration method for heterogeneous black-box programs according to claim 1, characterized in that, The process of triggering task flow based on the task triggering or degradation decision, creating a temporary working directory for the current task node, and generating a temporary configuration file based on the global variables, the data pointer, and the cross-node status signal includes: If no local degradation occurs during the task triggering or degradation decision, a globally unique universal identifier is generated for the current task node; a temporary working directory for the current task node is created based on the globally unique universal identifier; the global variables, the data pointer, and the cross-node status signals are used to replace the original configuration placeholders in the static configuration file preset template to generate a temporary configuration file. If a local degradation occurs during task triggering or degradation decision-making, it indicates that a meteorological observation station is missing. The geographical coverage of the missing meteorological observation station is obtained; a globally unique universal identifier is generated for the current task node; a temporary working directory for the current task node is created based on the globally unique universal identifier; an effective coverage mask and boundary transition weight file are generated based on the geographical coverage; a degradation control parameter package containing the effective coverage mask and the boundary transition weight file is constructed; the degradation control parameter package, the global variables, the data pointer, and the cross-node status signal replace the original configuration placeholders in the static configuration file preset template to generate a temporary configuration file.

3. The non-intrusive concurrent orchestration method for heterogeneous black-box programs according to claim 2, characterized in that, The step of generating an effective overlay mask and boundary transition weight file based on the geographical coverage area includes: Set the grid values ​​within the geographic coverage area to invalid values, and set the other grid values ​​except the invalid values ​​to valid values ​​to generate an effective coverage mask; Calculate the minimum vertical distance between any grid point in the planar coordinate system and the effective coverage boundary of the missing meteorological observation station; The physical boundary of the geographic coverage area is extended outward by the minimum vertical distance to generate a boundary transition weight file.

4. The non-intrusive concurrent orchestration method for heterogeneous black-box programs according to claim 2, characterized in that, After generating the temporary configuration file, the method further includes: Read the temporary configuration file; calculate the fusion weight using a smooth transition function based on the degradation control parameter package; the degradation control parameter package also includes available site identifiers and degradation alternative data sources, the available site identifiers are used to mark meteorological observation sites as available, and the degradation alternative data sources are data generated in the previous time segment or neighboring interpolated data; Data on available meteorological observation stations can be obtained through the available station identifiers; The data from the available meteorological observation stations and the data from the downgraded alternative data sources are weighted and fused using the fusion weights to obtain continuous physical field products.

5. The non-intrusive concurrent orchestration method for heterogeneous black-box programs according to claim 1, characterized in that, After starting the target subprocess in the black-box program, the method further includes: Construct an active task mapping table containing the unique identifier of the target subprocess, the temporary working directory, and the startup timestamp; The active task mapping table is traversed at a fixed period. If the target subprocess entity corresponding to the unique identifier of the target subprocess in the active task mapping table does not exist, and the lifespan of the temporary working directory associated with the target subprocess entity exceeds a preset aging time threshold, then the temporary working directory is determined to be an abnormal residual and is deleted.

6. The non-intrusive concurrent orchestration method for heterogeneous black-box programs according to claim 1, characterized in that, The step of diverting the log stream of the target subprocess to an asynchronous log pump in a non-blocking queue, and then using the asynchronous log pump to extract the target data pointer and target cross-node status signal from the extracted log stream, includes: Bind non-blocking read descriptors to the standard output and standard error data streams in the log stream; Based on the non-blocking read descriptor, data in the log stream is read with the highest priority through the asynchronous log pump and pushed into memory in a first-in-first-out queue; When the length of the first-in-first-out queue reaches a preset threshold, a backpressure strategy is triggered; the backpressure strategy discards ordinary computation process logs that do not contain a preset control prefix. Perform single-line length validation on the log stream that is not discarded, and retain the remaining log stream with a character length less than or equal to the preset length; The remaining log stream is subjected to fixed prefix matching to obtain candidate lines; Extract the target data pointer and the target cross-node status signal from the candidate rows.

7. The non-intrusive concurrent orchestration method for heterogeneous black-box programs according to claim 1, characterized in that, The step of determining the success or failure status after completing the current task node based on the target cross-node status signal and the target subprocess exit status signal includes: Construct a multi-level state arbitration tree that includes priority adjudication of business semantics, verification of system exceptions and interruptions, and a fallback system exit code; Based on the target cross-node status signal and the target subprocess exit status signal, the success or failure status after the current task node is completed is determined by the multi-level status arbitration tree.

8. A non-intrusive concurrent orchestration system for heterogeneous black-box programs, characterized in that, The system includes: The data parsing module is used to parse the spatiotemporal dependencies of each task node in the meteorological business pipeline, and to mark the data in the upstream data source in the spatiotemporal dependencies as critical dependencies and non-critical dependencies. The data acquisition module is used to acquire global variables, data pointers and cross-node status signals generated during the workflow of each task in the meteorological business pipeline. The cross-node status signals include business status identifiers or degradation mode identifiers. The decision execution module is used to execute task triggering or degradation decisions based on the critical dependency identifier and the non-critical dependency identifier under strict time window constraints. The file generation module is used to trigger task flow based on the task triggering or degradation decision, create a temporary working directory for the current task node, and generate a temporary configuration file based on the global variables, the data pointer, and the cross-node status signal. The subprocess startup module is used to inject the path of the temporary configuration file into the startup environment of the black-box program by redirecting the temporary working directory, so as to start the target subprocess in the black-box program. The data extraction module is used to divert the log stream of the target subprocess to the asynchronous log pump in a non-blocking queue, and extract the target data pointer and target cross-node status signal from the extracted log stream through the asynchronous log pump; the log stream is a data stream containing the standard output and standard error of the target subprocess, and the asynchronous log pump is an independent asynchronous reading coroutine for reading the standard output and standard error data stream; The status judgment module is used to determine the success or failure status after the current task node is completed based on the target cross-node status signal and the target subprocess exit status signal. The orchestration completion module is used to drive the execution of downstream tasks by reporting the success or failure status and the target data pointer, while destroying the temporary working directory and temporary configuration file of the current task node, so as to complete the non-intrusive concurrent orchestration of the black-box program.

9. An electronic device, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enable the at least one control processor to perform the non-intrusive concurrent orchestration method for heterogeneous black-box programs as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the non-intrusive concurrent orchestration method for heterogeneous black-box programs as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Android platform performance detection method and device

    CN120578566A

  • Natural-social coupling mode public innovation service system

    CN120631549A