An operator data discovery method and system for reducing computer resource consumption
Patent Information
- Application Number
- CN202611013081.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]为解决上述现有技术中存在的计算机资源消耗高、算子及函数缺乏自适应发现与注册机制、且无法与大型语言模型代理进行高效交互式创作的技术问题,本发明在如下的多个方面中提供方案
[0010]因此,本发明实现了算子的自动化发现、标准化注册与动态化查询的全生命周期管理,在保证系统运行期灵活性的同时,大幅降低了新算子接入的门槛与维护成本,尤其适用于大规模数据管道系统中需要频繁扩展算子类型且要求即时生效的复杂业务场景。
Smart Images

Figure CN122816728A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology. More specifically, this invention relates to an operator data discovery method and system for reducing computer resource consumption. Background Technology
[0002] With the widespread deployment of Java applications in development, distributed systems, and cloud platform environments, ensuring both code delivery security and application stability has become a dual challenge for the industry. The ease with which Java bytecode can be decompiled makes source code protection particularly urgent, while the frequent expansion of dynamic arrays (such as ArrayList) in the JVM heap memory also poses a risk of OutOfMemory (OOM). Existing technologies address these issues primarily suffer from the following technical shortcomings.
[0003] Firstly, taking the Chinese patent with publication number CN117828555B as an example, this scheme first encrypts the business class files and then decrypts them at runtime using a custom CustomerClassLoader. However, this scheme explicitly acknowledges that since the custom class loader itself is written in unencrypted Java, it becomes the first weak point that attackers can exploit. Once CustomerClassLoader is decompiled, the AES symmetric key and decryption logic stored within it will be completely exposed, allowing attackers to decrypt all protected class files. Furthermore, if this scheme uses a simple XOR operation to encrypt CustomerClassLoader, its resistance to cracking is limited; even with more complex encryption, the corresponding decryption logic still needs to rely on the JVMTI Agent to deliver it independently as a C++ dynamic link library. Essentially, it still exposes the core decryption logic outside the Java runtime, failing to form a closed-loop protection from the class loader to the business class, and undermining Java's "compile once, run anywhere" cross-platform feature.
[0004] Secondly, regarding JVM memory warnings, existing solutions are mostly passive responses after memory overflow occurs, lacking proactive warning capabilities based on object granularity and resizing behavior. Taking Chinese patent CN117312109B as an example, this solution calculates the remaining JVM memory during ArrayList resizing via a Java proxy and compares it with a preset threshold, achieving a certain degree of pre-warning. However, its drawback is that the warning trigger point is only set after the resizing operation, i.e., "resizing first, then judging." When the remaining JVM memory is already below the threshold, a memory shortage has already occurred, making it impossible to intercept or suppress it before the resizing action or before the resizing space is requested. Furthermore, this solution requires traversing the entire array and calculating the object header, instance data, and padding space for each object when calculating the memory usage of each object. For dynamic arrays storing millions of objects, the traversal overhead is enormous, potentially causing new performance bottlenecks or even triggering GC pauses due to the monitoring logic itself. Furthermore, its threshold configuration is a static preset value, which cannot be adaptively adjusted according to the current dynamic memory distribution of the JVM (such as the ratio of young generation / old generation occupancy and object promotion rate). In multi-tenant or cloud-native environments, it is difficult to balance the warning sensitivity and false alarm rate.
[0005] Therefore, existing technologies suffer from a combination of defects, such as incomplete Java source code protection links, the vulnerability of custom class loaders to decompilation leading to key leaks, memory warnings lagging behind resizing actions, and excessive overhead in object traversal monitoring. These defects make it difficult to achieve fine-grained, low-overhead proactive warnings for JVM memory while ensuring delivery security. Summary of the Invention
[0006] To address the technical problems of high computer resource consumption, lack of adaptive discovery and registration mechanisms for operators and functions, and inability to perform efficient interactive creation with large language model agents in the prior art, the present invention provides solutions in the following aspects.
[0007] In a first aspect, the present invention provides an operator data discovery method for reducing computer resource consumption, characterized by comprising the following steps: S10: Obtain the preset operator code path; perform classpath scanning based on the operator code path to obtain candidate operator classes; S20: Determine the target operator implementation class from the candidate operator classes according to the preset operator recognition rules; S30: Parse the implementation class of the target operator and generate the corresponding operator metadata; S40: Register the operator metadata to the operator registry; S50: In response to an operator query request, output the corresponding operator metadata from the operator registry for use in pipeline construction or operator invocation.
[0008] Compared with existing technologies, this invention is based on an automatic operator discovery and registration mechanism. Through a decoupled process of preset code path scanning, identification rule filtering, metadata parsing and registry query, it realizes zero-configuration dynamic registration and on-demand acquisition of operators, providing a unified metadata service foundation for pipeline construction and operator invocation.
[0009] Starting with class path scanning of preset operator code paths, this replaces the hard-coding method of listing operator classes one by one in configuration files or code in existing technologies. This allows new operators to be automatically discovered by the system without modifying any existing code, significantly reducing the manual cost of operator extension and the risk of configuration errors. Based on preset operator identification rules, target operator implementation classes are selected from candidate operator classes. Compared to the inefficient mode of relying on manual judgment or full loading of all classes in existing technologies, this avoids the overhead of loading and parsing invalid classes and supports flexible adjustment of identification rules to adapt to operator libraries with different naming conventions or annotation conventions. The target operator implementation class is parsed and corresponding operator metadata is generated. The operator metadata is registered to the operator registry. The combination of these two forms a standardized output and centralized storage of operator metadata, which solves the problem of operator type, input and output declaration, configurable attributes and other information being scattered in multiple parts of the code and difficult to manage in a unified manner in existing technologies. It provides a programmable metadata foundation for upper-level pipeline orchestration tools. In response to operator query requests, the corresponding operator metadata is output from the operator registry for use in pipeline construction or operator invocation. Compared to existing technologies that require real-time parsing of classes through reflection or rely on hard-coded static mapping tables, this enables runtime on-demand acquisition of operator information, which not only improves query efficiency, but also allows pipeline configuration information to be dynamically generated according to the actual registered operators, avoiding runtime anomalies caused by missing operators or inconsistent versions.
[0010] Therefore, this invention realizes the full lifecycle management of automated operator discovery, standardized registration, and dynamic querying. While ensuring the flexibility of system operation, it significantly reduces the threshold and maintenance cost of new operator access, and is especially suitable for complex business scenarios in large-scale data pipeline systems that require frequent expansion of operator types and immediate effect.
[0011] Preferably, step S20 includes: Determine whether the candidate operator class meets preset conditions, the preset conditions including: Inherit from or implement a predefined operator base interface; Candidate operator classes that meet the conditions are determined as the target operator implementation classes.
[0012] The beneficial effects of this invention are that, based on the type hierarchy judgment mechanism of class inheritance or interface implementation, it realizes the automatic and accurate screening of candidate operator classes. This avoids the risk of runtime type conversion exceptions caused by mistakenly registering non-operator classes to the operator registry, and also eliminates the additional memory overhead caused by loading and parsing a large number of irrelevant classes. Thus, while ensuring the accuracy of operator discovery, it further reduces the computational resource consumption of the classpath scanning stage and improves the overall efficiency of system startup and operator registration.
[0013] Preferably, step S30 includes: The operator type is determined based on the code path where the target operator implementation class is located. The operator types include data input type operators, data transformation type operators, and data output type operators.
[0014] Preferably, the operator metadata includes at least the operator identifier, operator name, operator type, parameter information, input / output information, and description information.
[0015] The beneficial effects of this invention are that it automatically infers the operator type based on the code path where the operator class is located, achieving zero-configuration type classification and avoiding type misjudgment caused by annotation omissions or configuration errors. At the same time, by integrating the operator's identifier, name, type, parameters, input / output, and description information into a standardized metadata structure, each record in the operator registry contains complete self-descriptive information, providing the pipeline builder with renderable data in one step. This allows the front end to obtain the full picture of the operator without repeatedly parsing the class through reflection, reducing the coupling between the UI layer and the bytecode layer, avoiding the performance loss caused by repeated reflection, and significantly improving the rendering response speed and user experience of the pipeline orchestration interface.
[0016] Preferably, step S40 specifically includes: Establish a corresponding type index based on the operator type in the operator metadata; Build a name index based on the operator name in the operator metadata; Store operator metadata in the operator registry, so that the operator registry can return the corresponding operator metadata based on the type index or name index.
[0017] Preferred options also include: A data processing flowchart is constructed based on the operator registry. Determine the target subgraph in the data processing flowchart based on the target operator; Execute the part of the process corresponding to the target subgraph to obtain the preview results; Some processes reuse the operator registry and operator instances from the production execution environment, and limit the amount of preview data and execution time.
[0018] The beneficial effects of this invention are as follows: by establishing a dual-dimensional retrieval mechanism of type index and name index, the pipeline builder can quickly filter candidate operators by operator type and accurately locate target operators by name, significantly improving query efficiency and UI rendering response speed under large-scale operator sets. At the same time, by building a data processing flowchart based on the same operator registry and executing target subgraph preview, and by directly reusing the operator registry and operator instances in the production environment during the preview process, compared with the existing technology that requires building an independent simulation environment or sandbox for preview, this not only eliminates the risk of inconsistency between preview results and production execution results caused by environmental differences, but also avoids the resource overhead caused by additional deployment of preview environment. Furthermore, the mandatory limitation on data volume and execution time during the preview process effectively eliminates the risk of system resource exhaustion caused by full data calculation triggered by preview operation, making the preview function a true native capability of the production environment while ensuring security.
[0019] In a second aspect, the present invention discloses an operator data discovery system for reducing computer resource consumption, specifically: The scanning module is used to obtain the preset operator code path; and to perform a classpath scan based on the preset operator code path to obtain candidate operator classes. The identification module is used to determine the target operator implementation class from the candidate operator classes according to the preset operator identification rules; The parsing module is used to parse the implementation class of the target operator and generate the corresponding operator metadata; The registration module is used to register operator metadata to the operator registry. The query module is used to respond to operator query requests and output the corresponding operator metadata from the operator registry.
[0020] Preferably, the registration module includes: The type index unit is used to create a corresponding type index in the operator registry based on the operator type in the operator metadata. The name index unit is used to create a corresponding name index in the operator registry based on the operator name in the operator metadata.
[0021] Preferably, it also includes a process construction module, specifically: The process construction module is used to generate a directed acyclic graph of operators based on operator metadata; It performs topological sorting and cycle detection based on the operator-directed acyclic graph.
[0022] Preferably, it also includes a preview execution module, specifically: The preview execution module is used to call the part of the operator process corresponding to the target operator; The preview execution module shares the operator registry, operator instances, and data processing context with the production execution module, thereby reducing the consumption of computing resources.
[0023] The beneficial effects of this invention are as follows: By separating the responsibilities and pipelined collaboration of the scanning, identification, parsing, registration, and query modules, the entire chain of operator discovery and use is automated. Each module performs its own function and is sequentially decoupled, so that when adding a new operator type, only the operator implementation itself needs to be considered, without modifying any existing registration logic, thus reducing system maintenance costs and expansion risks. Simultaneously, the registration module establishes a two-dimensional retrieval mechanism through type and name index units, enabling the query module to quickly filter by type or accurately locate operator metadata by name, avoiding the overhead of full traversal and improving query performance under large-scale operator sets. Based on this, the process construction module automatically generates a directed acyclic graph using the registered operator metadata and performs topological sorting and cycle detection, organically connecting operator discovery with process orchestration, eliminating the need for users to manually maintain dependencies. The preview execution module directly reuses the operator registry, operator instances, and data processing context of the production execution module. Compared to the existing technology that requires independent deployment of a sandbox environment for verification, this eliminates the risk of inconsistent results due to environmental differences and avoids the memory and CPU overhead caused by repeatedly loading operator instances.
[0024] In summary, this system significantly reduces the computational resource consumption and manual configuration burden in operator expansion, process orchestration, and preview verification stages by standardizing interfaces between modules and sharing runtime resources across paths, while ensuring functional integrity. Attached Figure Description
[0025] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein: Figure 1 This is a flowchart of the operator data discovery method for reducing computer resource consumption in Embodiment 1; Figure 2 This is an architecture block diagram of the operator data discovery system used to reduce computer resource consumption in Embodiment 2. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0028] In a specific application scenario, the operator data discovery method provided in this embodiment can be applied to an advertising data query and analysis system.
[0029] For example, when operators need to query metrics such as spending, impressions, and clicks of an advertising account within a specific time range, the system needs to pull raw data from multiple data sources (such as Facebook Ads, Google Ads, TikTok Ads, etc.), and process the data by cleaning, transforming, and aggregating it, and finally output structured reports for front-end display.
[0030] In such scenarios, the data processing flow typically involves multiple steps, including data source access operators, field mapping and transformation operators, expression calculation operators, data aggregation operators, and result output operators. Furthermore, the data formats and field names vary across different advertising platforms. If the logic of each operator is implemented using the traditional hard-coding method, the existing code must be modified and redeployed whenever a new platform is added or the processing logic is adjusted, resulting in poor scalability and high maintenance costs.
[0031] To this end, this embodiment achieves automatic operator discovery and registration during the startup phase. This ensures that all types of operators involved in the data processing flow (including data input, data transformation, and data output operators) are automatically scanned, identified, and registered in the operator registry upon system startup, forming a set of operator metadata available for subsequent pipeline construction and execution. Simultaneously, the system scans classes with custom namespace annotations, extracts all methods with custom expression function annotations within those classes, registers these methods as namespace functions, and stores the namespace function's metadata, including function name, description, parameter types, and usage examples. These namespace functions are called by the JEXL expression engine during operator execution to complete data transformation logic such as field mapping and expression calculation. When operators configure data processing tasks, the system responds to operator query requests, returning the required operator metadata from the operator registry for pipeline construction. This allows for flexible combination of operators to form a directed acyclic graph and execute data processing tasks according to actual business needs, adapting to the personalized data processing requirements of different advertising platforms without code modification.
[0032] The following sections will use this advertising data query scenario as an example to explain in detail the specific implementation process of operator automatic discovery, registration, pipeline construction, and execution.
[0033] Example 1 like Figure 1 As shown, this embodiment discloses an operator data discovery method for reducing computer resource consumption, including: S10: Obtain the preset operator code path; perform classpath scanning based on the operator code path to obtain candidate operator classes.
[0034] In this embodiment, the system obtains a preset operator code path. This path predefines the range of base packages containing the Java classes to be scanned, typically specified in a configuration file or system parameter. This limits the scan to a specific directory containing the operator implementation classes, avoiding startup delays and memory waste caused by a full scan. Subsequently, based on this path, the system performs a classpath scan using the Java class loading mechanism and reflection API, traversing all .class files under this path, reading bytecode, and parsing class structure information to identify candidate operator classes.
[0035] Specifically, this scanning process can utilize the features provided by the Spring framework. The `ClassPathScanningCandidateComponentProvider` or a custom class scanner implementation uses inclusion filters to initially identify all non-abstract, public classes, and encapsulates their metadata into a candidate operator class set for further identification and filtering in subsequent steps. This approach enables configurable operator code paths and precise control over the scanning scope.
[0036] S20: Determine the target operator implementation class from the candidate operator classes according to the preset operator recognition rules.
[0037] Preferably, in order to automatically and accurately select the implementation classes with operator execution capabilities from the candidate operator classes, avoid registering non-operator classes to the operator registry and causing runtime type conversion exceptions, and eliminate the additional memory overhead caused by loading and parsing irrelevant classes, step S20 includes: Step S21: Determine whether the candidate operator class meets the preset conditions, the preset conditions including, Step S22: Inherit a preset operator basic interface or implement a preset operator basic interface; Step S23: Determine the candidate operator classes that meet the conditions as the target operator implementation classes.
[0038] In this embodiment, the core judgment logic of the preset identification rule is to determine whether the candidate operator class meets the preset conditions, which include whether the candidate operator class inherits the preset operator basic interface or implements the preset operator basic interface.
[0039] Specifically, in advertising data query and analysis scenarios, all operator implementation classes must adhere to a unified interface specification, that is, inherit or implement the system's predefined operator base interface. This base interface declares the core methods of the operator, such as data reading methods, data processing methods, or data writing methods. Different operator implementation classes implement specific business logic under the constraints of this base interface. The system iterates through the acquired set of candidate operator classes, using Java reflection or type checking tools to determine the type hierarchy of each candidate class to see if it meets the above conditions. Candidate operator classes that meet the conditions are identified as target operator implementation classes, while candidate classes that do not meet the conditions, such as ordinary utility classes, configuration classes, or other non-operator classes, are filtered out. By leveraging the type hierarchy judgment mechanism based on interface inheritance, the system achieves automated and accurate filtering of operator implementation classes, ensuring that all classes ultimately registered in the operator registry have operator execution capabilities, while eliminating the risk of additional memory overhead and type conversion anomalies caused by the misidentification and loading of irrelevant classes.
[0040] S30: Parse the implementation class of the target operator and generate the corresponding operator metadata; Preferably, in order to automatically classify operator types without additional annotations or configuration files, and to avoid type misjudgment due to missing annotations or configuration errors, step S30 includes: The operator type is determined based on the code path where the target operator implementation class is located. The operator types include data input type operators, data transformation type operators, and data output type operators.
[0041] Preferably, the operator metadata includes at least the operator identifier, operator name, operator type, parameter information, input / output information, and description information.
[0042] In this embodiment, the system parses the determined target operator implementation class to generate the corresponding operator metadata.
[0043] Specifically, in the scenario of advertising data query and analysis, the system reads the bytecode information of the target operator implementation class through the Java reflection mechanism, extracts structural information such as class name, method signature, annotation, generic parameters and fields, and converts this information into a standardized operator metadata structure.
[0044] The parsing process includes determining the operator type based on the code path where the target operator implementation class is located: the system automatically determines the type based on the package name fragment in the package name string. If it contains a preset source type identifier (such as ".source"), it is determined to be a data input type operator; if it contains a preset output type identifier (such as ".sink"), it is determined to be a data output type operator; if it does not contain either, it is determined to be a data conversion type operator.
[0045] The three types mentioned above correspond to the data reading stage, data transformation and processing stage, and data writing stage in the data processing flow, respectively. In the advertising data query and analysis scenario, they are specifically manifested as follows: the data input type operator is responsible for pulling raw data from the APIs of various advertising platforms, the data transformation type operator is responsible for executing processing logic such as field mapping, expression calculation, data cleaning and aggregation, and the data output type operator is responsible for writing the processed data into the target storage.
[0046] Furthermore, the system extracts parameter information, input / output information, and descriptive information by traversing the methods and annotations of the operator class. The resulting operator metadata includes at least the operator identifier, operator name, operator type, parameter information, input / output information, and descriptive information. This metadata collectively constitutes a complete self-describing information set for the operator, providing a standardized data source for subsequent pipeline construction, front-end rendering, and query retrieval. By transforming the original Java class bytecode into clearly structured and complete operator metadata, a standardized mapping of operator information from the code layer to the data layer is achieved.
[0047] In addition, the system scans classes with custom namespace annotations. These annotations identify a function namespace through name and description attributes, such as string processing, time processing, or JSON processing namespaces. For each class, the system iterates through all its methods, extracting methods with custom expression function annotations. These annotations contain descriptions, usage examples, and parameter descriptions. The system registers the extracted methods as callable functions under the corresponding namespaces, storing the function name, description, parameter types, and usage examples as function metadata. It also establishes a mapping between namespace names and class instances containing these methods, allowing the JEXL expression engine to call them at runtime using the syntax "namespace:function name(parameters)," enabling the expression evaluation logic to be dynamically expanded with zero configuration.
[0048] S40: Register the operator metadata to the operator registry.
[0049] Preferably, step S40 specifically includes: Step S41: Establish the corresponding type index based on the operator type in the operator metadata; Step S42: Build a name index based on the operator name in the operator metadata; Step S43: Store the operator metadata in the operator registry, so that the operator registry returns the corresponding operator metadata based on the type index or name index.
[0050] Preferably, in order to verify the correctness of the conversion logic in the intermediate links before the pipeline is officially deployed, the following is also included: A data processing flowchart is constructed based on the operator registry. Determine the target subgraph in the data processing flowchart based on the target operator; Execute the part of the process corresponding to the target subgraph to obtain the preview results; Some processes reuse the operator registry and operator instances from the production execution environment, and limit the amount of preview data and execution time.
[0051] In this embodiment, the generated operator metadata is registered to the operator registry, which serves as a centralized metadata storage center, providing a unified data source for subsequent pipeline construction, operator querying, and process orchestration.
[0052] Specifically, in advertising data query and analysis scenarios, once the metadata of each type of operator (data input type operator, data transformation type operator, and data output type operator) has been registered, the system establishes a corresponding type index based on the operator type in the operator metadata, and simultaneously establishes a name index based on the operator name, forming a two-dimensional retrieval mechanism. The operator metadata is then stored in the operator registry, enabling subsequent quick filtering of candidate operators by category through the type index, or precise location of the target operator through the name index. This avoids the overhead of full traversal and significantly improves query performance under large-scale operator sets.
[0053] Furthermore, the system can construct a directed acyclic graph based on the operator registry. This graph consists of registered operators as nodes and upstream dependencies as edges, and the execution order is determined through topological sorting. During the pipeline creation or debugging phase, the system determines a subgraph from the starting operator to the target operator based on the specified target operator, and executes the corresponding part of the process to obtain a preview result. This preview process directly reuses the operator registry and registered operator instances from the production environment, eliminating the risk of inconsistencies between preview and production results caused by environmental differences, and avoiding the memory and CPU overhead caused by repeatedly loading operator instances. During the production execution phase, multiple pipelines reuse the same Java Virtual Machine process. The system dynamically allocates the number of threads in the thread pool according to the concurrency requirements of each pipeline, rather than allocating fixed resources independently for each pipeline. This avoids the fixed overhead of approximately 0.5 CPU and hundreds of MB of memory caused by the need to independently request containers for each pipeline in traditional solutions, significantly reducing the overall resource consumption in multi-pipeline concurrent scenarios.
[0054] The preview execution process also enforces restrictions on the amount of preview data and execution time. Specifically, the number of output data rows shall not exceed 100 rows and the execution timeout shall be 90 seconds. When the execution timeout occurs, the execution of the current subgraph shall be canceled and a timeout prompt shall be returned, thereby effectively preventing the risk of system resource exhaustion caused by full data calculation or logical errors triggered by the preview operation.
[0055] S50: In response to an operator query request, output the corresponding operator metadata from the operator registry for use in pipeline construction or operator invocation.
[0056] In this embodiment, in response to an operator query request, the corresponding operator metadata is output from the operator registry for use in pipeline construction or operator invocation.
[0057] In addition, the system scans classes with custom namespace annotations through the function registration module, extracts methods with custom expression function annotations, registers these methods as namespace functions, and stores their metadata, including function name, description, parameter types, and usage examples. At the same time, it exposes this metadata to the outside world through the metadata application interface, allowing the large language model agent to call it to obtain all available namespace functions and their parameter descriptions.
[0058] Specifically, in advertising data query and analysis scenarios, when operations personnel configure data processing pipelines through the front-end interface, the front-end page initiates an operator query request. This request can include query conditions such as filtering by type or exact matching by name. Based on the query parameters in the request, the system quickly retrieves and returns matching operator metadata using the type or name index in the operator registry, including operator identifier, name, type, parameter information, input / output information, and description. The front-end page then renders the operator selection list and displays parameter configuration items to assist operations personnel in completing pipeline configuration. When the pipeline is submitted and execution is triggered, the system also retrieves the required operator metadata from the registry through a query request, then instantiates the operator and executes the data processing task according to the topology. The operator registry provides retrieval services externally through a standardized interface, enabling dynamic acquisition of available operator information during the pipeline construction phase and correct instantiation and execution of operators during the operator invocation phase. This achieves seamless integration between the operator discovery mechanism and pipeline orchestration and task execution, avoiding the problems of poor scalability and high maintenance costs caused by hard coding.
[0059] Preferred options also include: A data processing flowchart is constructed based on the operator registry. Determine the target subgraph in the data processing flowchart based on the target operator; Execute the part of the process corresponding to the target subgraph to obtain the preview results; Some processes reuse the operator registry and operator instances from the production execution environment, and limit the amount of preview data and execution time.
[0060] In this embodiment, a preview operation can also be performed based on the operator registry.
[0061] Specifically, in advertising data query and analysis scenarios, when operators configure data processing pipelines through the front-end interface, they often need to verify whether the conversion logic of intermediate links is correct. For example, they need to confirm whether the mapping of a certain field or the calculation of an expression meets expectations. If verification is performed after the pipeline is fully deployed, the debugging cycle is long and the problem is difficult to locate.
[0062] To this end, the system first constructs a data processing flowchart based on the operator metadata registered in the operator registry. This flowchart forms a directed acyclic graph with operators as nodes and upstream dependencies as edges, and determines the execution order of each operator through topological sorting. When an operator specifies a target operator and initiates a preview request on the front-end interface, the system determines the target subgraph in the data processing flowchart based on the target operator, which is the subsequence from the starting operator to the target operator.
[0063] Subsequently, the system executes the corresponding part of the process in the target subgraph to obtain the preview results. This preview execution process directly reuses the operator registry and registered operator instances in the production execution environment, rather than building an independent simulation environment, thereby ensuring the consistency between the preview results and the production execution results.
[0064] When execution times out, the current subgraph execution is canceled and a timeout message is returned. This preview mechanism allows operations personnel to quickly verify intermediate transformation logic before the pipeline is officially deployed, avoiding the resource consumption and waiting time associated with full pipeline execution.
[0065] Example 2 like Figure 2 The diagram shown is the architecture of this system. This embodiment discloses an operator-based data discovery system for reducing computer resource consumption, specifically: The scanning module is used to obtain the preset operator code path; and to perform a classpath scan based on the operator code path to obtain candidate operator classes. The identification module is used to determine the target operator implementation class from the candidate operator classes according to the preset operator identification rules; The parsing module is used to parse the implementation class of the target operator and generate the corresponding operator metadata; The registration module is used to register operator metadata to the operator registry. The query module is used to respond to operator query requests and output the corresponding operator metadata from the operator registry.
[0066] Preferably, the registration module includes: The type index unit is used to create a corresponding type index in the operator registry based on the operator type in the operator metadata. The name index unit is used to create a corresponding name index in the operator registry based on the operator name in the operator metadata.
[0067] Preferably, in order to automatically orchestrate the registered operator metadata into an executable logical flow, a flow construction module is also included, specifically: The process construction module is used to generate a directed acyclic graph of operators based on operator metadata; It performs topological sorting and cycle detection based on the operator-directed acyclic graph.
[0068] Preferably, in order to verify the correctness of the operator logic before the pipeline is officially deployed, a preview execution module is also included, specifically: The preview execution module is used to call the part of the operator process corresponding to the target operator; The preview execution module shares the operator registry, operator instances, and data processing context with the production execution module, thereby reducing the consumption of computing resources.
[0069] The specific operating logic of this embodiment is as follows: The system begins with the scanning module obtaining the preset operator code path and performing a class path scan upon system startup to acquire candidate operator classes. The identification module, based on preset operator identification rules, filters out target operator implementation classes from the candidate operator classes that meet the conditions of inheriting or implementing the basic operator interface. The parsing module parses the target operator implementation classes, determines the operator type based on the package name, and extracts the operator identifier, name, parameter information, input / output information, and description information to generate standardized operator metadata. The registration module stores the operator metadata in the operator registry, and establishes dual-dimensional indexes using both the type index unit and the name index unit to support fast retrieval. During the pipeline construction phase, the process construction module generates a directed acyclic graph based on the metadata in the operator registry and performs topological sorting and cycle detection. The query module responds to operator query requests, outputting the corresponding operator metadata through the type index or name index for pipeline configuration or operator invocation. During the pipeline preview phase, the preview execution module determines the target subgraph in the directed acyclic graph based on the target operator and executes part of the process. This preview execution module shares the operator registry, operator instance, and data processing context with the production execution module to avoid repeated loading and instantiation overhead, while limiting the amount of preview data and execution time to prevent resource exhaustion.
[0070] The system also includes other components well known to those skilled in the art, such as communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.
[0071] In this invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc., or any other medium that can be used to store desired information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device. Any application or module described in this invention can be implemented using computer-readable / executable instructions that can be stored or otherwise maintained by such a computer-readable medium.
[0072] In the description of this specification, "multiple" means at least two, such as two, three or more, etc., unless otherwise expressly and specifically defined.
[0073] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.
Claims
1. A method for operator data discovery to reduce computer resource consumption, characterized in that, include: S10: Obtain the preset operator code path; Perform a classpath scan based on the operator code path to obtain candidate operator classes; S20: Determine the target operator implementation class from the candidate operator classes according to the preset operator identification rules; S30: Parse the target operator implementation class and generate the corresponding operator metadata; S40: Register the operator metadata to the operator registry; S50: In response to an operator query request, output the corresponding operator metadata from the operator registry for use in pipeline construction or operator invocation.
2. The operator data discovery method for reducing computer resource consumption according to claim 1, characterized in that, Step S20 includes: Determine whether the candidate operator class meets preset conditions, the preset conditions including: Inherit from or implement the preset operator basic interface; The candidate operator classes that meet the conditions are determined as the target operator implementation classes.
3. The operator data discovery method for reducing computer resource consumption according to claim 1, characterized in that, Step S30 includes: The operator type is determined based on the code path where the target operator implementation class is located. The operator types include data input type operators, data transformation type operators, and data output type operators.
4. The operator data discovery method for reducing computer resource consumption according to claim 3, characterized in that, The operator metadata includes at least the operator identifier, operator name, operator type, parameter information, input / output information, and description information.
5. The operator data discovery method for reducing computer resource consumption according to claim 1, characterized in that, Step S40 specifically involves: Establish a corresponding type index based on the operator type in the operator metadata; A name index is created based on the operator name in the operator metadata. The operator metadata is stored in the operator registry, and the operator registry returns the corresponding operator metadata based on the type index or the name index.
6. The operator data discovery method for reducing computer resource consumption according to claim 5, characterized in that, Also includes: A data processing flowchart is constructed based on the operator registry; The target subgraph in the data processing flowchart is determined based on the target operator; Execute the local process corresponding to the target subgraph to obtain the preview result; The local process execution process reuses the operator registry and operator instance in the production execution environment, and limits the amount of preview data and execution time.
7. An operator-based data discovery system for reducing computer resource consumption, characterized in that, The operator data discovery method for reducing computer resource consumption as described in any one of claims 1-6 includes: The scanning module is used to obtain the preset operator code path; and to perform a classpath scan based on the operator code path to obtain candidate operator classes. The identification module is used to determine the target operator implementation class from the candidate operator classes according to the preset operator identification rules; The parsing module is used to parse the implementation class of the target operator and generate the corresponding operator metadata; The registration module is used to register operator metadata to the operator registry. The query module is used to respond to operator query requests and output the corresponding operator metadata from the operator registry.
8. The operator data discovery system for reducing computer resource consumption according to claim 7, characterized in that, The registration module includes: A type index unit is used to establish a corresponding type index in the operator registry based on the operator type in the operator metadata. The name indexing unit is used to establish a corresponding name index in the operator registry based on the operator name in the operator metadata.
9. The operator data discovery system for reducing computer resource consumption according to claim 7, characterized in that, It also includes a process building module, specifically: The process construction module is used to generate an operator directed acyclic graph based on the operator metadata; And perform topological sorting and cycle detection on the directed acyclic graph based on the operator.
10. The operator data discovery system for reducing computer resource consumption according to claim 7, characterized in that, It also includes a preview execution module, specifically: The preview execution module is used to call the partial operator process corresponding to the target operator; The preview execution module shares the operator registry, operator instance, and data processing context with the production execution module, thereby reducing the consumption of computing resources.
Citation Information
Patent Citations
A memory warning method for Java dynamic arrays
CN117312109B
A low-cost Java source code protection method and device
CN117828555B