Container cluster-oriented big data job intelligent submission method and system
Through field mapping and deep decision tree combined with intelligent routing strategies, the configuration complexity and low resource utilization in big data job submission are solved, intelligent task submission and optimization decisions are realized, and task execution efficiency is improved.
Patent Information
- Application Number
- CN202510875739.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing big data job submission technology has complex configurations and lacks unified standards, high user usage complexity, low resource utilization, lack of intelligent decision-making capabilities, and cannot form a closed-loop optimization system.
Through field mapping and configuration standardized conversion technology, task configuration information from different sources is converted into standard intermediate formats, combined with deep decision trees and intelligent routing strategies, dynamically select the optimal submission method, and monitor task status optimization decisions in real time.
It significantly improves system integration efficiency and data conversion accuracy, reduces resource scheduling conflict rate, improves cluster resource utilization and task execution efficiency, simplifies user operation processes, and improves task execution efficiency by more than 30%.
Smart Images

Figure CN120371487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly to an intelligent submission method and system for big data jobs for container clusters. Background Art
[0002] With the rapid development of big data technology, various big data computing frameworks such as Hadoop, Spark, Flink, etc. are widely used in data processing and analysis scenarios. At the same time, the rise of container technology has made deploying and managing big data applications on container clusters such as Kubernetes the mainstream solution. In this environment, the submission and resource allocation of big data jobs become particularly important, which directly affects the system performance, resource utilization rate, and user experience.
[0003] Traditional big data job submission methods usually rely on command-line tools or scripts of specific computing frameworks, and users need to master different submission methods and parameter settings for different computing frameworks. With the popularization of containerized deployment, the submission methods of big data jobs have become more diverse. They can be submitted either through traditional computing framework tools or through resource definition files of container orchestration platforms. Although this diversity provides flexibility, it also increases the usage complexity and operation and maintenance difficulty for users.
[0004] However, the current big data job submission technology has the following defects and deficiencies: the configuration of big data job submission is complex and lacks a unified standard, and there are differences in configuration parameters between different frameworks and versions. Users need to face cumbersome parameter settings and version compatibility issues, which increase the usage threshold and error probability; the existing job submission methods lack intelligent decision-making capabilities and cannot automatically select the optimal submission method and resource configuration according to the cluster resource status, job characteristics, and historical operation data, resulting in low resource utilization or poor task execution efficiency; the existing technology lacks an effective feedback mechanism after job submission and cannot use data such as status information and resource usage during task execution to optimize subsequent task submission decisions and form a closed-loop optimization system, making the system unable to continuously self-improve and enhance with the usage process. Summary of the Invention
[0005] Embodiments of the present invention provide an intelligent submission method and system for big data jobs for container clusters, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention, an intelligent submission method for big data jobs for container clusters is provided, including: Receiving task configuration information submitted by a user; Performing field mapping and configuration standardization conversion based on field feature similarity calculation and historical call relationships, and combining version compatibility evaluation to generate standard intermediate configuration data; Input the standard intermediate configuration data into the deep decision tree. Obtain the cluster resource status and load metrics through the computing resource evaluation node, and obtain the task running characteristics through the task feature analysis node. Calculate and determine the final submission method based on the intelligent routing policy. Based on the template mapping mechanism, select the corresponding template according to the final submission method, and combine the historical task data to perform intelligent filling and optimization of the template parameters to generate a task submission instruction. According to the task submission instruction, when the final submission method is the first submission method, generate a resource definition configuration file and submit the task through the container cluster command; when it is the second submission method, generate a command-line tool for the computing framework to submit the task. Monitor the execution status of the task in real time, collect status information, resource information, and running logs, and feedback them to the intelligent routing policy for optimizing the subsequent task submission decision.
[0007] In an alternative embodiment, based on the field feature similarity calculation and historical call relationships, perform field mapping and configuration standardization conversion, and combine the version compatibility evaluation to generate the standard intermediate configuration data, including: Perform feature vector construction processing on the source field and the target field respectively. The feature vector construction processing uses a preset tokenizer to tokenize the field name, generates a basic word vector using the log-likelihood function of the central word and the context words, combines multi-dimensional semantic enhancement to generate a word vector, and at the same time extracts the data type identifier and the boundary values of the value range of the field to construct a field feature mapping table. Calculate the mapping correlation degree between the source field and the target field based on the field feature mapping table, extract the field call data from the historical task records, establish a field call relationship table, and count the co-occurrence frequency between fields to obtain the call weight; perform weighted combination of the mapping correlation degree and the call weight to generate a field mapping score table, and select the target field with the highest score as the initial mapping field set. Extract the computing engine fields from the initial mapping field set, parse the version information and the corresponding configuration items, and construct a version configuration association table; based on the version configuration association table, analyze the configuration item differences of each version to generate a configuration migration rule set; perform version compatibility evaluation on the configuration items according to the configuration migration rule set to obtain a version-compatible configuration set. Perform configuration item standardization conversion based on the version-compatible configuration set to generate standard configuration items; perform integrity verification on the standard configuration items, complete the configuration dependency relationship, and generate the standard intermediate configuration data.
[0008] In an alternative embodiment, use a preset tokenizer to tokenize the field name, generate a basic word vector using the log-likelihood function of the central word and the context words, and combine multi-dimensional semantic enhancement to generate a word vector, including: Vectorize the word segmentation results based on the central word and context words, maximize the log-likelihood function of the central word and the word sequence within the corresponding context window, calculate the conditional probability between word vectors through normalized exponentiation, and generate basic word vectors; extract the character sequence of each word in the basic word vectors, use a sliding window to scan to obtain local character combination features, and construct character-level feature vectors; Identify the root word and affix structure based on the character-level feature vectors, combine the corresponding word form change rules with the basic word vectors, and generate word form vectors; analyze the syntactic dependency relationship of the word form vectors to construct a syntactic tree, extract syntactic dependency features and fuse them with the word form vectors to obtain semantically enhanced vectors; Extract concept nodes and relationship edges from a preset dictionary, calculate the similarity between the semantically enhanced vectors and the concept nodes to establish entity mappings; extract concept hierarchy relationships and attribute constraint information based on the entity mappings, and combine them with the semantically enhanced vectors to obtain knowledge-associated vectors; Calculate the dynamic adjustment factor of the basic word vectors and the knowledge-associated vectors based on vector quality evaluation metrics, and the vector quality evaluation metrics include vector clustering metrics, vector discrimination metrics, and vector coverage metrics; Adaptively fuse the basic word vectors and the knowledge-associated vectors according to the dynamic adjustment factor to obtain the final word vectors.
[0009] In an alternative embodiment, input standard intermediate configuration data into a deep decision tree, obtain the cluster resource status and load metrics through a computing resource evaluation node, obtain task running features through a task feature analysis node, and calculate and determine the final submission method based on an intelligent routing strategy, including: Receive standard intermediate configuration data, generate a deep decision tree model, and construct a resource evaluation node, a task feature analysis node, and a historical task analysis node; Collect cluster resource status metric data, perform normalization processing to obtain standardized resource metrics, input them into a neural network trained based on historical resource data, and generate a resource evaluation score as the first scoring metric; Parse task-related data from the standard intermediate configuration data, obtain task feature data through static analysis, construct a task feature vector based on the task feature data, input the task feature vector into a pre-trained task classification model, and generate a task feature score as the second scoring metric; Retrieve similar task records in the historical task database based on the task feature vector, extract task execution data from the similar task records, and calculate and generate a historical task score as the third scoring metric; Calculate the weight coefficients through the backpropagation algorithm, perform weighted fusion on the first scoring metric, the second scoring metric, and the third scoring metric, and generate a routing comprehensive score; Calculate the dynamic routing threshold based on the cluster resource status and task distribution, compare the routing comprehensive score with the dynamic routing threshold, and determine the final submission method; Calculate the predicted deviation value between the final submission method and the actual execution result, and perform online updates on the neural network, task classification model, and weight coefficients according to the predicted deviation value.
[0010] In an alternative embodiment, generating the resource evaluation score includes: Convert the standardized resource metrics into a feature matrix, perform a convolution operation on the feature matrix to extract resource status features, input the resource status features into a fully connected layer, and obtain the resource evaluation score through sigmoid function mapping; Generating the task feature score includes: Input the task feature vector into the attention layer to extract key features, perform temporal analysis on the extracted key features based on the long short-term memory network, calculate the probability distribution of different task types using the softmax function, and calculate the task feature score through weighted averaging according to the probability distribution; Generating the historical task score includes: Calculate the similarity between the task feature vector and the historical task record based on cosine similarity, select the historical task record with the highest similarity according to the preset selection ratio, extract the corresponding execution duration, resource utilization rate, and success rate, and calculate the historical task score through weighted averaging.
[0011] In an alternative embodiment, based on the template mapping mechanism, select the corresponding template according to the final submission method, and combine the historical task data to perform intelligent filling and optimization of the template parameters to generate the task submission instruction, including: Establish a two-layer template mapping mechanism. The first layer selects the corresponding template type based on the final submission method, and the second layer constructs a resource requirement index for the template type. The resource requirement index divides the task resource requirements into different grades, and each grade corresponds to a template parameter configuration scheme; Identify the critical path of the parameter configuration from the historical task data. The critical path is determined by analyzing the influence degree of parameter changes on the task execution effect. Construct the parameter configuration combinations on the critical path into a parameter pattern library, where each parameter pattern contains a complete parameter dependency chain; Determine the basic template in the two-layer template mapping mechanism according to the final submission method, match the parameter configuration scheme in the resource requirement index based on the resource requirements of the current task, and select the parameter pattern that adapts to the parameter configuration scheme from the parameter pattern library; Apply the parameter pattern to the base template through iterative replacement, perform parameter consistency verification after each replacement, finally generate a task submission instruction, and feedback the task execution result after the task is executed to update the parameter pattern library.
[0012] In an alternative embodiment, identifying the critical path of parameter configuration includes: Construct a task execution directed graph, where nodes are determined based on the execution stage, and edges between nodes are determined based on the data flow relationship between execution stages; Mark configurable parameters on each node of the task execution directed graph, count the change records of the configurable parameters in the historical task data, and calculate the influence degree of the configurable parameters on adjacent execution stages; Adopt a message passing algorithm to perform iterative calculations along the edges of the task execution directed graph, obtain the influence propagation range of each configurable parameter, and calculate the cumulative influence value of the configurable parameter based on the influence propagation range; Sort the configurable parameters according to the cumulative influence value, select the parameter sequence with the largest cumulative influence value, and determine the critical path; Construct a parameter pattern library, where each parameter pattern in the parameter pattern library contains the complete parameter dependency chain on the critical path; Convert the critical path into a parameter pattern and store it in the parameter pattern library, establish a parameter pattern score based on the usage frequency and task success rate, delete the parameter patterns with scores lower than the preset score lower threshold, and update the parameter pattern library.
[0013] In the second aspect of the embodiments of the present invention, a big data job intelligent submission system for a container cluster is provided, including: A first unit for receiving task configuration information submitted by a user; A second unit for performing field mapping and configuration standardization conversion based on field feature similarity calculation and historical call relationships, and generating standard intermediate configuration data in combination with version compatibility evaluation; A third unit for inputting the standard intermediate configuration data into a deep decision tree, obtaining the cluster resource status and load metrics through a computing resource evaluation node, obtaining task running characteristics through a task feature analysis node, and calculating and determining the final submission method based on an intelligent routing strategy; A fourth unit for selecting a corresponding template according to the final submission method based on a template mapping mechanism, and performing intelligent filling and optimization of the template parameters in combination with historical task data to generate a task submission instruction; A fifth unit for generating a resource definition configuration file and submitting a task through a container cluster command according to the task submission instruction when the final submission method is the first submission method; and generating a command line tool for the computing framework to submit a task when the final submission method is the second submission method. The sixth unit is used to monitor the execution status of real-time tasks, collect status information, resource information, and operation logs, and feedback them to the intelligent routing policy for optimizing the submission decisions of subsequent tasks.
[0014] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0015] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0016] In the embodiments of the present invention, for the intelligent submission method of big data jobs for container clusters, through field mapping and configuration standardization conversion technologies, task configuration information from different sources is converted into a standard intermediate format, effectively solving the configuration compatibility problem between heterogeneous systems, significantly improving the system integration efficiency and data conversion accuracy rate; adopting a deep decision tree combined with an intelligent routing policy, realizing the automated decision-making of task submission methods, dynamically selecting the optimal submission path according to the cluster resource status and task characteristics, greatly reducing the resource scheduling conflict rate, improving the overall resource utilization rate of the cluster, and at the same time reducing the task queuing waiting time; the adaptation mechanism and intelligent parameter filling technology can automatically optimize submission parameters according to historical task data, form a closed-loop optimization system in combination with real-time monitoring feedback, not only simplify the user operation process, reduce the configuration error rate, but also improve the system's adaptability to various tasks through continuous learning, making the task execution efficiency increase by more than 30% on average. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of the intelligent submission method of big data jobs for container clusters in the embodiments of the present invention; Figure 2 It is a schematic flowchart of generating word vectors; Figure 3 It is a framework diagram for uniformly submitting Spark jobs on a Kubernetes cluster. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0020] Figure 1 It is a schematic flowchart of a method for intelligent submission of big data jobs for a container cluster according to an embodiment of the present invention. As Figure 1 shown, the method includes: Receiving task configuration information submitted by a user; Based on field feature similarity calculation and historical call relationships, performing field mapping and configuration standardization conversion, and combining version compatibility evaluation to generate standard intermediate configuration data; Inputting the standard intermediate configuration data into a deep decision tree, obtaining the cluster resource status and load metrics through a computing resource evaluation node, obtaining the task running characteristics through a task feature analysis node, and determining the final submission method based on intelligent routing policy calculation; Based on a template mapping mechanism, selecting a corresponding template according to the final submission method, and combining historical task data to perform intelligent filling and optimization of template parameters to generate a task submission instruction; According to the task submission instruction, when the final submission method is the first submission method, generating a resource definition configuration file and submitting the task through a container cluster command; when it is the second submission method, generating a command-line tool for a computing framework to submit the task; Real-time monitoring the execution status of the task, collecting status information, resource information, and running logs, and feeding them back to the intelligent routing policy for optimizing subsequent task submission decisions.
[0021] In a specific implementation, task configuration information submitted by the user is received. This information contains multiple fields, such as task name "data_processing_job", computing resource request "cpu: 8 cores, memory: 16GB", task dependency "input_dataset: hdfs: / / data / input", etc. A field feature library is maintained, which stores the field mapping relationships under different versions and different computing frameworks. For the task configuration submitted by the user, standardization conversion is performed by calculating the field similarity. The field similarity calculation uses the term frequency-inverse document frequency (TF-IDF) technology to compare the current field with the standard fields in the feature library. For example, when the "num_workers" field in the user configuration is detected, its similarity with each field in the standard field library is calculated, and it is found that the similarity with the standard field "worker_count" is 0.92, exceeding the threshold of 0.85. Therefore, "num_workers: 5" is mapped to the standard field "worker_count: 5".
[0022] Analyze the historical task call relationships and construct a field association graph. Taking the task dependency configuration as an example, it is found that "depends_on: job_A" configured by the user co-occurs with the standard field "prerequisites" in 87% of the cases in historical calls. Therefore, it is mapped to "prerequisites: job_A". For version compatibility evaluation, a compatibility matrix for different versions is maintained. For example, for the user configuration "spark.executor.memory: 4g", check that the running Spark version in the current cluster environment is 3.2.1. Through the compatibility matrix, it is confirmed that this configuration is valid in the current version, and the original configuration is retained. Through the above processing, the user's heterogeneous configuration information is converted into standard intermediate configuration data, such as {"job_name": "data_processing_job", "worker_count": 5, "prerequisites": ["job_A"], "resource_request": {"cpu": 8, "memory": "16GB"}}.
[0023] After the standard intermediate configuration data is generated, it is input into a deep decision tree for intelligent routing analysis. The resource evaluation node of the decision tree first obtains the real-time resource status from the cluster monitoring system, including indicators such as the CPU usage rate, memory occupancy, and network bandwidth of each node. Taking the submission of a certain task as an example, the overall resource status of the cluster is obtained as follows: the average CPU usage rate is 65%, the total available memory is 45GB, and the proportion of high-load nodes is 28%. The task feature analysis node of the decision tree will extract the key features of the current task, such as compute-intensive, I / O-intensive, or hybrid. For the data processing task submitted by the user, the system analyzes its configuration information and extracts the feature vector [0.82, 0.45, 0.23], representing the compute intensity, I / O intensity, and memory intensity.
[0024] Based on the cluster resource status and task features, the final submission method is determined through the calculation of an intelligent routing strategy. The intelligent routing strategy includes three dimensions: resource matching score, historical performance data, and task priority. The calculation method of the resource matching score is the matching degree between the resource requirements of the current task and the available resources of each node, and the score range is 0-100. Taking the data processing task in this example as an example, the resource matching degree for submitting by container is calculated to be 78 points, and the matching degree for submitting by the computing framework command line is 62 points. The historical performance database will also be queried, and it is found that the average completion time of tasks with similar features in the container environment is 45 minutes, and it is 52 minutes in the direct command line submission method. Combining the task priority "normal", the system finally decides to select the first submission method, that is, to submit the task through the container cluster command.
[0025] After the submission method is determined, specific task submission instructions are generated based on the template mapping mechanism. A template library is maintained, and the corresponding template definitions are stored for different submission methods. For the container cluster submission method, the container configuration file template is selected.
[0026] Based on the historical task data, intelligent filling and optimization of the template parameters are performed. For the CPU resource request, analyzing the historical data finds that for tasks with similar features, setting the CPU request to 1.2 times the actual requirement can obtain more stable performance. Therefore, the original request "cpu: 8 cores" is optimized to "cpu: 10 cores". For the container image selection, according to the task type and historical success rate, the most suitable image version "data-processor:v2.3" is selected from the image library. This image has a 98% success rate in similar tasks, and the final container configuration file is generated.
[0027] Submit this configuration file using the container cluster command-line tool: kubectl apply -f job_config.yaml to complete the task submission. If the decision result is the second submission method, commands will be generated based on the computing framework command-line template. For Spark tasks, an exemplary command is "spark-submit --master yarn --deploy-mode cluster --executor-memory 16g --num-executors 5 --class ProcessorMain processor.jar hdfs: / / data / input hdfs: / / data / output".
[0028] After the task is submitted, monitor the task execution status in real time. The monitoring module collects task status information every 30 seconds by integrating the API of the cluster management system, including the execution stage, completed percentage, resource usage, etc. For tasks running in the container cluster, collect container-level metrics such as CPU usage, memory occupancy, and network I / O. For example, the average CPU usage of the data_processing_job task during runtime is 78%, the peak memory occupancy is 14.2GB, and the network read speed is 125MB / s. Also collect the running logs of the task and identify potential performance bottlenecks or error patterns through log analysis.
[0029] These monitoring data are stored in the performance database and fed back to the intelligent routing policy to optimize the subsequent task submission decisions. Use a sliding window mechanism to focus on the execution data of the last 100 similar tasks, and continuously optimize the parameters of the decision model through incremental learning. Through the closed-loop optimization mechanism, it can adapt to changes in the cluster environment and evolution of task characteristics, and continuously improve the submission efficiency and execution performance of big data jobs.
[0030] In an alternative embodiment, based on field feature similarity calculation and historical call relationships, perform field mapping and configuration standardization transformation, combined with version compatibility evaluation, to generate standard intermediate configuration data including:[[]]END]] Construct feature vectors for the source field and the target field respectively. The feature vector construction process uses a preset tokenizer to tokenize the field name, generates a basic word vector using the log-likelihood function of the central word and context words, combines multi-dimensional semantic enhancement to generate a word vector, and at the same time extracts the data type identifier and boundary values of the value range of the field to construct a field feature mapping table; Calculate the mapping correlation degree between the source field and the target field based on the field feature mapping table, extract the field call data from the historical task records, establish a field call relationship table, and count the co-occurrence frequency between fields to obtain the call weight; perform a weighted combination of the mapping correlation degree and the call weight to generate a field mapping score table, and select the target field with the highest score as the initial mapping field set; Extract the calculation engine fields from the initial mapping field set, parse the version information and the corresponding configuration items, and construct a version configuration association table; based on the version configuration association table, analyze the differences in the configuration items of each version to generate a configuration migration rule set; perform a version compatibility evaluation on the configuration items according to the configuration migration rule set to obtain a version-compatible configuration set; Perform a standardized conversion of the configuration items based on the version-compatible configuration set to generate standard configuration items; perform an integrity verification on the standard configuration items, supplement the configuration dependency relationship, and generate standard intermediate configuration data.
[0031] In a specific implementation, perform feature vector construction processing on the source field and the target field respectively. Taking the resource configuration field common in big data task configuration as an example, the source field "executor_cores" needs to be mapped to the corresponding field in the target environment. The system uses a preset word segmenter to segment the field name. The word segmenter adopts the bidirectional maximum matching algorithm and segments "executor_cores" into ["executor", "cores"]. The segmentation result is used as the basic corpus, and the logarithmic likelihood function of the central word and the context words is used to generate the basic word vector. Specifically, a special vocabulary for the big data field is maintained, which contains about 5,000 common technical terms, and a 128-dimensional basic vector is pre-calculated for each term. For the word "executor", the values of the first 5 dimensions in its basic vector are [0.35, -0.12, 0.47, 0.22, -0.31]; for the word "cores", the first 5 dimensions of its basic vector are [0.28, 0.43, -0.15, 0.37, 0.19].
[0032] To enhance semantic expression ability, the final word vectors are generated by combining multi-dimensional semantic enhancement technology. Through the context-aware mechanism, the basic vector of a word is fused with its context information in the big data domain corpus. The context related to "executor" and "cores" is extracted from the preprocessed 100TB big data configuration corpus, such as they often co-occur with words like "memory", "instances", etc. After fusion, the first 5 dimensions of the enhanced word vector of "executor" become [0.42, -0.09, 0.51, 0.25, -0.28], and the enhanced word vector of "cores" becomes [0.31, 0.46, -0.12, 0.40, 0.22]. At the same time, the data type identifier and the boundary values of the value range of the field are extracted. For "executor_cores", its data type is integer (INT), and the value range is [1, 32]. Combining the word vectors and type information, a field feature mapping table is constructed, which stores 256-dimensional feature vectors and type information for each field.
[0033] Based on the field feature mapping table, the mapping correlation degree between the source field and the target field is calculated. The calculation method is a weighted combination of the cosine similarity of the feature vectors and the type matching degree. There are multiple candidate fields in the target environment, such as "num_cores", "core_count", "vcpu_num", etc. The cosine similarity of the feature vectors between "executor_cores" and each candidate field is calculated, and the result is ["num_cores": 0.87, "core_count": 0.82, "vcpu_num": 0.76]. In terms of type matching degree, both "num_cores" and "core_count" are integers and their value ranges cover [1, 32], so the matching degree is 1.0; while "vcpu_num" is an integer but its value range is [2, 16], so the matching degree is 0.8. The comprehensive calculated mapping correlation degree is ["num_cores": 0.91, "core_count": 0.88, "vcpu_num": 0.77].
[0034] Extract field call data from historical task records and establish a field call relationship table. Among the 50,000 tasks successfully executed in the past 6 months, the system counts the co-occurrence frequencies of the source field "executor_cores" and each target field: it co-occurs with "num_cores" 8,500 times, with "core_count" 6,200 times, and with "vcpu_num" 2,300 times. Calculate the call weights based on the co-occurrence frequencies to obtain ["num_cores": 0.5, "core_count": 0.36, "vcpu_num": 0.14]. Perform a weighted combination of the mapping correlation degree and the call weights with weight coefficients of 0.6 and 0.4 respectively to generate a field mapping score table: ["num_cores": 0.746, "core_count": 0.672, "vcpu_num": 0.514]. According to the score table, select the "num_cores" with the highest score as the initial mapping field to form the initial mapping field set.
[0035] Extract the fields related to the computing engine from the initial mapping field set. Taking the Spark engine configuration as an example, the source configuration contains the field "spark.executor.memory = 4g". Parse the version information to identify the source environment as Spark 2.4.5 and the target environment as Spark 3.2.1. Build a version configuration association table that contains the corresponding relationships of configuration items between different versions. The association table shows that "spark.executor.memory" remains the same in both versions, but there are differences in the unit representation: the 2.x version supports the units "g" and "m", and the 3.x version is normalized to "GB" and "MB".
[0036] Based on the version configuration association table, parse the configuration item differences to generate a configuration migration rule set. For "spark.executor.memory = 4g", the migration rule specifies converting the unit "g" to "GB", and after conversion, it becomes "spark.executor.memory = 4GB". Perform a version compatibility evaluation on the configuration item to check whether the configuration item exists in the target version, whether the parameter types are compatible, and whether the value ranges are legal. The evaluation result shows that this configuration item is fully compatible and is added to the version-compatible configuration set.
[0037] Perform standardized conversion of configuration items based on the version-compatible configuration set. The standardization process includes name normalization, value format normalization, and unit unification. For "spark.executor.memory=4GB", the standardized configuration item is "memory_per_executor=4GB", which is the standard intermediate format defined by the system. Configuration related to containers is also converted, such as standardizing "container.memory_limit=8g" to "container_memory_limit=8GB".
[0038] Performing integrity verification on standard configuration items is a crucial step to ensure configuration validity. Check whether required fields exist, whether related configuration items conflict, and whether dependencies are met. For example, when "memory_per_executor=4GB" is configured, it is detected that the "executor_count" configuration is missing. According to empirical rules and resource estimation, "executor_count=5" is automatically supplemented. At the same time, the system analyzes the dependency relationship between configuration items. For example, the product of "memory_per_executor" and "executor_count" should not exceed the available memory of the cluster nodes. During the verification process, it is found that the total memory requirement is 20GB, which meets the limit of 32GB memory of the cluster nodes. Finally, standard intermediate configuration data is generated.
[0039] This standard intermediate configuration data becomes the basis for subsequent intelligent routing decisions and task submissions. Through technical means such as word vector construction, mapping correlation calculation, historical call analysis, and version compatibility evaluation, the system realizes accurate mapping and standardized conversion of configuration fields in heterogeneous environments, improving the submission success rate and running efficiency of big data jobs in the container cluster environment.
[0040] In an alternative implementation, a preset tokenizer is used to tokenize the field name, and a basic word vector is generated using the log-likelihood function of the central word and context words. Combining multi-dimensional semantic enhancement, the generated word vector includes: Perform vectorization processing on the tokenization result based on the central word and context words, maximize the log-likelihood function of the central word and the word sequence within the corresponding context window, calculate the conditional probability between word vectors through normalized exponentiation, and generate a basic word vector; extract the character sequence of each word in the basic word vector, use a sliding window to scan to obtain local character combination features, and construct a character-level feature vector; Based on the character-level feature vector, identify the root and affix structure, combine the corresponding word form change rules with the basic word vector to generate a word form vector; analyze the syntactic dependency relationship of the word form vector to construct a syntactic tree, extract syntactic dependency features and fuse them with the word form vector to obtain a semantic enhancement vector; Extracting concept nodes and relationship edges from a preset dictionary, calculating the similarity between the semantic enhancement vector and the concept node to establish entity mapping; extracting concept hierarchical relationships and attribute constraint information based on the entity mapping, and combining them with the semantic enhancement vector to obtain a knowledge association vector; The dynamic adjustment factors of the basic word vector and the knowledge association vector are calculated based on the vector quality evaluation index, and the vector quality evaluation index includes a vector clustering index, a vector discrimination index and a vector coverage index; the basic word vector and the knowledge association vector are adaptively fused according to the dynamic adjustment factor to obtain the final word vector.
[0041] Figure 2 The flowchart for generating word vectors is shown in a specific implementation. In a specific implementation, the word vector generation process first performs word segmentation preprocessing on the text. Word segmentation can be performed using word segmentation tools such as Jieba to segment the original text and obtain a word sequence. Taking "deep learning technology is widely used" as an example, the word segmentation result is "deep learning / technology / application / widely". When the word segmentation result is vectorized, the center word and the words in the context window are selected for processing. For example, "technology" is selected as the center word, and the context window size is set to 2, then the context words are "deep learning" and "application". The log-likelihood function of the center word and the context word is calculated to maximize the function value. In specific implementation, for the center word w and the context word c, its conditional probability P(c|w) can be calculated by vector dot product and normalized exponential function. The vector dimension is set to 300 dimensions, and the initial value of the learning rate is set to 0.025, which decreases linearly with the training process. The training corpus selects 100GB of Chinese text and iterates 5 times to obtain the initial basic word vector.
[0042] In the character-level feature extraction stage, the character sequence of each word is processed. Taking "学" as an example, the character sequence is "学" and "习". The character sequence is scanned by sliding window, with the window size set to 3 and the step size set to 1. For words of insufficient length, special padding symbols are added. For each character combination in the window, a local sensitive hash code is used to represent it and generate a 128-dimensional character-level feature vector. The character-level feature vector is concatenated with the basic word vector and then mapped back to the original dimension through the fully connected layer to maintain the consistency of the vector dimension.
[0043] In the morphological structure analysis stage, the root and affix structure is recognized based on character-level feature vectors. During implementation, a pre-constructed Chinese morpheme library is used, which contains 5,000 common roots and 200 common affixes. The components of a word are recognized through the forward maximum matching algorithm. For example, "re-study" can be recognized as the prefix "re-" + the root "study". The recognized morphological change rules include root mapping, prefix conversion, suffix change, etc., and a total of 50 transformation types are designed. The morphological change rules are represented as discrete features and combined with the basic word vectors in a residual connection manner to obtain a morphological vector with a dimension of still 300.
[0044] When performing syntactic dependency analysis, a transition-based dependency parser is used to construct a syntactic tree. The parser adopts a buffer-stack structure, and dependency relationships are established through operations such as left arc, right arc, and shift. The types of dependency relationships include subject-predicate relationship, attributive-head relationship, verb-object relationship, etc., a total of 30 types. The extracted syntactic features include information such as the type of dependency relationship of the current word, the dependency parent node, and the sibling nodes. These features are fused with the morphological vector through an attention mechanism to obtain a semantically enhanced vector. During the fusion process, the attention weights are dynamically adjusted according to the importance of the dependency relationships, and the key syntactic relationships obtain higher weights.
[0045] In the knowledge enhancement stage, a knowledge graph is constructed using a preset dictionary. The dictionary contains 100,000 core concept nodes and 1.5 million relationship edges. Entity mapping is established by calculating the cosine similarity between the semantically enhanced vector and the concept nodes, and the similarity threshold is set to 0.75. For example, the "neural network" represented by the vector can be mapped to the concept node of "artificial intelligence / deep learning / neural network" in the knowledge graph. The hierarchical relationships (such as "superordinate concept", "subordinate concept") and attribute constraint information (such as "field", "usage") of the concepts are extracted from the knowledge graph, and these knowledge information are represented as 100-dimensional vectors, concatenated with the semantically enhanced vector and passed through a non-linear transformation to obtain a knowledge-associated vector.
[0046] In the vector quality evaluation stage, a dynamic adjustment factor is calculated to balance the basic word vector and the knowledge-associated vector. The evaluation metrics include: the vector clustering metric, which calculates the ratio of the intra-class distance to the inter-class distance through K-means clustering, and the smaller the value, the better the clustering; the vector discrimination metric, which calculates the entropy value of the cosine similarity distribution between word vectors, and the larger the value, the better the discrimination; the vector coverage metric, which statistics the proportion of language phenomena that the vector can express, such as the polysemy coverage rate reaching more than 85%. The dynamic adjustment factor α is calculated according to the weighted values of these three metrics, and its value range is [0.3, 0.7], with an initial value set to 0.5 and updated every 500 batches of data.
[0047] When generating the final word vector, the basic word vector and the knowledge association vector are combined through an adaptive fusion method. When the quality of the basic word vector is high, the value of α tends to 0.7; when the knowledge association vector provides more semantic information, the value of α tends to 0.3. The fusion formula is: the final word vector = α × the basic word vector + (1 - α) × the knowledge association vector. The fused vector is subjected to L2 normalization to ensure that the modulus length of all word vectors is 1, which is convenient for subsequent similarity calculation.
[0048] Existing word vector technologies are mainly constructed based on the distributional hypothesis, such as Word2Vec, GloVe, etc., which only utilize word co-occurrence information and ignore the internal structure of words and external knowledge. The method of this embodiment addresses the limitations of the existing technology. Starting from the perspective of multi-level feature fusion, it introduces character-level features, word form structures, syntactic dependency relationships, and knowledge graph information to construct a complete word vector enhancement framework. Compared with the traditional method that only relies on co-occurrence statistics, the method of this embodiment improves in word sense discrimination and semantic relevance evaluation.
[0049] In an alternative embodiment, the standard intermediate configuration data is input into a deep decision tree. The cluster resource status and load metrics are obtained through a computing resource evaluation node, and the task running characteristics are obtained through a task feature analysis node. Based on the intelligent routing strategy calculation, the final submission method is determined as follows: Receive the standard intermediate configuration data, generate a deep decision tree model, and construct a resource evaluation node, a task feature analysis node, and a historical task analysis node; Collect the cluster resource status metric data, perform normalization processing to obtain standardized resource metrics, input them into a neural network trained based on historical resource data, and generate a resource evaluation score as the first scoring metric; Parse the task-related data from the standard intermediate configuration data, obtain the task feature data through static analysis, construct a task feature vector based on the task feature data, input the task feature vector into a pre-trained task classification model, and generate a task feature score as the second scoring metric; Retrieve similar task records in the historical task database based on the task feature vector, extract task execution data from the similar task records, and calculate and generate a historical task score as the third scoring metric; Calculate the weight coefficients through the backpropagation algorithm, perform weighted fusion on the first scoring metric, the second scoring metric, and the third scoring metric, and generate a routing comprehensive score; Calculate the dynamic routing threshold based on the cluster resource status and task distribution, compare the routing comprehensive score with the dynamic routing threshold, and determine the final submission method; Calculate the prediction deviation value between the final submission method and the actual execution result, and perform online updates on the neural network, the task classification model, and the weight coefficients according to the prediction deviation value.
[0050] In a specific embodiment, the resource evaluation node is responsible for real-time collection of cluster resource status indicator data, including CPU utilization rate, memory occupancy rate, network throughput, and storage I / O status, etc. The collected raw data is converted into standardized resource indicators between 0 and 1 through maximum-minimum normalization processing. For example, when the CPU utilization rate of a certain computing node is 75%, the maximum recorded value is 95%, and the minimum recorded value is 20%, the normalized value is (75 - 20) / (95 - 20) = 0.73. These standardized indicators are then input into a five-layer feedforward neural network pre-trained based on historical resource data. The network contains 64 input neurons, three hidden layers (with 128, 64, and 32 neurons respectively), and 16 output neurons. Through forward propagation calculation, a resource evaluation score is generated as the first scoring indicator, with a value range of 0 to 100. The higher the score, the more suitable the resource status is for receiving new tasks.
[0051] The task feature analysis node parses task-related data from the standard intermediate configuration data, including algorithm type, data scale, parallelism requirement, etc. Task feature data is obtained through static analysis. For example, for a data analysis task, information such as data input volume (e.g., 50GB), length of the processing operator chain (e.g., 8), and distribution of operator complexity (e.g., 80% are high-complexity operators) is extracted. Based on these data, a 128-dimensional task feature vector is constructed, with each dimension representing a feature attribute. This feature vector is input into a pre-trained task classification model. The model adopts a random forest structure and contains 100 decision trees, each with a depth of 12 layers. The model outputs a task feature score as the second scoring indicator, also with a value range of 0 to 100. The score reflects the resource consumption characteristics and execution difficulty of the task.
[0052] The historical task analysis node retrieves similar task records in the historical task database based on the task feature vector. The system uses a vector similarity retrieval method and determines tasks as similar when the cosine similarity of the feature vectors exceeds 0.85. For example, if the similarities between the current big data analysis task and 5 task records in the history are 0.92, 0.88, 0.87, 0.83, and 0.79 respectively, the first three records are extracted as similar tasks. Execution data such as execution time, resource utilization rate, and completion status is extracted from these task records, and a historical task score is calculated as the third scoring indicator, with a value range also of 0 to 100, reflecting the historical performance of the task in a specific execution environment.
[0053] Calculate the weight coefficients of the three scoring metrics through the backpropagation algorithm. The initial weights are set to 0.4 (resource evaluation), 0.35 (task characteristics), and 0.25 (historical tasks) respectively. After each task execution is completed, the system calculates the difference between the predicted result and the actual execution effect, and adjusts the weights through the gradient descent method. For example, after a certain execution, it is found that the resource evaluation contributes more to the result prediction, and the weights may be adjusted to 0.45, 0.32, 0.23. The adjusted weights are applied to subsequent tasks, and the three scoring metrics are weighted and fused to generate a comprehensive routing score. The calculation method is to sum the scores of each metric multiplied by the corresponding weights. For example, if the scores of a certain task in the three items are 80, 65, and 90 respectively, the comprehensive routing score calculated using the above weights is 0.45×80 + 0.32×65 + 0.23×90 = 77.35.
[0054] Calculate the dynamic routing threshold based on the cluster resource status and task distribution. This threshold is dynamically adjusted according to the current cluster load. When the cluster resources are sufficient (such as the average load is lower than 50%), the threshold is set lower (such as 60 points); when the resources are tense (such as the average load exceeds 80%), the threshold is increased (such as 85 points). Compare the comprehensive routing score with the dynamic routing threshold. When the score is higher than the threshold, select a high-performance execution method (such as a dedicated computing node); when the score is lower than the threshold, select a normal execution method (such as a shared resource pool). For example, when the comprehensive routing score is 77.35 and the current dynamic threshold is 75, the system selects a high-performance execution method.
[0055] After the task execution is completed, calculate the prediction deviation value between the final submission method and the actual execution result. The deviation value is calculated by comparing the difference between the expected execution time and the actual execution time. For example, if the expected time is 2 hours and the actual time used is 2.5 hours, the deviation rate is 25%. Online update the neural network, task classification model, and weight coefficients according to the prediction deviation value. For the neural network, use the actual resource consumption data for incremental training; for the task classification model, add the new task characteristics and execution results to the training set; for the weight coefficients, adjust the weight distribution according to the prediction accuracy of each metric. This continuous learning mechanism ensures that the system can adapt to the changing workload and resource environment.
[0056] In an alternative embodiment, generating the resource evaluation score includes: Convert the standardized resource metrics into a feature matrix, perform a convolution operation on the feature matrix to extract resource status features, input the resource status features into a fully connected layer, and obtain the resource evaluation score through the sigmoid function mapping; Generating the task characteristic score includes: Input the task feature vector into the attention layer to extract key features, conduct temporal analysis on the extracted key features based on the long short-term memory network, calculate the probability distribution of different task types using the softmax function, and calculate the task feature score through weighted calculation according to the probability distribution. The generation of the historical task score includes: Calculate the similarity between the task feature vector and the historical task records based on cosine similarity, select the historical task record with the highest similarity according to the preset selection ratio, extract the corresponding execution duration, resource utilization rate, and success rate, and calculate the historical task score through weighted average.
[0057] In a specific implementation, in the link of generating the resource evaluation score, collect and standardize various resource indicators. Taking a computing node as an example, possible resource indicators include CPU usage rate (0.65), memory occupancy rate (0.78), network bandwidth usage rate (0.42), storage I / O load (0.55), etc. These values are transformed into the 0-1 interval through min-max normalization to form standardized resource indicators. Organize these standardized indicators into a feature matrix, such as a 5×4 matrix, where the rows represent different computing nodes and the columns represent different resource dimensions. Use a 3×3 convolutional kernel to perform convolution operations on the feature matrix to extract resource status features. During the convolution operation, the sliding window moves on the feature matrix, and the product of the elements in the window and the corresponding elements of the convolutional kernel is calculated and summed each time to obtain a new feature representation. After being processed by the convolutional layer, the feature dimension may change from the original 5×4 to 3×2, representing the extracted high-level resource status features. These features are then flattened and input into a fully connected layer with 64 neurons, and linear transformation is performed through the weight matrix and bias term. Finally, apply the sigmoid function to map the output to between 0 and 1 to obtain the final resource evaluation score of 0.82, indicating the applicability of the current resource status.
[0058] In the task feature score generation phase, the input task feature vector is processed. Assume that the task feature vector contains 10 features such as task type, estimated execution time, resource requirements, etc., with values [2, 45, 8, 3, 0, 1, 5, 9, 7, 4]. This vector is input into the attention layer, which extracts key features by calculating the importance weights of each feature. For example, it is identified that the task type (weight 0.3), estimated execution time (weight 0.25), and resource requirements (weight 0.2) are the most important features. The attention mechanism generates a new feature representation [1.8, 42.75, 7.6, 2.4, 0, 0.9, 4.5, 7.2, 5.6, 3.2] by weighting the input features. These extracted key features are then input into a long short-term memory network (LSTM) for temporal analysis. The LSTM contains three control units: an input gate, a forget gate, and an output gate, and can capture the temporal dependencies between task features. Through LSTM processing, the system can understand the temporal patterns of different task features, such as the typical execution time patterns of specific types of tasks. The output of the LSTM is then passed through the softmax function to calculate the probability distribution of different task types, such as computational tasks (0.6), storage tasks (0.1), and network tasks (0.3). Based on the preset weights corresponding to these probability values (0.8, 0.5, 0.7 respectively), a weighted average is calculated to obtain a task feature score of 0.73.
[0059] In the historical task score generation phase, the cosine similarity is used to calculate the similarity between the current task and historical tasks. Assume that the current task feature vector is [2, 45, 8, 3, 0, 1, 5, 9, 7, 4], and there are five records in the historical database, with their feature vectors being [2, 40, 8, 3, 1, 1, 5, 8, 7, 4], [3, 45, 7, 3, 0, 2, 5, 9, 6, 4], [2, 47, 8, 2, 0, 1, 4, 9, 7, 3], [1, 30, 6, 4, 2, 1, 6, 7, 5, 3], and [4, 50, 9, 2, 1, 0, 4, 10, 8, 5]. The similarities calculated by the cosine similarity are 0.95, 0.92, 0.94, 0.80, and 0.85 respectively. The selection ratio is set to 60%, and the three records with the highest similarities (0.95, 0.92, 0.94) are selected. From these historical records, the corresponding execution durations (42 minutes, 48 minutes, 45 minutes respectively), resource utilization rates (0.75, 0.70, 0.72), and task success rates (0.98, 0.95, 0.97) are extracted. By setting weights (execution duration 0.4, resource utilization rate 0.3, success rate 0.3), the weighted average is calculated. For the execution duration, it is first normalized (the smaller the value, the better), and the historical task score obtained is 0.88.
[0060] The resource evaluation score (0.82), task feature score (0.73), and historical task score (0.88) are weighted and averaged according to preset weights (0.35, 0.35, and 0.3 respectively), and the comprehensive score of 0.807 is calculated. This score is used as the decision basis for task resource allocation, and it can be determined whether to accept the task or allocate what resources according to the score level. Through this multi-dimensional scoring mechanism, the resource status, task characteristics, and historical execution conditions can be comprehensively considered, improving the accuracy and efficiency of resource allocation.
[0061] In an alternative implementation, based on the template mapping mechanism, the corresponding template is selected according to the final submission method, and combined with historical task data, intelligent filling and optimization of the template parameters are performed to generate a task submission instruction including: A two-layer template mapping mechanism is established. The first layer selects the corresponding template type based on the final submission method, and the second layer constructs a resource requirement index for the template type. The resource requirement index divides the task resource requirements into different grades, and each grade corresponds to a template parameter configuration scheme; Identify the critical path of parameter configuration from historical task data. The critical path is determined by analyzing the impact degree of parameter changes on task execution effects. The parameter configuration combinations on the critical path are constructed into a parameter pattern library, where each parameter pattern contains a complete parameter dependency chain; Determine the basic template in the two-layer template mapping mechanism according to the final submission method, match the parameter configuration scheme in the resource requirement index based on the resource requirements of the current task, and select the parameter pattern adapted to the parameter configuration scheme from the parameter pattern library; Through iterative substitution, apply the parameter pattern to the basic template, perform parameter consistency verification after each substitution, finally generate a task submission instruction, and feedback the task execution result after the task is executed to update the parameter pattern library.
[0062] In a specific implementation, the establishment process of the double - layer template mapping mechanism involves the construction of mapping relationships at two levels. The first - layer mapping selects the corresponding template type based on the final submission method. For example, when the submission method is a distributed computing framework, a distributed computing template is selected; when the submission method is containerized deployment, a container template is selected; when the submission method is single - machine execution, a local execution template is selected. The second - layer mapping constructs a resource requirement index for the selected template type, which divides the task resource requirements into different grades. Taking the distributed computing template as an example, the computing resource requirements can be divided into three grades: the low grade corresponds to 1 - 4 CPU cores and 4 - 8 GB of memory; the medium grade corresponds to 5 - 16 CPU cores and 9 - 32 GB of memory; the high grade corresponds to 17 - 64 CPU cores and 33 - 128 GB of memory. Each grade is associated with a set of preset template parameter configuration schemes, including specific values of key parameters such as the number of executors, task parallelism, and memory allocation ratio.
[0063] The construction process of the parameter pattern library is completed by analyzing historical task data. The execution records of 1000 completed tasks are collected, including parameter configurations and execution effect data. By calculating the influence degree of parameter changes on task completion time, resource utilization rate, and execution success rate, the critical path of parameter configuration is identified. For example, in a certain type of data - processing task, it is found that the changes in the three parameters of executor memory configuration, shuffle parallelism, and data compression ratio have the greatest impact on the execution effect and constitute the critical path. The system constructs the parameter configuration combinations on these critical paths into the parameter pattern library. A typical parameter pattern contains a complete parameter dependency chain. For example, a pattern for processing a large amount of data may include a series of interdependent parameter combinations such as 8 GB of executor memory, 4 cores per executor, a dynamic allocation ratio of 0.6, a shuffle parallelism of 200, a compression algorithm of Snappy, and a compression ratio of 0.5.
[0064] Determine the base template in the double - layer template mapping mechanism according to the final submission method selected by the user. Suppose the user selects the distributed computing framework as the submission method and selects the distributed computing template as the base template. Then, analyze the resource requirement characteristics of the current task, such as the data volume is 500 GB and the expected processing time is 1 hour. According to these characteristics, match the parameter configuration scheme of the "medium grade" in the resource requirement index, and this scheme recommends using a configuration of 12 CPU cores and 24 GB of memory.
[0065] Select a parameter pattern from the parameter pattern library that is adapted to this parameter configuration scheme. By calculating the matching degree between the task characteristics and each parameter pattern in the pattern library, find the parameter pattern that best suits the current task. For example, a parameter pattern optimized for medium data volume is selected for the above task, and this pattern includes parameter configurations such as the number of executors 10, memory per executor 2.4GB, dynamic resource allocation ratio 0.5, and task parallelism 120.
[0066] Through an iterative replacement process, apply the selected parameter pattern to the base template. The iterative replacement is carried out in the order of parameter dependencies. First, replace the basic resource configuration parameters, then replace the execution parameters that depend on the basic configuration, and finally replace the optimization parameters. After each replacement, perform a parameter consistency check to ensure that there are no conflicts between parameters. For example, check whether the total memory allocation exceeds the system limit, and whether the parallelism setting matches the number of executors. If a conflict is found, the parameter value will be adjusted according to the preset rules to achieve consistency. Suppose it is detected that the parallelism does not match the number of executors, the system will automatically adjust the parallelism to an integer multiple of the number of executors to ensure load balancing.
[0067] After multiple rounds of replacement and verification, finally generate a complete task submission instruction, including all necessary configuration parameters and execution commands.
[0068] After the task is executed, collect the task execution result data, including indicators such as the actual execution time, resource utilization rate, and execution success rate. Suppose the actual execution time of this task is 45 minutes, the average CPU utilization rate is 85%, the memory utilization rate is 70%, and the execution is successful. Compare these results with the expected effect, calculate the effectiveness score of the parameter configuration, and update the parameter pattern library based on this. If the execution effect is better than the historical average level, increase the weight of this parameter pattern; if it is found that certain parameter combinations are particularly efficient, a new parameter pattern will also be created and added to the library. Through this feedback mechanism, the parameter pattern library is continuously optimized to improve the accuracy of future task configurations.
[0069] In an alternative embodiment, identifying the critical path of the parameter configuration includes: Construct a task execution directed graph, where nodes are determined based on the execution phases, and edges between nodes are determined based on the data flow relationship between the execution phases; Mark the configurable parameters on each node of the task execution directed graph, count the change records of the configurable parameters in the historical task data, and calculate the influence degree of the configurable parameters on the adjacent execution phases; Use the message passing algorithm to perform iterative calculations along the edges of the task execution directed graph, obtain the influence propagation range of each configurable parameter, and calculate the cumulative influence value of the configurable parameters based on the influence propagation range; Sort the configurable parameters according to the cumulative impact value, select the parameter sequence with the largest cumulative impact value, and determine the critical path; Construct a parameter pattern library, where each parameter pattern in the parameter pattern library contains a complete parameter dependency chain on the critical path; Convert the critical path into a parameter pattern and store it in the parameter pattern library, establish a parameter pattern score based on the usage frequency and task success rate, delete the parameter patterns with scores lower than the preset score lower threshold, and update the parameter pattern library.
[0070] In a specific implementation, the big data job submission process first constructs a task execution directed graph. Taking the Spark big data processing task as an example, according to its execution stages, it is divided into four main stages: data input, data transformation, aggregation calculation, and result output. Each stage is used as a node in the directed graph. The edges between the nodes represent the data flow relationship between the execution stages. For example, there is a data flow between the data input node and the data transformation node, so a directed edge pointing from the data input to the data transformation is established in the directed graph. For complex distributed computing tasks, there may be multiple parallel transformation and aggregation stages, forming a more complex graph structure. In practical applications, a typical Hadoop MapReduce task can be constructed as a directed graph with five nodes: input, Map, Shuffle, Reduce, and output; while a Spark SQL query task can be constructed as a directed graph with five nodes: parsing, optimization, physical plan generation, task distribution, and execution.
[0071] On the constructed task execution directed graph, mark the configurable parameters for each node. For Spark jobs, the parameters marked on the data input node include spark.sql.files.maxPartitionBytes, spark.sql.shuffle.partitions, etc.; the parameters marked on the data transformation node include spark.default.parallelism, spark.serializer, etc.; the parameters marked on the aggregation calculation node include spark.memory.fraction, spark.memory.storageFraction, etc.; the parameters marked on the result output node include spark.hadoop.mapreduce.fileoutputcommitter.algorithm.version, etc. By analyzing the task records of 500 successful executions in the historical task database, the change situation of each parameter is statistically analyzed. For example, the value range of the spark.sql.shuffle.partitions parameter in the historical records is from 200 to 2000, the average value is 800, and the standard deviation is 300, which indicates that this parameter has large volatility.
[0072] When calculating the influence degree of configurable parameters on adjacent execution stages, the parameter sensitivity analysis method is adopted. By the method of controlling variables, while keeping other parameters unchanged, the value of a single parameter is adjusted, and the change of task execution performance is observed. Taking the spark.sql.shuffle.partitions parameter as an example, when this parameter is adjusted from the default value of 200 to 800, the execution time of the data conversion stage is reduced by 25%, but the memory usage increases by 40%; when adjusted to 2000, the execution time is reduced by 5%, but the memory usage increases by 120%. Based on these observation results, the influence coefficient of the parameter on the adjacent stage is calculated. The value range of the influence coefficient is from 0 to 1, where 0 means no influence and 1 means the greatest influence. The influence coefficient of spark.sql.shuffle.partitions on the data conversion stage is 0.85, and the influence coefficient on the aggregation calculation stage is 0.72.
[0073] The message passing algorithm is used to calculate the propagation range of parameter influence. This algorithm starts from each parameter node and performs iterative propagation calculations along the edges of the directed graph. The number of iterations is set to 5. In each iteration, the node multiplies its own influence value by the influence coefficient and passes it to the adjacent node. For example, the initial influence value of the spark.sql.shuffle.partitions parameter at the data conversion node is 1.0. After the first iteration, the influence value passed to the aggregation calculation node is 1.0×0.72 = 0.72; after the second iteration, the influence value passed to the result output node is 0.72×0.45 = 0.324. After the iteration is completed, the influence values of each parameter on all nodes are accumulated to obtain the cumulative influence value of the parameter. The cumulative influence value of spark.sql.shuffle.partitions is 2.52, and the cumulative influence value of spark.memory.fraction is 1.85.
[0074] The parameters are sorted according to the cumulative influence values, and the top 10 parameters with the largest influence values are selected to form a sequence to determine the critical path. In the experimental analysis, the critical path includes parameters such as spark.sql.shuffle.partitions(2.52), spark.default.parallelism(2.35), spark.memory.fraction(1.85), spark.executor.cores(1.76), spark.executor.memory(1.73), etc. These parameters constitute the main influence factor chain in the task execution process, that is, the parameter dependency chain.
[0075] When building the parameter pattern library, the parameter dependency chain on the critical path is regarded as a complete parameter pattern. A typical parameter pattern contains the following structure: {taskType: "spark-sql", dataSize: "large", complexity: "high", paramChain: [{param: "spark.sql.shuffle.partitions", value: 1000}, {param: "spark.default.parallelism", value: 200}, {param: "spark.memory.fraction", value: 0.7}]}. The initial capacity of the parameter pattern library is set to 100 and it is continuously expanded and optimized as the system runs.
[0076] The scoring mechanism of the parameter pattern is calculated based on the usage frequency and task success rate. The usage frequency is defined as the number of times this pattern has been applied in the last 30 days divided by the average number of applications of all patterns; the task success rate is defined as the proportion of tasks using this pattern that have been successfully completed. The final score is the weighted sum of the usage frequency and the task success rate, with weights of 0.4 and 0.6 respectively. The scoring range is from 0 to 10, and the preset lower threshold of the score is 4.5. For example, a certain parameter pattern has been applied 15 times in the last 30 days, and the average number of applications of all patterns is 10, so the usage frequency is 1.5; the task success rate of this pattern is 85%, then the score is 0.4×1.5 + 0.6×8.5 = 5.7, which is higher than the lower threshold and is retained in the pattern library. The pattern library is maintained once a week, deleting patterns with scores lower than the threshold, and at the same time adding newly identified efficient parameter patterns to the library.
[0077] The size of the parameter pattern library is adjusted dynamically. The initial capacity is 100 patterns, and it can grow to 500 patterns as the application scenario expands. To prevent the pattern library from expanding excessively, a dual control mechanism of the maximum capacity limit and the minimum score requirement is set. When the library capacity is close to the upper limit, the lower threshold of the score is automatically increased by 0.5 to ensure that only the highest-quality parameter patterns are retained.
[0078] The existing big data job submission technologies mainly adopt fixed parameter templates or simple resource estimation methods, which are difficult to adapt to complex and changeable computing environments. Traditional methods such as rule-based configuration systems lack in-depth understanding of the dependency relationships between parameters, resulting in unreasonable resource allocation and low task execution efficiency. The method of this embodiment identifies and understands the complex dependency relationships between task parameters, and establishes a parameter critical path identification mechanism through constructing a task execution directed graph and impact propagation analysis. Compared with the existing technologies, the method of this embodiment improves the parameter optimization accuracy, increases the task execution success rate, and raises the average resource utilization rate. Especially in the scenarios of processing large-scale data and complex computing logics, the task completion time is shortened, greatly improving the resource efficiency and task throughput of the container cluster.
[0079] In an optional specific implementation manner, there are currently two mainstream ways to run Apache Spark jobs on a Kubernetes cluster: one is to submit jobs using the spark-submit command-line tool, and the other is to manage Spark jobs through the Kubernetes Operator mode. These two ways have their own advantages and disadvantages, but the configurations are incompatible. Users need to learn two different configuration systems, and it is very troublesome to switch between the two submission methods. It is necessary to modify configurations, parameters, submission commands, etc., and seamless switching cannot be achieved.
[0080] The intelligent submission method of this embodiment receives the task configuration information submitted by the user. The configuration information usually includes basic information such as task ID, computing engine type, engine version, task name, and possible task parameters. After receiving this information, it enters the configuration standardization processing stage.
[0081] In the configuration standardization processing stage, based on the field feature similarity calculation and historical call relationships, field mapping and configuration standardization conversion are performed. Specifically, the source fields and target fields are respectively processed for constructing feature vectors. In the process of constructing feature vectors, a preset tokenizer is used to tokenize the field names, the logarithmic likelihood function of the central word and the context words is used to generate basic word vectors, and multi-dimensional semantic enhancement is combined to generate word vectors. At the same time, the data type identifiers of the fields and the boundary values of the value ranges are extracted to construct a field feature mapping table.
[0082] During the construction of word vectors, the word segmentation results are vectorized based on the central word and context words. The log-likelihood function of the central word and the word sequence within the corresponding context window is maximized, and the conditional probability between word vectors is calculated through the normalized exponential to generate the basic word vectors. The character sequence of each word in the basic word vectors is also extracted, and local character combination features are obtained by scanning with a sliding window to construct character-level feature vectors. Based on the character-level feature vectors, the root word and affix structure are identified to generate word form vectors, and their syntactic dependency relationships are analyzed to construct a syntactic tree, finally obtaining the semantic enhanced vectors. Concept nodes and relationship edges are also extracted from the preset dictionary to establish entity mappings, forming knowledge association vectors, and adaptive fusion is performed through a dynamic adjustment factor to obtain the final word vectors.
[0083] Based on the field feature mapping table, calculate the mapping correlation degree between the source field and the target field, extract the field call data from the historical task records, establish a field call relationship table, and count the co-occurrence frequency between fields to obtain the call weight. The mapping correlation degree and the call weight are weighted and combined to generate a field mapping score table, and the target field with the highest score is selected as the initial mapping field set.
[0084] Extract the computing engine fields from the initial mapping field set, parse the version information and the corresponding configuration items, construct a version configuration association table, and based on this, analyze the differences in configuration items for each version to generate a configuration migration rule set. Perform a version compatibility assessment on the configuration items according to the configuration migration rule set to obtain a version-compatible configuration set, and then perform a standardized conversion of the configuration items to generate standard configuration items. Finally, complete the integrity verification, supplement the configuration dependency relationships, and generate standard intermediate configuration data.
[0085] After obtaining the standard intermediate configuration data, input it into the deep decision tree. Obtain the cluster resource status and load metrics through the computing resource assessment node, and obtain the task running characteristics through the task feature analysis node. Calculate and determine the final submission method based on the intelligent routing strategy. Specifically, construct a deep decision tree model, including a resource assessment node, a task feature analysis node, and a historical task analysis node.
[0086] Collect the cluster resource status metric data, perform normalization processing to obtain the standardized resource metrics, input them into the neural network trained based on historical resource data, and generate a resource assessment score as the first scoring metric. Parse the task-related data from the standard intermediate configuration data, obtain the task feature data through static analysis, construct a task feature vector, input it into the pre-trained task classification model, and generate a task feature score as the second scoring metric. Retrieve similar task records in the historical task database based on the task feature vector, extract the task execution data, and calculate and generate a historical task score as the third scoring metric.
[0087] The calculation process of the resource evaluation score is to convert the standardized resource indicators into a feature matrix, perform a convolution operation on the feature matrix to extract the resource status features, input the resource status features into the fully connected layer, and obtain the resource evaluation score through the mapping of the sigmoid function. The calculation process of the task feature score is to input the task feature vector into the attention layer to extract the key features, perform a temporal analysis on the extracted key features based on the long short-term memory network, calculate the probability distribution of different task types using the softmax function, and calculate the task feature score through weighted calculation according to the probability distribution. The historical task score is calculated based on the cosine similarity of the task feature vector and the historical task records, select the historical task record with the highest similarity, extract the corresponding execution duration, resource utilization rate, and success rate, and calculate through weighted average.
[0088] Calculate the weight coefficients through the backpropagation algorithm, perform weighted fusion on the three scoring indicators, and generate the comprehensive routing score. Calculate the dynamic routing threshold based on the cluster resource status and task distribution, compare the comprehensive routing score with the dynamic routing threshold, and determine the final submission method, that is, select the spark-submit method or the spark-operator method. The prediction deviation value between the final submission method and the actual execution result will also be calculated, and the neural network, task classification model, and weight coefficients will be updated online according to the deviation value to continuously optimize the decision-making accuracy.
[0089] After determining the final submission method, based on the dynamic adaptation mechanism, select the corresponding command template according to the final submission method, and combine the historical task data to perform intelligent filling and optimization of the template parameters to generate the task submission instruction. Specifically, establish a two-layer template mapping mechanism. The first layer selects the corresponding template type based on the final submission method, and the second layer constructs a resource requirement index for the template type, divides the task resource requirements into different grades, and each grade corresponds to a template parameter configuration scheme.
[0090] Identify the critical path of parameter configuration from the historical task data, which is determined by analyzing the influence degree of parameter changes on the task execution effect. Construct a parameter pattern library with the parameter configuration combinations on the critical path, and each parameter pattern contains a complete parameter dependency chain. The process of identifying the critical path is to construct a task execution directed graph, where the nodes are determined based on the execution stages, the edges between the nodes are determined based on the data flow relationship between the execution stages, mark the configurable parameters on each node, count their change records in the historical task data, and calculate the influence degree on the adjacent execution stages. Through the message passing algorithm, perform iterative calculations along the edges of the task execution directed graph to obtain the influence propagation range of each configurable parameter, calculate the cumulative influence value, and select the parameter sequence with the largest cumulative influence value according to the influence value ranking to determine the critical path.
[0091] According to the final submission method, determine the base template in the double-layer template mapping mechanism, match the parameter configuration scheme in the resource requirement index based on the resource requirements of the current task, select the appropriate parameter mode from the parameter mode library, apply the parameter mode to the base template through iterative replacement, and perform parameter consistency verification after each replacement, finally generating a task submission instruction.
[0092] When the final submission method is the first submission method (spark-operator), generate a resource definition configuration file and submit the task through the container cluster command; when it is the second submission method (spark-submit), generate a command-line tool for the computing framework to submit the task. The execution status of the task will be monitored in real time, status information, resource information, and running logs will be collected, and fed back to the intelligent routing policy for optimizing the subsequent task submission decision, forming a closed-loop optimization mechanism.
[0093] Through the above method, intelligent submission of big data jobs for container clusters is realized. Using a unified configuration method and submission entry, the optimal submission path is automatically selected, greatly simplifying the user operation process, improving the task submission efficiency and resource utilization rate, and providing a more intelligent solution for the operation of big data jobs in the container environment.
[0094] As Figure 3 shown, it shows the overall architecture and working process of the intelligent submission system for big data jobs for container clusters. Starting from the left, the user submits task configuration information through the Java standard interface (step 1 marked in the figure), and this information is sent to the conversion layer for processing (step 2). The conversion layer is the core component of the system, which depends on the historical task records and resource information stored in the database. Inside the conversion layer, there are two parallel processing paths: above is the Operator adapter (step 3), responsible for generating the SparkApplication CRD resource definition; below is the Spark-submit adapter (step 4), responsible for generating the spark-submit command. Based on the intelligent decision-making algorithm, the system selects the optimal submission method routing, and forwards the processed configuration information to the execution environment through the corresponding adapter. In the Kubernetes cluster environment on the right, the system starts the Spark task (step 5), executes the computing process, and outputs the running logs to the storage system. The whole process forms a complete closed loop, and the task execution data will be collected and stored in the database for optimizing the subsequent task submission decision. This architecture design effectively solves the problem of incompatibility between the two submission methods of Spark in the Kubernetes environment. Through the unified configuration layer and intelligent routing mechanism, it realizes the standardized processing of configuration and the selection of the optimal submission path, enabling users to not care about the underlying implementation details, greatly simplifying the operation process, and improving the system efficiency.
[0095] The intelligent big data job submission system for container clusters according to the embodiments of the present invention includes: A first unit, configured to receive task configuration information submitted by a user; A second unit, configured to perform field mapping and configuration standardization conversion based on field feature similarity calculation and historical call relationships, and generate standard intermediate configuration data in combination with version compatibility evaluation; A third unit, configured to input the standard intermediate configuration data into a deep decision tree, obtain the cluster resource status and load metrics through a computing resource evaluation node, obtain the task running characteristics through a task feature analysis node, and determine the final submission method based on intelligent routing strategy calculation; A fourth unit, based on a template mapping mechanism, selects a corresponding template according to the final submission method, and performs intelligent filling and optimization on the template parameters in combination with historical task data to generate a task submission instruction; A fifth unit, configured to, according to the task submission instruction, when the final submission method is the first submission method, generate a resource definition configuration file and submit the task through a container cluster command; when it is the second submission method, generate a command-line tool for a computing framework to submit the task; A sixth unit, configured to monitor the execution status of the task in real time, collect status information, resource information, and operation logs, and feedback them to the intelligent routing strategy for optimizing subsequent task submission decisions.
[0096] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0097] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0098] The present invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent submission of big data jobs for container clusters, characterized in that, Including: Receiving task configuration information submitted by a user; Based on field feature similarity calculation and historical call relationships, performing field mapping and configuration standardization conversion, and combining version compatibility evaluation to generate standard intermediate configuration data; Inputting the standard intermediate configuration data into a deep decision tree, obtaining the cluster resource status and load metrics through a computing resource evaluation node, obtaining task running characteristics through a task feature analysis node, and determining the final submission method based on intelligent routing policy calculation; Based on a template mapping mechanism, selecting a corresponding template according to the final submission method, and combining historical task data to perform intelligent filling and optimization of template parameters to generate a task submission instruction; According to the task submission instruction, when the final submission method is the first submission method, generating a resource definition configuration file and submitting the task through container cluster commands; When it is the second submission method, generating a command-line tool for the computing framework to submit the task; Real-time monitoring the execution status of the task, collecting status information, resource information, and running logs, and feeding them back to the intelligent routing policy for optimizing subsequent task submission decisions.
2. The method according to claim 1, characterized in that, Based on field feature similarity calculation and historical call relationships, performing field mapping and configuration standardization conversion, and combining version compatibility evaluation to generate standard intermediate configuration data including: Performing feature vector construction processing on the source field and the target field respectively. The feature vector construction processing uses a preset tokenizer to tokenize the field name, generates a basic word vector using the log-likelihood function of the central word and context words, combines multi-dimensional semantic enhancement to generate a word vector, and at the same time extracts the data type identifier of the field and the boundary values of the value range to construct a field feature mapping table; Calculating the mapping correlation degree between the source field and the target field based on the field feature mapping table, extracting field call data from historical task records, establishing a field call relationship table, and statistically obtaining the co-occurrence frequency between fields to obtain a call weight; performing weighted combination of the mapping correlation degree and the call weight to generate a field mapping score table, and selecting the target field with the highest score as the initial mapping field set; Extracting computing engine fields from the initial mapping field set, parsing the version information and corresponding configuration items, and constructing a version configuration association table; parsing the configuration item differences of each version based on the version configuration association table to generate a configuration migration rule set; performing version compatibility evaluation on the configuration items according to the configuration migration rule set to obtain a version-compatible configuration set; Performing configuration item standardization conversion based on the version-compatible configuration set to generate standard configuration items; performing integrity verification on the standard configuration items, complementing configuration dependency relationships, and generating standard intermediate configuration data.
3. The method according to claim 2, wherein Using a preset tokenizer to tokenize the field name, generating a basic word vector using the log-likelihood function of the central word and context words, and combining multi-dimensional semantic enhancement to generate a word vector including: Performing vectorization processing on the tokenization result based on the central word and context words, maximizing the log-likelihood function of the central word and the word sequence within the corresponding context window, calculating the conditional probability between word vectors through normalized exponentiation to generate a basic word vector; extracting the character sequence of each word in the basic word vector, using a sliding window scan to obtain local character combination features, and constructing a character-level feature vector; Identify the root and affix structures based on the character-level feature vectors, combine the corresponding word form change rules with the basic word vectors to generate word form vectors; analyze the syntactic dependency relationships of the word form vectors to construct a syntactic tree, extract syntactic dependency features and fuse them with the word form vectors to obtain semantically enhanced vectors; Extract concept nodes and relationship edges from a preset dictionary, calculate the similarity between the semantically enhanced vectors and the concept nodes to establish entity mappings; extract concept hierarchy relationships and attribute constraint information based on the entity mappings, and combine them with the semantically enhanced vectors to obtain knowledge-associated vectors; Calculate the dynamic adjustment factors of the basic word vectors and the knowledge-associated vectors based on vector quality evaluation metrics, where the vector quality evaluation metrics include vector clustering metrics, vector discrimination metrics, and vector coverage metrics; Perform adaptive fusion of the basic word vectors and the knowledge-associated vectors according to the dynamic adjustment factors to obtain the final word vectors.
4. The method according to claim 1, wherein Input the standard intermediate configuration data into a deep decision tree, obtain the cluster resource status and load metrics through a computing resource evaluation node, obtain the task running characteristics through a task feature analysis node, and calculate and determine the final submission method based on an intelligent routing strategy, including: Receive the standard intermediate configuration data, generate a deep decision tree model, and construct a resource evaluation node, a task feature analysis node, and a historical task analysis node; Collect cluster resource status metric data, perform normalization processing to obtain standardized resource metrics, input them into a neural network trained based on historical resource data, and generate a resource evaluation score as the first scoring metric; Parse task-related data from the standard intermediate configuration data, obtain task feature data through static analysis, construct a task feature vector based on the task feature data, input the task feature vector into a pre-trained task classification model, and generate a task feature score as the second scoring metric; Retrieve similar task records in the historical task database based on the task feature vector, extract task execution data from the similar task records, and calculate and generate a historical task score as the third scoring metric; Calculate the weight coefficients through the backpropagation algorithm, perform weighted fusion on the first scoring metric, the second scoring metric, and the third scoring metric to generate a routing comprehensive score; Calculate the dynamic routing threshold based on the cluster resource status and task distribution, compare the routing comprehensive score with the dynamic routing threshold, and determine the final submission method; Calculate the prediction deviation value between the final submission method and the actual execution result, and perform online updates on the neural network, the task classification model, and the weight coefficients according to the prediction deviation value.
5. The method according to claim 4, characterized in that, Generating the resource evaluation score includes: Convert the standardized resource metrics into a feature matrix, perform a convolution operation on the feature matrix to extract resource status features, input the resource status features into a fully connected layer, and obtain the resource evaluation score through a sigmoid function mapping; Generating the task feature score includes: Input the task feature vector into an attention layer to extract key features, perform temporal analysis on the extracted key features based on a long short-term memory network, calculate the probability distribution of different task types using a softmax function, and calculate the task feature score based on the probability distribution with weighting; Generating historical task scores includes: Calculating the similarity between the task feature vector and historical task records based on cosine similarity, selecting the historical task record with the highest similarity according to a preset selection ratio, extracting the corresponding execution duration, resource utilization rate, and success rate, and calculating the historical task score through weighted average.
6. The method according to claim 1, wherein Based on the template mapping mechanism, select the corresponding template according to the final submission method, and combine historical task data to intelligently fill and optimize the template parameters. The generated task submission instructions include: Establish a two-layer template mapping mechanism. The first layer selects the corresponding template type based on the final submission method, and the second layer constructs a resource requirement index for the template type. The resource requirement index divides the task resource requirements into different grades, and each grade corresponds to a template parameter configuration scheme; Identify the critical path of parameter configuration from historical task data. The critical path is determined by analyzing the impact degree of parameter changes on task execution effects. Construct the parameter configuration combinations on the critical path into a parameter pattern library, where each parameter pattern contains a complete parameter dependency chain; Determine the basic template in the two-layer template mapping mechanism according to the final submission method, match the parameter configuration scheme in the resource requirement index based on the resource requirements of the current task, and select the parameter pattern adapted to the parameter configuration scheme from the parameter pattern library; Through iterative replacement, apply the parameter pattern to the basic template, perform parameter consistency verification after each replacement, finally generate the task submission instruction, and feedback the task execution result after the task is executed to update the parameter pattern library.
7. The method according to claim 6, wherein Identifying the critical path of parameter configuration includes: Construct a task execution directed graph, where nodes are determined based on the execution stage, and edges between nodes are determined based on the data flow relationship between execution stages; Mark the configurable parameters on each node of the task execution directed graph, count the change records of the configurable parameters in historical task data, and calculate the impact degree of the configurable parameters on adjacent execution stages; Use the message passing algorithm to perform iterative calculations along the edges of the task execution directed graph to obtain the influence propagation range of each configurable parameter, and calculate the cumulative influence value of the configurable parameter based on the influence propagation range; Sort the configurable parameters according to the cumulative influence value, select the parameter sequence with the largest cumulative influence value, and determine the critical path; Construct a parameter pattern library, where each parameter pattern in the parameter pattern library contains the complete parameter dependency chain on the critical path; Convert the critical path into a parameter pattern and store it in the parameter pattern library, establish a parameter pattern score based on the usage frequency and task success rate, delete the parameter patterns with scores lower than the preset score lower threshold, and update the parameter pattern library.
8. A big data job intelligent submission system for a container cluster, which is used to implement the method described in any one of the foregoing claims 1-7, and is characterized in that Including: The first unit is used to receive the task configuration information submitted by the user; The second unit is used to perform field mapping and configuration standardization conversion based on field feature similarity calculation and historical call relationships, and generate standard intermediate configuration data in combination with version compatibility evaluation; The third unit is used to input standard intermediate configuration data into a deep decision tree, obtain the cluster resource status and load metrics through a computing resource evaluation node, obtain the task running characteristics through a task feature analysis node, and calculate and determine the final submission method based on an intelligent routing policy; The fourth unit is used to select a corresponding template based on the final submission method according to a template mapping mechanism, and intelligently fill and optimize the template parameters in combination with historical task data to generate a task submission instruction; The fifth unit is used to generate a resource definition configuration file and submit a task through a container cluster command according to the task submission instruction when the final submission method is the first submission method; When it is the second submission method, generate a command-line tool for the computing framework to submit a task; The sixth unit is used to monitor the execution status of a task in real time, collect status information, resource information and running logs, and feedback them to the intelligent routing policy for optimizing the submission decision of subsequent tasks.
9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-source computing power data integration and intelligent scheduling system and method
CN118916147A
Intelligent traffic data processing method and system based on edge calculation
CN119479316A
Data asset multi-dimensional evaluation method and system based on block chain and machine learning
CN119494575A
Big data platform scheduling task and data collaborative smooth migration method and system
CN119576506A
AI-based big data distributed computing task automatic optimization method and system
CN119576507A