Data engineering task allocation method and system based on multi-dimensional dynamic capability prediction

By constructing an explicit and implicit skills database and semantic field matching, the problem of inaccurate task allocation in data engineering was solved, achieving precise matching between engineers and tasks and efficient utilization of resources.

CN122367009APending Publication Date: 2026-07-10CHONGQING KAIYUAN GONGCHUANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING KAIYUAN GONGCHUANG TECH CO LTD
Filing Date
2026-04-15
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing technologies, data engineering task allocation relies on subjective human judgment and lacks systematic management methods, resulting in inaccurate allocation, inefficient resource utilization, and difficulty in meeting the precise control requirements in complex scenarios.

Method used

By collecting historical task records of data engineers, a database of explicit and implicit skills is constructed. The capabilities of explicit and implicit skills are predicted. Combined with semantic field construction and neurosemantic coupling matching, the most suitable engineers are selected for task allocation.

Benefits of technology

It achieves precise matching of data engineering tasks with engineers, balances team workload, and improves resource utilization efficiency and project execution quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122367009A_ABST
    Figure CN122367009A_ABST
Patent Text Reader

Abstract

This invention discloses a data engineering task allocation method and system based on multi-dimensional dynamic capability prediction, belonging to the field of intelligent task management and scheduling technology. The method includes: collecting historical task records of all data engineers, modeling explicit and implicit skills and constructing a database; constructing a semantic field for the data engineering tasks to be allocated, extracting explicit requirements and quantifying uncertainties; screening candidates through neural semantic coupling matching, and then optimizing the screening based on workload to determine the final task allocator, achieving scientific allocation. This invention solves the technical problem that traditional data engineering task allocation struggles to fully grasp the actual capabilities of personnel and the core requirements of tasks, resulting in inaccurate allocation and inefficient resource utilization. It achieves precise matching of data engineering tasks with suitable engineers, balances team workload, and improves resource utilization efficiency and project execution quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent task management and scheduling technology, and in particular to a data engineering task allocation method and system based on multi-dimensional dynamic capability prediction. Background Technology

[0002] The rational allocation of data engineering tasks is crucial for the efficient advancement of projects, and its scientific management methods are essential for resource utilization, task progress, and deliverable quality. Currently, data engineering task allocation largely relies on subjective human judgment and experience, lacking systematic management methods and depending solely on limited explicit skill information. While these management methods may be effective in simple scenarios, as data engineering becomes more complex, and given the dynamic changes in engineers' capabilities and the diversity of task requirements, traditional management methods easily lead to imbalances and inaccurate matching, failing to meet the needs of precise control and optimization in data engineering projects. Summary of the Invention

[0003] This application provides a data engineering task allocation method and system based on multi-dimensional dynamic capability prediction, which solves the technical problem that traditional data engineering task allocation is difficult to fully control the actual capabilities of personnel and the core needs of tasks, resulting in inaccurate allocation and inefficient resource utilization.

[0004] The first aspect of this application provides a data engineering task allocation method based on multi-dimensional dynamic capability prediction. The method includes: collecting historical task records of all data engineers, performing capability prediction modeling of explicit and implicit skills, and establishing an explicit and implicit skills database; constructing a semantic field for the data engineering tasks to be allocated, including extracting explicit task requirements and quantifying requirement uncertainties; performing neural semantic coupling matching between the explicit and implicit skill data of each data engineer in the explicit and implicit skills database and the semantic field to filter a set of candidate engineers; performing workload-based filtering based on the set of candidate engineers, and finally selecting the selected engineers as the allocators of the data engineering tasks.

[0005] A second aspect of this application provides a data engineering task allocation system based on multi-dimensional dynamic capability prediction. The system includes: a module for constructing an explicit and implicit skills database, used to collect historical task records of all data engineers, perform capability prediction modeling of explicit and implicit skills, and establish an explicit and implicit skills database; a semantic field construction and execution module, used to construct a semantic field for the data engineering tasks to be allocated, including extracting explicit task requirements and quantifying requirement uncertainties; a candidate engineer set acquisition module, used to perform neural semantic coupling matching between the explicit and implicit skill data of each data engineer in the explicit and implicit skills database and the semantic field, and filter the candidate engineer set; and a final engineer selection module, which performs workload-based selection based on the candidate engineer set, and selects the final selected engineers as the assigners of the data engineering tasks.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application constructs an explicit and implicit skills database by collecting historical task records of data engineers, constructs a semantic field for the data engineering tasks to be assigned to extract explicit requirements and quantify uncertainties, and performs neural semantic coupling matching between explicit and implicit skills data and semantic field to screen candidate engineers. Finally, it combines workload to make final screening and adjustment, achieving the technical effect of accurately matching data engineering tasks with suitable engineers, balancing team workload, and improving resource utilization efficiency and project execution quality. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 This is a flowchart illustrating the data engineering task allocation method based on multi-dimensional dynamic capability prediction provided in an embodiment of this application.

[0009] Figure 2 This is a schematic diagram of the structure of the data engineering task allocation system based on multi-dimensional dynamic capability prediction provided in the embodiments of this application.

[0010] Figure labeling: 1. Explicit and implicit skills database construction module; 2. Semantic field construction and execution module; 3. Candidate engineer set acquisition module; 4. Final engineer selection acquisition module. Detailed Implementation

[0011] This application provides a data engineering task allocation method and system based on multi-dimensional dynamic capability prediction, which solves the technical problem that traditional data engineering task allocation is difficult to fully control the actual capabilities of personnel and the core needs of tasks, resulting in inaccurate allocation and inefficient resource utilization.

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0013] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0014] Example 1, as Figure 1 As shown, a data engineering task allocation method based on multi-dimensional dynamic capability prediction is described, wherein the method includes: Collect all historical task records of data engineers, perform predictive modeling of explicit and implicit skills, and establish a database of explicit and implicit skills.

[0015] In this embodiment, a data engineer is a professional technician responsible for data engineering tasks such as data extraction, transformation, and loading (ETL), possessing relevant technical skills and experience in task execution. Explicit skills are the hard technical abilities that a data engineer can directly demonstrate and are work-related, such as the ability to use specific technology stacks, frameworks, and corresponding versions. Implicit skills are the comprehensive practical abilities of a data engineer to efficiently complete tasks in uncertain scenarios such as changing task requirements or ambiguous descriptions, and are indirectly reflected through historical task performance.

[0016] Specifically, the process begins by collecting a set of schedulable engineers. For each engineer in the set, a historical task record set containing historical ETL task code submission records and task description change records is collected. By analyzing the code submission records, an explicit skill tag set is constructed. Based on the task description change records, the code submission records are classified, and the deterministic probability of task execution under different levels of uncertainty is identified to establish implicit skill tags. The two types of tags are then associated with the corresponding engineers to generate explicit and implicit skill data, which is then incorporated into the explicit and implicit skill database. This step will be explained in detail later.

[0017] Semantic field construction is performed on the data engineering tasks to be assigned, including extracting explicit task requirements and quantifying requirement uncertainties.

[0018] In this application embodiment, the data engineering task is a professional technical task that involves extracting, transforming, loading, and processing data to meet the needs of business analysis, system operation, or data application.

[0019] Optionally, natural language processing technology is used to parse the data engineering task description to extract explicit technical requirements and business objectives. A counterfactual reasoning engine is used to identify ambiguities in the task description. Historical technical requirements are retrieved according to explicit business objectives to optimize and generate target explicit requirements. Based on the attributes of ambiguities, historical task execution records are retrieved to calculate the probability distribution of requirement uncertainty. Finally, a semantic field is constructed with the target explicit requirements and the probability distribution. This step will be explained in detail in the following content.

[0020] The explicit and implicit skill data of each data engineer in the explicit and implicit skill database are matched with the semantic field through neural semantic coupling to filter the candidate engineer set.

[0021] In one embodiment of this application, the explicit and implicit skill data of each data engineer in the explicit and implicit skill database are matched with the semantic field through neural semantic coupling. First, the explicit skill labels of engineers are matched with the target explicit requirements of the semantic field to generate an explicit matching engineer set. Then, the implicit skill labels of engineers are matched with the uncertainty probability distribution of the semantic field to generate an implicit matching engineer set. Finally, the intersection of the two sets is taken to generate a candidate engineer set. This step will be described in detail in the following content.

[0022] Based on the set of candidate engineers, a workload-based screening is performed, and the final selected engineers are used as the assigners of the data engineering tasks.

[0023] Specifically, firstly, we collect basic workload data for each data engineer in the candidate engineer pool. Then, we integrate with the enterprise's internal project management system or task tracking tools, such as Jira or Trello, to extract core workload metrics for each candidate engineer: including the number of assigned but incomplete tasks, the estimated remaining time for each incomplete task, and task priority distribution, categorized into high, medium, and low levels. The estimated remaining time is directly obtained from the task tracking tool; if the tool does not record it, it is calculated using the method of "total time for similar historical tasks - time already consumed." Task priorities are extracted according to project-preset rules to ensure the data source is authentic and traceable.

[0024] Secondly, a weighted summation method is used to quantify the workload and generate a workload score for each engineer. Weights are assigned to each indicator: estimated remaining working hours (0.6 as the core influencing factor), number of unfinished tasks (0.3 as the parallel task pressure), and percentage of high-priority tasks (0.1 as the energy consumption coefficient). The specific calculation steps are as follows: First, the estimated remaining working hours are standardized to remaining workdays, converted to an 8-hour workday, i.e., remaining workdays = estimated remaining working hours / 8; then, the percentage of high-priority tasks is calculated as: number of unfinished high-priority tasks / total number of unfinished tasks, where the percentage is 0 when there are no unfinished tasks; finally, the results are substituted into the formula: Workload Score = Remaining Workdays × 0.6 + Number of Unfinished Tasks × 0.3 + Percentage of High-Priority Tasks × 0.1, yielding a quantified workload score for each engineer. A higher score indicates a heavier workload.

[0025] Next, a workload screening threshold is set to filter engineers whose workload is within a reasonable range. The threshold is determined through historical data statistics: workload score data of the candidate engineer's team over the past three months is collected, and the arithmetic mean and standard deviation of all data are calculated. The upper limit threshold is set as "mean + 10% × mean". This threshold setting references the team's average workload level while reserving a 10% flexibility to avoid excessive restriction. The workload score of each candidate engineer is iterated through, retaining engineers whose scores are below or equal to the upper limit threshold, and removing overloaded engineers whose scores exceed the threshold, forming a subset of workload-suitable engineers.

[0026] Then, if the load matching engineer subset contains multiple engineers, a comprehensive ranking based on skill matching degree is required to determine the optimal assignee. Extract the explicit and implicit skill matching details for each engineer: Explicit matching scores are calculated as 1 point for complete coverage and 0 points for otherwise. Since explicit matching in the candidate engineer set already satisfies complete coverage, the score is 1 for each engineer. Implicit matching scores are the average of the deterministic probabilities corresponding to the uncertainty levels of each objective. For example, if the deterministic probabilities of uncertainty for two objectives are 0.65 and 0.63 respectively, then the implicit matching score = (0.65 + 0.63) / 2 = 0.64. Construct a comprehensive scoring formula: Comprehensive Score = Implicit Matching Score × 0.8 + (Threshold - Workload Score) / Threshold × 0.2, where the threshold refers to the upper limit of the load, and (Threshold - Workload Score) / Threshold converts the workload score into a load advantage score in the 0-1 range; the lower the load, the higher the advantage score. Sort the load matching engineer subset from highest to lowest comprehensive score, and select the engineer ranked first as the priority assignee.

[0027] Finally, the final assigner is confirmed and the task allocation is completed. If the load balancing engineer subset contains only one engineer, that engineer is directly identified as the final assigner. If there are engineers with the same overall score after sorting, the engineer with the higher historical task completion rate is prioritized. The historical completion rate is extracted from the project management system and calculated based on the number of completed assigned tasks divided by the total number of assigned tasks. The information of the final selected data engineers is associated with the tasks to be assigned, and a task allocation notification is sent through the project management tool. The workload data of the data engineers is updated synchronously, completing the entire allocation process.

[0028] By employing methods such as data collection using task management tools, weighted summation to quantify load, setting thresholds based on historical data, and comprehensive matching degree and load ranking, we achieved accurate screening based on workload balancing. Ultimately, we identified task assigners who met both skill matching requirements and had no risk of overload, ensuring task execution efficiency and optimized utilization of team resources.

[0029] Furthermore, the method provided in this application embodiment includes: A set of schedulable engineers is collected; for the first engineer in the engineer set, a first set of historical task records is collected, wherein each set of records in the first set of historical task records includes code commit records and task description change records in historical ETL tasks; the explicit skill items used in the code commit records are analyzed to construct a first set of explicit skill tags; based on the task description change records, the code commit records are classified according to the degree of uncertainty, and the deterministic probability of task execution under various degrees of uncertainty is identified to establish a first set of implicit skill tags; the first set of explicit skill tags, the first set of implicit skill tags, and the first engineer are associated and marked to generate first explicit and implicit skill data; the first explicit and implicit skill data is added to the explicit and implicit skill database.

[0030] In the implementation of this application, the first engineer refers to any one engineer in the set of engineers.

[0031] Specifically, firstly, by using the company's commonly used human resource management system or project management platform, data engineers who are currently available for task assignment are selected. The unique identification information of these engineers, such as their employee ID, internal system unique account ID, and project management platform association identifier, is collected to form a set of dispatchable engineers. This ensures that the set can fully cover all personnel who meet the current task assignment conditions, thus defining a clear scope for subsequent skills data collection.

[0032] Next, using version control systems and task management tools, we connected to the relevant data storage platform for the first engineer's historical ETL tasks and extracted the first engineer's historical task record set. Code commit records, which can be obtained from version control tools like Git, contain technical implementation details during task execution; task description change records, which can be extracted from task management tools like Jira, record adjustments, additions, and corrections to ambiguous descriptions of task requirements, ensuring that the collected records can fully support subsequent skills analysis. Furthermore, after completing the first engineer's historical task record set, the above operations were performed on each engineer in the engineer set to obtain their corresponding historical task record set.

[0033] Next, the collected code commit records are parsed. By identifying the technical elements involved in the code, such as programming languages, data processing frameworks, and database systems, and using version analysis tools to confirm the specific versions corresponding to each technical element, these technical stack, framework, and version distribution information are organized into explicit skill tags according to a standardized tag format. This ultimately forms the first set of explicit skill tags for the first engineer. Similarly, the above operation is performed on each engineer in the engineer set to obtain the corresponding set of explicit skill tags, thereby accurately reflecting the engineer's hard technical capabilities.

[0034] Next, the uncertainty level of the task description change records is identified to generate the corresponding task uncertainty index. Then, based on the index, the code submission records are clustered to obtain deterministic code submission records and code submission records corresponding to multiple uncertainty clusters. Finally, by analyzing the differences in submission quality and speed of code pairs of similar difficulty in the above records, the deterministic probability under multiple uncertainty levels is obtained, thereby generating the first implicit skill tag. This step will be explained in detail in the following content.

[0035] Subsequently, using the engineer's unique identifier as the core index, the constructed first set of explicit skill tags is bound to the first set of implicit skill tags to generate structured explicit and implicit skill data containing the engineer's identity information, explicit skill characteristics, and implicit skill characteristics, ensuring that the two types of skill information form a unique correspondence with the engineer subject.

[0036] Finally, a structured database of explicit and implicit skills is built. Through the database write interface, each engineer's explicit and implicit skill data is stored in the database in batches or one by one according to the preset data storage structure. Relevant indexes are also established to improve the efficiency of subsequent data query and matching, thereby completing the construction of the database of explicit and implicit skills.

[0037] By integrating system integration, data parsing, text analysis, and database construction, we gradually completed the collection, analysis, tagging, and storage of engineer skill data, building a comprehensive explicit and implicit skills database that reflects engineers' capabilities, providing a reliable data foundation for the accurate allocation of subsequent data engineering tasks.

[0038] Furthermore, the method provided in this application embodiment includes: The explicit skills items include technology stacks, frameworks, and version distributions.

[0039] Optionally, the technology stack refers to the set of foundational technologies that data engineers rely on to complete data engineering tasks. It encompasses core technical elements such as programming languages, database systems, and data storage technologies. Its purpose is to define the scope of an engineer's basic technical capabilities, ensuring that they can meet the core technical implementation requirements of the task. A framework is a standardized set of development tools built upon the technology stack, such as Spark and Flink frameworks in data processing, and Talend framework in data integration. These are used to simplify development processes and improve task execution efficiency. Specific tasks often have specific requirements for particular frameworks, and framework proficiency directly affects the quality of task completion. Version distribution refers to the specific version information of the technology stack and frameworks that the engineer is proficient in. Different versions differ in functionality and compatibility. Its purpose is to accurately match the specific requirements of the task for the technology or framework version, avoiding task execution anomalies due to version incompatibility.

[0040] When extracting technology stack information, a combination of code keyword extraction and syntax analysis is used. Regular expression matching tools are used to identify keywords in the code that indicate programming languages, such as Python, Java, and Scala. At the same time, syntax analysis tools are used to parse database connection statements and data storage operation-related code segments in the code to extract identifiers for databases and storage systems such as MySQL, PostgreSQL, Hadoop, and HBase. This information is then integrated to form a technology stack list for engineers.

[0041] For extracting framework information, code dependency analysis tools are used to interface with code project files in the version control system to identify framework-related import statements or configuration information in the code. For example, by parsing the `import org.apache.spark` statements in Java code, the use of the Spark framework is confirmed; by analyzing the `import flink` statements in Python code, the use of the Flink framework is confirmed. After systematically reviewing these statements, a framework mastery checklist for engineers is created.

[0042] When obtaining version distribution information, a text parsing method is used to process dependency configuration files in the code project. For Python projects, the library names and corresponding version numbers in the requirements.txt file are parsed; for Java projects, the version information under the dependency node in the pom.xml file is parsed; for scenarios without explicit configuration files, the version identifiers of the technology stack and frameworks are extracted by combining the comment information of code commit records in the version control system and dependency package download logs to form version distribution information corresponding to each technology and framework.

[0043] Finally, the extracted technology stack, framework, and version distribution information are categorized and organized according to a standardized tag format, such as programming language - Python - 3.9 and data processing framework - Spark - 3.2, to construct a structured set of explicit skill tags. This ensures that the tag information is clear and standardized, making it easier to match and compare with task requirements in the future.

[0044] By performing keyword extraction, syntactic analysis, and text parsing, we accurately identify and integrate engineers' technology stack, frameworks, and version distribution information, and construct a set of explicit skill tags that can comprehensively reflect engineers' hard technical capabilities. This provides an accurate and reliable capability reference for matching the explicit requirements of subsequent tasks.

[0045] Furthermore, the method provided in this application embodiment includes: Identify the uncertainty level for the task description change records, and generate the task uncertainty indicators for each group of records; cluster the code submission records based on the task uncertainty indicators to obtain the deterministic code submission records and multiple groups of code submission records corresponding to the uncertainty clustering set indicators; analyze and extract the differences in the submission quality and speed of equivalent difficulty code pairs from the deterministic code submission records and the multiple groups of code submission records, obtain the deterministic probabilities under various uncertainty levels, and generate the first implicit skill label.

[0046] In the embodiment of the present application, the uncertainty level is the task change degree, that is, the difference degree between the initial task description text and the final task description text.

[0047] Specifically, first, extract the initial task description text and the final task description text corresponding to each group of task description change records from the task management tool according to the foregoing steps, ensuring that the obtained text completely includes the initial expression of the task requirements and all the finally determined content. Adopt text preprocessing methods to perform word segmentation and stop word removal operations on the two types of texts respectively. For example, use the jieba word segmentation tool to complete Chinese word segmentation, and filter out meaningless words such as "de" and "he" through the general stop word list to obtain pure text corpora, laying a foundation for subsequent difference analysis.

[0048] Next, use the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm to extract keywords from the preprocessed initial and final task description texts. The specific steps are as follows: First, the preprocessed initial task description text and the final task description text are merged to construct a joint vocabulary containing all unique words from both types of texts. This serves as the unified dimensional benchmark for subsequent feature vectors. For both the initial and final task description texts, the frequency of each word in the corresponding text is counted. This count is divided by the total number of words in the text to obtain the term frequency (TF) of each word in the corresponding text. Based on the document set composed of the initial and final task description texts, the number of documents in which each word appears in the set is counted. The inverse document frequency (IDF) of each word is calculated using the formula: IDF = log(total number of documents in the set / (number of documents containing the word + 1)). The addition of 1 in this formula represents the initial document frequency. To avoid cases where the denominator is 0, the term frequency (TF) of each word is multiplied by the inverse document frequency (IDF) to obtain the TF-IDF weight of that word in the corresponding text. The higher the weight, the stronger the word's ability to represent the core meaning of the text. Based on this, the core expressive words and their corresponding weights in the two types of text are selected. Finally, based on the above joint lexicon, the dimensions of the feature vectors are determined, with each dimension uniquely corresponding to a word. For the initial task description text and the final task description text, the TF-IDF weight of each word is filled into the position of its corresponding dimension. Words that do not appear in the text are filled with 0 in the corresponding dimension, thereby transforming the two types of text into structured text feature vectors, including the initial task description feature vector and the final task description feature vector.

[0049] Subsequently, the cosine similarity algorithm was used to calculate the similarity value, thus clarifying the initial task description feature vector as follows: The final task description feature vector is Both are built on the same joint thesaurus and have the same number of dimensions. The position of each dimension corresponds one-to-one with the words in the joint thesaurus, and the dimension value is the TF-IDF weight of the corresponding word. Then, the calculation is performed. and The dot product is calculated by multiplying the corresponding dimensions of two vectors one by one and then summing the results. Then, each vector is calculated separately. The modulus and The modulus is the square root of the sum of the squares of the values ​​in each dimension. Finally, divide the dot product result by... and The product of the moduli is used to obtain the cosine similarity value, which ranges from 0 to 1. The closer the value is to 1, the higher the overlap of the core meaning of the initial and final task descriptions and the smaller the change. The closer the value is to 0, the greater the difference in the core meaning and the more significant the change. Next, the calculated similarity value is converted into the task difference degree, i.e., task difference degree = 1 - similarity value. Then, the difference degree is normalized to the numerical range of 0-1 through a linear mapping method to generate a task uncertainty index for each set of records, where 0 represents no change, i.e., complete certainty, and 1 represents maximum change, i.e., high uncertainty.

[0050] Then, the K-means clustering algorithm is used to classify the code submission records. First, the task uncertainty index corresponding to all code submission records is collected. Using the value of a single index as a one-dimensional feature, a feature vector is constructed for each code submission record. The Min-Max normalization method is used to map the values ​​of all feature vectors to the data range of [0,1] to eliminate the influence of units and ensure the accuracy of the clustering results. Next, the number of clusters K is determined by the elbow rule. Usually, the value of K is set to 3-5, which can effectively distinguish different levels of uncertainty and avoid over-refinement of clusters. K cluster centers are initialized. K feature vectors can be randomly selected as initial centers. Then, the Euclidean distance from the feature vector of each code submission record to each cluster center is calculated iteratively. Each record is assigned to the category of the nearest cluster center, and the mean of each category is recalculated as the new cluster center. Repeat the above iterative process until one of the following two termination conditions is met: First, in two consecutive iterations, the change in the Euclidean distance of all cluster centers is less than the preset distance threshold of 0.001, that is, it is determined that the cluster centers no longer change significantly; Second, the number of iterations reaches the preset upper limit of 100 iterations to ensure that the algorithm converges within a reasonable time and avoids infinite iteration. At this point, regardless of whether there are still slight changes in the cluster centers, the iteration process is terminated and the clustering operation is completed.

[0051] The clustering results are then divided, and the cluster with the smallest uncertainty index value is identified as a deterministic code submission record. The remaining clusters are identified as multi-class code submission records corresponding to different uncertainty cluster indices, thus classifying code submission records according to the degree of uncertainty.

[0052] Finally, the code difficulty is evaluated for each code record in the deterministic code submission records and the multiple types of uncertain code submission records. Based on the evaluation results, the multiple types of uncertain code submission records are combined with records in the deterministic code submission records whose difficulty difference is less than a preset value to establish multiple sets of comparison codes. Finally, the differences in submission quality and speed between uncertain codes and deterministic codes in the comparison codes are analyzed to obtain the deterministic probability under multiple degrees of uncertainty that are inversely proportional to the degree of difference, and the first implicit skill tag is generated. This step will be explained in detail in the following content.

[0053] By employing methods such as text preprocessing, TF-IDF keyword extraction, cosine similarity calculation, and K-means clustering, the quantification of task uncertainty and the classification of code submission records were completed, providing a clear and standardized data source for subsequent calculation of the deterministic probability of task execution under various degrees of uncertainty.

[0054] Furthermore, the method provided in this application embodiment includes: For each code record in the deterministic code submission record and the multi-type code submission record, a code difficulty evaluation is performed to obtain a difficulty evaluation result. Based on the difficulty evaluation result, the multi-type code submission records are combined with the code records in the deterministic code submission records whose difficulty difference is less than a preset difficulty difference to establish multiple sets of comparison codes corresponding to multiple uncertainty clustering indicators. For the multiple sets of comparison codes, the degree of difference in submission quality and speed between uncertain codes and deterministic codes is analyzed to obtain the deterministic probability under multiple degrees of uncertainty, wherein the deterministic probability is inversely proportional to the degree of difference.

[0055] Specifically, firstly, a difficulty assessment is performed on each code record. The structure of the code record is analyzed using code syntax analysis tools such as ANTLR, and core difficulty indicators are statistically determined: the total number of lines of code is counted directly using text analysis tools; the nesting levels of loops and conditional statements are identified and accumulated through syntax tree traversal; the number of data processing operation types involved (such as data cleaning, join queries, aggregation calculations, etc.) is statistically analyzed against a pre-defined list of operation types; and the scale of the processed data is extracted from the corresponding task description, including the number of data entries and file size, and converted into standardized values ​​(e.g., 100,000 data entries are assigned 0.3, 1 million data entries are assigned 0.6, etc.). Clear weights are assigned to each indicator: total lines of code account for 0.2, nesting levels account for 0.3, number of operation types account for 0.3, and data scale accounts for 0.2. The standardized values ​​of each indicator are multiplied by their corresponding weights and summed to obtain a difficulty score for each code record, with a score ranging from 0 to 1, where a higher score indicates greater difficulty, thus forming a complete difficulty assessment result.

[0056] Secondly, a method based on historical data statistics is used to determine the preset difficulty difference. Code record pairs of similar difficulty levels, confirmed by the project leader, are collected from similar data engineering tasks completed within the past six months, and at least 50 valid samples are selected. The difficulty score difference between the two code records in each sample is calculated, and the arithmetic mean and standard deviation of these differences are calculated. The "mean + 1 standard deviation" is used as the preset difficulty difference. For example, if the mean difficulty difference of the 50 samples is 0.08 and the standard deviation is 0.02, then the preset difficulty difference is set to 0.10.

[0057] Next, comparison codes are constructed based on the difficulty assessment results and the preset difficulty difference. For each type of uncertain cluster, each code submission record is extracted and paired with all code records in the deterministic code submission records. The difficulty score difference of each pair of code records is calculated. If the difference is less than the preset difficulty difference, the two code records are identified as a pair of comparison codes, where the code from the uncertain cluster is the uncertain code and the code from the deterministic set is the deterministic code. The above logic is used to pair all types of uncertain code records with deterministic code records, forming multiple sets of comparison codes corresponding to different levels of uncertainty. This ensures that the difficulty base of each set of comparison codes is consistent, eliminating difficulty interference for subsequent difference analysis.

[0058] Next, the differences in code submission quality and speed are analyzed and compared. For submission quality, code style checking tools such as SonarQube are used to check the syntax error rate and style compliance. Combined with test records, the first-pass rate and number of bugs are obtained: 1 point for a syntax error rate of 0% and style compliance ≥ 95%, 1 point for a 100% first-pass rate, and 1 point for no functional bugs. A quality score is calculated for each line of code on a 3-point scale. Then, the ratio of the uncertain code quality score to the deterministic code quality score is calculated. If the ratio is less than 1, the quality difference is 1 - the ratio; if the ratio is ≥ 1, the quality difference is 0. For submission speed, the "task allocation time" and "final code submission time" corresponding to each code record are extracted from the task management tool. The time difference between the two is calculated as the task time. The speed difference is obtained by subtracting the deterministic code time from the uncertain code time and dividing by the deterministic code time. If the difference is negative, the speed difference is 0. The arithmetic mean of the quality difference and speed difference is taken to obtain the total degree of difference of the comparison code. Since the probability of certainty is inversely proportional to the degree of difference, the probability of certainty = 1 - the total degree of difference. Based on this, the probability of certainty under each type of uncertainty is calculated.

[0059] Finally, the deterministic probabilities corresponding to different levels of uncertainty are organized in a standardized format to form the first implicit skill tag for each engineer. The tag clearly marks each level of uncertainty, such as low, medium, and high, and its corresponding deterministic probability value, which intuitively presents the engineer's comprehensive ability to cope with different levels of task uncertainty.

[0060] By employing methods such as code feature quantification, historical data statistics, pairwise comparison, and dual-dimensional evaluation of quality and speed, the system completed the construction and difference analysis of the comparison code, and finally generated skill tags that can accurately reflect the implicit abilities of engineers, providing a reliable implicit ability reference for the accurate matching of subsequent tasks and engineers.

[0061] Furthermore, the method provided in this application embodiment includes: The task description of the data engineering task is parsed using natural language processing technology to extract explicit technical requirements and explicit business objectives. Ambiguities in the task description are identified using a counterfactual reasoning engine. Historical technical requirements are retrieved based on the explicit business objectives, and these requirements are optimized to generate target explicit requirements. Historical task execution records are retrieved based on the fuzzy attributes of the ambiguous points, and the uncertainty probability distribution of the task requirements is calculated according to these records. The semantic field is constructed using the target explicit requirements and the uncertainty probability distribution.

[0062] Specifically, firstly, a natural language processing (NLP) workflow is used to parse the task description of the data engineering task. Specifically: The first step involves text preprocessing. The task description text is segmented using the jieba word segmentation tool, and meaningless words are filtered out using a general stop word list. Then, spaCy is used for part-of-speech tagging and named entity recognition to identify core semantic components such as nouns, verbs, and technical terms. The second step involves constructing two thesauruses: a technical requirement thesaurus containing keywords common in the data engineering field, such as programming languages, frameworks, database systems, and data processing operations (e.g., data extraction, cleaning, and aggregation); and a business objective thesaurus containing terms related to business scenarios such as data visualization, report generation, real-time synchronization, and batch processing. The third step uses keyword matching and semantic association analysis to extract words matching the technical requirement thesaurus from the preprocessed text and integrate them into explicit technical requirements, such as "using Python and the Spark framework to process tens of millions of data points"; and extracts words and related semantic expressions matching the business objective thesaurus to define explicit business objectives, such as "generating monthly sales data reports and supporting multi-dimensional queries."

[0063] Then, a counterfactual reasoning engine is built for data engineering task scenarios to identify ambiguities in task descriptions. The engine consists of four modular units: a text preprocessing unit that reuses the word segmentation, stop word filtering, and part-of-speech tagging tools from the previous steps to output clean core corpus; a ambiguity candidate extraction unit that constructs a data engineering task-specific fuzzy rule library, with rules including: absence of key constraint words such as data range, time window, output format, and cleaning standards; inclusion of vague modifiers such as "approximately," "as far as possible," and "similar"; mention of referencing historical projects without specifying the project name or number; and vague descriptions of data volume, such as large amounts of data without specifying concrete values. Candidate ambiguities are extracted through keyword matching and rule validation; and a counterfactual verification unit that uses a hybrid approach of rule judgment and a lightweight text classification model to perform counterfactual verification on the candidate ambiguities. The system makes real assumptions, such as whether supplementing data within the last 6 months would change the task execution path, or whether specifying CSV as the output format would reduce execution ambiguity. The rule base pre-sets that "if the absence of key constraints leads to ≥2 different execution paths, it is judged as a true ambiguity point." It can also use publicly available pre-trained text classification models, such as a lightweight BERT-based model, as input. The input consists of the original task description and a counterfactual hypothesis description; the model outputs a judgment result indicating whether there is an execution divergence. The combination of these two methods verifies the authenticity of candidate ambiguities. The ambiguity point output unit labels verified true ambiguities with ambiguity attributes, such as constraint missing type, expression ambiguity type, and reference object ambiguity type, forming a complete ambiguity point identification result.

[0064] Next, using explicit business objectives as the core of the retrieval process, explicit technical requirements are optimized to generate target explicit requirements. Explicit business objectives are transformed into search keyword vectors, and the TF-IDF algorithm is used to calculate the weight of each keyword, constructing a structured search vector. Then, the process connects to a historical task database, which stores information such as business objectives, technical requirements, and execution results of previously completed data engineering tasks. The cosine similarity algorithm is used to calculate the similarity between the current business objective search vector and the historical task business objective vectors, filtering out highly relevant historical tasks with a similarity ≥ 0.8. The TF-IDF and cosine similarity algorithms described above are similar to the steps mentioned earlier for generating task uncertainty indicators, and will not be elaborated here. Then, the explicit technical requirements of highly relevant historical tasks are extracted and compared with the currently extracted explicit technical requirements. Necessary technical elements missing from the current requirements are supplemented. For example, if the current explicit technical requirement is "using the Spark framework to process data," and the historical related task technical requirements include "using Spark SQL for multi-table join queries and storing intermediate data in Parquet format," these supplementary elements are integrated into the current explicit technical requirements, ultimately generating comprehensive and accurate target explicit requirements.

[0065] Then, based on the fuzzy attributes of fuzzy points, historical task execution records are retrieved to calculate the uncertainty probability distribution of task requirements. The first step involves classifying and retrieving historical task execution records according to the fuzzy attributes of the fuzzy points, extracting all task records corresponding to the same fuzzy attribute from the historical task database, ensuring a minimum of 30 samples for statistical validity. The second step involves statistically analyzing the retrieved historical task execution status, with key indicators including: the number of reworks caused by fuzzy points, the communication costs incurred due to fuzzy points (quantified by additional communication time), the proportion of task delays, and the degree of deviation between the execution results and expectations. The third step calculates the uncertainty probability: for each type of fuzzy attribute, (number of reworks + number of delays) is divided by the total number of samples to obtain the basic uncertainty probability corresponding to that fuzzy attribute. This is then adjusted based on communication costs and the degree of result deviation; for example, if communication time exceeds the average level by 20%, the probability increases by 5%. Finally, the uncertainty probabilities corresponding to each type of fuzzy attribute are formed and integrated into the uncertainty probability distribution of task requirements.

[0066] Finally, a structured semantic field is constructed to integrate and store explicit target requirements with uncertainty probability distributions. Explicit target requirements are presented in list form, clearly defining the technology stack, framework, data processing steps, business output requirements, and other specific details. Uncertainty probability distributions are stored in key-value pairs, where the key is a fuzzy attribute type and the value is the corresponding uncertainty probability. The semantic field uses common structured data formats such as JSON to ensure compatibility and parsability when matching explicit and implicit skill data later.

[0067] By completing the extraction of explicit target requirements, identification of ambiguities, optimization of requirements, and calculation of uncertainty probability distribution, a semantic field that fully reflects the characteristics of the task is finally constructed, providing a structured and quantifiable task feature foundation for the subsequent accurate matching of explicit and implicit skill data of data engineers with tasks.

[0068] Furthermore, the method provided in this application embodiment includes: The explicit requirements of the target in the semantic field are matched with the explicit skill tags in the explicit and implicit skill data of each data engineer to generate an explicit matching engineer set; the uncertainty probability distribution in the semantic field is matched with the implicit skill tags in the explicit and implicit skill data of each data engineer to generate an implicit matching engineer set; the intersection of the explicit matching engineer set and the implicit matching engineer set is taken to generate the candidate engineer set.

[0069] In one embodiment, the explicit target requirements in the semantic field are first structurally decomposed to form a clear list of matching items. The explicit target requirements already include core elements such as technology stack, framework, version distribution, data processing operations, and output format. They are decomposed into independent matching items according to the format of "technology type-specific content-version / standard". For example, "process order data using Python language and Spark framework version 3.2 to output CSV format reports" is decomposed into four matching items: programming language - Python, data processing framework - Spark - 3.2, data object - order data, and output format - CSV, ensuring that each matching item has a clear judgment criterion.

[0070] Next, the broken-down target explicit requirements are matched with the explicit skill tags of each data engineer. The explicit skill tag set for each engineer is retrieved from the explicit and implicit skill database. This set stores information such as technology stack, framework, and version distribution in a standardized format. All matching items for the target explicit requirements are iterated through, and each engineer's explicit skill tag set is checked to see if it contains every matching item. If an engineer's tag set completely covers all matching items (i.e., no matching items are missing or mismatched), then that engineer is included in the explicit matching engineer set, strictly adhering to the matching rule that explicit skill tags completely cover the target explicit requirements.

[0071] Then, the target uncertainty degree with a probability greater than a preset threshold is extracted from the uncertainty probability distribution of the semantic field. Then, the matching certainty probability of the target uncertainty degree in the implicit skill tags of each data engineer is matched. Finally, engineers with a matching certainty probability greater than a preset probability threshold are selected to form an implicit matching engineer set. This step will be explained in detail in the following content.

[0072] Finally, a candidate engineer set is generated through set intersection operations. The explicit matching engineer set and the implicit matching engineer set are transformed into a set data structure with the unique identifier information of engineers as elements. The common unique identifier information of engineers in the two sets is extracted. The engineers corresponding to these identifiers are the candidates that simultaneously meet the explicit requirement coverage and implicit ability adaptation, and finally form a structured candidate engineer set, which includes basic information such as engineer identifier information, explicit skill matching results, and implicit ability matching results.

[0073] By accurately coupling and matching engineers' explicit and implicit skills with the semantic field of the task, a set of candidate engineers who simultaneously meet the explicit technical requirements of the task and the ability to cope with uncertainty is selected, providing a high-quality candidate base for the subsequent final selection based on workload.

[0074] Furthermore, the method provided in this application embodiment includes: The explicit matching engineer set includes data engineers whose explicit skill tags fully cover the explicit requirements of the target.

[0075] Optionally, firstly, the explicit skill tags for data engineers should be standardized in format. Extract the set of explicit skill tags for each engineer from the explicit and implicit skill database. This set is stored in a standardized format of "Technology Type - Specific Content - Version / Standard". Using a data formatting tool, all engineers' explicit skill tags are uniformly converted into a structure consistent with the target explicit requirements. This ensures that the dimensions and expression of the matching items are completely aligned with the skill tags, avoiding matching deviations caused by format differences. For example, "Python3.9" in the engineer tag is standardized to "Python-3.9" to maintain consistency with the format of the requirement matching items.

[0076] Next, a subset-based matching method is used to perform full coverage matching. All matching items for the target explicit requirements are grouped into a requirement set S, and the explicit skill tags of each engineer are grouped into a skill set T. Using built-in set operation functions in the programming language or multi-condition matching queries in the database, it is determined whether the requirement set S is a subset of the skill set T. Specifically, each matching item in the requirement set S is traversed, and each matching item is verified to exist in the engineer's skill set T: if all matching items can be found with a completely identical counterpart in the skill set T, then the engineer's explicit skill tags are determined to fully cover the target explicit requirements; if any matching item cannot be found with a completely identical counterpart in the skill set T, it is determined to be uncovered and not included in the set of engineers explicitly matched.

[0077] Finally, information on all engineers meeting the full coverage criteria is integrated to generate an explicit matching engineer set. Unique identifiers of engineers who pass the matching verification are collected, such as employee IDs and system account IDs, and their explicit skill tags are linked to the matching details between the target explicit requirements, such as the number of covered matches and a description of no omissions. This data is stored in structured data formats such as tables and JSON arrays to form a complete explicit matching engineer set. This ensures that each engineer in the set meets the core requirement that explicit skill tags fully cover the target explicit requirements, providing a standardized and usable data foundation for the intersection operation of the implicit matching engineer set.

[0078] By accurately selecting data engineers whose explicit skills are fully adapted to the task requirements, a set of explicit matching engineers is generated, providing a qualified candidate base of explicit personnel for the complete execution of subsequent neurosemantic coupling matching.

[0079] Furthermore, the method provided in this application embodiment includes: Based on the uncertainty probability distribution, the target uncertainty degree with a probability threshold greater than a preset threshold is extracted, and the matching certainty probability corresponding to the target uncertainty degree is matched in the implicit skill tags of each data engineer; data engineers with a probability greater than the preset probability threshold are selected according to the matching certainty probability to form the implicit matching engineer set.

[0080] In one embodiment, firstly, a preset threshold for the uncertainty probability distribution is determined. This threshold is derived from historical task execution data statistics to ensure its rationality and practicality. Uncertainty probability distribution data for data engineering tasks completed within the past 12 months are collected. Cases where uncertainty leads to task execution risks (such as rework, delays, and result deviations) exceeding acceptable limits are identified. The minimum uncertainty probability corresponding to these cases is statistically analyzed, and this minimum value is set as the preset threshold. This defines the high-risk uncertainty levels that require focused adaptation.

[0081] Next, the target uncertainty level is extracted based on a preset threshold. From the uncertainty probability distribution of the semantic field, all fuzzy attributes with probability values ​​greater than the preset threshold and their corresponding probabilities are extracted; these fuzzy attributes constitute the target uncertainty level. The uncertainty probability distribution is stored in key-value pair format, for example, missing time window: 0.35, fuzzy output format: 0.28, fuzzy data magnitude: 0.32, etc. By comparing the probability value in each key-value pair with the preset threshold one by one, if the probability value is greater than the preset threshold, the corresponding fuzzy attribute is included in the target uncertainty level set, thus forming the target uncertainty level set.

[0082] Next, a preset probability threshold for matching deterministic probabilities is determined. This threshold is used to select engineers with sufficient implicit capabilities to cope with target uncertainty. Also based on historical data statistics, deterministic probability data of engineers who have successfully completed tasks involving target uncertainty are collected, and the arithmetic mean of these data is calculated. This mean is set as the default preset probability threshold to ensure that selected engineers have a high probability of successfully coping with target uncertainty.

[0083] Subsequently, the implicit skill tags are matched with the target uncertainty levels. Implicit skill tags for each engineer are extracted from the explicit and implicit skill database; these tags contain the probability of certainty corresponding to different levels of uncertainty. For each fuzzy attribute in the target uncertainty set, a query is performed to see if a corresponding probability of certainty exists in the engineer's implicit skill tags: if the engineer's tags contain the probability of certainty for all target uncertainties, and each probability of certainty is greater than a preset probability threshold, then the engineer's implicit ability is determined to be suitable for the task requirements; if any target uncertainty has no corresponding probability of certainty, or the corresponding probability is less than the preset probability threshold, then it is determined to be unsuitable.

[0084] Finally, all suitable engineer information is integrated to generate a set of implicitly matched engineers. Unique identifiers of engineers who pass the matching verification are collected, such as employee IDs and system account IDs, and linked to the deterministic probability details corresponding to the degree of uncertainty in their implicit skill tags. This information is stored in structured data formats such as JSON arrays and database tables to form a complete set of implicitly matched engineers. This ensures that each engineer in the set possesses sufficient implicit capabilities to cope with task uncertainty, providing standardized data support for finding the intersection of the explicitly matched engineer sets and generating a candidate engineer set.

[0085] By employing historical data statistical methods to determine dual preset thresholds, iterative comparison, and key-value pair query matching, data engineers with implicit capabilities adapted to the uncertain requirements of tasks are accurately screened, generating a set of implicit matching engineers. This provides a suitable pool of implicit personnel for the complete execution of neurosemantic coupling matching.

[0086] In summary, the data engineering task allocation method based on multi-dimensional dynamic capability prediction provided in this application has the following technical effects: This application collects historical task records of data engineers, models implicit and explicit skills to build a database, constructs a semantic field for tasks to be assigned, and filters candidates through neural semantic coupling matching. Combined with workload optimization, it accurately matches suitable engineers, achieving the scientific allocation of data engineering tasks. This results in a precise match between data engineering tasks and suitable engineers, balancing team workload, and improving resource utilization efficiency and project execution quality.

[0087] Example 2, as Figure 2 As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a data engineering task allocation system based on multi-dimensional dynamic capability prediction, the system comprising: The explicit and implicit skills database construction module 1 is used to collect all the historical task records of all data engineers, perform ability prediction modeling of explicit and implicit skills, and establish an explicit and implicit skills database.

[0088] Semantic field construction execution module 2 is used to construct the semantic field for the data engineering task to be assigned, including extracting explicit requirements of the task and quantifying the uncertainty of the requirements.

[0089] The candidate engineer set acquisition module 3 is used to perform neural semantic coupling matching between the explicit and implicit skill data of each data engineer in the explicit and implicit skill database and the semantic field to filter the candidate engineer set.

[0090] The final engineer selection module 4 performs workload-based selection based on the candidate engineer set, and selects the final selected engineers as the assigners of the data engineering tasks.

[0091] Furthermore, the explicit / implicit skill database construction module 1 is used to perform the following steps: A set of schedulable engineers is collected; for the first engineer in the engineer set, a first set of historical task records is collected, wherein each set of records in the first set of historical task records includes code commit records and task description change records in historical ETL tasks; the explicit skill items used in the code commit records are analyzed to construct a first set of explicit skill tags; based on the task description change records, the code commit records are classified according to the degree of uncertainty, and the deterministic probability of task execution under various degrees of uncertainty is identified to establish a first set of implicit skill tags; the first set of explicit skill tags, the first set of implicit skill tags, and the first engineer are associated and marked to generate first explicit and implicit skill data; the first explicit and implicit skill data is added to the explicit and implicit skill database.

[0092] Furthermore, the explicit / implicit skill database construction module 1 is used to perform the following steps: The explicit skills items include technology stacks, frameworks, and version distributions.

[0093] Furthermore, the explicit / implicit skill database construction module 1 is used to perform the following steps: Uncertainty level identification is performed on the task description change records to generate a task uncertainty index for each group of records; the code submission records are clustered based on the task uncertainty index to obtain deterministic code submission records and multiple types of code submission records corresponding to multiple uncertainty cluster set indices; the submission quality and speed differences of code pairs of equal difficulty are analyzed and extracted from the deterministic code submission records and the multiple types of code submission records to obtain deterministic probabilities under multiple uncertainty levels, and the first implicit skill tag is generated.

[0094] Furthermore, the explicit / implicit skill database construction module 1 is used to perform the following steps: For each code record in the deterministic code submission record and the multi-type code submission record, a code difficulty evaluation is performed to obtain a difficulty evaluation result. Based on the difficulty evaluation result, the multi-type code submission records are combined with the code records in the deterministic code submission records whose difficulty difference is less than a preset difficulty difference to establish multiple sets of comparison codes corresponding to multiple uncertainty clustering indicators. For the multiple sets of comparison codes, the degree of difference in submission quality and speed between uncertain codes and deterministic codes is analyzed to obtain the deterministic probability under multiple degrees of uncertainty, wherein the deterministic probability is inversely proportional to the degree of difference.

[0095] Furthermore, the semantic field construction execution module 2 is used to perform the following steps: The task description of the data engineering task is parsed using natural language processing technology to extract explicit technical requirements and explicit business objectives. Ambiguities in the task description are identified using a counterfactual reasoning engine. Historical technical requirements are retrieved based on the explicit business objectives, and these requirements are optimized to generate target explicit requirements. Historical task execution records are retrieved based on the fuzzy attributes of the ambiguous points, and the uncertainty probability distribution of the task requirements is calculated according to these records. The semantic field is constructed using the target explicit requirements and the uncertainty probability distribution.

[0096] Furthermore, the candidate engineer set acquisition module 3 is used to perform the following steps: The explicit requirements of the target in the semantic field are matched with the explicit skill tags in the explicit and implicit skill data of each data engineer to generate an explicit matching engineer set; the uncertainty probability distribution in the semantic field is matched with the implicit skill tags in the explicit and implicit skill data of each data engineer to generate an implicit matching engineer set; the intersection of the explicit matching engineer set and the implicit matching engineer set is taken to generate the candidate engineer set.

[0097] Furthermore, the candidate engineer set acquisition module 3 is used to perform the following steps: The explicit matching engineer set includes data engineers whose explicit skill tags fully cover the explicit requirements of the target.

[0098] Furthermore, the candidate engineer set acquisition module 3 is used to perform the following steps: Based on the uncertainty probability distribution, the target uncertainty degree with a probability threshold greater than a preset threshold is extracted, and the matching certainty probability corresponding to the target uncertainty degree is matched in the implicit skill tags of each data engineer; data engineers with a probability greater than the preset probability threshold are selected according to the matching certainty probability to form the implicit matching engineer set.

[0099] The data engineering task allocation system based on multi-dimensional dynamic capability prediction provided in the embodiments of the present invention can execute the data engineering task allocation method based on multi-dimensional dynamic capability prediction provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0100] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.

[0101] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A data engineering task allocation method based on multi-dimensional dynamic capability prediction, characterized in that, include: Collect all historical task records of data engineers, perform predictive modeling of explicit and implicit skills, and establish a database of explicit and implicit skills; Semantic field construction is performed on the data engineering tasks to be assigned, including extracting explicit task requirements and quantifying requirement uncertainties; The explicit and implicit skill data of each data engineer in the explicit and implicit skill database are matched with the semantic field through neurosemantic coupling to filter the candidate engineer set; Based on the set of candidate engineers, a workload-based screening is performed, and the final selected engineers are used as the assigners of the data engineering tasks.

2. The data engineering task allocation method based on multi-dimensional dynamic capability prediction as described in claim 1, characterized in that, Collect all historical task records of data engineers, perform predictive modeling of explicit and implicit skills, and establish a database of explicit and implicit skills, including: Collect a set of schedulable engineers; For the first engineer in the engineer set, a first historical task record set is collected, wherein each set of records in the first historical task record set includes code commit records and task description change records in historical ETL tasks; Analyze the explicit skill items used in the code commit records to construct a first set of explicit skill tags; Based on the task description change records, code submission records are classified according to the degree of uncertainty, the deterministic probability of task execution under various degrees of uncertainty is identified, and a first implicit skill tag is established. Associate the first set of explicit skill tags, the first set of implicit skill tags, and the first engineer with the tags to generate first explicit and implicit skill data; Add the first reveal / conceal skill data to the reveal / conceal skill database.

3. The data engineering task allocation method based on multi-dimensional dynamic capability prediction as described in claim 2, characterized in that, The explicit skills items include technology stacks, frameworks, and version distributions.

4. The data engineering task allocation method based on multi-dimensional dynamic capability prediction as described in claim 2, characterized in that, Based on the task description change records, code commit records are categorized according to their degree of uncertainty. The deterministic probability of task execution under various degrees of uncertainty is identified, and a first implicit skill tag is established, including: For the task description change records, the degree of uncertainty is identified, and a task uncertainty index is generated for each group of records; Based on the task uncertainty index, the code submission records are clustered to obtain deterministic code submission records and multiple types of code submission records corresponding to multiple uncertainty clustering indices. The quality and speed differences of code pairs of equal difficulty are analyzed and extracted from the deterministic code submission records and the multi-type code submission records to obtain deterministic probabilities under various degrees of uncertainty, and the first implicit skill tag is generated.

5. The data engineering task allocation method based on multi-dimensional dynamic capability prediction as described in claim 4, characterized in that, The commit quality and speed differences of code pairs of equal difficulty are analyzed and extracted from the deterministic code commit records and the multi-type code commit records to obtain deterministic probabilities under various degrees of uncertainty, including: The code difficulty is evaluated for each code record in the deterministic code submission record and the multi-type code submission record to obtain the difficulty evaluation result. Based on the difficulty evaluation results, the code records with the difficulty difference of the extraction of the multiple types of code submission records are combined with the code records with the determination code submission records whose extraction difficulty difference is less than the preset difficulty difference, and multiple sets of comparison codes corresponding to multiple uncertainty clustering central indicators are established. For the multiple sets of comparison codes, the differences in submission quality and speed between uncertain codes and deterministic codes are analyzed to obtain deterministic probabilities under various degrees of uncertainty, wherein the deterministic probability is inversely proportional to the degree of difference.

6. The data engineering task allocation method based on multi-dimensional dynamic capability prediction as described in claim 1, characterized in that, The semantic field of the data engineering tasks to be assigned is constructed, including extracting explicit task requirements and quantifying the uncertainty of those requirements, including: The task description of the data engineering task is analyzed using natural language processing technology to extract explicit technical requirements and explicit business objectives. Identify ambiguities in the task description using a counterfactual reasoning engine; Historical technical requirements are retrieved based on the explicit business objectives, and the explicit technical requirements are optimized to generate target explicit requirements. Historical task execution records are retrieved based on the fuzzy attributes of the fuzzy points, and the uncertainty probability distribution of task requirements is calculated based on the historical task execution records. The semantic field is constructed using the explicit target requirement and the uncertainty probability distribution.

7. The data engineering task allocation method based on multi-dimensional dynamic capability prediction as described in claim 6, characterized in that, The explicit and implicit skill data of each data engineer in the explicit and implicit skill database are matched with the semantic field using neural semantic coupling to filter the candidate engineer set, including: The explicit requirements of the target in the semantic field are matched with the explicit skill tags in the explicit and implicit skill data of each data engineer to generate a set of explicitly matched engineers. The uncertainty probability distribution in the semantic field is matched with the implicit skill tags in the explicit and implicit skill data of each data engineer to generate a set of implicitly matched engineers. The candidate engineer set is generated by taking the intersection of the explicit matching engineer set and the implicit matching engineer set.

8. The data engineering task allocation method based on multi-dimensional dynamic capability prediction as described in claim 7, characterized in that, The explicit matching engineer set includes data engineers whose explicit skill tags fully cover the explicit requirements of the target.

9. The data engineering task allocation method based on multi-dimensional dynamic capability prediction as described in claim 7, characterized in that, The uncertainty probability distribution in the semantic field is matched with the implicit skill tags in the explicit and implicit skill data of each data engineer to generate a set of implicitly matched engineers, including: Based on the probability distribution of uncertainty, the degree of uncertainty of the target is extracted when the probability threshold is greater than a preset threshold, and the degree of uncertainty of the target is matched with the probability of certainty in the implicit skill tags of each data engineer. The implicit matching engineer set is formed by selecting data engineers whose probability of matching certainty is greater than a preset probability threshold.

10. A data engineering task allocation system based on multi-dimensional dynamic capability prediction, characterized in that, The system is used to implement the data engineering task allocation method based on multi-dimensional dynamic capability prediction as described in any one of claims 1-9, the system comprising: The explicit and implicit skills database construction module is used to collect all the historical task records of data engineers, perform ability prediction modeling of explicit and implicit skills, and establish an explicit and implicit skills database. The semantic field construction and execution module is used to construct semantic fields for the data engineering tasks to be assigned, including extracting explicit requirements of the tasks and quantifying the uncertainty of the requirements; The candidate engineer set acquisition module is used to perform neural semantic coupling matching between the explicit and implicit skill data of each data engineer in the explicit and implicit skill database and the semantic field to filter the candidate engineer set. The final engineer selection module performs workload-based selection based on the candidate engineer set, and selects the final selected engineers as the assigners of the data engineering tasks.