Data classification method based on dynamic tasks
By building a task tree and a classification model to work collaboratively with the agent, the data classification rules are dynamically adjusted, and the problem of single data classification methods in the existing technology is solved, and efficient and flexible data processing and rule optimization are achieved.
Patent Information
- Application Number
- CN202510386618.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing data classification method is single and cannot dynamically optimize the classification rules, which makes it impossible to meet actual needs when facing data changes or new application scenarios, and consumes a lot of time and manpower.
By building a task tree, using the classification model and agent to work together, dynamically adjust the task distribution strategy, call appropriate data classification tools, realize data attribute alignment and classification, support the addition of new tasks and modification of existing tasks, and optimize the rule base.
While ensuring the quality of data classification, it saves time and labor costs, improves data processing efficiency, and quickly adapts to data changes and new application scenarios.
Smart Images

Figure CN120372378A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a data classification method based on dynamic tasks. Background Art
[0002] With the rapid development of information technology, the scale and complexity of data have increased exponentially, and data classification technology has emerged and evolved continuously. With the development of technologies such as artificial intelligence and big data, data classification also plays a key role in data mining, machine learning, knowledge graph construction, etc., providing a basis for extracting valuable information from massive data and building intelligent applications.
[0003] Many existing data classification methods rely on manually predefined rules and standards. Once these rules are formulated, it is difficult to flexibly adjust them in the face of data changes or new application scenarios. For example, in some traditional industries, data classification standards are formulated based on industry experience and historical data. When new business models appear or data information changes, the original classification standards cannot be effectively adapted. This results in the inability to obtain classification results that meet actual needs when dealing with dynamically changing data; moreover, adding and modifying rules, as well as validating the rules, all require a large amount of time and manpower.
[0004] Although some methods use deep learning networks or large language models for data classification, they do not consider that the data characteristics and class distributions of different classification tasks are different. Using the same classification method, these differences cannot be fully exploited and adapted, resulting in poor classification effects. Moreover, different classification tasks may have specific business requirements and goals. Using a unified method may not be able to optimize for these specific task requirements, thus unable to meet the requirements of actual application scenarios. Summary of the Invention
[0005] In view of the above analysis, embodiments of the present invention aim to provide a data classification method based on dynamic tasks to solve the problems of poor classification effects caused by the single existing data classification method and the inability to dynamically optimize classification rules.
[0006] Embodiments of the present invention provide a data classification method based on dynamic tasks, including the following steps:
[0007] S1. Obtain multiple tasks from the rule library according to the received task identifier to construct a task tree, and display it on the interface end;
[0008] S2. After adjusting the tasks in the task tree on the interface end, pass them into the classification large model, select the root task as the to-do task and distribute it to the classification intelligent agent; the classification large model and the classification intelligent agent are large model applications based on LangChain.
[0009] S3. The classification agent identifies whether the to-do task has been aligned with attributes. For the to-do task with aligned attributes, it calls the data classification tool to execute the to-do task and feeds back the execution result to the classification large model. The classification large model selects new to-do tasks from the task tree according to the execution result and distributes them to the classification agent, repeating step S3 until there are no new to-do tasks, thus completing data classification.
[0010] As a further improvement based on the above method, the method further includes: for the to-do task with unaligned data attributes, feeding back the abnormal execution result to the classification large model; the classification large model saves the newly added and modified to-do tasks with normal execution results to the rule library.
[0011] As a further improvement based on the above method, obtaining multiple tasks from the rule library according to the received task identifier is to take the task corresponding to the task identifier as the root task, and traversing layer by layer from the rule library according to the subtask identifiers of the root task to retrieve all subtasks until the subtask identifier is empty.
[0012] As a further improvement based on the above method, the to-do task includes: task identifier, parent task identifier, subtask identifier, alignment identifier, task name, data source, data description, task objective, task description, task requirements, task parameters, tool model, and tool hyperparameters; the task parameters include a similarity threshold and a sample sampling quantity.
[0013] As a further improvement based on the above method, the classification agent identifying whether the to-do task has been aligned with data attributes includes:
[0014] Identifying the type of the to-do task according to the data source, task identifier, and alignment identifier of the to-do task. If it is a newly added and modified structured data classification task, it is regarded as a task to be aligned; otherwise, the to-do task has been aligned with data attributes;
[0015] Obtaining the metadata information of each field in the data table according to the data source of the task to be aligned; comparing each attribute in the data description of the task to be aligned with the metadata information of each field, and screening out the attributes that do not exist in the metadata information as the attributes to be aligned;
[0016] Obtaining the similarity threshold from the task parameters of the task to be aligned, and identifying whether fields can be selected from each field based on the similarity threshold. If so, after replacing the attributes to be aligned with the selected fields, the to-do task has been aligned with data attributes; otherwise, the to-do task has not been aligned with data attributes.
[0017] Based on further improvements to the above method, it is determined whether a field can be selected from each field based on a similarity threshold, including: respectively obtaining the semantic vectors of the attribute to be aligned and each field, and if there is a field whose similarity to the semantic vector of the attribute to be aligned exceeds the similarity threshold, it is placed in the candidate set, and the field corresponding to the maximum similarity in the candidate set is selected to replace the attribute to be aligned; otherwise, the similarity threshold is gradually decreased proportionally but not less than the minimum similarity threshold, and the fields greater than the decreased similarity threshold are placed in the candidate set, and the field corresponding to the maximum similarity in the candidate set is selected to replace the attribute to be aligned; when the maximum similarity is less than the minimum similarity threshold, a field cannot be selected.
[0018] Based on further improvements to the above method, for the to-do tasks of the already aligned data attributes, a data classification tool is called to execute the to-do tasks, including:
[0019] Using the data extraction tool bound to the second LLM model corresponding to the classification agent, obtain the data to be classified from the data source of the to-do task;
[0020] When the tool model and tool hyperparameters in the to-do task are not empty, extract the corresponding tool model from the data classification tools bound to the second LLM model, obtain the hyperparameters from the tool hyperparameters and set them to the tool model, and the second LLM model calls the tool model to perform the data classification task on the data to be classified according to the task objective in the to-do task to obtain the execution result; otherwise, construct a classification prompt for the to-do task, and the second LLM model performs the data classification task according to the classification prompt to obtain the execution result.
[0021] Based on further improvements to the above method, obtaining the data to be classified from the data source of the to-do task includes: when the sample sampling quantity in the task parameters of the to-do task is not empty, obtain the data to be classified according to the sample sampling quantity, and when the execution result is normal, select all the data to be classified and execute again.
[0022] Based on further improvements to the above method, when the execution result is abnormal and the number of fields in the candidate set is greater than 1, select the field with the second largest similarity in the candidate set to replace the attribute to be aligned, and call the data classification tool again to execute the to-do task to obtain the final execution result.
[0023] Based on further improvements to the above method, the classification large model selects a new to-do task from the task tree according to the execution result, including:
[0024] When the execution result is normal, take the next task from the task tree in breadth-first order as the new to-do task;
[0025] When the execution result is a timeout and the number of executions of the current to-do task does not exceed the threshold, the current to-do task is used as the new to-do task again;
[0026] When the execution result is an exception, or the number of executions of the current to-do task exceeds the threshold, the current to-do task and all subtasks with the current to-do task as the parent task will no longer be executed, and the next task will be retrieved from the task tree in a breadth-first manner as the new to-do task.
[0027] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:
[0028] 1. Utilize the classification large model to distribute the to-do tasks submitted by the interface end. During the iterative execution process, use the classification agent to call the appropriate data classification tool to complete data classification, and feedback the execution result to the classification large model. The classification large model dynamically determines the distribution strategy according to the execution result, and the two parties cooperate to complete it, saving time cost and labor cost while ensuring the data classification quality and improving the data processing efficiency.
[0029] 2. Support the addition of new tasks and the modification of existing tasks during classification, realizing the automatic expansion and optimization of task rules in the rule library, and quickly adapting to data changes and new application scenarios.
[0030] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent specification, and some advantages can be made obvious from the specification, or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings are only for the purpose of showing specific embodiments, and are not considered as limiting the present invention. Throughout the drawings, the same reference signs denote the same components;
[0032] Figure 1 It is a flowchart of a data classification method based on dynamic tasks in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings, where the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, and are not used to limit the scope of the present invention.
[0034] A specific embodiment of the present invention discloses a data classification method based on dynamic tasks, which is the basis for an enterprise to perform data asset management and build intelligent applications. As Figure 1 shown, it includes the following steps:
[0035] S1. Obtain multiple tasks from the rule library according to the received task identifier to construct a task tree, and display it on the interface end;
[0036] S2. After adjusting the tasks in the task tree at the interface end, input them into the classification large model, select the root task as the to-do task and distribute it to the classification intelligent agent; the classification large model and the classification intelligent agent are large model applications built based on LangChain.
[0037] S3. The classification intelligent agent identifies whether the to-do task has been aligned with the data attributes. For the to-do task with aligned data attributes, call the data classification tool to execute the to-do task, and feedback the execution result to the classification large model; the classification large model selects a new to-do task from the task tree according to the execution result and distributes it to the classification intelligent agent, and repeat step S3 until there is no new to-do task, and the data classification is completed.
[0038] During implementation, use the classification large model to distribute the to-do tasks submitted at the interface end. During the iterative execution process, use the classification intelligent agent to call the appropriate data classification tool to complete the data classification, and feedback the execution result to the classification large model. The classification large model dynamically determines the distribution strategy according to the execution result, and the two parties cooperate to complete it, which not only ensures the quality of data classification, but also saves time cost and labor cost, and improves the data processing efficiency.
[0039] It should be noted that the rule library in step S1 includes multiple classification task rules, referred to as tasks for short; each task includes: task identifier, parent task identifier, sub-task identifier, alignment identifier, task name, data source, data description, task objective, task description, task requirements, task parameters, tool model, and tool hyperparameters; the task identifier is automatically generated when saved to the rule library, and the parent task identifier and sub-task identifier are automatically associated according to the hierarchical relationship of the tasks; the alignment identifier is used to indicate whether the data attributes need to be aligned before classification; the data description is used to describe the data information from multiple dimensions such as attribute name, Chinese name of the attribute, attribute type, length, or precision; the task objective is used to describe the basis and objective of data classification; the task requirements are used to indicate the special processing requirements of this task, for example, the date format needs to be unified, or there are no special requirements; the task parameters are used to indicate the parameters required by the classification intelligent agent during the task execution process, including but not limited to: similarity threshold and sample sampling quantity; the tool model is used to indicate the tool model used for processing this task, and the tool model comes from the tool library. The tool types in the tool library include: general tools and data classification tools; general tools include but not limited to: data extraction tools, data vectorization tools, and image processing tools; data classification tools include but not limited to: text-CNN classification model, text-SVM classification model, text-decision tree classification model, and image-CNN classification model; the tool hyperparameters are the paths of the configuration files used to indicate the hyperparameters required by the tool model.
[0040] Exemplarily, the task of "department employee classification" in the rule library is shown in Table 1.
[0041] Task of Table 1 "Department Employee Classification"
[0042]
[0043] It should be noted that before data classification, data information from different data sources is collected, including structured data and unstructured data. For example, structured data in the telecommunications industry includes master data, transaction data, and analytical data, and unstructured data includes text data and multimedia data.
[0044] Furthermore, the collected data is preprocessed, including: extraction of data attributes and data cleaning. Among them, the extraction of data attributes is to obtain the data attributes of each data table and the association relationships between data tables for structured data; data cleaning includes: default value processing, outlier processing, data format conversion, etc.
[0045] Furthermore, the data attributes are vectorized and stored in a vector database. For example, the Word2Vec model is used to obtain the vector embedding representation of data attributes and store it in the Chroma or Pinecone vector database for convenient subsequent retrieval and analysis.
[0046] The task identifier received in step S1 can be transmitted through an internal interface or method, or it can come from the interface end; for example, tasks in the rule library are displayed on the interface end, or the corresponding task identifier is obtained by retrieving tasks in the rule library through the keyword of the task name. Obtaining multiple tasks from the rule library according to the received task identifier is based on the task corresponding to the task identifier as the root task, and all subtasks are retrieved layer by layer from the rule library according to the subtask identifiers of the root task until the subtask identifier is empty.
[0047] Furthermore, the constructed task tree is displayed on the interface end, and detailed information of each task node on the task tree is obtained from the rule library according to its corresponding task identifier; adjusting tasks in the task tree on the interface end includes: adding new tasks and detailed information of tasks, modifying and deleting existing tasks / or detailed information of existing tasks.
[0048] Tasks without a task identifier are newly added tasks. If the data sources of the newly added tasks and the modified tasks are structured database tables, these tasks are structured data classification tasks, and an alignment identifier of 1 is set for these tasks, indicating that data attributes need to be aligned before data classification to avoid inconsistency with the data attributes in the data source.
[0049] It can be understood that if the database table is upgraded on a large scale, an alignment identifier of 1 is set for all tasks in the task tree.
[0050] In step S2, the adjusted task tree is used as the task tree to be processed corresponding to the current operator and is passed into the classification large model. Initially, the root task is taken from the task tree as the to-do task and distributed to the classification agent.
[0051] It should be noted that the classification large model and the classification agent are large model applications based on LangChain that include a series of processing processes. Among them, the first LLM (Large Language Model) model is used in the classification large model, and the second LLM model is used in the classification agent. The first / second LLM model includes, but is not limited to: GPT series models, LLaMA models, and Bert models.
[0052] In step S3, the classification agent identifies whether the to-do task has been aligned with the data attributes, including:
[0053] Identify the type of the to-do task according to the data source, task identifier, and alignment identifier of the to-do task. If it is a new or modified structured data classification task, it is used as the task to be aligned; otherwise, the to-do task has been aligned with the data attributes;
[0054] Obtain the metadata information of each field in the data table according to the data source of the task to be aligned; compare each attribute in the data description of the task to be aligned with the metadata information of each field, and filter out the attributes that do not exist in the metadata information as the attributes to be aligned; exemplarily, perform a preliminary filter according to the attribute name.
[0055] Obtain the similarity threshold from the task parameters of the task to be aligned, and identify whether a field can be selected from each field based on the similarity threshold. If so, after replacing the attribute to be aligned with the selected field, the to-do task has been aligned with the data attributes; otherwise, the to-do task has not been aligned with the data attributes.
[0056] Furthermore, identifying whether a field can be selected from each field based on the similarity threshold includes:
[0057] Obtain the semantic vectors of the attribute to be aligned and each field respectively. If there is a field whose similarity to the semantic vector of the attribute to be aligned exceeds the similarity threshold, it is put into the candidate set, and the field corresponding to the maximum similarity in the candidate set is selected to replace the attribute to be aligned; otherwise, gradually reduce the similarity threshold proportionally but not less than the minimum similarity threshold, put the fields greater than the reduced similarity threshold into the candidate set, and select the field corresponding to the maximum similarity in the candidate set to replace the attribute to be aligned; when the maximum similarity is less than the minimum similarity threshold, no field can be selected.
[0058] Exemplarily, obtain the semantic vector of the attribute to be aligned according to the attribute name and the Chinese description of the attribute, obtain the semantic vector of the field according to the field name and the field description, and calculate the cosine similarity between the two semantic vectors.
[0059] Furthermore, for the to-do tasks of the classification agent for unaligned data attributes, the classification agent feeds back the abnormal execution results to the classification large model; for the to-do tasks of aligned data attributes, the classification agent calls the data classification tool to execute the to-do tasks, including:
[0060] ① Use the data extraction tool bound to the second LLM model corresponding to the classification agent to obtain the data to be classified from the data source of the to-do task.
[0061] It should be noted that the data extraction tool is a general tool in the tool library and can be bound to multiple LLM models. The data extraction tool includes but is not limited to: connecting to a relational database and obtaining table data; accessing the file system and obtaining file data.
[0062] Obtaining the data to be classified from the data source of the to-do task includes: when the sample sampling quantity in the task parameters of the to-do task is not empty, obtaining the data to be classified according to the sample sampling quantity, and when the execution result is normal, selecting all the data to be classified and re-executing.
[0063] ② When the tool model and tool hyperparameters in the to-do task are not empty, extract the corresponding tool model from the data classification tools bound to the second LLM model, obtain the hyperparameters from the tool hyperparameters and set them to the tool model, and the second LLM model calls the tool model to perform the data classification task on the data to be classified according to the task objective in the to-do task to obtain the execution result; otherwise, construct a classification prompt for the to-do task, and the second LLM model performs the data classification task according to the classification prompt to obtain the execution result.
[0064] It should be noted that in the classification prompt, the role of the second LLM model is defined as a data classification expert, the data description in the to-do task is used as the context of the data to be classified; the task objective in the to-do task is used as the objective of the second LLM model; the task requirements in the to-do task are used as the classification constraints of the second LLM model, and it is required that the second LLM model automatically selects a suitable tool model from the bound data classification tools for classification by analyzing the data to be classified and the task objective, and organizes the execution result in a set format.
[0065] Preferably, considering that when aligning data attributes, in some cases, the similarities of multiple fields are all greater than the similarity threshold and the differences in similarities are relatively small. At this time, the number of candidates is greater than 1. First, preferentially select the attribute with the largest similarity to replace the attribute to be aligned. The data extraction tool samples data from the data source according to the sample sampling quantity in the task parameters for data classification. If the execution result is abnormal, the classification agent selects the field with the second largest similarity from the candidate set to replace the attribute to be aligned, repeats the above process, and calls the data classification tool again to execute the to-do task to obtain the final execution result.
[0066] The classification agent feeds back the execution result to the classification large model; the classification large model selects new to-do tasks from the task tree according to the execution result, including:
[0067] When the execution result is normal, the next task is taken from the task tree in breadth-first order as the new to-do task;
[0068] When the execution result is a timeout and the execution count of the current to-do task does not exceed the threshold, the current to-do task is again used as the new to-do task;
[0069] When the execution result is abnormal, or the execution count of the current to-do task exceeds the threshold, the current to-do task and all its subtasks with the current to-do task as the parent task are no longer executed, and the next task is taken from the task tree in breadth-first order as the new to-do task.
[0070] Preferably, the execution result fed back by the classification agent to the classification large model also includes: the detailed information of the to-do task and the classification result. For the execution result without anomalies, the classification large model constructs a prompt word according to the detailed information of the to-do task, sets a series of rules and constraints in the prompt word for secondary reasoning, and verifies whether the result of the secondary reasoning matches the classification result. If not, new to-do tasks are selected according to the abnormal execution result.
[0071] The classification large model distributes the selected new to-do tasks to the classification agent, and repeats step S3 until there are no new to-do tasks, completing data classification.
[0072] Finally, the classification large model saves the newly added and modified to-do tasks with normal execution results to the rule library, including: generating a task identifier for the newly added to-do task, and automatically filling in the parent task identifier and subtask identifier according to its hierarchical relationship in the task tree; overwriting the existing information of the task in the rule library with the information of the modified to-do task, such as the reduced similarity threshold.
[0073] It should be noted that the classification results are all part of the data assets, and like the originally collected data, they are vectorized and stored to facilitate the retrieval and management of the data assets, forming an important data list and a sensitive data asset map.
[0074] Preferably, in order to improve the accuracy of the data assets, the classification large model is periodically used to execute the tasks in the rule library, update the data in each category in a timely manner, and continuously optimize the rule library.
[0075] Compared with the prior art, a data classification method based on dynamic tasks provided in this embodiment utilizes a classification large model to distribute the to-do tasks submitted by the interface end. During the iterative execution process, a classification agent is used to call appropriate data classification tools to complete data classification, and the execution results are fed back to the classification large model. The classification large model dynamically determines the distribution strategy according to the execution results, and the two parties cooperate to complete it. While ensuring the quality of data classification, time costs and labor costs are saved, and the data processing efficiency is improved. When classifying, it supports the addition of new tasks and the modification of existing tasks, realizes the automatic expansion and optimization of task rules in the rule library, and quickly adapts to data changes and new application scenarios.
[0076] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disk, a read-only memory, or a random access memory, etc.
[0077] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A data classification method based on dynamic tasks, characterized in that, It includes the following steps: S1. Obtain multiple tasks from the rule base according to the received task identifier, construct them into a task tree, and display it on the interface side; S2. After adjusting the tasks in the task tree on the interface side, input them into the classification large model, select the root task as the to-do task, and distribute it to the classification intelligent agent; the classification large model and the classification intelligent agent are large model applications based on LangChain; S3. The classification intelligent agent identifies whether the to-do task has been aligned with the data attributes. For the to-do task with aligned data attributes, call the data classification tool to execute the to-do task, and feedback the execution result to the classification large model; the classification large model selects a new to-do task from the task tree according to the execution result and distributes it to the classification intelligent agent, and repeat step S3 until there is no new to-do task, and the data classification is completed.
2. The data classification method based on dynamic tasks according to claim 1, wherein The method further includes: for the to-do task with unaligned data attributes, feedback the abnormal execution result to the classification large model; the classification large model saves the newly added and modified to-do tasks with normal execution results to the rule base.
3. The data classification method based on dynamic tasks according to claim 1, wherein The obtaining of multiple tasks from the rule base according to the received task identifier is to take the task corresponding to the task identifier as the root task, and traverse and extract all sub-tasks from the rule base layer by layer according to the sub-task identifier of the root task until the sub-task identifier is empty.
4. The data classification method based on dynamic tasks according to claim 1, characterized in that The to-do task includes: task identifier, parent task identifier, sub-task identifier, alignment identifier, task name, data source, data description, task objective, task description, task requirements, task parameters, tool model, and tool hyperparameters; the task parameters include a similarity threshold and a sample sampling quantity.
5. The data classification method based on dynamic tasks according to claim 4, characterized in that The classification intelligent agent identifies whether the to-do task has been aligned with the data attributes, including: Identifying the type of the to-do task according to the data source, task identifier, and alignment identifier of the to-do task. If it is a new or modified structured data classification task, it is regarded as a task to be aligned; otherwise, the to-do task has been aligned with the data attributes; Obtaining the metadata information of each field in the data table according to the data source of the task to be aligned; comparing each attribute in the data description of the task to be aligned with the metadata information of each field, and screening out the attributes that do not exist in the metadata information as the attributes to be aligned; Obtaining the similarity threshold from the task parameters of the task to be aligned, and identifying whether a field can be selected from each field based on the similarity threshold. If so, after replacing the attribute to be aligned with the selected field, the to-do task has been aligned with the data attributes; otherwise, the to-do task has not been aligned with the data attributes.
6. The data classification method based on dynamic tasks according to claim 5, wherein, The identifying whether a field can be selected from each field based on the similarity threshold includes: respectively obtaining the semantic vectors of the attribute to be aligned and each field. If there is a field whose similarity with the semantic vector of the attribute to be aligned exceeds the similarity threshold, put it into the candidate set, and select the field corresponding to the maximum similarity from the candidate set to replace the attribute to be aligned; otherwise, gradually reduce the similarity threshold proportionally but not less than the minimum similarity threshold, put the fields greater than the reduced similarity threshold into the candidate set, and select the field corresponding to the maximum similarity from the candidate set to replace the attribute to be aligned; when the maximum similarity is less than the minimum similarity threshold, a field cannot be selected.
7. The data classification method based on dynamic tasks according to claim 4, wherein Invoking a data classification tool for the to-do task of the aligned data attribute to execute the to-do task includes: Using a data extraction tool bound to the second LLM model corresponding to the classification agent to obtain the data to be classified from the data source of the to-do task; When the tool model and tool hyperparameters in the to-do task are not empty, extract the corresponding tool model from the data classification tool bound to the second LLM model, obtain the hyperparameter settings from the tool hyperparameters and set them for the tool model, and the second LLM model invokes the tool model to perform a data classification task on the data to be classified according to the task objective in the to-do task to obtain an execution result; otherwise, construct a classification prompt word for the to-do task, and the second LLM model performs a data classification task according to the classification prompt word to obtain an execution result.
8. The data classification method based on dynamic tasks according to claim 7, wherein The obtaining the data to be classified from the data source of the to-do task includes: when the sample sampling quantity in the task parameters of the to-do task is not empty, obtain the data to be classified according to the sample sampling quantity, and when the execution result is normal, select all the data to be classified and execute again.
9. The data classification method based on dynamic tasks according to claim 7, wherein When the execution result is abnormal and the number of fields in the candidate set is greater than 1, select the field with the second largest similarity from the candidate set to replace the attribute to be aligned, and call the data classification tool again to execute the to-do task to obtain the final execution result.
10. The data classification method based on dynamic tasks according to claim 1 or 7, characterized in that, The classification large model selects a new to-do task from the task tree according to the execution result, including: When the execution result is normal, take out the next task from the task tree in a breadth-first manner as the new to-do task; When the execution result is a timeout and the number of executions of the current to-do task does not exceed the threshold, the current to-do task is used as the new to-do task again; When the execution result is abnormal, or the number of executions of the current to-do task exceeds the threshold, the current to-do task and all its subtasks with the current to-do task as the parent task are no longer executed, and the next task is taken out from the task tree in a breadth-first manner as the new to-do task.
Citation Information
Patent Citations
Dynamic structured data classification method, device and system and storage medium
CN119128146A
Data classification processing dynamic optimization method suitable for OA system and related products
CN119416004A
Information processing method, device, equipment and storage medium based on large language model
US20250005018A1
Automatic database enrichment and curation using large language models
US20250045256A1
Knowledge fusion method and apparatus based on data relationship analysis, and computer device and storage medium
WO2021051630A1